Neighboring edge requantization for trisoup
The TriSoup scheme with dynamic OBUF entropy coding addresses the data size challenge of point clouds by achieving efficient compression and decoding, suitable for applications in augmented reality and autonomous driving.
Patent Information
- Application Number
- PCT/US2025/012056
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
The large data size of point clouds, comprising millions or billions of points with geometry and attribute information, poses challenges for efficient storage and transmission, necessitating effective compression techniques to reduce data volume while maintaining visual quality or ensuring lossless integrity.
The use of a TriSoup scheme for lossy compression, where TriSoup nodes are modeled with triangles and centroid residuals, combined with dynamic OBUF for entropy coding, to efficiently encode and decode point cloud data, achieving compression rates of 0.7 bits per point for dense point clouds.
This approach achieves high compression efficiency, reducing data size by up to 25% while maintaining visual fidelity and enabling practical deployment of point cloud technologies in applications like AR, VR, and autonomous driving.
Smart Images

Figure US2025012056_24072025_PF_FP_ABST
Abstract
Description
Docket No.: 24-2010PCT TITLE Neighboring Edge Requantization for TriSoup CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No.63 / 623,169, filed January 19, 2024, which is hereby incorporated by reference in its entirety. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Examples of several of the various embodiments of the present disclosure are described herein with reference to the drawings.
[0003] FIG.1 illustrates an exemplary point cloud coding / decoding system in which embodiments of the present disclosure may be implemented.
[0004] FIG.2 illustrates the Morton order of eight sub-cuboids split from a cuboid.
[0005] FIG.3 illustrates an example processing or scanning order for the first three level of an occupancy tree.
[0006] FIG.4 illustrates an example of already-coded occupancies of cuboids that may be used to code the occupancy of a current child cuboid.
[0007] FIG.5 illustrates an example of a dynamic reduction function DR that may be used in dynamic OBUF.
[0008] FIG.6 illustrates a flowchart of an example method for coding the occupancy (e.g., as indicated by a single bit) of a current child cuboid using dynamic OBUF.
[0009] FIG.7 illustrates an example of an occupied cube of size NxNxN (where N > 1) that corresponds to a TriSoup node of an occupancy tree.
[0010] FIG.8A illustrates an example cube corresponding to a TriSoup node with a number K of TriSoup vertices Vk.
[0011] FIG.8B illustrates an example refinement to the TriSoup model by coding a centroid residual vector Cresinto the bitstream such as to use C+Cres instead of C as pivoting vertex for the triangles.
[0012] FIG.8C illustrates an example of coding a centroid residual vector Cresin / from the bitstream such that an adjusted centroid C+Cres is used instead of centroid C for generating TriSoup triangles of a cuboid corresponding to a portion of a point cloud, according to some embodiments.
[0013] FIG.9A and FIG.9B illustrate examples of voxelization.
[0014] FIGS.10A-B illustrate 12 cuboids with volumes that intersect a current edge E being entropy coded.
[0015] FIGS.11A-C illustrate the at most five edges that may be used to entropy code a current edge E in at least one implementation.
[0016] FIGS.12A-C illustrate examples of neighboring already-coded edges of an edge.
[0017] FIGS.13A-C illustrates examples of neighboring already-coded edges of an edge.
[0018] FIGS.14A-C illustrate a set of 18 neighboring edges constituting a causal neighborhood of an edge.
[0019] FIGS.15A-B illustrates a quantizer when the TriSoup node size B as well as the size of the quantization steps are both powers of two.Docket No.: 24-2010PCT
[0020] FIG.16 illustrates an example encoding process of presence flag and position of TriSoup vertex of a current edge of a TriSoup node in a bitstream.
[0021] FIG.17 illustrates an example decoding process of encoded presence flag and position of TriSoup vertex of a current edge of a TriSoup node from a bitstream.
[0022] FIGS.18A-B illustrate a quantizing process, according to some embodiments.
[0023] FIGS.19A-C illustrate a quantizing process, according to some embodiments.
[0024] FIGS.20A-C illustrate examples of dequantizing distances determined between the middle position of the current edge and the centers of quantization intervals.
[0025] FIG.21 illustrates an example encoding process for encoding, in a bitstream, a presence flag and a current position of a TriSoup vertex on a current edge of a cuboid, corresponding to a TriSoup node, according to some embodiments.
[0026] FIG.22 illustrates an example decoding process for decoding, from a bitstream, an encoded presence flag and quantized position (e.g., current position) of a TriSoup vertex on a current edge of a cuboid, corresponding to a TriSoup node, according to some embodiments.
[0027] FIG.23 illustrates an example of obtaining TriSoup edge information of at least one neighboring already- coded edge, according to some embodiments.
[0028] FIGS.24A-B illustrate examples of re-quantizing a distance indicated by a neighborhood codeword, according to some embodiments.
[0029] FIG.25 illustrates an example encoding process for encoding, in a bitstream, a presence flag and a current position of a TriSoup vertex of a current edge of a cuboid containing a portion of a point cloud, according to some embodiments.
[0030] FIG.26 illustrates an example decoding process for decoding, from the bitstream, an encoded presence flag and a current position of a TriSoup vertex of a current edge of a cuboid containing a portion of a point cloud, according to some embodiments.
[0031] FIG.27 illustrates a flowchart of an example method for encoding, in a bitstream, a current position of a TriSoup vertex for a point cloud, according to some embodiments.
[0032] FIG.28 illustrates a flowchart of an example method for decoding, from a bitstream, a current position of a TriSoup vertex for a point cloud, according to some embodiments.
[0033] FIG.29 illustrates a flowchart of an example method for coding a current position of a TriSoup vertex on a current edge based on re-quantizing the quantized position of a neighboring TriSoup vertex on a neighboring already- coded edge of the current edge, according to some embodiments.
[0034] FIG.30 illustrates a block diagram of an example computer system in which embodiments of the present disclosure may be implemented.Docket No.: 24-2010PCT DETAILED DESCRIPTION
[0035] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the disclosure. However, it will be apparent to those skilled in the art that the disclosure, including structures, systems, and methods, may be practiced without these specific details. The description and representation herein are the common means used by those experienced or skilled in the art to most effectively convey the substance of their work to others skilled in the art. In other instances, well-known methods, procedures, components, and circuitry have not been described in detail to avoid unnecessarily obscuring aspects of the disclosure.
[0036] References in the specification to “one embodiment,” “an embodiment,” “an example embodiment,” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0037] Also, it is noted that individual embodiments may be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.
[0038] The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.Docket No.: 24-2010PCT
[0039] Furthermore, embodiments may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks.
[0040] Traditional visual data describes an object or scene using a series of points that each comprise a position in two dimensions (x and y) and one or more optional attributes like color. Volumetric visual data adds another positional dimension to this traditional visual data. Volumetric visual data describes an object or scene using a series of points that each comprise a position in three dimensions (x, y, and z) and one or more optional attributes like color, reflectance, time stamp, etc. Compared to traditional visual data, volumetric visual data may provide a more immersive way to experience visual data.
[0041] For example, an object or scene described by volumetric visual data may be viewed from any (or multiple) angles, whereas traditional visual data may generally only be viewed from the angle in which it was captured or rendered. Volumetric visual data may be used in many applications, including Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR). Sparse volumetric visual data may be used in the automotive industry for the representation of 3D maps (cartography) or as input to assisted driving systems. In the latter use case, volumetric visual data is typically input to driving decision algorithms. In another example, volumetric visual data may be used to store valuable objects in digital form. In applications for preserving cultural heritage, the goal is to keep a representation of objects that may be threatened by natural disasters. For example, statues, vases, and temples may be entirely scanned and stored as volumetric visual data having several billions of samples. This use case for volumetric visual data may be particularly relevant for valuable objects in locations where earthquakes, tsunamis, and typhoons are frequent. Volumetric visual data may be in the form of a volumetric frame that describes an object or scene captured at a particular time instance or in the form of a sequence of volumetric frames (referred to as a volumetric sequence or volumetric video) that describes an object or scene captured at multiple different time instances.
[0042] One format for storing volumetric visual data is point clouds. A point cloud comprises a collection of points in three-dimensional (3D) space. Each point in a point cloud may comprise geometry information that indicates the point’s position in 3D space. For example, the geometry information may indicate the point’s position in 3D space using three Cartesian coordinates (x, y, and z) or using spherical coordinates (r, phi, theta) (e.g., when acquired by a rotating sensor). The positions of points in a point cloud may be quantized according to a space precision, which may be the same or different in each dimension. The quantization process may create a grid in 3D space. One or more points residing within each sub-grid volume may be mapped to the sub-grid center coordinates, referred to as voxels. A voxel (also referred to as a volumetric pixel) may be considered as a 3D extension of pixels corresponding to the 2D image grid coordinates. For example, similar to a pixel being the smallest unit when dividing the 2D space (or 2D image) into discrete, uniform (e.g., equally sized) regions, a voxel may be the smallest unit of volume when dividing 3D space into discrete, uniform regions. The sub-grid center coordinates (which correspond to voxels) may be referred to as aDocket No.: 24-2010PCT voxelized grid. A point in a point cloud may further comprise one or more types of attribute information. Attribute information may indicate a property of a point’s visual appearance. For example, attribute information may indicate a texture (e.g., color) of the point, a material type of the point, transparency information of the point, reflectance information of the point, a normal vector to a surface of the point, a velocity at the point, an acceleration at the point, a time stamp indicating when the point was captured, or a modality indicating how the point was captured (e.g., running, walking, or flying). In another example, a point in a point cloud may comprise light field data in the form of multiple view- dependent texture information. Light field data may be another type of optional attribute information.
[0043] The points in a point cloud may describe an object or a scene. For example, the points in a point cloud may describe the external surface and / or the internal structure of an object or scene. The object or scene may be synthetically generated by a computer or may be generated from the capture of a real-world object or scene. The geometry information of a real-world object or scene may be obtained by 3D scanning and / or photogrammetry.3D scanning may include laser scanning, structured light scanning, and / or modulated light scanning.3D scanning may obtain geometry information by moving one or more laser heads, structured light cameras, and / or modulated light cameras relative to an object or scene being scanned. Photogrammetry may obtain geometry information by triangulating the same feature or point in different spatially shifted 2D photographs. Point cloud data may be in the form of a point cloud frame that describes an object or scene captured at a particular time instance or in the form of a sequence of point cloud frames (referred to as a point cloud sequence or point cloud video) that describes an object or scene captured at multiple different time instances.
[0044] The data size of a point cloud frame or sequence may be too large for storage and / or transmission in many applications. For example, a single point cloud may comprise over a million points or even billions of points, where each point may comprise geometry information and one or more optional types of attribute information. The geometry information of each point may comprise three Cartesian coordinates (x, y, and z) or spherical coordinates (r, phi, theta) that are each represented, for example, using at least 10 bits per component or 30 bits in total. The attribute information of each point may comprise a texture corresponding to three color components (e.g., R, G, and B color components) that are each represented, for example, using 8-10 bits per component or 24-30 bits in total. A single point therefore comprises at least 54 bits of information in this example, with at least 30 bits of geometry information and at least 24 bits of texture. If a point cloud frame includes a million such points, each point cloud frame would require 54 million bits or 54 megabits to represent. In case of dynamic point clouds that change over time, at a frame rate of 30 frames per second, a data rate of 1.62 gigabits per second would be required to transmit the points of the point cloud sequence. Therefore, raw representations of point clouds may require a large amount of data and the practical deployment of point-cloud-based technologies may need compression technologies that enable the storage and distribution of point clouds with reasonable cost.
[0045] Encoding may be used to compress and / or reduce the data size of a point cloud frame or sequence to provide for more efficient storage and / or transmission. Decoding may be used to decompress a compressed point cloud frameDocket No.: 24-2010PCT or sequence for display and / or other forms of consumption (e.g., by a machine learning-based device, neural network- based device, artificial intelligence-based device, or other forms of consumption by other types of machine-based processing algorithms and / or devices). Compression of point clouds may be lossy (introducing differences relative to the original data) for the distribution to and visualization by an end-user, for example, on AR or VR glasses or any other 3D-capable device. Lossy compression may allow for a high ratio of compression but may imply a trade-off between compression and visual quality perceived by an end-user. Other frameworks, like medical applications or autonomous driving, may require lossless compression to avoid altering the results of a decision obtained based on the analysis of the transmitted and decompressed point cloud frame.
[0046] FIG.1 illustrates an exemplary point cloud coding system 100 in which embodiments of the present disclosure may be implemented. Point cloud coding system 100 comprises a source device 102, a transmission medium 104, and a destination device 106. Source device 102 encodes a point cloud sequence 108 into a bitstream 110 for more efficient storage and / or transmission. Source device 102 may store and / or transmit bitstream 110 to destination device 106 via transmission medium 104. Destination device 106 decodes bitstream 110 to display point cloud sequence 108 or for other forms of consumption. Destination device 106 may receive bitstream 110 from source device 102 via a storage medium or transmission medium 104. Source device 102 and destination device 106 may be any one of a number of different devices, including a cluster of interconnected computer systems acting as a pool of seamless resources (also referred to as a cloud of computers or cloud computer), a server, a desktop computer, a laptop computer, a tablet computer, a smart phone, a wearable device, a television, a camera, a video gaming console, a set- top box, a video streaming device, an autonomous vehicle, or a head mounted display. A head mounted display may allow a user to view a VR, AR, or MR scene and adjust the view of the scene based on movement of the user’s head. A head mounted display may be tethered to a processing device (e.g., a server, desktop computer, set-top box, or video gaming counsel) or may be fully self-contained.
[0047] To encode point cloud sequence 108 into bitstream 110, source device 102 may comprise a point cloud source 112, an encoder 114, and an output interface 116. Point cloud source 112 may provide or generate point cloud sequence 108 from a capture of a natural scene and / or a synthetically generated scene. A synthetically generated scene may be a scene comprising computer generated graphics. Point cloud source 112 may comprise one or more point cloud capture devices (e.g., one or more laser scanning devices, structured light scanning devices, modulated light scanning devices, and / or passive scanning devices), a point cloud archive comprising previously captured natural scenes and / or synthetically generated scenes, a point cloud feed interface to receive captured natural scenes and / or synthetically generated scenes from a point cloud content provider, and / or a processor to generate synthetic point cloud scenes.
[0048] As shown in FIG.1, a point cloud sequence 108 may comprise a series of point cloud frames 124. A point cloud frame may describe an object or scene captured at a particular time instance. Point cloud sequence 108 may achieve the impression of motion when a constant or variable time is used to successively present point cloud framesDocket No.: 24-2010PCT 124 of point cloud sequence 108. A point cloud frame may comprise a collection of points 126 in 3D space. Each of points 126 may comprise geometry information that indicates the point’s position in 3D space. For example, the geometry information may indicate the point’s position in 3D space using three Cartesian coordinates (x, y, and z). One or more of points 126 may further comprise one or more types of attribute information. Attribute information may indicate a property of a point’s visual appearance. For example, attribute information may indicate a texture (e.g., color) of a point, a material type of a point, transparency information of a point, reflectance information of a point, a normal vector to a surface of a point, a velocity at a point, an acceleration at a point, a time stamp indicating when a point was captured, a modality indicating how a point was captured (e.g., running, walking, or flying). In another example, one or more of points 126 may comprise light field data in the form of multiple view-dependent texture information. Light field data may be another type of optional attribute information. Color attribute information of one or more of points 126 may comprise a luminance value and two chrominance values. The luminance value may represent the brightness (or luma component, Y) of the point. The chrominance values may respectively represent the blue and red components of the point (or chroma components, Cb and Cr) separate from the brightness. Other color attribute values are possible based on different color schemes (e.g., an RGB or monochrome color scheme).
[0049] Encoder 114 may encode point cloud sequence 108 into bitstream 110. To encode point cloud sequence 108, encoder 114 may apply one or more lossy compression techniques and / or prediction techniques to reduce redundant information in point cloud sequence 108. Redundant information is information that may be predicted at a decoder and therefore may not be needed to be transmitted to the decoder for accurate decoding of point cloud sequence 108. For example, Motion Picture Expert Group (MPEG) introduced a geometry-based point cloud compression (G-PCC) standard (ISO / IEC standard 23090-9: Geometry-based point cloud compression). G-PCC specifies the encoded bitstream syntax and semantics for transmission and / or storage of a compressed point cloud frame and the decoder operation for reconstructing the compressed point cloud frame from the bitstream. During standardization of G-PCC, a reference software (ISO / IEC standard 23090-21: Reference Software for G-PCC) was developed to encode the geometry and attribute information of a point cloud frame. To encode geometry information of a point cloud frame, the G-PCC reference software encoder may perform voxelization by quantizing positions of points in a point cloud, which creates a grid in 3D space. The G-PCC reference software encoder may map the points to the center coordinates of the sub-grid volume (or voxel) that their quantized locations reside. The G-PCC reference software encoder may perform geometry analysis using an occupancy tree to compress the geometry information. The G-PCC reference software encoder may entropy encode the result of the geometry analysis to further compress the geometry information. To encode attribute information of a point cloud, the G-PCC reference software encoder may apply a transform tool, such as Region Adaptive Hierarchical Transform (RAHT), the Predicting Transform, and / or the Lifting Transform. The Lifting Transform may be built on top of the Predicting Transform but with an extra update / lifting step. Consequently, these two transforms may be referred to as Predicting / Lifting Transform or pred lift. Encoder 114 may operate in a same or similar manner to an encoder provided by the G-PCC reference software.Docket No.: 24-2010PCT
[0050] Output interface 116 may be configured to write and / or store bitstream 110 onto transmission medium 104 for transmission to destination device 106. In addition, or alternatively, output interface 116 may be configured to transmit, upload, and / or stream bitstream 110 to destination device 106 via transmission medium 104. Output interface 116 may comprise a wired and / or wireless transmitter configured to transmit, upload, and / or stream bitstream 110 according to one or more proprietary and / or standardized communication protocols, such as Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, 3rd Generation Partnership Project (3GPP) standards, Institute of Electrical and Electronics Engineers (IEEE) standards, Internet Protocol (IP) standards, and Wireless Application Protocol (WAP) standards.
[0051] Transmission medium 104 may comprise a wireless, wired, and / or computer readable medium. For example, transmission medium 104 may comprise one or more wires, cables, air interfaces, optical discs, flash memory, and / or magnetic memory. In addition or alternatively, transmission medium 104 may comprise one more networks (e.g., the Internet) or file servers configured to store and / or transmit encoded video data.
[0052] To decode bitstream 110 into point cloud sequence 108 for display or other forms of consumption, destination device 106 may comprise an input interface 118, a decoder 120, and a point cloud display 122. Input interface 118 may be configured to read bitstream 110 stored on transmission medium 104 by source device 102. In addition, or alternatively, input interface 118 may be configured to receive, download, and / or stream bitstream 110 from source device 102 via transmission medium 104. Input interface 118 may comprise a wired and / or wireless receiver configured to receive, download, and / or stream bitstream 110 according to one or more proprietary and / or standardized communication protocols, such as those mentioned above.
[0053] Decoder 120 may decode point cloud sequence 108 from encoded bitstream 110. For example, decoder 120 may operate in a same or similar manner to a decoder provided by G-PCC reference software. In some examples, decoder 120 may decode a point cloud sequence that approximates point cloud sequence 108 due to, for example, lossy compression of point cloud sequence 108 by encoder 114 and / or errors introduced into encoded bitstream 110 during transmission to destination device 106.
[0054] Point cloud display 122 may display point cloud sequence 108 to a user. Point cloud display 122 may comprise a cathode rate tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light emitting diode (LED) display, a 3D display, a holographic display, a head mounted display, or any other display device suitable for displaying point cloud sequence 108.
[0055] It should be noted that point cloud coding / decoding system 100 is presented by way of example and not limitation. In the example of FIG.1, point cloud coding / decoding system 100 may have other components and / or arrangements. For example, point cloud source 112 may be external to source device 102. Similarly, point cloud display 122 may be external to destination device 106 or omitted altogether where point cloud sequence is intended for consumption by a machine and / or storage device. In another example, source device 102 may further comprise a pointDocket No.: 24-2010PCT cloud decoder and destination device 106 may comprise a point cloud encoder. In such an example, source device 102 may be configured to further receive an encoded bit stream from destination device 106 to support two-way point cloud transmission between the devices.
[0056] As mentioned above, an encoder may quantize the positions of points in a point cloud according to a space precision, which may be the same or different in each dimension of the points. The quantization process may create a grid in 3D space. The encoder may map any points residing within each sub-grid volume to the sub-grid center coordinates, referred to as a voxel (or a volumetric pixel). A voxel may be considered as a 3D extension of pixels corresponding to 2D image grid coordinates.
[0057] The encoder may represent or code the point cloud using an occupancy tree. For example, the encoder may split the initial volume or cuboid (also referred to as a bounding box) containing the point cloud into sub-cuboids. The encoder may then recursively split each sub-cuboid that contains at least one point of the point cloud. The encoder may not further split sub-cuboids that do not contain at least one point of the point cloud. A sub-cuboid that contains at least one point of the point cloud may be referred to as an occupied sub-cuboid. A sub-cuboid that does not contain at least one point of the point cloud may be referred to as an unoccupied sub-cuboid. The encoder may split an occupied cuboid into, for example, two sub-cuboids (to form a binary tree), four sub-cuboids (to form a quadtree), or eight sub- cuboids (to form an octree). The encoder may split an occupied cuboid to obtain sub-cuboids all with the same size and shape at a given depth level of the occupancy tree by splitting following a plane passing through the middle of edges of the cuboid.
[0058] The initial volume or cuboid containing the point cloud may correspond to the root node of the occupancy tree. Each occupied sub-cuboid, split from the initial volume / cuboid, may correspond to a node (of the root node) in a second level of the occupancy tree. Each occupied sub-cuboid, split from an occupied sub-cuboid in the second level, may correspond to a node (off the occupied sub-cuboid in the second level from which it was split) in a third level of the occupancy tree. The occupancy tree structure may continue to form in this manner for each recursive split iteration until, for example, a maximum depth level of the occupancy tree is reached or each occupied sub-cuboid has a volume corresponding to one voxel.
[0059] Each non-leaf node of the occupancy tree may comprise or be associated with an occupancy word representing an occupancy state of the cuboid corresponding to the node. For example, a node of the occupancy tree corresponding to a cuboid that is split into 8 sub-cuboids may comprise or be associated with a 1-byte occupancy word. Each bit (referred to as an occupancy bit) of the 1-byte occupancy word may represent or indicate the occupancy of a different one of the eight sub-cuboids. Occupied sub-cuboids may be represented or indicated by a binary value of 1 in the 1-byte occupancy word and unoccupied sub-cuboids may be represented or indicated by a binary value of 0 in the 1-byte occupancy word. In other examples, occupied and un-occupied sub-cuboids may be represented or indicated by opposite 1-bit binary values in the 1-byte occupancy word.Docket No.: 24-2010PCT
[0060] Each bit of an occupancy word may represent or indicate the occupancy of a different one of the eight sub- cuboids following the so-called Morton order. For example, the least significant bit of an occupancy word may represent or indicate the occupancy of a first one of the eight sub-cuboids following the Morton order, the second least significant bit of an occupancy word may represent or indicate the occupancy of a second one of the eight sub-cuboids following the Morton order, etc.
[0061] FIG.2 illustrates the Morton order of eight sub-cuboids 202-216 split from a cuboid 200. Sub-cuboids 202-216 are labeled based on their Morton order, with child node 202 being the first in Morton order and child node 216 being the last in Morton order. The Morton order for sub-cuboids 202-216 is a local lexicographic order in xyz.
[0062] The geometry of the point cloud is represented by, and therefore may be determined from, the initial volume and the occupancy words of the nodes in the occupancy tree. The encoder may therefore transmit the initial volume and the occupancy words of the nodes in the occupancy tree in a bitstream to a decoder for reconstructing the point cloud. Before transmitting the initial volume and the occupancy words of the nodes in the occupancy tree, the encoder may entropy encode the occupancy words. For example, the encoder may encode an occupancy bit of an occupancy word of a node corresponding to a cuboid, based on one or more occupancy bits of occupancy words of other nodes corresponding to cuboids that are adjacent or spatially close to the cuboid of the occupancy bit being encoded.
[0063] An encoder and / or decoder may code occupancy bits of occupancy words in sequence of a scan order. For example, an encoder and / or decoder may scan an occupancy tree in breadth-first order: all the occupancy words of the nodes of a given depth (or level) within the occupancy tree may be scanned before scanning the occupancy words of the nodes of the next depth (or level). Within a depth, the encoder and / or decoder may scan the occupancy words of nodes in the Morton order. Within a node, the encoder and / or decoder may scan the occupancy bits of the occupancy word of the node further in the Morton order.
[0064] FIG.3 illustrates an example of this scanning order for the first three levels of an occupancy tree 300. At each level of occupancy tree 300, a plurality of cuboids (e.g., cubes) are generated. In FIG.3, a cube 302 corresponding to the root node of occupancy tree 300 is divided into eight sub-cubes. Two sub-cubes 304 and 306 of the eight sub- cubes are occupied, while the other six sub-cubes are unoccupied. Following the Morton order, a first eight-bit occupancy word occW1,1is constructed to represent the occupancy word of the root node. The least significant occupancy bit of the first eight-bit occupancy word occW1,1 represents or indicates the occupancy of the first sub-cube of the eight sub-cubes in Morton order, the second least significant occupancy bit of the first eight-bit occupancy word occW1,1represents or indicates the occupancy of the second sub-cube of the eight sub-cubes in Morton order, etc.
[0065] Each of the two occupied sub-cubes 304 and 306 corresponds to a node off the root node in a second level of occupancy tree 300. The two occupied sub-cubes 304 and 306 are each further split into eight sub-cubes. One of the sub-cubes 308 of the eight sub-cubes split from sub-cube 304 is occupied, while the other seven sub-cubes are unoccupied. Three of the sub-cubes 310, 312, and 314 of the eight sub-cubes split from sub-cube 306 are occupied, while the other five sub-cubes of the eight sub-cubes split from sub-cube 306 are unoccupied. Two second eight-bitDocket No.: 24-2010PCT occupancy words occW2,1 and occW2,2 are constructed in this order to respectively represent the occupancy word of the node corresponding to sub-cube 304 and the occupancy word of the node corresponding to sub-cube 306.
[0066] Each of the four occupied sub-cubes 308, 310, 312, and 314 corresponds to a node in a third level of occupancy tree 300. The four occupied sub-cubes 308, 310, 312, and 314 are each further split into eight sub-cubes or 32 sub-cubes in total. Four third eight-bit occupancy words occW3,1, occW3,2, occW3,3 and occW3,4 are constructed in this order to respectively represent the occupancy word of the node corresponding to sub-cube 308, the occupancy word of the node corresponding to sub-cube 310, the occupancy word of the node corresponding to sub-cube 312, and the occupancy word of the node corresponding to sub-cube 314.
[0067] Following the scanning order discussed above, the occupancy words of this exemplary occupancy tree 300 may be entropy coded (e.g., entropy encoded by an encoder and entropy decoded by a decoder) as the succession of the seven occupancy words occW1,1to occW3,4. As a consequence of the breadth-first scanning order, when entropy coding the occupancy word of a current child node belonging to a current parent node, the occupancy words of all nodes having the same depth (or level) as the current parent node have already been entropy coded. In addition, the occupancy words of all nodes having the same depth (or level) as the current child node and having a lower Morton order than the current child node have also already been entropy coded. Part of these already coded occupancy words may be used to entropy code the occupancy word of the current child node. For example, the already coded occupancy words of neighboring parent and child nodes may be used to entropy code the occupancy word of the current child node. When entropy coding a particular occupancy bit of the occupancy word of the current child node, the occupancy bits of the occupancy word having a lower Morton order than the particular occupancy bit have also already been entropy coded and may be used to code the occupancy bit of the occupancy word of the current child node.
[0068] FIG.4 illustrates an example neighborhood of cuboids with already-coded occupancy bits that may be used to entropy code the occupancy bit of a current child cuboid 400. The neighborhood of cuboids with already-coded occupancy bits may be determined based on the scanning order of an occupancy tree representing the geometry of the cuboids in FIG.4 as discussed above. As illustrated in FIG.4, current child cuboid 400 belongs to a current parent cuboid 402. Following the scanning order of the occupancy words and occupancy bits of nodes of the occupancy tree, the occupancy bits of four child cuboids 404, 406, 408, and 410, belonging to the same current parent cuboid 402, have already been coded. Also, the occupancy bit of child cuboids 412 of preceding parent cuboids have already been coded. Furthermore, the occupancy bits of parent cuboids 414, for which the occupancy bits of child cuboids have not already been coded, have already been coded. Therefore, the already-coded occupancy bits of cuboids 404, 406, 408, 410, 412, and 414 may be used to code the occupancy bit of the current child cuboid 400.
[0069] The number of possible occupancy configurations for a neighborhood of a current child cuboid may be 2N, where N is the number of cuboids in the neighborhood of the current child cuboid with already-coded occupancy bits. The neighborhood of the current child cuboid may comprise several dozens of cuboids, among them the 26 adjacent parent cuboids sharing a face, an edge, or a vertex with the parent cuboid of the current child cuboid and also severalDocket No.: 24-2010PCT adjacent child cuboids (with occupancy bits already coded) sharing a face, an edge, or a vertex with the current child cuboid. Even limited to a subset of the adjacent cuboids, the occupancy configuration for a neighborhood of the current child cuboid may have billions of possible occupancy configurations making its direct use impractical. The occupancy configuration for a neighborhood of the current child cuboid may be used by an encoder and / or decoder to select the context (or equivalently the probability model), among a set of contexts, of a binary entropy coder (e.g., binary arithmetic coder) that codes the occupancy bit of the current child cuboid. The context-based binary entropy coding may be similar to the Context Adaptive Binary Arithmetic Coder (CABAC) used in MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC)).
[0070] Several methods may be used by an encoder and / or decoder to reduce the occupancy configurations for a neighborhood of a current child cuboid being coded to a practical number of reduced occupancy configurations. Firstly, the 26or 64 occupancy configurations of the six adjacent parent cuboids sharing a face with the parent cuboid of the current child cuboid may be reduced to 9 occupancy configurations by using geometry invariance. Secondly, an occupancy score for the current child cuboid may be obtained from the 226occupancy configurations of the 26 adjacent parent cuboids. The score may be further reduced into a ternary occupancy prediction (“predicted occupied”, “unsure”, “predicted unoccupied”) by applying score thresholds. Thirdly, the number of occupied and the number of unoccupied adjacent child cuboids may be used instead of the individual occupancies of these child cuboids.
[0071] An encoder and / or decoder employing one or more of the above methods may reduce the number of possible occupancy configurations for a neighborhood of a current child cuboid to a more manageable number (e.g., a few thousands). However, it has been observed that instead of associating a reduced number of contexts (or probability models) directly to the reduced occupancy configurations, another mechanism may be used, namely Optimal Binary Coders with Update on the Fly (OBUF). An encoder and / or decoder may implement OBUF to limit the number of contexts to a lower number (e.g., 32 contexts).
[0072] OBUF may use a limited number (e.g., 32) of contexts that may be fixed. These contexts may be ordered, referred to by a context index (e.g., a context index in the range of 0 to 31), and associated from a lowest virtual probability to a highest virtual probability to code a 1. A Look-Up Table (LUT) of context indices may be initialized at the beginning of a point cloud coding process. For example, the LUT may initially point to a context (e.g., context with context index 15), among the limited number of contexts, with the median virtual probability to code a 1 for all input. This LUT may take an occupancy configuration for a neighborhood of current child cuboid as input and output the context index associated with the occupancy configuration. Consequently, the LUT may have as many entries as reduced occupancy configurations (e.g., around a few thousand). The coding of the occupancy bit of a current child cuboid may follow the steps of determining the reduced occupancy configuration of the current child node, obtaining a context index by applying the reduced occupancy configuration as an entry to the LUT, coding the occupancy bit of the current child cuboid by using the context pointed to (or indicated) by the context index, and finally updating the LUT entry corresponding to the reduced occupancy configuration depending on the value of the coded occupancy bit of theDocket No.: 24-2010PCT current child cuboid. If a binary 0 (e.g., indicating the current child cuboid is unoccupied) is coded, the LUT entry may be decreased to a lower context index value, and if a binary 1 (e.g., indicating the current child cuboid is occupied) is coded, the LUT entry may be increased to a higher context index value. The update process of the context index may be based on a theoretical model of optimal distribution for virtual probabilities associated with the limited number of contexts. This virtual probability for a context may be fixed by a model and may be different from the internal probability of the context that evolves during the coding of bits of data. The evolution of the internal context may follow a well- known process similar to the process in CABAC.
[0073] An encoder and / or decoder may implement a “dynamic OBUF” scheme that may handle a much larger number of occupancy configurations for a neighborhood of a current child cuboid than can be handled by general OBUF, while maintaining complexity within reasonable bounds. The use of a larger number of occupancy configurations for a neighborhood of a current child cuboid may lead to improved compression capabilities. By using an occupancy tree compressed by OBUF, an encoder and / or decoder may reach a lossless compression performance as good as 1 bit per point (bpp) for coding the geometry of dense point clouds. An encoder and / or decoder may implement dynamic OBUF to potentially further reduce the bitrate by more than 25% to 0.7 bpp.
[0074] OBUF may not take as input a large variety of reduced occupancy configurations for a neighborhood of a current child cuboid, thus potentially leading to a loss of useful correlation. The size of the LUT of context indices may be increased to handle more various occupancy configurations for a neighborhood of a current child cuboid as input. However, by doing so, statistics may be diluted, and compression performance may be reduced. For example, if the LUT has millions of entries and the point cloud has a hundred thousand points, then most of the entries are never visited. Worse yet, many entries may be visited only a few times and their associated context indices may not be updated enough times to reflect any meaningful correlation between the occupancy configuration value and the probability of occupancy of the current child cuboid. Dynamic OBUF may be implemented to mitigate the dilution of statistics due to the increase in the number of occupancy configurations for a neighborhood of a current child cuboid. This mitigation is performed by a “dynamic reduction” of occupancy configurations in dynamic OBUF.
[0075] Dynamic OBUF may add an extra step of reduction of occupancy configurations for a neighborhood of a current child cuboid before applying the LUT of context indices. This step may be called a dynamic reduction because it evolves based on the progress of the coding of the point cloud or, more precisely, based on already visited occupancy configurations.
[0076] As discussed above, many possible occupancy configurations for a neighborhood of a current child cuboid are potentially involved but only a subset may be visited during the coding of a point cloud. This subset may characterize the type of the point cloud. For example, when coding AR or VR dense point clouds, most of the visited occupancy configurations may exhibit occupied adjacent cuboids of a current child cuboid. On the other hand, when coding sensor-acquired sparse point clouds, most of the visited occupancy configurations may exhibit only a few occupied adjacent cuboids of a current child cuboid. The role of the dynamic reduction may be to obtain a more preciseDocket No.: 24-2010PCT correlation based on the most visited occupancy configuration while putting aside (or reducing aggressively) other occupancy configurations that are much less visited. The dynamic reduction may be updated on-the-fly, as detailed below, after each visit of an occupancy configuration during the coding of occupancy data.
[0077] FIG.5 illustrates an example of a dynamic reduction function DR that may be used in dynamic OBUF. The dynamic reduction function DR may be obtained by masking bits βj of occupancy configurations 500: β = β1… βKmade of K bits. The size of the mask may decrease when occupancy configurations are visited a certain number of times. The initial dynamic reduction function DR0may mask all bits for all occupancy configurations such that it is a constant function DR0(β) = 0 for all occupancy configurations β. After each coding of an occupancy bit, the dynamic reduction function may evolve from a function DRnto an updated function DRn+1. The function may be defined by: β’ = DRn(β) = β1… βkn(β)where kn(β) 510 is the number of non-masked bits. The initialization of DR0may correspond to k0(β)=0, and the natural evolution of the reduction function towards finer statistics may lead to an increasing number of non-masked bits kn(β) ≤ kn+1(β). The dynamic reduction function may be entirely determined by the values of knfor all occupancy configurations β.
[0078] The visits to occupancy configurations may be tracked by a variable NV(β’) for all dynamically reduced occupancy configurations β’= DRn(β). After the coding of an occupancy bit based on an occupancy configuration βV, the corresponding number of visits NV(βV’) may be increased by one. If this number of visits NV(βV’) is greater than a threshold thV, NV(βV’) > thVthen the number of unmasked bits kn(β) may be increased by one for all occupancy configurations β being dynamically reduced to βV’. Practically, this corresponds to replacing the dynamically reduced occupancy configuration βV’ by the two new dynamically reduced occupancy configurations β0’ and β1’ defined by β0’ = βV’0 = βV1 … βVkn(β)0 and β1’ = βV’1 = βV1 … βVkn(β)1. In other words, the number of unmasked bits has been increased by one kn+1(β) = kn(β) + 1 for all occupancy configurations β such that DRn(β) = βV’. The number of visits of the two new dynamically reduced occupancy configurations may then be initialized to zero: NV(β0’) = NV(β1’) = 0. (I) At the start of the coding, the initial number of visits for the initial dynamic reduction function DR0may be set to NV(DR0(β)) = NV(0) = 0, and the evolution of NV on dynamically reduced occupancy configurations may now be entirely defined.
[0079] When a dynamically reduced occupancy configuration βV’ is replaced by the two new dynamically reduced occupancy configurations β0’ and β1’, the corresponding LUT entry LUT[βV’] may be replaced by the two new entries LUT[β0’] and LUT[β1’] that are initialized by the context index associated with βV’,Docket No.: 24-2010PCT LUT[β0’] = LUT[β1’] = LUT[βV’], (II) and then evolve separately. The evolution of the LUT of context indices on dynamically reduced occupancy configurations may thus be entirely defined.
[0080] The reduction function DRnmay be modeled by a series of growing binary trees Tn520 whose leaf nodes 530 are the reduced occupancy configurations β’ = DRn(β). The initial tree may be the single root node associated with 0 = DR0(β). The replacement of the dynamically reduced to βV’ by β0’ and β1’ corresponds to growing the tree Tnfrom the leaf node associated with βV’ by attaching to it two new nodes associated with β0’ and β1’. The tree Tn+1may be obtained by this growth. The number of visits NV and the LUT of context indices may be defined on the leaf nodes and evolve with the growth of the tree through equations (I) and (II).
[0081] In some examples, dynamic OBUF may be practically implemented by storage of the array NV[β’] and the LUT[β’] of context indices, as well as the trees Tn520. An alternative to the storage of the trees may be to store the array kn[β] 510 of the number of non-masked bits.
[0082] A limitation for implementing dynamic OBUF may be its memory footprint. In some applications, a few million occupancy configurations may be practically handled, leading to about 20 bits βiconstituting an entry configuration β to the reduction function DR. Each bit βimay correspond to the occupancy status of a neighboring cuboid of a current child cuboid or a set of neighboring cuboids of a current child cuboid.
[0083] Higher bits βi(e.g., β0, β1, etc.) may be the first bits to be unmasked during the evolution of the dynamic reduction function DR. Therefore, the order of neighbor-based information put in the bits βi may impact the compression performance. In some examples, neighboring information may be ordered from highest priority to lower priority and put in this order into the bits βi, from higher to lower weight. For example, the priority may be, from the most important to the least important, occupancy of sets of adjacent neighboring child cuboids, then occupancy of adjacent neighboring child cuboids, then occupancy of adjacent neighboring parent cuboids, then occupancy of non-adjacent neighboring child nodes, and finally occupancy of non-adjacent neighboring parent nodes. Adjacent nodes sharing a face with the current child node may also have higher priority than adjacent nodes sharing an edge or, worse, only a vertex with the current child node.
[0084] FIG.6 illustrates a flowchart of an exemplary method for coding the occupancy bit of a current child cuboid using dynamic OBUF. The method of the flowchart begins at block 602. At block 602, an encoder and / or decoder may determine the occupancy configuration β of already-coded cuboids in a neighborhood of the current child cuboid. At block 604, the encoder and / or decoder may dynamically reduce the occupancy configuration β into a reduced occupancy configuration β’ = DRn(β). At block 606, the encoder and / or decoder may lookup context index LUT[β’] in the LUT of the dynamic OBUF. At block 608, the encoder and / or decoder may select the context (or probability model) pointed to by the context index. At block 610, the encoder and / or decoder may entropy code (e.g., arithmetic code) the occupancy bit of the current child cuboid based on the context. Thus, the occupancy bit of the current child cuboid may be coded based on occupancy bits of the already-coded cuboids neighboring the current child cuboid.Docket No.: 24-2010PCT
[0085] Although not shown in FIG.6, the encoder and / or decoder may further update the reduction function DRninto DRn+1and update the context index LUT[β’] based on the occupancy bit of the current child cuboid. In addition, the method of FIG.6 may be repeated for additional or all child cuboids of parent cuboids corresponding to nodes of the occupancy tree in a scan order, such as the scan order discussed above with respect to FIG.3.
[0086] In general, the occupancy tree is a lossless compression technique. The occupancy tree may be adapted to provide lossy compression by modifying the point cloud on the encoder side (e.g., down-sampling, removing points, moving points, etc.) but the lossy compression performance may be reduced / weak. However, the use of the occupancy tree as a lossless compression technique may be very useful for dense point clouds.
[0087] One approach to lossy compression for point cloud geometry may be to set the maximum depth of the occupancy tree to not reach the smallest volume size of one voxel but instead to stop at a bigger volume size (e.g., NxNxN cubes, where N > 1). The geometry of the points belonging to each occupied leaf node associated with the bigger volumes may then be modeled. This approach may be particularly suited for dense and smooth point clouds that may be locally modeled by smooth functions like planes or polynomials. The coding cost may become the cost of the occupancy tree plus the cost of the local model in each of the occupied leaf nodes.
[0088] A scheme for modeling the geometry of the points belonging to each occupied leaf node, associated with a volume size larger than one voxel, may use sets of triangles as local models. This scheme may be referred to as the “TriSoup” scheme. TriSoup is short for “Triangle Soup” because the connectivity between triangles may not be part of the models. An occupied leaf node, of an occupancy tree, that corresponds to a cuboid with a volume greater than one voxel may be referred to as a TriSoup node. An edge belonging to at least one cuboid corresponding to a TriSoup node may be referred to as a TriSoup edge. A TriSoup node may comprise a presence flag (sk) for each TriSoup edge of its corresponding occupied cuboid. A presence flag (sk) of a TriSoup edge may indicate (a presence of or) whether a TriSoup vertex (Vk) is present or not on the TriSoup edge. At most one TriSoup vertex (Vk) may be present on a TriSoup edge. For each vertex (Vk) present on a TriSoup edge of an occupied cuboid, the TriSoup node corresponding to the occupied cuboid may further comprise a position (pk) of the vertex (Vk) along the TriSoup edge.
[0089] In addition to the occupancy words of an occupancy tree, an encoder may entropy encode, for each TriSoup node of the occupancy tree, a TriSoup vertex presence flag (and a position of a TriSoup vertex, if present, along a TriSoup edge) of each TriSoup edge belonging to the TriSoup node. A decoder may similarly entropy decode the TriSoup vertex presence flags and positions of each TriSoup vertex along a respective TriSoup edge belonging to a TriSoup node of the occupancy tree, in addition to the occupancy words of the occupancy tree.
[0090] FIG.7 illustrates an example of an occupied cube 700 of size NxNxN (where N > 1) that corresponds to a TriSoup node of an occupancy tree. Occupied cube 700 comprises TriSoup edges 710-721. The TriSoup node, corresponding to occupied cube 700, comprises a presence flag (sk) for each TriSoup edge of TriSoup edges 710-721. The presence flag of TriSoup edge 714 indicates that a TriSoup vertex V1 is present on TriSoup edge 714. The presence flag of TriSoup edge 715 indicates that a TriSoup vertex V2is present on TriSoup edge 715. The presenceDocket No.: 24-2010PCT flag of TriSoup edge 716 indicates that a TriSoup vertex V3 is present on TriSoup edge 716. The presence flag of TriSoup edge 717 indicates that a TriSoup vertex V4 is present on TriSoup edge 718. The presence flags of the remaining TriSoup edges each indicates that a TriSoup vertex is not present on their corresponding TriSoup edge. The TriSoup node, corresponding to occupied cube 700, further comprises a position (pk) for each TriSoup Vertex present along one of its TriSoup edges 710-721. More specifically, the TriSoup node (corresponding to occupied cube 700) further comprises a position p1for TriSoup vertex V1, a position p2for TriSoup vertex V2, a position p3for TriSoup vertex V3, and a position p4 for TriSoup vertex V4. The TriSoup vertices may be shared among TriSoup nodes along TriSoup edge(s) in common.
[0091] In some examples, a presence flag (sk) and, if the presence flag (sk) indicates the presence of a vertex, a position (pk) (the presence flag (sk) and position (pk) individually or collectively referred to as vertex information) of the vertex along a current TriSoup edge may be entropy coded based on already-coded presence flags and positions (of present TriSoup vertices) of TriSoup edges that neighbor the current TriSoup edge. A presence flag (sk) and, if the presence flag (sk) indicates the presence of a vertex, a position (pk) on (e.g., indicating a position of the vertex along) a current TriSoup edge may be additionally or alternatively entropy coded based on occupancies of cuboids that neighbor the current TriSoup edge. Similar to the entropy coding of the occupancy bits of the occupancy tree, a configuration βTS for a neighborhood (also referred to as a neighborhood configuration βTS) of a current TriSoup edge may be obtained. A neighborhood configuration βTS is derived based on presence flags (sk) and positions (pk) of vertices of neighboring already-coded edges of the current (TriSoup) edge and / or the occupancy of corresponding neighboring leaf nodes. The neighborhood configuration βTS may be dynamically reduced into a reduced configuration βTS’ = DRn(βTS) by using a dynamic OBUF scheme for TriSoup. A context index LUT[βTS’] may be obtained from the OBUF LUT and at least a part of the vertex information of the current TriSoup edge may be entropy coded using the context (or probability model) pointed to by the context index.
[0092] In order to use a binary entropy coder to entropy code at least part of the vertex information of the current TriSoup edge, the TriSoup vertex position (pk) (if present) along the current TriSoup edge may be binarized. A number of bits Nb may be set for the quantization of the TriSoup vertex position (pk) along the current TriSoup edge of length N that is uniformly partitioned into 2Nbquantization intervals. By doing so, the TriSoup vertex position (pk) may be represented by Nb bits (pkj, j=1,…,Nb) that may be individually coded by the dynamic OBUF scheme as well as the bit corresponding to the presence flag (sk) of the vertex on the current TriSoup edge. The neighborhood configuration βTS, the OBUF reduction function DRn, and thus the context index may depend on the nature / characteristic / property of the coded bit (presence flag (sk), highest position bit (pk1), second highest position bit (pk2), etc.). Therefore, there may be several dynamic OBUF schemes implemented, with each dedicated to a specific bit of information (presence flag (sk) or position bit (pkj)) of the vertex information.
[0093] FIG.8A illustrates a cuboid 800 (e.g., a cube) corresponding to a TriSoup node with a number K of TriSoup vertices Vk. Within cuboid 800, TriSoup triangles may be constructed from the TriSoup vertices Vkif at least three (K≥3)Docket No.: 24-2010PCT TriSoup vertices are present on the TriSoup edges of cuboid 800. In the example of FIG.8A, 4 TriSoup vertices are present and therefore TriSoup triangles are constructed. The TriSoup triangles may be constructed around the centroid vertex C defined as the mean of the TriSoup vertices Vk. In some examples, to construct the TriSoup triangles, a dominant direction may first be determined, then vertices Vk may be ordered by turning around this direction, and finally the following K TriSoup triangles (listed as triples of vertices) are constructed: V1V2C, V2V3C, …, VKV1C. The dominant direction may be chosen among the three directions parallel to the axis of the 3D space to increase or maximize the 2D surface of the triangles when projected along the dominant direction. By doing so, the dominant direction may be somewhat perpendicular to a local surface defined by the points of the point cloud belonging to the TriSoup node.
[0094] FIG.8B illustrates a refinement to the TriSoup model by coding a centroid residual vector Cresinto the bitstream such as to use C+Cres instead of C as a pivoting vertex for constructing / generating the triangles. By doing so, the vertex C+Cresmay be closer to the points of the point cloud than the centroid C used to model the points, which reduces the reconstruction error and leads to lower distortion at the cost of a small increase in bitrate needed for coding Cres.
[0095] FIG.8C illustrates a more detailed example of coding a centroid residual vector Cresin / from the bitstream such that an adjusted centroid C+Cresis used instead of centroid C for generating TriSoup triangles of a cuboid 800 (corresponding to a TriSoup node) corresponding to a portion of a point cloud, according to some embodiments. For example, the triangles may be generated based on adjusted centroid C+Cresand adjacent pairs of vertices of an ordering of the vertices V1-V4, determined as described above with respect to FIG.8A. Further, as described above, the TriSoup triangles of the cuboid may be voxelized at the decoder to generate voxels representing (or modeling) the portion, of the point cloud, corresponding to the cuboid. A unit vector ^⃗ (i.e., also referred to as a normalized vector) may be determined as a normalized mean vector of normal vectors to the triangles (V1V2C, V2V3C, …, VKV1C) constructed by centroid C and pairs of the vertices of the cuboid by pivoting around the centroid C (e.g., as described in FIG.8A). For example, the unit vector ^⃗ may be determined as the normalized vector based on a mean of cross-products representing areas of the trianglesFor example, theunit vector ^⃗ may be determined by dividing the mean vector (n) by the norm (or length) of the mean vector (i.e., ^⃗ = n / ||n||).
[0096] A value resulting from each cross product is equal to an area of a parallelogram formed by the two vectors in the cross product. Therefore, the value may be representative of an area of a triangle formed by the two vectors because the area of the triangle is equal to half of the value. Accordingly, since the vector ^⃗ indicates a direction of the triangles (e.g., TriSoup triangles) representing (e.g., modeling) the portion of the point cloud, the vector ^⃗ may be indicative of the direction normal to a local surface representative of the portion of the point cloud. In some examples, to maximize the effect of the centroid residual while minimizing its coding cost, a one-component residual αresalong the line (C, ^)⃗ 810 may be coded instead of a 3D residual vector.Docket No.: 24-2010PCT The residual value αres may be determined by the encoder as the intersection between the current point cloud and the line (C, ^)⃗, which is along the same direction of the normalized vector ^.⃗ For example, a set of points, of the portion of the point cloud, closest (e.g., within a threshold distance, a threshold number of points) to the line may be determined. The set of points may be projected on the line and the residual value αresmay be determined as the mean component along the line of the projected points. In some examples, the mean may be determined as a weighted mean whose weights depend on the distance of the set of points from the line. For example, a point from the set closer to the line may have a higher weight than another point from the set farther from the line.
[0097] In some examples, the residual value αres may be quantized. For example, it may be quantized by a uniform quantization function having quantization step similar to the quantization precision of the TriSoup vertices Vk. By doing so, the quantization error may be maintained to be uniform over all vertices Vkand C+Cressuch that the local surface is uniformly approximated.
[0098] In some examples, the residual value αres may be binarized and entropy coded into the bitstream, e.g., by using a unary-based coding scheme. In some examples, the residual value αresmay be coded using a set of flags. For example, a flag f0 may be coded to indicate if the residual value αres is equal to zero. If the flag f0 indicates the residual value αres is zero, no further syntax elements may be needed. If the flag f0 indicates the residual value αres is not zero, a sign bit indicating a sign may be coded and the residual magnitude |αres|-1 may be coded using an entropy code. For example, the residual magnitude may be coded using a unary coding scheme that codes successive flags fi (i≥1) indicating if the residual value magnitude |αres| is equal to ‘i’. A binary entropy coder may binarize the residual value αresinto the flags fi(i≥0) and entropy code the binarized residual value as well as the sign bit.
[0099] In some examples, compression of the residual value αres may be improved by determining bounds as shown in FIG.8C. As shown, the line (C, ^)⃗ 810 intersects the current cuboid 800 (corresponding to a TriSoup node) at two bounding points 820 and 821 and the encoder may impose that the adjusted centroid vertex C+Cres is located between the two bounding points 820 and 821. These bounding points 820 and 821 also bounds the residual value αres(which may be quantized) as belonging to an integral interval [m, M] where m ≤ 0 ≤ M. By doing so, some bits of the binarized residual value αres may be inferred. For example, if m=M=0, then residual value αres is necessarily equal to zero. In another example, if m=0<M, then the sign bit is necessarily positive. More generally, if the residual value αresis not equal to zero and its sign is known, its magnitude |αres| may be determined to be bounded by either |m| or M such that the magnitude may be coded by a truncated unary coding scheme that may infer the value of the last of successive flags fi(i≥1).
[0100] In some examples, the binary entropy coder used to code the binarized residual value αres may be a context- adaptive binary arithmetic coder (CABAC) such that the probability model (also referred to as a context or an entropy coder) used to code at least one bit (e.g., fior sign bit) of the binarized residual value αresare updated depending on precedingly coded bits. In some examples, the probability model of the binary entropy coder may be determined based on contextual information such as the values of the bounds m and M, the position of vertices Vk, or the size of theDocket No.: 24-2010PCT cuboid. In some examples, the selection of the probability model (i.e., also referred equivalently as an entropy coder or context) may be performed by a dynamic OBUF scheme with the contextual information described above as inputs.
[0101] The reconstruction of a decoded point cloud from the set of TriSoup triangles may be referred to as “voxelization” and may be performed, e.g., by ray tracing or rasterization, for each triangle individually before duplicate voxels from the voxelized triangles are removed.
[0102] FIG.9A illustrates an example of voxelization using ray tracing, according to some embodiments. For example, ray-triangle intersection algorithms, such as the Möller-Trumbore algorithm, rely on launching rays to determine whether rays intersect with TriSoup triangles and if so, at what points of the TriSoup triangles. Rays may be launched from integral coordinates that correspond to the centers of voxels. As illustrated by FIG.9A, rays such as ray 900 may be launched parallel to one of the three coordinate axes of the 3D space, starting from integral coordinates (sometimes referred to as integer coordinates) such as an origin point 905 (shown as origin or starting point Pstart).
[0103] An intersection point 904 (shown as Pint), if any, between ray 900 and a TriSoup triangle 901 belonging to a cube 902, corresponding to a TriSoup node, may be rounded (e.g., quantized) to obtain a decoded point corresponding to a voxel. For example, a ray, launched parallel to a coordinate axis in 3D space, may intersect a TriSoup triangle if and only if the projection, along the ray direction, of the center of a voxel belongs to the TriSoup triangle. In other words, the ray may be determined to intersect the TriSoup triangle if the point of intersection corresponds to the center of the voxel. In some examples, this intersection may be determined by applying a ray-triangle intersection algorithm (e.g., tracing or ray casting technique) such as the Möller-Trumbore algorithm to generate voxels representing the triangle.
[0104] Ray tracing techniques such as the Möller-Trumbore algorithm is based on generating, with respect to a triangle, barycentric coordinates of points of intersection between rays and a plane of the triangle. Then, points of the triangle may be determined from the barycentric coordinates.
[0105] FIG.9B illustrates an example of voxelization using barycentric coordinates (u, v, w) of a point 912 (P) relative to a TriSoup triangle 910 having vertices labeled A, B, and C in the 3D space, according to some embodiments. In some examples, point 912 may be determined as an intersection between a ray and a plane of TriSoup triangle 910 (e.g., containing or passing through the three vertices A, B, and C of TriSoup triangle 910). For example, the ray may be launched parallel to one of the three coordinate axes in 3D space. In some examples, this intersection point 912 may be uniquely represented as a sum of the three vertices of TriSoup triangle 910: P= uA + vB + wC under the condition u + v + w = 1. Therefore, any point P of the plane (containing TriSoup triangle 910) has unique coordinates (u,v,w) in the barycentric coordinate system. A point with barycentric coordinates (u,v,w) includes an ordered triple of numbers u, v, and w. A point with barycentric coordinates (u,v,w) that sum to 1 (i.e., u + v + w = 1) is known as homogeneous barycentric coordinates or normalized barycentric coordinates. The barycentric coordinates ofDocket No.: 24-2010PCT the intersection point with respect to TriSoup triangle 910 may be determined using, e.g., the well-known Möller- Trumbore algorithm.
[0106] By converting points with Cartesian coordinates in 3D space to homogeneous barycentric coordinates, the three vertices A, B, C of TriSoup triangle 910 have respective barycentric coordinates A(1,0,0), B(0,1,0) and C(0,0,1). In some examples, the convex hull (i.e., TriSoup triangle 910) of the three vertices A, B, and C is equal to the set of all points such that the barycentric coordinates u, v, and w is each greater than or equal to zero: 0 ≤ u, v, w Therefore, in some examples, the intersection point may be determined to belong to TriSoup triangle 910 based on the intersection point having barycentric coordinates with an ordered triple of values that is each greater than or equal to zero. Relatedly, if at least one of barycentric coordinates (i.e., one of u, v, or w) is negative or less than 0, then the intersection point may be determined to not belong to TriSoup triangle because it will be on the plane, but not on an edge or within the TriSoup triangle. In some examples, a point determined to belong to TriSoup triangle 910 may be the ray intersecting TriSoup triangle 910 (e.g., within or at an edge of TriSoup triangle 910).
[0107] FIG.10A and FIG.10B illustrate 12 cuboids 1000-1003, 1010-1013, and 1020-1023 with volumes that intersect a current (TriSoup) edge E being entropy coded. Current edge E is an edge of cuboids 1000-1003. The start point of current edge E intersects cuboids 1010-1013. The end point of current edge E intersects cuboids 1020-1023. The occupancy bits of one or more of the 12 cuboids 1000-1003, 1010-1013, and 1020-1023 is used to determine neighborhood configuration βTS for current edge E.
[0108] Edges may be oriented from a start point to an end point following the orientation of one of the three axes of the 3D space they are parallel to. A global ordering of the edges may then be defined as the lexicographic order over the couple (start point, end point). TriSoup vertex information of the edges may be coded following the edge ordering. Consequently, a causal neighborhood of a current edge may be obtained from the neighboring already-coded edges of the current edge.
[0109] Depending on the direction of current edge E, either two (FIG.11C for direction z), three (FIG.11B for direction y), or four (FIG.11A for direction x) of the four perpendicular edges have been already coded and their TriSoup vertex information may be used to construct neighborhood configuration βTSfor coding the TriSoup vertex information of the current edge E. The edge E’ has been already coded for each direction of the current edge E and its TriSoup vertex information may therefore be used to construct the neighborhood configuration βTS for current edge E independent of its direction.
[0110] In at least one implementation, the neighborhood configuration βTS for a current edge E may be obtained from one or more of the 12 occupancy bits of the 12 cuboids illustrated in FIG.10A and FIG.10B and from the vertex information of the at most five neighboring already-coded edges (E’ and E’’) illustrated in FIGS.11A-C.
[0111] In the implementation discussed above with respect to FIGS.11A-C, the TriSoup vertex information of at most five edges may be used to entropy code a current edge E. More particularly, in the implementation discussed aboveDocket No.: 24-2010PCT with respect to FIGS.11A-C, the TriSoup vertex information of at most five edges may be used to determine the neighborhood configuration βTS of the current edge E. The neighborhood configuration βTS may be dynamically reduced into a reduced configuration βTS’ = DRn(βTS) by using a dynamic OBUF scheme as discussed above. A context index LUT[βTS’] may be obtained from the OBUF LUT and at least a part of the TriSoup vertex information of the current edge E may be entropy coded using the context (or probability model) pointed to by the context index. Using the TriSoup vertex information of the at most five edges to entropy code the current edge E may result in a weak correlation between the neighborhood configuration βTS and the TriSoup vertex information of current edge E. The dynamic OBUF scheme may therefore provide a context index for entropy coding the current edge E with coding probabilities that are further weakly correlated with the TriSoup vertex information of the current edge E. Because of the weak correlation, the TriSoup vertex information of the current edge E may not be effectively compressed.
[0112] Beyond the at most five neighboring already-coded edges E’ and E’’ as illustrated in FIGS.11A-C, there are other neighboring already-coded edges belonging to the causal neighborhood of the current edge that may be used for coding the TriSoup vertex information of a current edge E. Using more neighboring already-coded edges leads to a neighborhood configuration βTSthat gets stronger correlation with the values of the TriSoup vertex information of the current edge E. Then, the dynamic OBUF schemes provide coder indices selecting entropy coders that have coding probabilities better correlated with the values of the bits representing the TriSoup vertex information of the current edge E. These bits are thus better compressed and the overall size of the bitstream is reduced.
[0113] FIGS.12A-C illustrate neighboring already-coded edges that, unlike the at most five edges illustrated in FIGS. 11A-C, do not intersect a start point of a current edge E being entropy coded. More particularly, FIGS.12A-C illustrate neighboring already-coded edges Eparthat are parallel to the current edge E and belong to a same TriSoup node as the current edge E.
[0114] FIG.12A illustrates already-coded parallel edges Epar that are available for coding a current edge E based on the current edge E being parallel to the x direction. FIG.12B illustrates already-coded parallel edges Eparthat are available for coding a current edge E based on the current edge E being parallel to the y direction. FIG.12C illustrates already-coded parallel edges Epar that are available for coding a current edge E based on the current edge E being parallel to the z direction. These parallel edges Eparare already coded according to the aforementioned lexicographic order that globally orders the set of edges.
[0115] An encoder and / or decoder may determine one or more symbols of a neighborhood configuration βTS of a current edge E based on one or more of the already-coded four parallel edges Eparillustrated in FIGS.12A-C that are available for coding the current edge E. The specific set of already-coded four parallel edges Epar illustrated in FIGS. 12A-C that are available for coding the current edge E may be determined based on the direction to which the current edge E is parallel as explained above. The TriSoup vertex information of the one or more of the already-coded four parallel edges Epar may be used to determine (e.g., in conjunction with one or more of the at least five edges illustrated in FIGS.11A-C) one or more symbols of a neighborhood configuration βTSof the current edge E. The encoder and / orDocket No.: 24-2010PCT decoder may select a context (or probability model) for coding the TriSoup vertex information of the current edge E based on the neighborhood configuration βTS. For example, the encoder and / or decoder may select the context for coding the TriSoup vertex information of the current edge E based on a reduced configuration βTS’ = DRn(βTS) representing a subset of the symbols of the neighborhood configuration βTS. The encoder and / or decoder may select the context based on an OBUF LUT that maps the neighborhood configuration βTS or the reduced configuration βTS’ to an index of the context. The encoder and / or decoder may entropy code (e.g., arithmetic code) the TriSoup vertex information of the current edge E based on the context.
[0116] FIGS.13A-C illustrate neighboring already-coded edges that, unlike the at most five edges illustrated in FIGS. 11A-C, do not intersect a start point of a current edge E being entropy coded. More particularly, FIGS.13A-C illustrate neighboring already-coded edges Eperp that are perpendicular to the current edge E and intersect the end point of the current edge E.
[0117] FIG.13A illustrates that no already-coded perpendicular edges Eperpare available for coding the TriSoup vertex information of a current edge E based on the current edge E being parallel to the x axis. FIG.13B illustrates that one already-coded perpendicular edge Eperpis available for coding the TriSoup vertex information of a current edge E based on the current edge E being parallel to the y axis. FIG.13C illustrates that two already-coded perpendicular edges Eperp are available for coding the TriSoup vertex information of a current edge E based on the current edge E being parallel to the z axis. These perpendicular edges Eperpare already coded according to the aforementioned lexicographic order that globally orders the set of edges.
[0118] An encoder and / or decoder may determine one or more symbols of a neighborhood configuration βTS of a current edge E based on one or more of the already-coded perpendicular edges Eperpillustrated in FIGS.13A-C that are available for coding the TriSoup vertex information of the current edge E. The specific set of already-coded perpendicular edges Eperp illustrated in FIGS.13A-C that are available for coding the TriSoup vertex information of the current edge E may be determined based on the direction to which the current edge E is parallel as explained above. The TriSoup vertex information of the one or more of the already-coded perpendicular edges Eperp may be used to determine (e.g., in conjunction with one or more of the at least five edges illustrated in FIGS.11A-C and / or in conjunction with one or more of the already-coded four parallel edges Eparillustrated in FIGS.12A-C) one or more symbols of a neighborhood configuration βTS of the current edge E. The encoder and / or decoder may select a context (or probability model) for coding the TriSoup vertex information of the current edge E based on the neighborhood configuration βTS. For example, the encoder and / or decoder may select the context for coding the TriSoup vertex information of the current edge E based on a reduced configuration βTS’ = DRn(βTS) representing a subset of the symbols of the neighborhood configuration βTS. The encoder and / or decoder may select the context based on an OBUF LUT that maps the neighborhood configuration βTSor the reduced configuration βTS’ to an index of the context. The encoder and / or decoder may entropy code (e.g., arithmetic code) the TriSoup vertex information of the current edge E based on the context.Docket No.: 24-2010PCT
[0119] Both parallel edges Epar and perpendicular edges Eperp are spatially close enough from the current edge E such that their TriSoup vertex information is well correlated with the TriSoup vertex information of the current edge E.
[0120] FIGS.14A-C illustrate neighboring already coded edges, taken from a spatial topology of 18 edges labelled from 0 to 17, of a current edge E which is parallel: to the x direction (FIG.14A), to the y direction (FIG.14B), or to the z direction (FIG.14C). Edge 0 corresponds to the unique edge (E’ in FIGS.11A-C) parallel to the current edge E and having its end point equal to the start point of the current edge E. Edges 1, 2, 3, 4 correspond to the at most four edges (E’’ in FIGS.11A-C) perpendicular to the current edge E and having a start or end point equal to the start point of the current edge E. Edges 14, 15, 16, 17 are edges (Eparin FIGS.12A-C) that are parallel to the current edge E and belong to a same TriSoup node as the current edge E. Edges 9, 10 are edges (Eperpin FIGS.13A-C) that are perpendicular to the current edge E and intersect the end point of the current edge E. Edges 5, 6, 7, 8 belong to a same TriSoup node as the current edge E and belong to the plane: perpendicular to the current edge E, and comprising the start point of the current edge E. Edges 10, 11, 12 belong to a same TriSoup node as the current edge E and belong to the plane: perpendicular to the current edge E, and comprising the end point of the current edge E.
[0121] An encoder and / or decoder may determine one or more symbols of a neighborhood configuration βTSof a current edge E based on one or more of the already-coded edges 0 to 17 illustrated in FIGS.14A-C that are available for coding the current edge E. The specific set of the already-coded edges 0 to 17 illustrated in FIGS.14A-C that are available for coding the current edge E may be determined based on the direction to which the current edge E is parallel. The TriSoup vertex information of the one or more of the already-coded edges 0 to 17 may be used to determine one or more symbols of a neighborhood configuration βTS of the current edge E. The encoder and / or decoder may select a context (or probability model) for coding the TriSoup vertex information of the current edge E based on the neighborhood configuration βTS. For example, the encoder and / or decoder may select the context for coding the TriSoup vertex information of the current edge E based on a reduced configuration βTS’ = DRn(βTS) representing a subset of the symbols of the neighborhood configuration βTS. The encoder and / or decoder may select the context based on an OBUF LUT that maps the neighborhood configuration βTS or the reduced configuration βTS’ to an index of the context. The encoder and / or decoder may entropy code (e.g., arithmetic code) the TriSoup vertex information of the current edge E based on the context.
[0122] The positions of TriSoup vertices along edges of TriSoup nodes may be quantized before being coded. Quantization of the positions of TriSoup vertices to a restricted number of possible positions along edges leads to a lower bitrate of coding of the positions with the drawback of higher distortion between the TriSoup model made of triangles based on the TriSoup vertices and the original point cloud geometry. Nevertheless, quantization of TriSoup vertices often leads to an improved tradeoff between bitrate and distortion, which results in improved compression capabilities.
[0123] In existing implementations, the TriSoup scheme (such as those defined for example in GPCCv2 and GeS-TM under development in MPEG SC29 / WG7) has limited quantization capabilities that impose that the size B of a TriSoupDocket No.: 24-2010PCT node as well as the size of the quantization steps used for quantizing the positions of TriSoup vertices along TriSoup edges to be both powers of two. This restriction relates to reducing computational complexity since powers of two can be efficiently processed.
[0124] FIGS.15A-B illustrate examples of quantizer when the TriSoup node size B as well as the size of the quantization steps are both powers of two. FIG.15A illustrates edges 1500 of length B=8 belonging to TriSoup nodes of size 8x8x8. FIG.15B illustrates edges 1501 of length B=4 belonging to TriSoup nodes of size 4x4x4. A parameter Nbsignals the number of quantization bits allowed for signaling the quantized positions of the TriSoup vertices along edges. Quantization is uniform over edges and quantization steps cannot be smaller than a voxel. Consequently, for TriSoup nodes having size B=2N, the number Nbof quantization bits must be smaller than N. FIG.15A illustrates the quantization process for Nb = 3, 2, 1 and 0 when B=8; and FIG.15B illustrates the quantization process for Nb = 2, 1 and 0 when B=4.
[0125] Edges are represented by segments [-0.5, B-0.5]. By definition, points 1510 (white circles) of the point cloud geometry have integral coordinates (xV,yV,zV) and are associated with 3D voxels being cubes of size 1x1x1 centered at (xV,yV,zV). The edges (1500, 1501) are divided uniformly into a number 2Nbof quantization intervals 1520 each of same length 2N-Nbfor Nb≤ N. The quantization intervals are associated with codewords 1530 belonging to the set {0; …; 2Nb- 1}. The position pk of a TriSoup vertex along an edge k (1500,1501) is a value in the segment [-0.5, B-0.5] that is quantized into a codeword pk,Nbassociated with the quantization interval this value belongs to. Dequantized positions 1540 (gray dots) obtained from codewords pk,Nb may be the middle of the quantization interval associated with the codewords pk,Nb.
[0126] FIG.16 illustrates an example encoding process 1600 of presence flag (sk) and position (pk) of TriSoup vertex of a current edge (k) of a TriSoup node in a bitstream, according to some embodiments.
[0127] At block 1605 the presence flag (sk) and position (pk) of a TriSoup vertex are determined based on, for example, a portion of the point cloud neighboring the current edge k. For example the presence flag (sk) may be determined by the encoder depending on the presence of at least one point of the point cloud closer than a threshold distance from the edge. If the presence flag (sk) is true, the position (pk) may be determined, for example, by the mean value of the projected positions onto the edge of the at least one point of the point cloud closer than a threshold distance from the edge.
[0128] At block 1610, if the presence flag (sk) is true, the position (pk) of the TriSoup vertex is quantized into codeword (pk,Nb) of Nbbits.
[0129] At block 1620, presence flag (snei) and if the presence flag (snei) is true, codeword (pnei,Nb) of Nb bits of at least one neighboring already-coded edge of the current edge k, are obtained from a set 1621 of already-coded edges. An edge is said already coded when its presence flag and, if the presence flag is true, position of TriSoup vertex along it, are coded before coding presence flag and, if the presence flag is true, position of TriSoup vertex of the current edge (k).Docket No.: 24-2010PCT
[0130] At block 1630, neighborhood configuration βTS is derived based on the presence flag (snei) and, if the presence flag (snei) is true, codeword (pnei,Nb) of Nb bits of the at least one neighboring already-coded edge of the current edge k.
[0131] At block 1640, the bits representing the presence flag (sk) and, if the presence flag (sk) is true, codeword (pk,Nb) of Nb bits of the current edge (k) are entropy encoded one by one in the bitstream 1690 based on the neighborhood information βTS.
[0132] For example, block 1640 may comprise block 1641 and block 1642.
[0133] At block 1641, an entropy encoder 1643 is selected based on the neighborhood configuration βTS.
[0134] For example, an OBUF or dynamic OBUF process takes neighborhood configuration βTSas input and selects the entropy encoder 1643.
[0135] At block 1642, the bits representing the presence flag (sk) and, if the presence flag (sk) is true, codeword (pk,Nb) of Nbbits of the current edge (k) are entropy encoded one by one in the bitstream 1690 based on the selected entropy encoder 1643. Several entropy encoders 1643 may be selected, each of them being dedicated to the encoding of one bit.
[0136] After encoding, the presence flag (sk) and codeword (pk,Nb) of Nbbits of the current edge (k) are added to the set 1621 of already coded edges.
[0137] FIG.17 illustrates an example decoding process 1700 of encoded presence flag (sk) and position (pk) of a TriSoup vertex of a current edge (k) of a TriSoup node from a bitstream. The process 1700 may include the same operations (shown as having the same labeled blocks) as those described in FIG.16. Different from the process 1600 of FIG.16, the process 1700 includes block 1710 (1711, 1712) and 1720.
[0138] At block 1710, the bit representing the presence flag (sk) and, if the presence flag (sk) is true, the bits representing codeword (pk,Nb) of Nbbits of the current edge (k) are entropy decoded one by one from the bitstream 1690 based on the neighborhood information βTS. The decoded presence flag (sk) and codeword (pk,Nb) of Nb bits are added to a set 1621 of already coded edges of the current edge (k).
[0139] In some embodiments, block 1710 may comprise block 1711 and block 1712.
[0140] At block 1711, an entropy decoder 1713 is selected based on the neighborhood information βTS.
[0141] In some embodiments, an OBUF or dynamic OBUF process takes neighborhood information βTSas input and selects the entropy decoder 1713.
[0142] At block 1712, the bit representing the presence flag (sk) and if the presence flag (sk) is true, the bits representing codeword (pk,Nb) of Nbbits of the current edge (k) are entropy decoded one by one from the bitstream 1690 based on the selected entropy decoder 1713. Several entropy decoders 1713 may be selected, each of them being dedicated to the decoding of one bit.
[0143] At step 1720, dequantized position pDQ,kof the TriSoup vertex is determined based on (decoded) codeword (pk,Nb) of Nb bits.Docket No.: 24-2010PCT
[0144] The TriSoup vertex information of TriSoup vertex along the current edge (k) is then made of the (decoded) presence flag sk and the dequantized position pDQ,k of TriSoup vertex.
[0145] The neighborhood configuration βTS(block 1630) is derived by combining presence flags (snei) and codeword (pnei,Nb) (representing positions) of TriSoup vertices along neighboring already-coded edges of the current edge (k). In particular, the location of neighboring TriSoup vertices relative to (close to or far from) the start point or the end point of the current edge is a strong indicator of the presence and / or position of a TriSoup vertex on the current edge.
[0146] Presence flags (snei) and codewords (pnei,Nb) of neighboring already-coded edges of the current TriSoup edge (k) are much information that must be reduced to obtain a neighborhood configuration βTSof practical sizes to be used as input to the OBUF process (1641, 1711). Practically, only the first one (or two) leading bits of the codewords pnei,Nbare used to decide if a TriSoup vertex along a neighboring already-coded edge is close or far from the current edge. The leading bit of the codewords (pnei,Nb) indicates to which half edge the position of the TriSoup vertex associated with a neighboring already-coded edge belongs to. The two leading bits of the codewords (pnei,Nb) indicate to which quarter of edge the position of the TriSoup vertex along a neighboring already-coded edge belongs to. Half edge or quarter of edge precision is sufficient to construct the neighborhood configuration βTS. Deriving (block 1630) neighborhood configuration βTShas thus been optimized for the use of the two leading bits of the codewords (pnei,Nb).
[0147] In existing technologies, the TriSoup scheme is limited to powers of two for both the size B of TriSoup node and the number 2Nbof quantization intervals over edges. This limitation is due to the specific quantizers illustrated by FIG.16.
[0148] This limitation leads to a coarse granularity of sizes of TriSoup node and quantizers that does not allow for a fine control of the overall geometry bitrate. Given a fixed parameter Nb, the geometry bitrate is typically multiplied by a factor three to four when decreasing the size of TriSoup node from B to B / 2. On the other hand, given a fixed size B of TriSoup node, increasing Nb by one leads typically to an increase of the geometry bitrate by several tens of precents. These massive changes in bitrate when changing the value of parameters B or Nbby the smallest amount possible imply that advanced rate control algorithms cannot be implemented.
[0149] This limitation further leads to local geometry quality that is mainly adjusted through the parameter Nb that drives the accuracy of the positions of the TriSoup triangles. Due to the very structure of the TriSoup nodes, the size B of TriSoup node cannot be adjusted locally and is the same for a whole slice (or brick or geometry data unit). Changing the value of Nb by one unit leads to massive changes in quality for both PSNR and visual quality. Even if the quality may be adjusted locally by signaling Nbat some sub-slice granularly (say at TriSoup node granularity), this adjustment is very rough. For comparison, video coders typically use a local integral Quality Parameter (QP) such that a decrease of 6 units for QP leads to twice smaller quantization steps for colors. Local quality of colors can be thus finely tuned. The current TriSoup scheme lacks intermediate qualities between Nb(corresponding to some QP) and Nb+1 (corresponding to some QP-6).Docket No.: 24-2010PCT
[0150] Embodiments of the present disclosure relates to a quantizing process of positions of TriSoup vertices along edges having any length (and thus allows for any size B of TriSoup node), e.g., for lengths that are not powers of two. The quantizing process according to some embodiments of the present disclosure provides uniform quantization of TriSoup edge into quantization intervals of any length. It allows a fine control of the overall point cloud geometry bitrate because the size B of TriSoup node may be any integer value and the size of quantization intervals can be finely controlled. Fine changes in bitrate may be possible by changing the value of parameters B and / or newly introduced GQP (replacing Nb). Therefore, advanced rate control algorithms can be implemented.
[0151] The quantizing process according to some embodiments of the present disclosure derives quantized position (pQP,k) of TriSoup vertex along a current edge (k) made of a bit (blr) and a codeword (Cp,k). The bit (blr) indicates whether the position (pk) of the TriSoup vertex is on a first half or a second half of the current edge (k). The codeword (Cp,k) indicates a distance (dp) between the middle position of the current edge (k) separating the first half and the second half and the position (pk) of the TriSoup vertex.
[0152] In some embodiments, the bit blr, the codeword (Cp,k) may be entropy encoded in a bitstream 1690.
[0153] FIGS.18A-B illustrate the quantizing process according to some embodiments. FIGS.19A-C illustrate the quantizing process according to some embodiments.
[0154] In FIG.18A, the edge (k) is represented by a thick line 1800. The edge (k) has length B that is any integral number not necessarily a power of two. Without loss of generality, it is assumed that the edge (k) is a segment [-0.5; B- 0.5] along the coordinate axis parallel to the edge (k). Points 1810 of the point cloud in a TriSoup node that neighbors the edge (k) (line 1800) have integral coordinate between 0 and B-1 along the axis. For example, in the case of cubic TriSoup nodes, points in the TriSoup node may belong to the grid {0; …; B-1}3within the cube [-0.5; B-0.5]3defining the volume of the TriSoup node. The middle position 1820 of the edge (k) (line 1800) has coordinate (B-1) / 2 in the segment [-0.5; B-0.5] defining the edge (k). The edge (k) (line1800) is thus split into two equal parts: a left half edge 1830 and a right half edge 1831. The left half edge 1830 is defined by the segment [-0.5; B / 2-0.5) starting from the lower bound (- 0,5) of the edge (k) (line 1800) and ending at the middle position (B / 2-0,5) of the edge (k) (line 1800). The right half edge 1831 is defined by the segment [B / 2-0.5; B-0.5] starting from the middle position (B / 2-0.5) of the edge (k) (line 1800) and ending at the upper bound (B-0.5) of the edge (k) (line 1800).
[0155] A TriSoup vertex located on the edge (k) (line 1800) has coordinate (pk) in the segment [-0.5; B-0.5]. Quantizing the position (pk) of TriSoup vertex according to some embodiments of the present disclosure firstly involves determining whether a position (pk) of the TriSoup vertex is on a first half or a second half of the current edge (k). Therefore, the bit (blr) indicates if the TriSoup vertex position (pk) belongs to either the left half edge 1830 (for example, blr = 0) or the right half edge 1831 (for example, blr = 1).
[0156] Secondly, a distance (dp) is determined between the middle position 1821 of the edge (k) (line 1800) as illustrated by FIG.18B and the position (pk) of the TriSoup vertex and the distance (dp) is quantized as illustrated in FIGS.19A-C.Docket No.: 24-2010PCT
[0157] In FIG.18B, the distance (dp) belongs to an interval of definition defined as a segment [0; B / 2] illustrated by a thick line 1840. The lower bound 0 represents the middle position 1821 of the edge.
[0158] In some embodiments, the distance dpmay be determined as the absolute difference between the middle position 1821 of the current edge (k) and the position (pk) of the TriSoup vertex: dp = | pk - (B-1) / 2|
[0159] Assuming the position (pk) of the TriSoup vertex has been determined by points of the point cloud belonging to TriSoup nodes intersecting entirely the edge (k) (line 1800), the distance (dp) is not higher than B / 2-0.5 corresponding to the farthest points 1811 from the middle position 1821 of the edge (k).
[0160] FIGS.19A-C illustrate examples of quantizing the distance (dp). The distance (dp) is quantized for three examples of quantization step Qs 1900 (FIG.19A), 1901 (FIG.19B) and 1902 (FIG.19C).
[0161] In some embodiments, the segment [0; B / 2] over which the distance (dp) is defined may be uniformly partitioned into quantization intervals and a respective codeword is associated with each quantization interval.
[0162] For example, the segment [0; B / 2] over which the distance (dp) may be uniformly partitioned into quantization intervals, starting from its lower bound 0 until a last interval of quantization 1910 (FIG.19A), 1911 (FIG.19B) and 1912 (FIG.19C) containing the upper bound B / 2-0.5. The distance (dp) between the middle position 1821 and the position (pk) of the TriSoup vertex may then be quantized into Ns quantization intervals [0, Qs), [Qs, 2Qs), …, [(Ns-1)Qs, NsQs). A codeword may be associated with each interval [kQs, (k+1)Qs).
[0163] In some embodiments, the codeword Cp,k indicating the distance (dp) from the middle position of the current edge (k) and the position (pk) of the TriSoup vertex is associated with the quantization interval the distance (dp) belongs to.
[0164] The distance (dp) may be quantized into the codeword Cp,kassociated with the quantization interval the distance (dp) belongs to.
[0165] The number Nsof possible codewords may thus be determined as the unique integer that fulfills the equalities (Ns-1)Qs ≤ dp < NsQs.
[0166] Described embodiments relate to dequantizing process of quantized positions of TriSoup vertices along edges having any length (and thus allows for any size B of TriSoup node), e.g., for lengths that are not powers of two. The dequantizing process according to some embodiments of the present disclosure provides uniform dequantization of TriSoup edge from quantization intervals of any length. It allows a fine control of the overall point cloud geometry bitrate because the size B of TriSoup node may be any integer value and the size of quantization intervals can be finely controlled. Fine changes in bitrate may be possible by changing the value of parameters B and / or newly introduced GQP (replacing Nb). Therefore, advanced rate control algorithms can be implemented.
[0167] The dequantizing process according to some embodiments of the present disclosure derives position of TriSoup vertex along a current edge (k) made of a presence flag (sk) and a position (pQP,k) of TriSoup vertex along the current edge (k) from a quantized position (pQP,k) of TriSoup vertex along the current edge (k) made of a bit (blr) and aDocket No.: 24-2010PCT codeword (Cp,k). The bit (blr) indicates whether the position (pk) of the TriSoup vertex is on a first half or a second half of the current edge (k). The codeword (Cp,k) indicates a distance (dp) between the middle position of the current edge (k), separating the first half and the second half, and the position (pk) of the TriSoup vertex.
[0168] In some embodiments, the bit blr and the codeword Cp,k may be entropy decoded from the bitstream 1690.
[0169] Firstly, interval of definition of a distance from the middle position of the current edge (k) is uniformly partitioned into quantization intervals and a respective codeword is associated with each quantization interval. A distance dDQ,k is determined based on the codeword (Cp,k) that indicates a quantization internal.
[0170] In some embodiments, the size of quantization intervals is derived based on a quantization step Qs.
[0171] In some embodiments, the distance dDQ,kmay be determined by a linear combination of the codeword (Cp,k) and the quantization step Qs such that, for example: dDQ,k(Cp,k, Qs) = (Cp,k+ 0.5) Qs
[0172] In some embodiments, the distance (dDQ,k) may be determined between the middle position of the current edge (k) and the center of a quantization interval indicated by the codeword (Cp,k).
[0173] FIG.20A illustrates an example of distances dDQ,k(2000) determined between the middle position of a current edge (k) and the centers of quantization intervals.
[0174] Secondly, the position (pDQ,k) of the TriSoup vertex along the current edge (k) is determined based on the bit (blr) and the distance (dDQ,k).
[0175] In some embodiments, the position (pDQ,k) of the TriSoup vertex along the current edge (k) may be determined by subtracting the distance (dDQ,k) from the middle position (B / 2-0,5) of the current edge (k) if the bit (blr) equals to a first value, e.g.0, to indicate the first half of the current edge (k), and by adding the distance (dDQ,k) to the middle position of the current edge (k) if the bit (blr) equals to a second value, e.g.1, different from the first value, to indicate the second half of the current edge (k).
[0176] The position pDQ,kof TriSoup vertex along the current edge (k) may be given by: pDQ,k = (B-1) / 2 - dDQ,k if blr = 0 pDQ,k = (B-1) / 2 + dDQ,k if blr = 1
[0177] Determining the distance dDQ,kassociated with the last quantization interval [(Ns-1)Qs, NsQs) based on the center of the last interval may not be optimal as illustrated in FIG.20B. The distance dp may only belong to the subinterval 2020 (vertical stripes) at the left of the coordinate B / 2 - 0.5 and not to the subinterval 2021 (horizontal stripes) at the right of the coordinate B / 2 - 0.5. The dequantized distance dDQ,k(2011) of the last quantization interval may thus belong to the left subinterval [(Ns-1) Qs, B / 2 - 0.5].
[0178] In some embodiments, the dequantized distance dDQ,k (2011) of the last quantization interval may be the middle position of the left subinterval [(Ns-1) Qs, B / 2 - 0.5], i.e. dDQ,k(Ns-1, Qs) = ((Ns-1) Qs + B / 2 - 0.5 ) / 2.Docket No.: 24-2010PCT
[0179] It has been observed from statistics on the positions of TriSoup vertices that the distribution of distances (dp) tends to peak near the bound B / 2 - 0.5. Consequently, the distance dDQ,k (2012, FIG.20C) of the last interval of quantization may be located at the right of the middle position of the left subinterval [(Ns-1) Qs, B / 2 - 0.5].
[0180] For example, the distance dDQ,k (2012) may be obtained by dDQ(Ns-1, Qs) = ( (Ns-1) Qs + 3*(B / 2 - 0.5) ) / 4.
[0181] In some embodiments, the quantization step Qsmay have sub-voxel precision such that the quantization step Qs is not necessarily a multiple of a voxel. Sub-voxel precision facilitates fine adjusting of the local quality and / or the rate allocation for the point cloud geometry.
[0182] For example, the quantization step Qsmay be represented by an integer value IQswith fixed number M of precision, e.g.8 bits, such that the quantization step Qs is equal to IQs / 2M. Internal computation of the quantizing and de-quantizing processes may be performed using this M-bit precision. In particular, positions pDQ,kmay have M-bit precision and the TriSoup triangles may be constructed based on TriSoup vertices positions with M-bit precision. The voxelizing process of TriSoup triangles comes back to integer precision when obtaining the decoded points of the point cloud.
[0183] In some embodiments, a geometry quality parameter (GQP) representative of values of the integer IQsmay be provided for example by an end-user.
[0184] For example, the geometry quality parameter GQP may equal the integer IQs.
[0185] For example, the geometry quality parameter GQP may mimic the well-known behavior of video quality parameters by having the quantization step Qs being proportional to power of two of a scaled GQP. The following formula may be used for positive integral quality parameter GQP Qs= 2(GQP-12) / 6(GQP ≥0)
[0186] When the geometry quality parameter GQP is equal to 0, the quantization step Qs has size of a quarter of a voxel. When the geometry quality parameter GQP is equal to 6, the quantization step Qshas size of half a voxel. When the geometry quality parameter GQP is equal to 12, the quantization step Qs has the size of a voxel. When the geometry quality parameter GQP is equal to 18, the quantization step Qs has size of two voxels, etc. A change of one unit in the integral geometry quality parameter GQP induces a change of the quantization step Qsof a factor 21 / 6≈ 1.12 (equivalently a change of about 12%), much below the factor 2 of existing technologies. This allows for finer adjusting of the geometry quality.
[0187] In some embodiments, the geometry quality parameter GQP may be signaled in the bitstream to signal to the decoder the size of the quantization step Qs.
[0188] For example, the geometry quality parameter GQP may be signaled in the high-level syntax of the bitstream, for example in the Geometry Parameter Set (GPS) or in the headers of the Geometry Data Units (GDU, aka slices or bricks).Docket No.: 24-2010PCT
[0189] For example, the geometry quality parameter GQP may be signaled at a more local level to allow for local quality adjustment. For example, the value of the geometry quality parameter GQP may be signaled at TriSoup node level or in Quality Units that partition Geometry Data Units.
[0190] FIG.21 illustrates an example encoding process 2100 for encoding, in a bitstream, a presence flag (sk) and a current position (pk) of a TriSoup vertex on a current edge (k) of a cuboid, corresponding to a TriSoup node, according to some embodiments. For example, the cuboid may contain a portion of a point cloud and one or more TriSoup vertices of the cuboid may be encoded to represent a geometry of the point cloud contained in the cuboid. The process 2100 may include the same operations (shown as having the same labeled blocks) as those described in FIG.16. Different from the process 1600 of FIG.16, the process 2100 includes block 2110 rather than block 1610 in FIG.16 to quantize the current position (pk) of the TriSoup vertex and block 2120 rather than block 1620 to obtain TriSoup edge information Ineiof at least one neighboring already-coded edge of the current edge (k).
[0191] At block 2110, the current position (pk) of the TriSoup vertex is quantized into a quantized position (pQP,k) of the TriSoup vertex. The quantized position (pQP,k) is indicated by (e.g., made of) a current bit (blr,k) indicating whether the current position (pk) of the TriSoup vertex is on a first half or a second half of the current edge (k) and a current codeword (Cp,k) indicating a distance between the middle position of the current edge (k) separating the first half and the second half of the current edge (k), and the current position (pk) of the TriSoup vertex as discussed in relation with FIGS.18 and 19.
[0192] Quantizing the current position (pk) of the TriSoup vertex may be performed according to a quantization step Qs that may be determined based on the geometry quality parameter GQP.
[0193] At block 2120, TriSoup edge information Ineiof the at least one neighboring already-coded edge of the current edge (k) is obtained.
[0194] At block 1630, neighborhood configuration βTS is derived based on the TriSoup edge information Inei of the at the least one neighboring already-coded edge of the current edge (k).
[0195] In some embodiments, the encoding process 2100 further comprises entropy encoding, in the bitstream, a bit representing a presence flag indicating the presence of the TriSoup vertex on the current edge (k).
[0196] FIG.22 illustrates an example decoding process for decoding, from a bitstream, an encoded presence flag and quantized position (e.g., current position) of a TriSoup vertex on a current edge of a cuboid, corresponding to a TriSoup node, according to some embodiments. For example, the cuboid may contain a portion of a point cloud and one or more TriSoup vertices of the cuboid may be encoded to represent a geometry of the point cloud contained in the cuboid. The process 2200 may include the same operations (shown as having the same labeled blocks) as those described in FIG.17. Different from the process 1700 of FIG.17, the process 2200 includes block 2210 to dequantize the quantized current position (pQP,k) of the TriSoup vertex and block 2120 rather than block 1620 to obtain TriSoup edge information Inei of at least one neighboring already-coded edge of the current edge (k).Docket No.: 24-2010PCT
[0197] At block 2120, TriSoup edge information Inei of the at least one neighboring already-coded edge of the current edge (k) is obtained.
[0198] At block 1630, neighborhood configuration βTSmay be derived based on TriSoup edge information Ineiof the at least one neighboring already-coded edge of the current edge (k). Blocks 2120 and 1630 are similarly performed in the encoding process of FIG.21.
[0199] At block 2210, the quantized current position (pQP,k) of a TriSoup vertex, indicated by (e.g., made of) a current bit (blr,k) and a current codeword (Cp,k), is dequantized to obtain a current position (pDQ,k) of the TriSoup vertex on the current edge (k), for example, as discussed hereinabove in relation with FIG.20.
[0200] A portion of a decoded point cloud geometry may then be obtained by constructing and voxelizing triangles obtained based on the TriSoup vertices associated with the positions (pDQ,k) of TriSoup vertices on edges of the cuboid containing the portion of the decoded point cloud geometry.
[0201] Dequantizing the quantized current position (pQP,k) of the TriSoup vertex may be performed according to a quantization step Qs that may be determined based on the geometry quality parameter GQP.
[0202] In some embodiments, the decoding process 2200 further comprises entropy decoding, from the bitstream, a bit representing a presence flag indicating the presence of the TriSoup vertex on the current edge (k).
[0203] In some embodiments, the entropy encoder or decoder (e.g., also referred to as a context or probability model) is selected based on OBUF or dynamic OBUF taking the neighborhood configuration βTSas input.
[0204] In some embodiments of block 2120, the TriSoup edge information Inei of the at least one neighboring already- coded edge of the current edge (k) is comprises (e.g., made of) a presence flag (snei) of the at least one neighboring already-coded edge, and if a presence flag (snei) of one neighboring already-coded edge is true, the TriSoup information Ineiis further comprises (e.g., made of) a bit (blr,nei) indicating whether the quantized position of the neighboring TriSoup vertex is on a first half or a second half of the neighboring already-coded edge and a codeword (Cp,nei) indicating a distance between the middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge, and the quantized position of the neighboring TriSoup vertex on the neighboring already-coded edge. For example, the quantized position of the neighboring TriSoup vertex may have been previously determined and encoded in bitstream 1690.
[0205] FIG.23 illustrates an example of obtaining TriSoup edge information of at least one neighboring already- coded edge (e.g., block 2120), according to some embodiments.
[0206] A quantized position (pQP,k) of a TriSoup vertex along a current edge (k) may be indicated by (e.g., represented by or made of) a current bit (blr,k) and a current codeword (Cp,k). A presence flag (sk) and possibly the quantized current position (pQP,k) of a TriSoup vertex along the current edge (k) may be added to a set of 2121 of already-coded edge as TriSoup edge information of the current edge (k).
[0207] At block 21201, a presence flag (snei) and possibly a bit (blr,nei) and a codeword (Cp,nei) may be obtained for at least one neighboring already-coded edge of the current edge (k) from the set 2121 of already-coded edge and, theDocket No.: 24-2010PCT TriSoup edge information Inei may comprise (e.g., include) at least one presence flag (snei), and if a presence flag (snei) of one neighboring already-coded edge is true, the TriSoup edge information Inei may comprise the presence flag (snei) of the neighboring already-coded edge and one bit (blr,nei) and a codeword (Cp,nei) of the neighboring already-coded edge.
[0208] In some embodiments, the bit representing the presence flag (sk) may be entropy encoded or decoded based on a first neighborhood configuration derived based on at least one bit representing at least one presence flag (snei) indicating the presence of at least one TriSoup vertex on one of the at least one neighboring already-coded edge, the bit (blr,k) may be entropy encoded or decoded based on a second neighborhood configuration derived based on at least one bit (blr,nei) indicating whether a position of at least one TriSoup vertex is on a first half or a second half of the at least one neighboring already-coded edge, and the bits representing the codeword (Cp,k) may be entropy coded based on a third neighborhood configuration derived based on bits representing at least one codeword (Cp,nei) indicating at least one distance between the middle position of the at least one neighboring already-coded edge and a position of at least one TriSoup vertex.
[0209] In some embodiments, only the leading bit of the codeword (Cp,k) may be entropy encoded or decoded based on a fourth neighborhood configuration derived based on leading bit representing at least one codeword (Cp,nei) indicating at least one distance between the middle position of the at least one neighboring already-coded edge and a position of at least one TriSoup vertex.
[0210] In some embodiments, the first, second, third, or fourth neighborhood configuration may be constructed based on precedingly coded bits representing the presence flag (sk), the coded bit (blr,k) and the coded bits representing the codeword (Cp,k). This is particularly efficient when coding the bits representing the codeword (Cp,k) because not all combinations of bits lead to a valid codeword.
[0211] In some implementations, the introduction of a quantizing block 2110 and dequantizing block 2210 impacts the derivation of neighborhood configuration βTSbecause the quantized positions (pQP,k) of the TriSoup vertices along the already-coded edges have changed to a quantizer that does not necessarily have a number of intervals equal to a power of two. For example, the quantized position (pQP,nei) of TriSoup vertex along a neighboring already-coded edge of the current edge (k) may be indicated by (e.g., represented by) a bit (blr,nei) indicating whether the position (pnei) of the TriSoup vertex is on a first half or a second half of the neighboring already-coded edge of the current edge (k), and a codeword (Cp,nei) indicating a distance between the middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge and the position (pnei) of the TriSoup vertex on the neighboring already-coded edge. The usage of the codeword may enable quantization to any number of intervals not limited to powers of two. For example, the leading bits of codewords (Cp,nei) may not represent anymore quarter edge position of TriSoup vertices along the neighboring already-coded edges. Derived neighborhood configurations βTSare thus likely to become much weaker due to inadequate input of bits from the codewords (Cp,nei). This would lead toDocket No.: 24-2010PCT degraded compression efficiency of the TriSoup vertex information and poorer compression capability of the overall TriSoup scheme.
[0212] Therefore, embodiments of the present disclosure relate to addressing the problem of degraded compression by generating TriSoup edge information Inei to comprise a neighborhood codeword (Cp,nei,2) that re-quantizes a quantize position of a neighboring TriSoup vertex according to a uniform quantizer. For example, the quantized position represented by the bit (blrnei) and the codeword (Cp,nei), as explained above, may be converted (e.g., transformed) into the neighborhood codeword. For example, the neighborhood codeword may indicate whether the quantized position of the neighboring TriSoup vertex is on a first, second, third, or fourth quarter of one neighboring already-coded edge of the current edge. In this example, the neighborhood codeword may be a 2-bit codeword Cp,nei,2with each possible value of the 2-bit codeword representing a respective quarter of the one neighboring already-coded edge.
[0213] The neighborhood configuration βTSmay then derived based on the neighborhood codeword Cp,nei,2(e.g., a 2- bit codeword) of at least one neighboring already-coded edge of the current edge (k). The derivation of the neighborhood configuration βTS is similar to the derivation as explained with respect to FIGS.16 and 17 with Nb =2. The neighborhood configuration βTScan thus be used without degrading compression capability.
[0214] In some embodiments, the first bit b1of the neighborhood codeword Cp,nei,2is associated with one of the at least one neighboring already-coded edge of the current edge (k) and may indicate whether a quantized position (pnei) of the neighboring TriSoup vertex is on a first half or a second half of the neighboring already-coded edge. For example, the first bit b1 of the neighborhood codeword may be the same as the bit (blrnei) used to represent the quantized position.
[0215] In some embodiments, the other bits (e.g., remaining bits and excluding the first bit) of the neighborhood codeword Cp,nei,2may indicate which of the power of two number of intervals in the half of the neighboring edge indicated by the first bit the quantized position (pnei) of the neighboring TriSoup vertex belongs. For example, when the neighborhood codeword Cp,nei,2is a 2-bit codeword, the second bit b2of the neighborhood codeword Cp,nei,2may indicate whether the position (pnei) of the TriSoup vertex is on a first half or the second half of the half of the neighboring already-coded edge indicated by the first bit b1 of the 2-bit codeword Cp,nei,2. In some embodiments, the second bit b2 of the 2-bit codeword Cp,nei,2may be flipped if the first bit b1is equal to 0. The 2-bit codeword may then be the concatenation b1b2 of the bit b1 and the (flipped) bit b2.
[0216] In some examples, the second bit b2 of the neighborhood codeword Cp,nei,2 may be derived based on a distance between the middle position of the at least one neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge and the quantized position (pnei) of the neighboring TriSoup vertex on the neighboring already-coded edge.
[0217] In some embodiments, the neighborhood codeword Cp,nei,2associated with a neighboring already-coded edge may be determined based on the quantized position (pQP,nei) of a TriSoup vertex along the neighboring already-coded edge.Docket No.: 24-2010PCT
[0218] Determining the neighborhood codeword Cp,nei,2 based on the quantized position (pQP,nei) may represent and result in a re-quantizing process introduced to preserve compression efficiency. The re-quantizing process is particularly advantageous in case the geometry quality parameter GQP is signaled locally and is not homogeneous over all TriSoup vertices. In this case, the bits (blr,nei) and the codewords (Cp,nei) associated with the neighboring already-coded TriSoup edges may have been quantized with several different geometry quality parameters GQP depending on the neighboring already-coded edge, which results in quantization intervals that may not be uniform across cuboids containing different portions of the point cloud. The re-quantizing process homogenizes the different quantization qualities to neighborhood codewords that may be combined to derive the neighborhood configuration βTS.Specifically, the neighborhood codewords may have the same number of bits (e.g., be 2-bit codewords) independent of number of quantization intervals used to quantize TriSoup vertices representing different portions of the point cloud. In practice, the neighborhood codewords may be 2-bit codewords that provides a fine enough granularity of neighborhood vertex information for encoding and decoding the current position of the TriSoup vertex. However, in other implementations, the neighborhood codewords may be fixed-bit codewords that is not 2, for example, they may be 3-bit codewords, 4-bit codewords, 6-bit codewords, etc.
[0219] For example, the first bit b1of the neighborhood codeword Cp,nei,2may equal to the bit (blr,nei) of the quantized position (pQP,nei). The quantized position (pQP,nei) may be obtained as discussed hereinabove in relation with FIG.23.
[0220] For example, the second bit b2of the neighborhood codeword (Cp,nei,2) may be derived based on the codeword (Cp,nei) of the quantized position (pQP,nei).
[0221] In some embodiments, the second bit b2 of the neighborhood codeword (Cp,nei,2) may be derived based on a distance (dDQ,k) indicated by the codeword (Cp,nei).
[0222] In some embodiments, the second bit b2of the neighborhood codeword (Cp,nei,2) may be derived based on the belonging of the distance (dDQ,k) to the first half or to the second half of the half edge indicated by the first bit of the neighborhood codeword (Cp,nei,2).
[0223] FIGS.24A-B illustrate examples of re-quantizing the distance indicated by the neighborhood codeword, according to some embodiments. For illustration purposes, the neighborhood codeword (Cp,nei,2) that represents a re- quantized value of a quantized position of a neighboring TriSoup vertex on an already-coded neighboring edge will be described with respect to a 2-bit codeword. However, it is to be understood the distance may be dequantized according to other lengths of the neighborhood codeword.
[0224] For example, the distance (dDQ,k) (2320) is determined as discussed above with respect to FIG.20 and the second bit b2 of the 2-bit codeword (Cp,nei,2) is derived based on the belonging of the distance (2320) to the first half (2300, b2 is set to 0) or to the second half (2301, b2 is set to 1) of the half edge the distance (dp) has been defined over as illustrated in FIG.24A.Docket No.: 24-2010PCT
[0225] For example, the second bit b2 of the 2-bit codeword (Cp,nei,2) may be derived based on the belonging of the distance (2320) to the first half (2310, b2 is set to 0) or to the second half (2311, b2 is set to 0) of the interval [0; B / 2-0.5] of the half edge the distance (dp) may necessarily belong to as illustrated in FIG.24B.
[0226] FIG.25 illustrates an example encoding process 2500 for encoding, in a bitstream, a presence flag and a current position of a TriSoup vertex of a current edge of a cuboid containing a portion of a point cloud, according to some embodiments.
[0227] The process 2500 may include many of the same operations (shown as having the same labeled blocks) as those described in FIG.21. Different from the process 2100 of FIG.21, the process 2500 includes block 2510 to re- quantize the quantized position (pQP,nei) of TriSoup edge information Ineiof at least one neighboring already-coded TriSoup edge of the current TriSoup edge k. The TriSoup edge information Inei may comprise (e.g., include or be made of) a presence flag (snei) of the at least one neighboring already-coded edge of the current edge (k), and if the presence flag (snei) is true, the TriSoup information Ineifurther comprises a bit (blr,nei) indicating whether the quantized position of a neighboring TriSoup vertex is on a first half or a second half of the neighboring already-coded edge, and a codeword (Cp,nei) indicating a distance between the middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge, and the quantized position of the neighboring TriSoup vertex on the neighboring already-coded edge.
[0228] At block 2510, a neighborhood codeword Cp,nei,2is derived by re-quantizing the quantized position (pQP,nei) of each of the at least one neighboring already-coded edge. For example, the neighborhood codeword may be a fixed-bit codeword such as a 2-bit codeword.
[0229] At block 1630, the neighborhood configuration βTSis derived based on the presence flags sneiand the neighborhood codeword Cp,nei,2of the at least one neighboring already-coded edge. For example, the neighborhood configuration βTS may comprise the neighborhood codeword such as at least two bits b1 and b2 of the neighborhood codeword.
[0230] FIG.26 illustrates an example decoding process 2600 for decoding, from the bitstream, an encoded presence flag and a current position of a TriSoup vertex of a current edge of a cuboid containing a portion of a point cloud, according to some embodiments. The process 2600 may include many of the same operations (shown as having the same labeled blocks) as those described in FIG.22.
[0231] At block 2510, a neighborhood codeword Cp,nei,2 is derived by re-quantizing the quantized position (pQP,nei) of each of the at least one neighboring already-coded edge. As explained with respect to FIG.25, the neighborhood codeword may be a fixed-length codeword such as a 2-bit codeword.
[0232] At block 1630, the neighborhood configuration βTS is derived based on the presence flags snei and the neighborhood codeword Cp,nei,2of the at least one neighboring already-coded edge. For example, the neighborhood configuration βTS may comprise the neighborhood codeword such as at least two bits b1 and b2 of the neighborhood codeword.Docket No.: 24-2010PCT
[0233] FIG.27 illustrates a flowchart 2700 of an example method for encoding in a bitstream a current position of a TriSoup vertex for a point cloud, according to some embodiments. The TriSoup vertex may be associated with a TriSoup scheme for representing geometry of the point cloud. For example, the method of flowchart 2700 may be performed by an encoder (e.g., encoder 114 of FIG.1).
[0234] At block 2710, an encoder determines, for a TriSoup vertex on a current edge (e.g., current edge k) of a cuboid containing a portion of a point cloud, a neighborhood codeword (e.g., Cp,nei,2) based on a quantized position of a neighboring TriSoup vertex on a neighboring already-coded edge of the current edge. The quantized position is indicated by a bit (e.g., blrnei) indicating whether the quantized position is on a first half or a second half of the neighboring already-coded edge; and a codeword (e.g., Cp,nei) indicating a distance between: the middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge, and the quantized position of the neighboring TriSoup vertex on the neighboring already-coded edge.
[0235] At block 2720, the encoder encodes (e.g., entropy encodes), in a bitstream and based on at least the neighborhood codeword, a current position (e.g., pQP,k) of the TriSoup vertex on the current edge.
[0236] In some examples, the encoder further encodes (e.g., entropy encodes), in the bitstream, a presence flag indicating the presence of the TriSoup vertex on the current edge. For example, the presence flag may be encoded as a one-bit indication.
[0237] FIG.28 illustrates a flowchart 2800 of an example method for decoding from a bitstream a current position of a TriSoup vertex for a point cloud, according to some embodiments. The TriSoup vertex may be associated with a TriSoup scheme for representing geometry of the point cloud. For example, the method of flowchart 2800 may be performed by a decoder (e.g., decoder 120 of FIG.1).
[0238] At block 2810, a decoder determines, for a TriSoup vertex on a current edge (e.g., current edge k) of a cuboid containing a portion of a point cloud, a neighborhood codeword (e.g., Cp,nei,2) based on a quantized position of a neighboring TriSoup vertex on a neighboring already-coded edge of the current edge. The quantized position is indicated by a bit (e.g., blrnei) indicating whether the quantized position is on a first half or a second half of the neighboring already-coded edge; and a codeword (e.g., Cp,nei) indicating a distance between: the middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge, and the quantized position of the neighboring TriSoup vertex on the neighboring already-coded edge. In some examples, operations performed by the decoder at block 2810 are the same as the operations performed by the encoder at block 2710 of FIG.27.
[0239] At block 2820, the decoder decodes (e.g., entropy decodes), from a bitstream and based on at least the neighborhood codeword, a current position (e.g., pQP,k) of the TriSoup vertex.
[0240] In some examples, the decoder further decodes (e.g., entropy decodes), from the bitstream, a presence flag indicating the presence of the TriSoup vertex on the current edge. For example, the presence flag may be decoded as a one-bit indication.Docket No.: 24-2010PCT
[0241] In some embodiments, the neighborhood codeword determined by the encoder and the decoder at blocks 2710 and 2810, respectively, converts the quantized position, indicated by the bit and the codeword, into a re-quantized position of a uniform quantizer. For example, the uniform quantizer may have uniform intervals on the neighboring already-coded edge. Generally, the quantity of uniform intervals may be a power of two. In these examples, the quantization scheme represented by the bit and the codeword may be converted to the uniform quantization scheme of the neighborhood codeword and the quantized position of the neighboring TriSoup vertex may be re-quantized into a re-quantized position indicated by the neighborhood codeword. Specifically the intervals of the quantization scheme may be converted to the uniform intervals (e.g., with a power of two number of intervals) associated with the neighborhood codeword.
[0242] For example, the uniform quantizer may be a 2-bit quantizer that converts (e.g., requantizes) the quantized position of the neighboring TriSoup vertex into one of four values. In this example, the neighborhood codeword may be a 2-bit codeword that indicates whether the quantized position of the neighboring TriSoup vertex is on a first, second, third, or fourth quarter of the neighboring already-coded edge. Other types of representation are possible. For example, the uniform quantizer may be a 4-bit quantizer with each bit indicating whether the quantized position is on a respective interval of the neighboring already-coded edge.
[0243] In some embodiments, the first bit of the neighborhood codeword is the same as the bit (blrnei), used as part of the representation for the quantized position, and indicates whether the quantized position of the neighboring TriSoup vertex is on the first half or the second half of the neighboring already-coded edge. In these embodiments, the remaining bits (e.g., the second bit) of the neighborhood codeword may be derived based on the codeword (Cp,nei) used in the representation of the quantized position. For example, at least the second bit of the neighborhood codeword may be derived based on the distance indicated by the codeword. For example, at least the second bit may be derived based on the belonging of the distance to the first half or to the second half of the half edge indicated by the first bit of the neighborhood codeword. For example, the second bit of the neighborhood codeword may be derived based on the belonging of the distance to the first half or to the second half of the half edge a second distance has been defined over, the second distance being defined between the middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge and the position of the TriSoup vertex on the neighboring already-coded edge. For example, the second bit of the neighborhood codeword may be derived based on the belonging of the distance to the first half or to the second half of the half edge a second distance may necessarily belong to, the second distance being defined between the middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge and the position of the TriSoup vertex on the neighboring already-coded edge.
[0244] In some embodiments, the encoder and decoder may identically determine the neighborhood codeword for respectively encoding and decoding the current position of the TriSoup vertex. For example, the neighborhood codeword may be used to generate a neighborhood configuration (βTS) or a reduced configuration (βTS’), which is usedDocket No.: 24-2010PCT to select a context for respectively encoding and decoding the current position of the TriSoup vertex. For example, the neighborhood codeword may be a portion of the neighborhood configuration.
[0245] In some embodiments, the current position may be encoded in and similarly decoded from the bitstream based on a plurality of neighborhood codewords including the neighborhood codeword. For example, a respective neighborhood codeword may be generated similarly as that described for blocks 2710 and 2810 to represent respective quantized positions of a plurality of neighboring TriSoup vertices of the TriSoup vertex. Examples of neighboring TriSoup vertices are described with respect to FIGS.11-14.
[0246] In some embodiments, the current position of the TriSoup vertex may be encoded and decoded using a non- uniform quantizer such as a quantizer that does not use a number of intervals that is a power of two. For example, the current position may be encoded in and decoded from a bitstream as a current bit indicating whether the current position of the current TriSoup vertex is on a first half or a second half of the current edge, and a current codeword (C_p,k) indicating the current position relative to the half of the current edge indicated by the current bit. For example, the current codeword may indicate a current distance between: the middle position of the current edge separating the first half and the second half of the current edge, and the current position of the TriSoup vertex on the current edge. By representing the current position using the bit and the current codeword, the current position may be represented by any number of intervals and allows for flexibility in quantization of the current position of the TriSoup vertex on the current edge because the current position is no longer restricted to be quantized according to a number of intervals that is a power of two.
[0247] FIG.29 illustrates a flowchart of an example process 2900 for coding a current position of a TriSoup vertex on a current edge based on re-quantizing the quantized position of a neighboring TriSoup vertex on a neighboring already- coded edge of the current edge for a point cloud, according to some embodiments. The TriSoup vertex may be associated with a TriSoup scheme for representing geometry of the point cloud. For example, the method of flowchart 2700 may be performed by an encoder (e.g., encoder 114 of FIG.1) or a decoder (e.g., decoder 120 of FIG.1).
[0248] At 2902, the process 2900 involves obtaining a quantized position of a neighboring TriSoup vertex on a neighboring already-coded edge of a current edge of a cuboid containing a portion of a point cloud. In some examples, the cuboid is associated with a TriSoup node. The cuboid is one of a plurality of cuboids partitioning a volume containing the point cloud. In some examples, the quantized position is indicated by a bit indicating whether the quantized position of the neighboring TriSoup vertex is on a first half or a second half of the neighboring already-coded edge and a codeword. The codeword indicates a distance between a middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge, and the quantized position of the neighboring TriSoup vertex on the neighboring already-coded edge.
[0249] At 2904, the process 2900 involves re-quantizing the quantized position of the neighboring TriSoup vertex to determine, for a TriSoup vertex on the current edge, a neighborhood codeword. In some examples, the neighborhood codeword indicates whether the quantized position of the neighboring TriSoup vertex is on a first, second, third, orDocket No.: 24-2010PCT fourth quarter of the neighboring already-coded edge. In some examples, the neighborhood codeword can include two bits. For instance, the neighborhood codeword converts the quantized position, indicated by the bit and the codeword, into the two bits. The first bit of the neighborhood codeword indicates whether the quantized position of the neighboring TriSoup vertex is on the first half or the second half of the neighboring already-coded edge. The second bit of the neighborhood codeword indicates whether the quantized position of the neighboring TriSoup vertex is on a first half or a second half of the half of the neighboring already-coded edge indicated by the first bit. In some examples, the second bit of the neighborhood codeword is flipped if the first bit equals to 0.
[0250] In some examples, the first bit of the neighborhood codeword is the same as the bit of the quantized position of the neighboring TriSoup. The second bit of the neighborhood codeword may be derived based on the codeword of the quantized position. The second bit of the neighborhood codeword may also be derived based on the distance indicated by the codeword. The second bit of the neighborhood codeword may further be derived based on the distance belonging to the first half or to the second half of the half edge indicated by the first bit of the neighborhood codeword. In further examples, the second bit of the neighborhood codeword may be derived based on the distance belonging to the first half or to the second half of the half edge a second distance has been defined over, and the second distance may be defined between the middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge and the position of the TriSoup vertex on the neighboring already- coded edge.
[0251] At 2906, the process 2900 involves entropy coding, based on at least the neighborhood codeword, a current position of the TriSoup vertex. In some examples, entropy coding the current position of the TriSoup vertex includes coding, for the TriSoup vertex a current bit indicating whether the current position of the current TriSoup vertex is on a first half or a second half of the current edge and a current codeword. The current codeword indicates a current distance between a middle position of the current edge separating the first half and the second half of the current edge and the current position of the TriSoup vertex on the current edge. For the decoder, the current position of the TriSoup vertex on the current edge can be determined based on the current bit and the current codeword.
[0252] In some examples, the current distance is defined over a segment of the current edge, the segment being uniformly partitioned into quantization intervals and a respective codeword is associated with each quantization interval. The current codeword indicating the distance from the middle position of the current edge and the position of the TriSoup vertex is associated with the quantization interval the current distance belongs to. The current distance may be determined between the middle position of the current edge and a center of the quantization interval indicated by the current codeword. The current distance can be an absolute difference between the middle position of the current edge and the current position of the TriSoup vertex. The current position of the TriSoup vertex may be determined by subtracting the distance from the middle position of the current TriSoup edge if the current bit indicates the first half edge and by adding the current distance to the middle position of the current edge if the current bit indicates the secondDocket No.: 24-2010PCT half edge. In some examples, the current distance of a last quantization interval is located to the right of the middle position of a left subinterval of the last quantization interval partitioned into the left subinterval and a right subinterval.
[0253] In some examples, the size of the quantization intervals is determined based on a quantization step. The quantization step may have sub-voxel precision. The quantization step may be represented by an integer value with a fixed number of bits of precision. The quantization step can be proportional to a power of two, and an exponent of the power is based on a scaled geometry quality parameter. The geometry quality parameter equals the integer value. In some examples, a geometry quality parameter is signaled in a bitstream to signal the size of the quantization step. For example, the geometry quality parameter is signaled in a high-level syntax of the bitstream. The geometry quality parameter may be signaled in a geometry parameter set or in headers of geometry data units. For instance, the geometry quality parameter is signaled at a TriSoup node level or in quality units that partition geometry data units.
[0254] For example, the decoder can determine the current distance based on the current bit, the current codeword, and the quantization step. The decoder can determine the current distance based on a linear combination of the current codeword and the quantization step.
[0255] In some examples, a presence flag indicating the presence of the TriSoup vertex on the current edge is also entropy coded (entropy encoded by the encoder or entropy decoded by the decoder). The presence flag, the current bit and bits of the current codeword of the current edge may be entropy coded (entropy encoded or decoded) based on neighborhood configuration derived based on information of at least one neighboring already-coded edge of the current edge. A bit representing the presence flag indicating the presence of the TriSoup vertex on the current edge can be entropy coded based on a first neighborhood configuration derived based on at least one bit representing at least one presence flag indicating the presence of at least one TriSoup vertex on one of the at least one neighboring already- coded edge. The current bit indicating whether the current position of a TriSoup vertex is on the first half or the second half of the current edge may be entropy coded based on a second neighborhood configuration derived based on at least one bit indicating whether a position of at least one TriSoup vertex is on a first half or a second half of the at least one neighboring already-coded edge. The bits representing the current codeword may be entropy coded based on a third neighborhood configuration derived based on bits representing at least one neighborhood codeword indicating at least one distance between the middle position of the at least one neighboring already-coded edge and a position of at least one TriSoup vertex. In some examples, only a leading bit of the current codeword is entropy coded based on a fourth neighborhood configuration derived based on at least one leading bit representing at least one neighborhood codeword indicating at least one distance between the middle position of the at least one neighboring already-coded edge and a position of at least one TriSoup vertex.
[0256] In some examples, entropy coding a bit based on neighborhood configuration includes selecting an entropy coder (entropy encoder or entropy decoder) based on the neighborhood configuration and entropy coding the current bit based on the selected entropy coder. For example, the entropy coder can be selected based on OBUF or dynamic OBUF taking the neighborhood configuration as input.Docket No.: 24-2010PCT
[0257] Embodiments of the present disclosure may be implemented in hardware using analog and / or digital circuits, in software, through the execution of instructions by one or more general purpose or special-purpose processors, or as a combination of hardware and software. Consequently, embodiments of the disclosure may be implemented in the environment of a computer system or other processing system. An example of such a computer system 3000 is shown in FIG.30. Blocks depicted in the figures above, such as the blocks in FIGS.1, 16-17, 21-23, and 25-29 may execute on one or more computer systems 3000. Furthermore, each of the steps of the flowcharts depicted in the present disclosure may be implemented on one or more computer systems 3000. When more than one computer system 3000 is used to implement embodiments of the present disclosure, the computer systems 3000 may be interconnected by one or more networks to form a cluster of computer systems that may act as a single pool of seamless resources. The interconnected computer systems 3000 may form a “cloud” of computers.
[0258] Computer system 3000 includes one or more processors, such as processor 3004. Processor 3004 may be, for example, a special purpose processor, general purpose processor, microprocessor, or digital signal processor. Processor 3004 may be connected to a communication infrastructure 3002 (for example, a bus or network). Computer system 3000 may also include a main memory 3006, such as random access memory (RAM), and may also include a secondary memory 3008.
[0259] Secondary memory 3008 may include, for example, a hard disk drive 3010 and / or a removable storage drive 3012, representing a magnetic tape drive, an optical disk drive, or the like. Removable storage drive 3012 may read from and / or write to a removable storage unit 3016 in a well-known manner. Removable storage unit 3016 represents a magnetic tape, optical disk, or the like, which is read by and written to by removable storage drive 3012. As will be appreciated by persons skilled in the relevant art(s), removable storage unit 3016 includes a computer usable storage medium having stored therein computer software and / or data.
[0260] In alternative implementations, secondary memory 3008 may include other similar means for allowing computer programs or other instructions to be loaded into computer system 3000. Such means may include, for example, a removable storage unit 3018 and an interface 3014. Examples of such means may include a program cartridge and cartridge interface (such as that found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a thumb drive and USB port, and other removable storage units 3018 and interfaces 3014 which allow software and data to be transferred from removable storage unit 3018 to computer system 3000.
[0261] Computer system 3000 may also include a communications interface 3020. Communications interface 3020 allows software and data to be transferred between computer system 3000 and external devices. Examples of communications interface 3020 may include a modem, a network interface (such as an Ethernet card), a communications port, etc. Software and data transferred via communications interface 3020 are in the form of signals which may be electronic, electromagnetic, optical, or other signals capable of being received by communications interface 3020. These signals are provided to communications interface 3020 via a communications path 3022.Docket No.: 24-2010PCT Communications path 3022 carries signals and may be implemented using wire or cable, fiber optics, a phone line, a cellular phone link, an RF link, and other communications channels.
[0262] Computer system 3000 may also include one or more sensor(s) 3024. Sensor(s) 3024 may measure or detect one or more physical quantities and convert the measured or detected physical quantities into an electrical signal in digital and / or analog form. For example, sensor(s) 3024 may include an eye tracking sensor to track the eye movement of a user. Based on the eye movement of a user, a display of a point cloud may be updated. In another example, sensor(s) 3024 may include a head tracking sensor to track the head movement of a user. Based on the head movement of a user, a display of a point cloud may be updated. In yet another example, sensor(s) 3024 may include a camera sensor for taking photographs and / or a 3D scanning device, like a laser scanning, structured light scanning, and / or modulated light scanning device.3D scanning devices may determine geometry information by moving one or more laser heads, structured light, and / or modulated light cameras relative to the object or scene being scanned. The geometry information may be used to construct a point cloud.
[0263] As used herein, the terms “computer program medium” and “computer readable medium” are used to refer to tangible storage media, such as removable storage units 3016 and 3018 or a hard disk installed in hard disk drive 3010. These computer program products are means for providing software to computer system 3000. Computer programs (also called computer control logic) may be stored in main memory 3006 and / or secondary memory 3008. Computer programs may also be received via communications interface 3020. Such computer programs, when executed, enable computer system 3000 to implement the present disclosure as discussed herein. In particular, the computer programs, when executed, enable processor 3004 to implement the processes of the present disclosure, such as any of the methods described herein. Accordingly, such computer programs represent controllers of the computer system 3050.
[0264] In another embodiment, features of the disclosure may be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. Implementation of a hardware state machine to perform the functions described herein will also be apparent to persons skilled in the relevant art(s).
Claims
Docket No.: 24-2010PCT CLAIMS What is claimed is:
1. A method comprising: obtaining a quantized position of a neighboring TriSoup vertex on a neighboring already-coded edge of a current edge of a cuboid containing a portion of a point cloud; re-quantizing the quantized position of the neighboring TriSoup vertex to determine, for a TriSoup vertex on the current edge, a neighborhood codeword, the neighborhood codeword indicating whether the quantized position of the neighboring TriSoup vertex is on a first, second, third, or fourth quarter of the neighboring already-coded edge; and entropy coding, based on at least the neighborhood codeword, a current position of the TriSoup vertex.
2. A method comprising: obtaining a quantized position of a neighboring TriSoup vertex on a neighboring already-coded edge of a current edge of a cuboid containing a portion of a point cloud; re-quantizing the quantized position of the neighboring TriSoup vertex to determine, for a TriSoup vertex on the current edge, a neighborhood codeword; and entropy coding, based on at least the neighborhood codeword, a current position of the TriSoup vertex.
3. The method according to any one of claims 1-2, further comprising entropy coding a presence flag indicating the presence of the TriSoup vertex on the current edge.
4. The method of any one of claims 1-3, wherein the entropy coding the current position of the TriSoup vertex comprises: coding, for the TriSoup vertex: a current bit indicating whether the current position of the current TriSoup vertex is on a first half or a second half of the current edge; and a current codeword indicating a current distance between: a middle position of the current edge separating the first half and the second half of the current edge; and the current position of the TriSoup vertex on the current edge.
5. The method of claim 4, further comprising determining the current position of the TriSoup vertex on the current edge based on the current bit and the current codeword.
6. The method of any one of claims 2-5, wherein the neighborhood codeword indicates whether the quantized position of the neighboring TriSoup vertex is on a first, second, third, or fourth quarter of the neighboring already- coded edge.
7. The method of any one of claims 1-6, wherein the neighborhood codeword comprises two bits.Docket No.: 24-2010PCT 8. The method of claim 7, wherein the neighborhood codeword converts the quantized position, indicated by the bit and the codeword, into the two bits.
9. The method of claim 7, wherein a first bit of the neighborhood codeword indicates whether the quantized position of the neighboring TriSoup vertex is on the first half or the second half of the neighboring already-coded edge.
10. The method of any one of claims 7-9, wherein the second bit of the neighborhood codeword indicates whether the quantized position of the neighboring TriSoup vertex is on a first half or a second half of the half of the neighboring already-coded edge indicated by the first bit.
11. The method of claim 10, wherein the second bit of the neighborhood codeword is flipped if the first bit equals to 0.
12. The method of any one of claims 1-11, wherein the quantized position is indicated by: a bit indicating whether the quantized position of the neighboring TriSoup vertex is on a first half or a second half of the neighboring already-coded edge; and a codeword indicating a distance between: a middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge, and the quantized position of the neighboring TriSoup vertex on the neighboring already-coded edge.
13. The method of claim 12, wherein a first bit of the neighborhood codeword is the same as the bit.
14. The method of any one of claims 12-13, wherein the second bit of the neighborhood codeword is derived based on the codeword of the quantized position.
15. The method of any one of claims 12-14, wherein the second bit of the neighborhood codeword is derived based on the distance indicated by the codeword.
16. The method of any one of claims 12-15, wherein the second bit of the neighborhood codeword is derived based on the distance belonging to the first half or to the second half of the half edge indicated by the first bit of the neighborhood codeword.
17. The method of any one of claims 12-15, wherein the second bit of the neighborhood codeword is derived based on the distance belonging to the first half or to the second half of the half edge a second distance has been defined over, the second distance being defined between the middle position of the neighboring already-coded edge separating the first half and the second half of the neighboring already-coded edge and the position of the TriSoup vertex on the neighboring already-coded edge.
18. The method of any one of claims 4-17, wherein the current distance is defined over a segment of the current edge, the segment being uniformly partitioned into quantization intervals and a respective codeword is associated with each quantization interval.
19. The method of claim 18, wherein the current codeword indicating the distance from the middle position of the current edge and the position of the TriSoup vertex is associated with the quantization interval the current distance belongs to.Docket No.: 24-2010PCT 20. The method of any one of claims 18-19, wherein the current distance is determined between the middle position of the current edge and a center of the quantization interval indicated by the current codeword.
21. The method of any one of claims 18-20, wherein the current distance is an absolute difference between the middle position of the current edge and the current position of the TriSoup vertex.
22. The method of any one of claims 18-21, wherein the current position of the TriSoup vertex is determined by subtracting the distance from the middle position of the current TriSoup edge if the current bit indicates the first half edge and by adding the current distance to the middle position of the current edge if the current bit indicates the second half edge.
23. The method of any one of claims 18-22, wherein the current distance of a last quantization interval is located to the right of the middle position of a left subinterval of the last quantization interval partitioned into the left subinterval and a right subinterval.
24. The method of any one of claims 18-23, wherein a size of the quantization intervals is determined based on a quantization step.
25. The method of claim 24, wherein the quantization step has sub-voxel precision.
26. The method of any one of claims 24-25, wherein the quantization step is represented by an integer value with a fixed number of bits of precision.
27. The method of any one of claims 24-26, wherein the quantization step is proportional to a power of two, and wherein an exponent of the power is based on a scaled geometry quality parameter.
28. The method of claim 27, wherein the scaled geometry quality parameter equals the integer value.
29. The method of any one of claims 24-28, wherein a geometry quality parameter is signaled in a bitstream to signal the size of the quantization step.
30. The method of claim 29, wherein the geometry quality parameter is signaled in a high-level syntax of the bitstream.
31. The method of claim 29, wherein the geometry quality parameter is signaled in a geometry parameter set or in headers of geometry data units.
32. The method of claim 29, wherein the geometry quality parameter is signaled at a TriSoup node level or in quality units that partition geometry data units.
33. The method of any one of claims 24-32, wherein the current distance is determined based on the current bit, the current codeword, and the quantization step.
34. The method of any one of claims 24-33, wherein the current distance is determined based on a linear combination of the current codeword and the quantization step.
35. The method according to any one of claims 1-34, wherein the cuboid is associated with a TriSoup node.
36. The method according to any one of claims 1-35, wherein the cuboid is one of a plurality of cuboids partitioning a volume containing the point cloud.Docket No.: 24-2010PCT 37. The method of any one of claims 4-36, wherein the presence flag, the current bit and bits of the current codeword of the current edge are entropy coded based on neighborhood configuration derived based on information of at least one neighboring already-coded edge of the current edge.
38. The method of any one of claims 3-37, wherein a bit representing the presence flag indicating the presence of the TriSoup vertex on the current edge is entropy coded based on a first neighborhood configuration derived based on at least one bit representing at least one presence flag indicating the presence of at least one TriSoup vertex on one of the at least one neighboring already-coded edge.
39. The method of any one of claims 4-38, wherein the current bit indicating whether the current position of a TriSoup vertex is on the first half or the second half of the current edge is entropy coded based on a second neighborhood configuration derived based on at least one bit indicating whether a position of at least one TriSoup vertex is on a first half or a second half of the at least one neighboring already-coded edge.
40. The method of any one of claims 4-39, wherein the bits representing the current codeword are entropy coded based on a third neighborhood configuration derived based on bits representing at least one neighborhood codeword indicating at least one distance between the middle position of the at least one neighboring already- coded edge and a position of at least one TriSoup vertex.
41. The method of any one of claims 4-40, wherein only a leading bit of the current codeword is entropy coded based on a fourth neighborhood configuration derived based on at least one leading bit representing at least one neighborhood codeword indicating at least one distance between the middle position of the at least one neighboring already-coded edge and a position of at least one TriSoup vertex.
42. The method according to any one of claims 38-41, wherein entropy coding a bit based on neighborhood configuration comprises selecting an entropy coder based on the neighborhood configuration and entropy coding the current bit based on the selected entropy coder.
43. The method of claim 42, wherein the entropy coder is selected based on OBUF or dynamic OBUF taking the neighborhood configuration as input.
44. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of an apparatus, cause the apparatus to perform the method of any one of claims 1-43.
45. An encoder comprising: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the encoder to perform the method of any one of claims 1-4, 6-32, or 35-43.
46. A non-transitory computer-readable recording medium storing a bitstream generated by the method for encoding a video according to any one of claims 1-4, 6-32, or 35-43.
47. A decoder comprising: one or more processors; andDocket No.: 24-2010PCT memory storing instructions that, when executed by the one or more processors, cause the decoder to perform the method of any one of claims 1-43.
48. A non-transitory computer readable medium storing a bitstream, which, when decoded by a decoder, causes the decoder to perform the method according to any one of claims 1-43.
49. A bitstream generated according to any one of claims 1-4, 6-32, or 35-43.
Citation Information
Patent Citations
Methods and apparatus for entropy coding a presence flag for a point cloud and data stream including the presence flag
EP4258213A1
Apparatus for coding vertex position for point cloud, and data stream including vertex position
WO2023193533A1