Planar patterns in point cloud encoding based on octree
By adopting a planar encoding mode in point cloud compression, identifying and utilizing the planarity of volume, the problem of low point cloud data compression efficiency in the prior art is solved, and more efficient storage and transmission are achieved.
Patent Information
- Application Number
- CN202080047369.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-28
- Filing Date
- 2020-06-04
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-06-04
AI Technical Summary
The prior art is difficult to efficiently compress point cloud data, resulting in waste of storage and transmission bandwidth, especially when dealing with point clouds in non-natural environments. Traditional tree structures such as octree and KD trees have inefficiency in signaling.
The plane encoding mode is used to identify the planarity of the volume and use the plane mode mark and the plane position mark to signal the plane characteristics of the volume, thereby optimizing the encoding and decoding process of point cloud data.
By utilizing planarity information, the number of bits that need to be encoded is reduced, the compression efficiency is improved, the bandwidth requirements for storage and transmission are reduced, and the operational performance of point cloud data is significantly improved.
Smart Images

Figure CN114073095B_ABST
Abstract
Description
Technical Field
[0001] The present application relates generally to point cloud compression, and in particular to methods and apparatus for improved compression of occupancy data in octree-based encoding of point clouds. Background Art
[0002] Data compression is used in communications and computer networks to efficiently store, send, and reproduce information. There is growing interest in the representation of three-dimensional objects or spaces, which can involve large data sets, and for which efficient and effective compression would be very useful and valuable. In some cases, a three-dimensional object or space can be represented using a point cloud, which is a collection of points that each have a three-dimensional coordinate position (X, Y, Z) and in some cases other attributes like color data (e.g., brightness and chromaticity), transparency, reflectivity, normal vectors, etc. Point clouds can be static (a fixed object, or a snapshot of an environment / object at a single point in time) or dynamic (a time-ordered sequence of point clouds).
[0003] Example applications for point clouds include topographic and mapping applications. Autonomous vehicles and other machine vision applications may rely on point cloud sensor data of the environment in the form of 3D scans (such as from a LiDAR scanner). Virtual reality simulations may rely on point clouds.
[0004] It will be appreciated that point clouds can involve large amounts of data, and that quickly and accurately compressing (encoding and decoding) this data is of significant interest. It would therefore be advantageous to provide methods and apparatus for more efficiently and / or effectively compressing data for point clouds. Such methods may result in savings in storage requirements (memory) through improved compression, or savings in bandwidth for transmitting the compressed data, resulting in improved operation of 3D vision systems (such as for vehicular applications), or improved operating speed and rendering of, for example, virtual reality systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Reference will now be made, by way of example, to the accompanying drawings which show example embodiments of the present application, and in which:
[0006] Figure 1 A simplified block diagram of an example point cloud encoder is shown;
[0007] Figure 2 A simplified block diagram of an example point cloud decoder is shown;
[0008] Figure 3 and Figure 4 illustrates an example of a volume that exhibits planarity within the child sub-volume it occupies;
[0009] Figure 5 An example method for encoding point cloud data using a planar encoding mode is shown in flowchart form;
[0010] Figure 6 An example method for decoding point cloud data using a planar coding mode is shown in flowchart form;
[0011] Figure 7 A portion showing one example of a process for encoding occupied bits based on planarity information;
[0012] Figure 8 A portion showing another example of a process for encoding occupied bits based on planarity information;
[0013] Fig. 9 Diagrammatically illustrating possible factors in determining a context for encoding a plane mode flag or a plane position flag;
[0014] Fig.10 An example mechanism for tracking the closest already encoded occupied node at the same depth and in a common plane is shown;
[0015] Fig.11 An example simplified block diagram showing an encoder; and
[0016] Fig.12 An example simplified block diagram of a decoder is shown.
[0017] Similar reference numerals may have been used in different drawings to denote similar components. DETAILED DESCRIPTION
[0018] The present application describes methods of encoding and decoding point clouds, as well as encoders and decoders for encoding and decoding point clouds.
[0019] In one aspect, the present application describes a method of encoding a point cloud to generate a bitstream of compressed point cloud data, the compressed point cloud data representing a three-dimensional position of an object, the point cloud being located within a volume space, the volume space being recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first group of child sub-volumes and a second group of child sub-volumes, the first group of child sub-volumes being located in a first plane and the second group of child sub-volumes being located in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points. The method may include: determining whether the volume is planar based on whether all child sub-volumes containing at least one point are located in the first group or the second group; encoding a planar mode flag in the bitstream to signal whether the volume is planar; in the bitstream, encoding occupancy bits for the child sub-volumes of the first group includes: for at least one occupancy bit, inferring a value of the at least one occupancy bit based on whether the volume is planar and not encoding the at least one occupancy bit in the bitstream; and outputting a bitstream of compressed point cloud data.
[0020] In another aspect, the present application describes a method of decoding a bitstream of compressed point cloud data to produce a reconstructed point cloud, the reconstructed point cloud representing a three-dimensional position of a physical object, the point cloud being located within a volumetric space, the volumetric space being recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first group of child sub-volumes and a second group of child sub-volumes, the first group of child sub-volumes being located in a first plane, and the second group of child sub-volumes being located in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points. The method may include: reconstructing the occupancy bits to reconstruct the points of the point cloud by: decoding a planar mode flag from the bitstream, the planar mode flag indicating whether the volume is planar, wherein the volume is planar if all child sub-volumes containing at least one point are located in the first group or the second group; and decoding the occupancy bits for the child sub-volumes of the first group from the bitstream includes: for at least one occupancy bit, inferring a value of the at least one occupancy bit based on whether the volume is planar and not decoding the at least one occupancy bit from the bitstream.
[0021] In some implementations, determining whether the volume is planar may include: determining that the volume is planar by determining that at least one of the child subvolumes in the first group contains at least one of the points and none of the child subvolumes in the second group contains any of the points, and the method may also include: encoding a plane position flag based on the volume being planar to signal that at least one of the child subvolumes is in the first plane. In such implementations, encoding the occupied bits in some cases includes: prohibiting encoding of occupied bits associated with the second group, and inferring a value for occupied bits associated with the second group based on the second group not containing points. Encoding the occupied bits may also include: inferring that a last occupied bit in the coding order of the occupied bits associated with the first group has a value indicating occupied based on determining that all other occupied bits in the coding order of the first group have values indicating unoccupied.
[0022] In some implementations, determining whether the volume is planar includes determining that the volume is not planar, and based on that, encoding the occupancy bit based on having a value indicating occupied, at least one of the occupancy bits in the first group, and at least one of the occupancy bits in the second group.
[0023] In some implementations, the point cloud is defined with respect to Cartesian axes in the volumetric space, the Cartesian axes having a vertically oriented z-axis perpendicular to the horizontal plane, and wherein the first plane and the second plane are parallel to the horizontal plane. In some implementations, the first plane and the second plane are orthogonal to the horizontal plane.
[0024] In some implementations, the method includes first determining that the volume is eligible ("eligible") for planar mode encoding. Determining that the volume is eligible for planar mode encoding may include determining a probability of planarity and determining that the probability of planarity is greater than a threshold eligibility value.
[0025] In some implementations, the coded plane mode flag may include: a coded horizontal plane mode flag and a coded vertical plane mode flag.
[0026] In yet another aspect, the present application describes a method of encoding a point cloud to generate a bitstream of compressed point cloud data representing a three-dimensional position of an object, the point cloud being located within a volumetric space that is recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first set of child sub-volumes and a second set of child sub-volumes, the first set of child sub-volumes being positioned in a first plane and the second set of child sub-volumes being positioned in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points, the first plane and the second plane both being orthogonal to an axis. The method may include: determining whether a volume is planar based on whether all child sub-volumes containing at least one point are positioned in a first group or a second group; entropy encoding a planar mode flag in a bitstream to signal whether the volume is planar, wherein the entropy encoding includes determining a context for encoding the planar mode flag based in part on one or more of: (a) whether a parent volume containing the volume is planar in occupancy, (b) occupancy of a neighboring volume at the parent's depth, the neighboring volume being adjacent to the volume and having a common face with the parent volume, or (c) a distance between the volume and a nearest already-encoded occupied volume at the same depth as the volume and having a position on the same axis as the volume; encoding occupancy bits for at least some of the child sub-volumes; and outputting a bitstream of compressed point cloud data.
[0027] In another aspect, the present application describes a method of decoding a bitstream of compressed point cloud data to produce a reconstructed point cloud representing a three-dimensional position of a physical object, the point cloud being located within a volumetric space that is recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first set of child sub-volumes and a second set of child sub-volumes, the first set of child sub-volumes being positioned in a first plane and the second set of child sub-volumes being positioned in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points, the first plane and the second plane both being orthogonal to an axis. The method may include reconstructing occupancy bits to reconstruct points of a point cloud by: entropy decoding a planar mode flag from a bitstream, the planar mode flag indicating whether a volume is planar, wherein the volume is planar if all child sub-volumes containing at least one point are positioned in a first group or a second group, wherein the entropy decoding includes determining a context for decoding the planar mode flag based in part on one or more of: (a) whether a parent volume containing the volume is planar in occupancy, (b) occupancy of a neighboring volume at a depth of the parent, the neighboring volume being adjacent to the volume and having a common face with the parent volume, or (c) a distance between the volume and a nearest already encoded occupied volume at the same depth as the volume and having a position on the same axis as the volume; and reconstructing the occupancy for the child sub-volumes.
[0028] In some implementations, if the parent planar mode flag indicates that the parent volume is planar, then the parent volume of the containing volume is planar in occupancy.
[0029] In some implementations, the distance is close or far and may be based on calculating a distance metric and comparing it to a threshold.
[0030] In some implementations, determining the context for the coding plane mode flag may be based on a combination of (a), (b), and (c).
[0031] In some implementations, determining whether the volume is planar includes determining that the volume is planar and thus entropy encoding a plane position flag to signal whether at least one point is located in the first group or the second group. Entropy encoding the plane position flag may include determining a context for encoding the plane position flag. Determining the context may be based in part on one or more of: (a') occupancy of neighboring volumes at a parent depth; (b') a distance between the volume and the nearest occupied volume that has been encoded; (c') a plane position of the nearest occupied volume that has been encoded (if any); or (d') a position of the volume within the parent volume. In some cases, determining the context for encoding the plane position flag may be based on a combination of three or more of (a'), (b'), (c'), and (d').
[0032] In some implementations, the distance is close, not too far, or far, and can be based on calculating a distance metric and comparing it to a first threshold and a second threshold.
[0033] In another aspect, the present application describes encoders and decoders configured to implement such encoding and decoding methods.
[0034] In yet another aspect, the present application describes a non-transitory computer-readable medium storing computer-executable program instructions that, when executed, cause one or more processors to perform the described encoding and / or decoding methods.
[0035] In yet another aspect, the present application describes a computer-readable signal containing program instructions which, when executed by a computer, cause the computer to perform the described encoding and / or decoding methods.
[0036] Notice Other aspects and features of the present application will be understood by those of ordinary skill in the art upon reviewing the following description of examples in conjunction with the accompanying drawings.
[0037] Any feature described in relation to one aspect or embodiment of the invention may also be used in relation to one or more other aspects / embodiments.These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described herein.
[0038] Sometimes in the description below, the terms "node", "volume", and "subvolume" may be used interchangeably. It will be appreciated that a node is associated with a volume or subvolume. A node is a specific point on a tree, which may be an internal node or a leaf node. A volume or subvolume is a bounded physical space represented by a node. The term "volume" may be used in some cases to refer to the largest bounded space defined to contain a point cloud. A volume may be recursively divided into subvolumes for the purpose of building out a tree structure of interconnected nodes for encoding point cloud data. The tree structure that partitions a volume into subvolumes may be referred to as a "father" and "child" relationship, where a subvolume is a child node or child subvolume for a parent node or parent volume. Subvolumes within the same volume may be referred to as sibling nodes or sibling subvolumes.
[0039] In the present application, the term “and / or” is intended to cover all possible combinations and subcombinations of the listed elements, including only any one element, any subcombination or all of the listed elements, and does not necessarily exclude additional elements.
[0040] In the present application, the phrase "at least one of... or..." is intended to cover any one or more of the listed elements, including only any one, any subcombination or all of the listed elements, without necessarily excluding any additional elements, and not necessarily requiring all elements.
[0041] A point cloud is a set of points in a three-dimensional coordinate system. Points are often intended to represent the external space of one or more objects. Each point has a location (position) in a three-dimensional coordinate system. The position can be represented by three coordinates (X, Y, Z), which can be a Cartesian or any other coordinate system. A point can have other associated attributes (such as color), which in some cases can also be three component values (such as R, G, B or Y, Cb, Cr). Depending on the desired application for the point cloud data, other associated attributes may include transparency, reflectivity, normal vectors, etc.
[0042] Point clouds can be static or dynamic. For example, a detailed scan or mapping of an object or terrain can be static point cloud data. A LiDAR-based scan of an environment for machine vision purposes can be dynamic in that the point cloud (at least potentially) changes over time, e.g., with each successive scan of a volume. A dynamic point cloud is thus a time-ordered sequence of point clouds.
[0043] Point cloud data can be used in many applications, including conservation (scanning of historical or cultural objects), mapping, machine vision (such as, autonomous or semi-autonomous vehicles), and virtual reality systems, to name a few examples. Dynamic point cloud data for applications like machine vision can be quite different from static point cloud data for applications like conservation purposes. Vehicle vision (e.g., typically involves relatively small resolution, non-color, highly dynamic point clouds obtained by lidar (or similar) sensors at high capture frequencies. Such point clouds are not intended for human consumption or viewing, but for machine object detection / classification in decision-making processes. As an example, a common lidar frame contains on the order of tens of thousands of points, while high-quality virtual reality applications require millions of points. It can be expected that there will be a demand for higher resolution data over time as computing speeds increase and new applications are discovered.
[0044] Despite the usefulness of point cloud data, the lack of effective and efficient compression (i.e., encoding and decoding processes) can hinder adoption and deployment. A particular challenge in encoding point clouds that does not arise with other data compression (like audio or video) is the encoding of the geometry of the point cloud. Point clouds tend to be sparsely populated, which makes it more challenging to efficiently encode the locations of the points.
[0045] One of the more common mechanisms for encoding point cloud data is by using a tree-based structure. In a tree-based structure, a bounded three-dimensional volume for a point cloud is recursively divided into sub-volumes. The nodes of the tree correspond to sub-volumes. The decision whether to further divide the sub-volume can be based on the resolution of the tree, and / or whether there are any points contained in the sub-volume. A node can have an occupancy flag that indicates whether its associated sub-volume contains points. A split flag can signal whether a node has child nodes (i.e., whether the current volume has been further split into sub-volumes). These flags can be entropy coded in some cases and predictive coding can be used in some cases.
[0046] A commonly used tree structure is an octree. In this structure, the volumes / subvolumes are all cubes and each split of a subvolume results in eight further subvolumes / subcubes. Another commonly used tree structure is a KD-tree, in which the volume (cube or rectangular cuboid) is recursively divided into two halves by a plane orthogonal to one of the axes. An octree is a special case of a KD-tree, in which the volume is divided by three planes, each of which is orthogonal to one of the three axes. The partitioning of a volume does not necessarily become two subvolumes (KD-tree) or eight subvolumes (octree), but can include other partitionings, including: partitioning into non-rectangular shapes or including non-adjacent subvolumes.
[0047] This application may refer to octrees for ease of description and because they are a popular candidate tree structure for autonomous driving applications, but it will be understood that the methods and apparatus described herein may be implemented using other tree structures.
[0048] Reference now Figure 1 , which shows a simplified block diagram of a point cloud encoder 10 according to aspects of the present application. The point cloud encoder 10 includes a tree construction module 12 for receiving point cloud data and generating a tree (in this example, an octree) that represents the geometry of a volume space containing a point cloud and indicates the location or position of points from the point cloud in the geometry.
[0049] In the case of a uniformly partitioned tree structure (like an octree), each node can be represented by a sequence of occupancy bits, where each occupancy bit corresponds to one of the subvolumes in the node and signals whether the subvolume contains at least one point. The occupied subvolumes are recursively split up to the maximum depth of the tree. This can be named as serialization or binarization of the tree. Figure 1 As shown in FIG. 1 , in this example, the point cloud encoder 10 includes a binarizer 14 for binarizing the octree to generate a bit stream representing binarized data of the tree.
[0050] This sequence of bits can then be encoded using an entropy encoder 16 to produce a compressed bit stream. The entropy encoder 16 can encode the sequence of bits using a context model 18, which specifies the probability of encoding the bits based on the context determined by the entropy encoder 16. After encoding each bit or a limited set of bits, the context model 18 can be adaptively updated. The entropy encoder 16 can be a binary arithmetic encoder in some cases. The binary arithmetic encoder can adopt context adaptive binary arithmetic coding (CABAC) in some implementations. In some implementations, an encoder other than an arithmetic encoder can be used.
[0051] In some cases, the entropy encoder 16 may not be a binary encoder, but may instead operate on non-binary data. The output octree data from the tree construction module 12 may not be evaluated in binary form, but may instead be encoded as non-binary data. For example, in the case of an octree, the eight flags (e.g., occupancy flags) within the subvolume in their scan order may be considered as 2 8 -1 bit number (e.g., an integer with a value between 1 and 255, since the value 0 cannot be used for a split subvolume, i.e., if it is completely unoccupied, then it will not have been split yet). In some implementations, this number can be encoded by the entropy encoder using a multi-symbol arithmetic encoder. Within a subvolume (e.g., a cube), the sequence of symbols defining this integer can be named a "pattern".
[0052] A convention commonly used in point cloud compression is that an occupancy bit value of 1 signals that the associated node or volume is "occupied", i.e., it contains at least one point, and an occupancy bit value of 0 signals that the associated node or volume is "unoccupied", i.e., it contains no points. More generally, an occupancy bit may have a value indicating occupied or a value indicating unoccupied. In the following description, for ease of illustration, example embodiments are described in which the convention of 1 = occupied and 0 = unoccupied is used; however, it will be understood that the present application is not limited to this convention.
[0053] A block diagram of an example point cloud decoder 50 corresponding to encoder 10 is shown at Figure 2 . The point cloud decoder 50 comprises an entropy decoder 52 that uses the same context model 54 used by the encoder 10. The entropy decoder 52 receives an input bit stream of compressed data and entropy decodes the data to produce an output sequence of decompressed bits. The sequence is then converted into reconstructed point cloud data by a tree reconstructor 56. The tree reconstructor 56 reconstructs the tree structure based on the decompressed data and knowledge of the scan order in which the tree data is binarized. The tree reconstructor 56 is therefore able to reconstruct the positions of the points from the point cloud (subject to the resolution of the tree encoding).
[0054] In European patent application no.18305037.6, the applicant describes a method and apparatus for selecting among available style distributions based on some occupancy information from previously encoded nodes near the specific node for use in encoding the occupancy style of a specific node. In one example implementation, the occupancy information is obtained from the occupancy style of the father of the specific node. In another example implementation, the occupancy information is obtained from one or more nodes adjacent to the specific node. The contents of European patent application no.18305037.6 are incorporated herein by reference. This is referred to as determining a "neighborhood configuration" and selecting a context (i.e., a style distribution) based at least in part on the neighborhood configuration.
[0055] In European patent application no. 18305415.4 the applicant describes a method and apparatus for binary entropy coding of occupancy patterns. The content of European patent application no. 18305415.4 is incorporated herein by reference.
[0056] Certain types of point cloud data tend to have strong directionality. Non-natural environments especially exhibit strong directionality because those environments tend to feature uniform surfaces. For example, in the case of LiDAR, the walls of roads and adjacent buildings are generally flat, either horizontally or vertically. In the case of interior scanning in a room, the floor, ceiling, and walls are all flat. LiDAR used for the purpose of vehicle vision and similar applications tends to be lower resolution and also needs to be compressed quickly and efficiently.
[0057] Octrees are efficient tree structures because they are based on a process of uniform partitioning of a cube into eight sub-cubes using three orthogonal planes in all cases, so signaling their structure is efficient. However, octrees using current signaling processes cannot take advantage of the efficiency available from recognizing the planar properties of some unnatural environments. However, KD trees are able to better customize the partitioning for the directionality of point clouds. This makes them a more efficient and effective structure for these types of environments. The disadvantage of KD trees is that the signaling of their structure requires significantly more data than octrees. The fact that KD trees are non-uniform means that some of the techniques used to improve octree compression are not available to KD trees or will be computationally difficult to implement.
[0058] Therefore, it would be advantageous to have a mechanism for representing unnatural environments using a uniform partitioning based tree structure in a manner that improves compression by exploiting horizontal and / or vertical directionality.
[0059] According to one aspect of the present application, an improved point cloud compression process and apparatus are characterized by a planar coding mode. Planar mode is signaled to indicate that a volume satisfies certain requirements for planarity in terms of its occupancy. Specifically, a volume is planar if all occupied subvolumes of the volume are positioned in or located in a common plane. The syntax used for signaling can indicate whether the volume is planar, and if so, the location of the common plane. By exploiting this knowledge of planarity, gains in compression can be achieved. Applying eligible criteria for enabling planar mode and a mechanism for context-adaptive coding of planar mode signaling helps improve compression performance.
[0060] In the following description, planarity is assumed to be with respect to the Cartesian axes aligned with the structure of the volume and subvolumes. That is, if all occupied subvolumes of the volume are positioned in a common plane orthogonal to one of the axes, the volume is planar. As a convention, the axis will assume that the z-axis is vertical, meaning that the (horizontal) plane is orthogonal to the z-axis. In many of the following examples, horizontal planarity will be used to illustrate the concept; however, it will be appreciated that the present application is not limited to horizontal planarity and may alternatively or additionally include vertical planarity with respect to the x-axis, the y-axis, or both the x and y-axis. In addition, in some examples, planarity is not necessarily aligned with the Cartesian axes by orthogonality. For illustration, in one example, diagonal vertical planarity can be defined, i.e., at a 45 degree angle with both the x and y axes.
[0061] Reference now Figure 3 and Figure 4 , each of which illustrates example volumes 300 and 400. In this example, horizontal planarity will be illustrated and discussed, but one of ordinary skill in the art will recognize that these concepts extend to vertical or other planarities.
[0062] Volume 300 is shown as being divided into eight subvolumes. Occupied subvolumes are indicated using shading, while unoccupied subvolumes are shown as empty. It will be noted that the lower four subvolumes (in the z-axis or vertical sense) in volume 300 are occupied. This occupancy pattern is horizontally planar; that is, each subvolume in the occupied subvolumes is in the same horizontal plane, i.e., has the same z position. Volume 400 shows another example of an occupied horizontal planar pattern. All occupied subvolumes of volume 400 are in the same horizontal plane. Volume 300 shows a situation where the lower plane is occupied. Volume 400 shows a situation where the upper plane is occupied. This can be named "plane position", where the plane position signals where the planar subvolume is located within the volume. In this case, it is a binary "above" or "below" signal.
[0063] The planarity of a volume is not limited to the case where all subvolumes of a plane (e.g., all subvolumes of the upper half of a 2x 2x 2 volume) are occupied. In some cases, only some of the subvolumes in the plane are occupied, provided that there are no occupied subvolumes outside the plane. In fact, as few as one occupied subvolume can be considered "planar". Volumes 302, 304, 306, 308, 402, 404, 406, and 408 each illustrate an example of horizontal plane occupancy. Note that with respect to volumes 308 and 408, they meet the requirement for being a horizontal plane because in each case, the upper or lower half of volumes 308 and 408 is empty, i.e., all occupied subvolumes (in these examples, one subvolume) are located in one horizontal half of volumes 308 and 408. It will also be appreciated that in these examples, a volume with a single occupied subvolume will also meet the requirements for vertical planarity about the y-axis and vertical planarity about the x-axis. That is, volumes 308 and 408 are planar in three directions.
[0064] Planarity can be signaled with respect to the volume by a plane mode flag (e.g., isPlanar). In the case where there are multiple possible plane modes (e.g., with respect to the z-axis, y-axis, and x-axis), there can be multiple flags: isZPlanar, isYPlanar, isXPlanar. In this example, for ease of illustration, it is assumed that only the horizontal plane mode is enabled.
[0065] The planeMode flag indicates whether the volume is planar. If it is planar, the second syntax element, the planePosition flag, can be used to signal the position of the plane within the volume. In this example, the planePosition flag signals whether the occupied subvolume of the plane is in the upper half or the lower half of the volume.
[0066] In more complex implementations involving non-orthogonal planes (eg, planes diagonal to one or more of the axes), more complex signaling syntaxes involving multiple flags or non-binary syntax elements may be used.
[0067] The plane mode flag and / or the plane position flag may be encoded in the bitstream using any suitable coding scheme. In some implementations, the flag may be entropy encoded using prediction and / or context adaptive coding to improve compression. Example techniques for determining the context for encoding the flag are discussed further below.
[0068] Occupancy coding and planar mode
[0069] By signaling planarity, the encoding of the occupancy bits can be altered, since the planarity information allows inferences to be made about the occupancy pattern as a shortcut to the signaling of occupancy. For example, if the volume is planar, the four subvolumes not pointed to by the plane position markers can be assumed to be empty, and their occupancy bits do not need to be encoded. Only up to four bits of the occupied plane need to be encoded. Furthermore, if the first three coded bits of the plane are zero (unoccupied), the last (fourth) bit in the coding order can be inferred to be one (occupied), since the plane signaling indicates that the plane is occupied. Additionally, if the plane signaling indicates that the volume is not planar, there must be at least one occupied subvolume in both planes, which allows additional inferred occupancy of the last bit of either plane if the first three occupancy bits of either plane are zero.
[0070] Therefore, signaling planar mode can provide efficiency in encoding of occupied data. However, planar mode adds syntax elements to the bitstream with signaling, and may not provide efficiency in all situations. For example, in dense point clouds and at certain depths, signaling planarity may not be advantageous because, by definition, any node with more than five occupied child nodes cannot be planar. Therefore, it may be further advantageous to have an eligibility criterion for enabling planar mode. It would be further advantageous to provide an eligibility criterion that adapts to local data.
[0071] In one example, eligibility may be biased towards use with sparse clouds. For example, the eligibility criteria may be based on a metric such as the average number of occupied child nodes of a volume. This running average may be determined for a particular depth of the tree and then used for eligibility for planar mode at the next lower depth. In one example, if the average number of occupied subvolumes is greater than 3, planar mode may be disabled. This technique has the advantage of simplicity, but lacks adaptability to local properties of the cloud.
[0072] In another example, a probability factor of operation may be determined. The probability factor may indicate the likelihood that a node is planar, i.e., the probability of planarity. If the probability is low, then it indicates that the cost of signaling planarity will be high for little potential gain. For a given volume / node, a threshold eligibility value may be set and planar mode may be enabled if the probability factor at that time is greater than the threshold:
[0073] p=prob(plane)≥threshold
[0074] As an example, the threshold selected may be 0.6; however, it will be appreciated that other values may be used for other situations. The probabilities are updated during the encoding process. The update process may be tuned for fast or slow updates for a particular application. Faster updates may give more weight or bias to recently encoded nodes.
[0075] An example probability update process can be expressed as:
[0076] p new =(Lp+δ(coded node)) / (L+1)
[0077] In the above expression, p is the current probability, p new is the probability of updating, δ(coded node) is the planar state of the current node, where 0 is for non-planar and 1 is for planar, and L is a weight factor used to tune how quickly the update occurs. The weight factor L can be set to a power of two minus one, such as 255, to allow a simple integer implementation; more specifically, the partitioning can be implemented using a simple shift operation. Note that the planar state does not necessarily mean that the planar mode flag is encoded, so that the probability tracks the planarity of whether the planar mode in the nearest node is enabled. For example, during the decoding process, the planar state of any node is known after decoding of the occupied bits associated with the node, regardless of whether the planar mode flag is decoded.
[0078] For example, the updating of the probability may occur when the node occupancy is encoded or decoded. In another example, the updating of the probability may occur when the node plane information is decoded for the current node. The updated probability is then used to determine the eligibility of the next node in the coding order for plane mode signaling.
[0079] As noted earlier, the planar mode can be signaled for the horizontal plane (z-axis) or the vertical plane (x-axis or y-axis) or for any two of the three. In the case where planar modes for more than one direction can be signaled, the eligibility criteria may be different. That is, for each additional planar mode signaled by the node, the benefit in terms of occupied signaling is half. With the first planar mode, half of the occupied bits can be inferred in the case of a plane. With the second planar mode, only two of the remaining four occupied bits can be inferred, and so on. Therefore, the threshold for signaling additional planar modes can be higher than the first planar mode.
[0080] The definition of "first", "second" and "third" plane modes can be based on probability and their order from most likely to least likely. Thresholds can then be applied to qualify, where the first threshold is lower than the second threshold, and so on. Example thresholds are 0.6, 0.77 and 0.88, although these are merely illustrative. In another embodiment, there is only one threshold and even if more than one of the plane modes meets the threshold, only the most likely plane mode is enabled.
[0081] Reference now Figure 5, which illustrates in flowchart form an example method 500 for encoding point cloud data using planar modes. The method 500 reflects a process for encoding occupancy information for a volume. In this example, the volume is evenly partitioned into eight sub-volumes, each with an occupancy bit, according to an octree-based encoding. For simplicity, the current example assumes that only one (e.g., horizontal) planar mode is used.
[0082] In operation 502, the encoder evaluates whether the volume is eligible for plane coding mode. As discussed above, in one example, eligibility can be based on cloud density, which can be evaluated using the average number of occupied child nodes. To improve local adaptation, eligibility can be tracked based on probabilistic factors. If the plane coding mode is not eligible, then the occupancy pattern for the volume is encoded without using the plane coding mode, as shown by operation 504.
[0083] If planar mode is enabled, then in operation 506 the encoder evaluates whether the volume is planar. If not, then in operation 508 it encodes a planar mode flag, e.g., isPlanar=0. In operation 510, the encoder then encodes an occupancy pattern based on the presence of at least one occupied subvolume per plane. That is, the occupancy pattern is encoded and if the first three bits encoded for any plane (upper or lower) are zero, then the last (fourth) bit for that plane is not encoded and is inferred to be one, since the corresponding subvolume must be occupied.
[0084] If planar mode is enabled and the volume is planar, then in operation 512 a planar mode flag is encoded, e.g., isPlanar=1. Because the volume is planar, the encoder then also encodes a plane position flag, planePosition. The plane position flag signals whether the occupied subvolume of the plane is in the upper or lower half of the volume. For example, planePosition=0 may correspond to the lower half (i.e., the lower z-axis position) and planePosition=1 may correspond to the upper half. The occupancy bits are then encoded based on knowledge of the planarity of the volume and the position of the occupied subvolume. That is, up to four bits are encoded because four can be inferred to be zero, and the fourth bit can be inferred to be one if the first three are encoded to be zero.
[0085] exist Figure 6 An example method 600 for decoding encoded point cloud data is shown in FIG. The example method 600 is implemented by a decoder that receives a bitstream of encoded data. For a current volume, in operation 602, the decoder determines whether the volume is eligible for planar mode. The eligibility evaluation is the same evaluation as performed at the encoder. If not eligible, the decoder occupies the pattern entropy decoding without adapting the planar mode signal, as shown by operation 604.
[0086] If planar mode is enabled, the decoder decodes the planar mode flag in operation 606. The decoded planar mode flag indicates whether the volume is planar, as shown by operation 608. If not planar, the decoder decodes the occupancy bits knowing that at least one subvolume in each plane is occupied. This may allow the decoder to infer one or two of the occupancy bits based on the values of the other bits decoded.
[0087] If the decoded plane mode flag indicates that the volume is planar, the decoder decodes the plane position flag in operation 612. The decoded plane position flag indicates whether the occupied subvolume is the upper half or the lower half of the volume. Based on this knowledge, the decoder then infers the value of the four occupancy bits in the unoccupied half to be zero, and it decodes up to four bits of the occupancy pattern for the occupied half, as shown by operation 614.
[0088] As noted earlier, encoding of the occupancy bits may include entropy encoding based in part on a neighbor configuration, where the neighbor configuration reflects the occupancy of volumes that share at least one face with the current volume. When evaluating the neighbor configuration, if the neighbor configuration (NC) is zero, meaning that no neighbor volumes are occupied, a flag may be encoded to signal whether the current volume has a single occupied subvolume.
[0089] The encoding of occupied bits using NC and individual node signaling can be based on Figure 7 The plane mode signal shown in is adjusted. Figure 7 An example method 700 for occupied bit encoding is shown. The portion of the flow chart showing method 700 reflects the effect of plane signaling on occupied bit encoding. Although not shown, it can be assumed that appropriate eligibility testing occurs and that plane flags and position flags are encoded / decoded if applicable. In this example, it will be assumed that only one plane mode is possible, but extensions to other modes or additional modes will be understood.
[0090] In operation 702, the encoder evaluates whether NC is zero, that is, whether all adjacent volumes are empty. If not, the encoder evaluates whether the volume is planar in operation 704. If not planar, the encoder encodes or decodes the eight occupied bits knowing that at least one bit in each plane is 1, as shown by operation 706. If planar, the encoder infers that the bits of the unoccupied planes are zero, and encodes the other four bits knowing that at least one of the occupied bits is 1.
[0091] If NC is zero, the encoder determines if there is a single occupied subvolume, which is indicated in the bitstream by a single node flag, as shown in operation 710. If single node is true, then data is bypass encoded in operation 712 to signal any remaining xyz position data about the position of the single node within the volume that is not already available from the encoded plane mode and plane position flags. For example, if the plane mode is enabled to encode horizontal plane information, then the x and y flags are bypass encoded to signal the single node position, but the z position is known from the plane position flag.
[0092] If a single node is false, then the occupancy bits are encoded knowing that at least two subvolumes are occupied, as indicated by operation 714. This may include determining whether the volume is planar, and if so, determining its plane positions, and then encoding the occupancy bits accordingly. For example, if planar then the unoccupied planes may be inferred to contain all zero bits, and the bits for the occupied planes may be encoded and up to two of them may be inferred based on the knowledge that at least two subvolumes are occupied.
[0093] Reference now Figure 8 , which shows a flow chart illustrating one example method 800 for encoding occupancy bits using three possible plane modes. The portion of the flow chart showing method 800 reflects the effect of plane signaling on the encoding of occupancy bits. Although not shown, it can be assumed that appropriate eligibility testing occurs and plane flags and position flags are encoded / decoded where applicable. The encoder first evaluates whether all three plane modes indicate that the volume is planar about all three axes, as shown at operation 802. If so, then they collectively indicate the location of a single occupied subvolume and all occupancy bits can be inferred, as shown by operation 804.
[0094] If not all planes, then the encoder evaluates whether the neighboring configuration is zero in operation 816. If NC is not zero, then the occupancy bits are encoded based on the plane mode signaling in operation 808. As discussed above, the occupancy bit encoding can be masked by the plane mode signaling, which allows multiple possible inferences to be made as a shortcut for occupancy encoding.
[0095] If NC is zero, then a single node flag may be encoded. The encoder first evaluates whether at least one of the plane mode flags indicates that the plane mode is false. If so, it would imply that it cannot be a single node situation because more than one subvolume is occupied. Therefore, if this is not the case in operation 810, i.e., plane is not false, then in operation 812 the single node flag is evaluated and encoded. If the single node flag is set, then the x, y, z bits of the single node position may be bypass encoded to the extent that they are not already inferred from the plane position data, as shown by operation 814.
[0096] If operation 810 determines that at least one planar mode flag indicates that the volume is non-planar, or if in operation 812 the signal node flag indicates that it is not a single node, then the encoder evaluates in operation 816 whether there are two planar mode flags indicating that the volume is planar in two directions, and if so, then all occupied bits may be inferred in operation 804. If not, then the occupied bits are encoded in operation 818 with knowledge of the planarity (if any), and at least two bits are non-zero.
[0097] Those skilled in the art will recognize that the feature employed in the current test mode for point cloud coding is the Inferred Direct Coding Mode (IDCM), which is used to handle very isolated points. Because there is little correlation with neighboring nodes to exploit, the position of the isolated point is encoded directly rather than encoding the occupancy information of the cascaded single child nodes. This mode is eligible under the condition of isolation of the node and where eligible, IDCM activation is signaled by the IDCM flag. In the case of activation, the local position of one or more points belonging to the node is encoded and the node then becomes a leaf node, effectively stopping the recursive segmentation and tree-based coding process for that node.
[0098] The process of signaling planarity herein can be incorporated into the encoding process with IDCM mode by signaling planarity (if eligible) before signaling IDCM. First, the eligibility of a node for IDCM can be affected by plane information. For example, if a node is not planar, then the node may become ineligible for IDCM. Second, in the case of IDCM activation, plane knowledge helps encode the position of points in the volume associated with the node. For example, the following rules can be applied:
[0099] If the node is x-planar, the position of the plane, planeXPosition, is known, so the highest bit of the point's x-coordinate is known from the plane position. This bit is not encoded in the IDCM; the decoder will infer it from the plane position.
[0100] If the node is y-planar, the position of the plane, planeYPosition, is known, so the highest bit of the point's y coordinate is known from the plane position. This bit is not encoded in the IDCM; the decoder will infer it from the plane position.
[0101] If the node is z-planar, the position of the plane, planeZPosition, is known, so the highest bit of the point's y coordinate is known from the plane position. This bit is not encoded in the IDCM; the decoder will infer it from the plane position.
[0102] In the case where a node is planar in multiple directions, the inference of the most significant bits of the xyz coordinates still holds. For example, if a node is x-planar and y-planar, then the most significant bits for both the x- and y-coordinates are inferred via planeXPosition and planeYPosition.
[0103] Entropy Coding of Planar Mode Syntax
[0104] Plane mode syntax such as plane mode flag or plane position flag may represent a significant portion of the bitstream. Therefore, in order to make plane mode effective in compressing point cloud data, it may be advantageous to ensure that the plane information is entropy encoded with effective context determination.
[0105] Recall that whether a node / volume is a plane is signaled using the plane mode flag isPlanar. In the discussion of the current example, it will be assumed that the plane mode applies to horizontal planarity, i.e., about the z-axis. In this example, the flag may be named isZPlanar. Entropy encoding of the flag may employ a binary arithmetic coder, such as a context adaptive binary arithmetic coder (CABAC). The context (or internal probability) may be determined using one or more predictors.
[0106] The planar mode flag for the current node or subvolume signals whether the child subvolumes within the subvolume are planar. The current node or subvolume exists within a parent volume. As an example, predictors for determining the context for encoding the planar mode flag may include one or more of the following:
[0107] (a) Father volume planarity;
[0108] (b) occupancy of adjacent neighboring volumes; and
[0109] (c) The distance to the closest, occupied, encoded node at the same depth and at the same z-axis position.
[0110] Fig. 9 Three example factors regarding the current node 900 within the parent node 902 are illustrated.
[0111] Factor (a) refers to whether the parent node 902 is planar. Regardless of whether it is encoded using planar mode, if the parent node 902 meets the criteria for planarity (in this case, horizontal planarity), then the parent node 902 is considered to be planar. Factor (a) is binary: "parent is planar" or "parent is not planar".
[0112] Factor (b) refers to the occupancy status of the neighbor volume 904 of the plane-aligned face adjacent to the parent volume at the parent depth. In the case of horizontal planarity, if the current node 900 is in the upper half of the parent volume 902, then the neighbor volume 904 is vertically above the parent node 902. If the current node 900 is in the lower half of the parent volume 902, then the neighbor volume 904 will be vertically below. In the case of vertical planarity, depending on the x-axis or y-axis planarity and position of the current node, the neighbor volume will be adjacent to one of the sides. Factor (b) is also binary: the neighbor is occupied or not occupied.
[0113] Factor (c) refers to how far the closest encoded node 906 is from the following conditions: the encoded node is located at the same depth as the current node 900 and is located in a common plane, that is, at the same z-axis position as the current node 900. The encoded node 906 is not necessarily in an adjacent volume and can be some distance away, depending on the density of the cloud. The encoder tracks the encoded nodes and identifies the closest node that meets these criteria. The distance d between the current node 900 and the encoded node 906 can be determined based on the relative positions of the nodes 900 and 906. In some embodiments, for ease of calculation, the L1 norm can be used to determine the distance, that is, the absolute value of delta-x plus the absolute value of delta-y. In some embodiments, the L2 norm can be used to determine the distance, that is, the square root of the sum of squares given by the square of delta-x and the square of delta-y.
[0114] In some implementations, the distance d may be discretized into two values, "near" and "far". The division between "near" d and "far" d may be chosen appropriately. Factor (c) is also binary by classifying the distance as near or far. It will be appreciated that in some implementations, the distance may be discretized into three or more values.
[0115] If all three example factors are used in context determination, then 2 x 2 x 2 = 8 separate contexts may be maintained for encoding of the planar mode flag.
[0116] If a plane mode flag is encoded for the current node 900 and the current node 900 is planar, then a plane position flag may be encoded, such as planeZPosition. The plane position flag signals which half of the current node 900 contains the occupied child subvolume. In the case of horizontal planarity, the plane position flag signals the lower half or the upper half.
[0117] The entropy coding of the plane position flags may also employ a binary arithmetic coder, such as CABAC. The context (or internal probability) may be determined using one or more predictors, possible examples of which include:
[0118] (a') occupancy of adjacent neighboring volume 904;
[0119] (b') distance to the closest occupied encoded node 906 at the same depth and at the same z-axis position;
[0120] (c') if the nearest occupied already encoded node 906 at the same depth and z-axis position is planar, its planar position; and
[0121] (d′) The position of the current node 900 within the parent node 902 .
[0122] Factor (a') is the same as factor (b) discussed above in the context of the planar mode flag. Factor (b') is the same as factor (c) discussed above in the context of the planar mode flag. In some example implementations, factor (b') may discretize the distance into three categories: "near", "not too far", and "far". As discussed above, the distance may be determined using the L1 norm or the L2 norm or any other appropriate metric.
[0123] Factor (c') refers to whether the closest, occupied, already encoded node 906 is planar, and if so, whether it is top planar or bottom planar, i.e., its planar position. It turns out that even a distant, already encoded node that is planar can be a strong predictor of the planarity or planar position of the current node. That is, factor (c') can have three results: non-planar, same planar position as the current node 900, different planar position than the current node 900. If the current node 900 and the closest, already encoded, occupied node have the same planar position, then their occupied child subvolumes are all aligned in a common horizontal plane at the same z-axis position.
[0124] The factor (d') refers to whether the current node 900 is located in the upper or lower half (in the case of horizontal planarity) of the parent node 902. Because if the current node 900 is planar, the parent is likely to be planar due to the eligibility requirement, and a planar position is slightly more likely to be "outside" the parent node 902 and not toward the middle. Therefore, the position of the current node 900 in its parent node 902 has a significant impact on the probability of a planar position within the current node 900.
[0125] In an implementation that combines all four factors, there may be 2 x 3 x 2 x 2 = 24 prediction combinations when the closest, occupied, coded node 906 at the same z and depth (as the current node 900) is planar; otherwise, when the closest coded node 906 at the same z and same depth is not planar, a specific context is used instead. Thus, 24 + 1 = 25 contexts may be used by the binary arithmetic encoder to encode the plane position flag in such an example.
[0126] Although the above examples involve three factors for context determination in the case of a planar mode flag, and four factors for context determination in the case of a planar position flag, it will be recognized that the present application includes the use of individual factors for context determination and all combinations and sub-combinations of such factors.
[0127] Reference now Fig.10 , which diagrammatically illustrates an example implementation of a mechanism for managing the determination of the closest occupied encoded nodes during a context determination process. In this example mechanism, an encoding device uses a memory (e.g., a volatile or persistent memory unit) to implement a buffer 1000 containing information about occupied encoded nodes. Specifically, buffer 1000 allocates space to track encoded nodes with the same z-axis position and depth in the tree. In this specific example, buffer 1000 tracks information associated with up to four encoded nodes with the same z-axis position and depth.
[0128] Each row of the example buffer 1000 corresponds to a z-axis position and a depth. The four columns correspond to the four most recently encoded occupied nodes having that z-axis position. For example, the example row 1002 contains data about four encoded, occupied nodes. The stored data for each encoded node may include: the x and y positions of the encoded, occupied node, whether the node is planar, and if so, the planar position.
[0129] In encoding the current node 1004, based on the example row 1002 being for the same z-axis position as the current node 1004, the encoding device accesses the buffer 1000 to identify the closest occupied, already encoded node from among the four stored nodes in the example row 1002. As discussed above, the distance metric can be based on the L1 norm, the L2 norm, or any other distance metric. The stored x and y positions for each node in the buffer 1000 help to easily make a determination of the closest node (especially in the case of the L1 norm).
[0130] Once the closest node is identified, such as the closest node 1006, its distance from the current node 100 and possibly its planarity and / or planar position are used in the context determination(s). The buffer 1000 is then updated by adding the current node 1004 to the first position 1008 of the buffer 1000, and shifting all other node data in this example row 1002 of the buffer 1000 to the right, resulting in the last item in the buffer 1000 being discarded. In some examples, based on the distance determination, it is possible that the identified closest node retains a higher potential relevance to the current encoding, so before adding the current node 1004 to the buffer 1000, the contents of the example row 1002 are first rearranged so that the closest node 1006 is moved to the first position 1008, and the nodes are shifted to the right to accommodate, such as in this example the node data in the first position 1008 and the second position are shifted to the second position and the third position, respectively. In this way, the encoding device avoids prematurely evacuating the most recently identified, closest node from the buffer 1000 .
[0131] It will be appreciated that the described buffer 1000 is one example implementation of a mechanism for managing data about the closest nodes, but the present application is not necessarily limited to this example, and many other mechanisms for tracking the closest node information may be used. Furthermore, it will be appreciated that retaining only a fixed number of recently encoded occupied nodes in the buffer 1000 means that there is a chance that the identified node is not actually the closest occupied encoded node, but is merely the closest encoded node available from the buffer 1000; however, even when the buffer is limited to four candidates as in the above example, the impact on performance is negligible.
[0132] Reference now Fig.11, which shows a simplified block diagram of an example embodiment of an encoder 1100. The encoder 1100 includes a processor 1102, a memory 1104, and an encoding application 1106. The encoding application 1106 may include a computer program or application stored in the memory 1104 and containing instructions, which when executed causes the processor 1102 to perform operations such as those described herein. For example, the encoding application 1106 can encode and output a bit stream encoded according to the process described herein. It will be understood that the encoding application 1106 can be stored on a non-transitory computer-readable medium, such as a compressed disk, a flash memory device, a random access memory, a hard drive, etc. When the instructions are executed, the processor 1102 performs the operations and functions specified in the instructions so as to be used as a dedicated processor to implement the process (s) described. In some examples, such a processor may be referred to as a "processor circuit" or "processor circuit".
[0133] Now also refer to Fig.12 , which shows a simplified block diagram of an example embodiment of a decoder 1200. The decoder 1200 includes a processor 1202, a memory 1204, and a decoding application 1206. The decoding application 1206 may include a computer program or application stored in the memory 1204 and containing instructions that, when executed, cause the processor 1202 to perform operations such as those described herein. It will be understood that the decoding application 1206 may be stored on a computer-readable medium, such as a compressed disk, a flash memory device, a random access memory, a hard drive, etc. When the instructions are executed, the processor 1202 performs the operations and functions specified in the instructions so as to function as a dedicated processor that implements the described process(es). In some examples, such a processor may be referred to as a "processor circuit" or "processor circuit."
[0134] It will be appreciated that decoders and / or encoders according to the present application can be implemented with a plurality of computing devices, including but not limited to servers, appropriately programmed general-purpose computers, machine vision systems, and mobile devices. A decoder or encoder can be implemented by software containing instructions for configuring a processor or processors to perform the functions described herein. Software instructions can be stored in any suitable non-transient computer-readable memory including CD, RAM, ROM, flash memory, etc.
[0135] It will be appreciated that the decoder and / or encoder described herein and the modules, routines, processes, threads or other software components implementing the described methods / processes for configuring the encoder or decoder can be implemented using standard computer programming techniques and languages. The application is not limited to specific processors, computer languages, computer programming conventions, data structures, other such implementation details. Those skilled in the art will appreciate that the described process can be implemented as a part of a computer executable code stored in a volatile or non-volatile memory, a part of an application specific integrated circuit (ASIC), etc.
[0136] The present application also provides a computer readable signal encoding data produced by application of an encoding process according to the present application.
[0137] Impact on compression performance
[0138] Tests using a planar pattern in three terms x, y and z using an example implementation have been performed using multiple example point clouds with different characteristics and compared to a Moving Picture Experts Group (MPEG) test model. The different types of point clouds used in the experiments include those related to outdoor scenes including urban built environments, indoor built environments, 3D maps, LiDAR road scans, and natural landscapes. Negligible conservative compression gains are seen in the case of natural landscapes. Compression gains of 2% to 4% are seen in the case of 3D maps, up to 10% are seen in the case of outdoor built scenes, 6 to 9% are seen in the case of LiDAR data, and up to over 50% are seen in some indoor built scenes.
[0139] Certain adaptations and modifications to the described embodiments may be made.Therefore, the embodiments discussed above are considered to be illustrative rather than restrictive.
Claims
1. A method of encoding a point cloud to generate a bitstream of compressed point cloud data, the compressed point cloud data representing a three-dimensional position of an object, the point cloud being located within a volumetric space, the volumetric space being recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first set of child sub-volumes and a second set of child sub-volumes, the first set of child sub-volumes being positioned in a first plane and the second set of child sub-volumes being positioned in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points, the method include: determining whether the volume is planar based on whether all child subvolumes containing at least one point are located in the first group or the second group; encoding a planar mode flag in the bitstream to signal whether the volume is planar; encoding, in the bitstream, occupancy bits for the child subvolume of the first group comprises: for at least one occupancy bit, speculating a value of the at least one occupancy bit based on whether the volume is planar and not encoding the at least one occupancy bit in the bitstream; as well as Outputs the bit stream of compressed point cloud data.
2. The method of claim 1, wherein determining whether the volume is planar include: The volume is determined to be planar by determining that at least one of the child subvolumes in the first group contains at least one of the points and none of the child subvolumes in the second group contains any of the points, and wherein the method further comprises: encoding a plane position flag based on the volume being planar to signal that the at least one of the child subvolumes is in the first plane.
3. The method according to claim 2, wherein the encoding occupies include: Encoding of occupied bits associated with the second group is prohibited, and values for the occupied bits associated with the second group are inferred based on the second group not including a point.
4. The method according to claim 3, wherein the encoding occupied bits are further include: A last occupied bit in the coding order of the occupied bits associated with the first group is inferred to have a value indicating occupied based on determining that all other occupied bits of the first group in coding order have values indicating unoccupied.
5. The method according to any one of claims 1 to 4, wherein determining whether the volume is planar include: It is determined that the volume is not planar, and based on this, occupancy bits are encoded based on at least one of the occupancy bits in the first group and at least one of the occupancy bits in the second group having a value indicating occupied.
6. The method of any one of claims 1-4, wherein the point cloud is defined with respect to Cartesian axes in the volumetric space, the Cartesian axes having a z-axis oriented vertically perpendicular to a horizontal plane, and wherein the first plane and the second plane are parallel to the horizontal plane.
7. A method according to any one of claims 1-4, wherein the point cloud is defined with respect to Cartesian axes in the volumetric space, the Cartesian axes having a vertically oriented z-axis perpendicular to a horizontal plane, and wherein the first plane and the second plane are orthogonal to the horizontal plane.
8. The method according to any one of claims 1 to 4, further comprising: include: First it is determined that the volume is eligible for planar mode encoding.
9. The method of claim 8, wherein determining the volume is eligible for planar mode encoding include: A probability of planarity is determined and the probability of planarity is determined to be greater than a threshold eligibility value.
10. The method according to any one of claims 1 to 4, wherein the planar mode flag is encoded include: Encoded horizontal plane mode flag and encoded vertical plane mode flag.
11. A method of decoding a bitstream of compressed point cloud data to produce a reconstructed point cloud, the reconstructed point cloud representing a three-dimensional position of a physical object, the point cloud being located within a volumetric space, the volumetric space being recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first set of child sub-volumes and a second set of child sub-volumes, the first set of child sub-volumes being positioned in a first plane, and the second set of child sub-volumes being positioned in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points, the method include: The occupied bits are reconstructed by the following process to reconstruct the points of the point cloud: decoding a planar mode flag from the bitstream, the planar mode flag indicating whether the volume is planar, wherein the volume is planar if all child subvolumes containing at least one point are located in the first group or the second group; as well as Decoding occupancy bits for the child subvolumes of the first group from the bitstream includes, for at least one occupancy bit, inferring a value of the at least one occupancy bit based on whether the volume is planar and not decoding the at least one occupancy bit from the bitstream.
12. The method of claim 11, wherein the planar mode flag indicates that the volume is planar, and based on the volume being planar, the method further include: A plane position flag is decoded, the plane position flag indicating that at least one of the child sub-volumes in the first group contains at least one of the points and none of the child sub-volumes in the second group contains any of the points.
13. The method according to claim 12, wherein decoding the occupied bits include: Decoding of occupied bits associated with the second group is inhibited, and values for the occupied bits associated with the second group are inferred based on the second group not including a point.
14. The method according to claim 13, wherein the decoding occupied bits are further include: A last occupied bit in coding order of the occupied bits associated with the first group is inferred to have a value indicating occupied based on determining that all other occupied bits of the first group in coding order have values indicating unoccupied.
15. A method according to any one of claims 11 to 14, wherein the decoded planar mode flag indicates that the volume is not planar, and on this basis, the occupancy bits are decoded based on at least one of the occupancy bits in the first group and at least one of the occupancy bits in the second group having a value indicating occupied.
16. The method of any one of claims 11 to 14, wherein the point cloud is defined with respect to Cartesian axes in the volumetric space, the Cartesian axes having a vertically oriented z-axis perpendicular to a horizontal plane, and wherein the first plane and the second plane are parallel to the horizontal plane.
17. A method according to any one of claims 11 to 14, wherein the point cloud is defined with respect to Cartesian axes in the volumetric space, the Cartesian axes having a vertically oriented z-axis perpendicular to a horizontal plane, and wherein the first plane and the second plane are orthogonal to the horizontal plane.
18. The method of any one of claims 11 to 14, further comprising first determining that the volume is eligible for planar mode encoding.
19. The method of claim 18, wherein determining the volume is eligible for planar mode encoding include: A probability of planarity is determined and the probability of planarity is determined to be greater than a threshold eligibility value.
20. The method according to any one of claims 11 to 14, wherein decoding the planar mode flag include: Decode horizontal plane mode flag and decode vertical plane mode flag.
21. An encoder for encoding a point cloud to generate a bitstream of compressed point cloud data, the compressed point cloud data representing a three-dimensional position of an object, the point cloud being located within a volumetric space, the volumetric space being recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first set of child sub-volumes and a second set of child sub-volumes, the first set of child sub-volumes being positioned in a first plane and the second set of child sub-volumes being positioned in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points, the encoder include: processor; Memory; as well as A coded application comprising instructions executable by the processor which, when executed, cause the processor to: determining whether the volume is planar based on whether all child subvolumes containing at least one point are located in the first group or the second group; encoding a planar mode flag in the bitstream to signal whether the volume is planar; encoding, in the bitstream, occupancy bits for the child subvolume of the first group comprises: for at least one occupancy bit, speculating a value of the at least one occupancy bit based on whether the volume is planar and not encoding the at least one occupancy bit in the bitstream; as well as Outputs the bit stream of compressed point cloud data.
22. A decoder for decoding a bitstream of compressed point cloud data to produce a reconstructed point cloud, the reconstructed point cloud representing a three-dimensional position of a physical object, the point cloud being located within a volumetric space, the volumetric space being recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first set of child sub-volumes and a second set of child sub-volumes, the first set of child sub-volumes being positioned in a first plane and the second set of child sub-volumes being positioned in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points, the decoder include: processor; Memory; as well as A decoding application comprising instructions executable by the processor, which, when executed, causes the processor to reconstruct the occupied bits by the following process to reconstruct the points of the point cloud: decoding a planar mode flag from the bitstream, the planar mode flag indicating whether the volume is planar, wherein the volume is planar if all child subvolumes containing at least one point are located in the first group or the second group; as well as Decoding occupancy bits for the child subvolume of the first group from the bitstream includes, for at least one occupancy bit, inferring a value of the at least one occupancy bit based on whether the volume is planar and not decoding the at least one occupancy bit from the bitstream.
23. A non-transitory processor-readable medium storing processor-executable instructions, the processor-executable instructions encoding a point cloud to generate a bitstream of compressed point cloud data, the compressed point cloud data representing three-dimensional positions of an object, the point cloud being located within a volumetric space, the volumetric space being recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first set of child sub-volumes and a second set of child sub-volumes, the first set of child sub-volumes being positioned in a first plane and the second set of child sub-volumes being positioned in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points, wherein the processor-executable instructions, when executed by a processor, cause the processor to: determining whether the volume is planar based on whether all child subvolumes containing at least one point are located in the first group or the second group; encoding a planar mode flag in the bitstream to signal whether the volume is planar; encoding, in the bitstream, occupancy bits for the child subvolume of the first group comprises: for at least one occupancy bit, speculating a value of the at least one occupancy bit based on whether the volume is planar and not encoding the at least one occupancy bit in the bitstream; as well as Outputs the bit stream of compressed point cloud data.
24. A non-transitory processor-readable medium storing processor-executable instructions, the processor-executable instructions decoding a bitstream of compressed point cloud data to produce a reconstructed point cloud, the reconstructed point cloud representing a three-dimensional position of a physical object, the point cloud being located within a volumetric space, the volumetric space being recursively split into sub-volumes and containing points of the point cloud, wherein the volume is partitioned into a first set of child sub-volumes and a second set of child sub-volumes, the first set of child sub-volumes being positioned in a first plane and the second set of child sub-volumes being positioned in a second plane parallel to the first plane, and wherein an occupancy bit associated with each respective child sub-volume indicates whether the respective child sub-volume contains at least one of the points, wherein the processor-executable instructions, when executed by a processor, cause the processor to: The occupied bits are reconstructed by the following process to reconstruct the points of the point cloud: decoding a planar mode flag from the bitstream, the planar mode flag indicating whether the volume is planar, wherein the volume is planar if all child subvolumes containing at least one point are located in the first group or the second group; and Decoding occupancy bits for the child subvolume of the first group from the bitstream include: For at least one occupied bit, a value of the at least one occupied bit is inferred based on whether the volume is planar and the at least one occupied bit is not decoded from the bitstream.
Citation Information
Patent Citations
Sign coding for blocks with transform skipped
CN105723714A
Coding of a spatial sampling of a two-dimensional information signal using sub-division
US20170134760A1