Method and device for multi-point direct encoding and decoding in point cloud compression
By sorting and pairing the coordinates of multiple points in the sub-volume using the direct codec mode in point cloud compression, the problem of low processing efficiency of isolated points is solved, and faster codec speed and better compression performance are achieved.
Patent Information
- Application Number
- CN202080097025.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-02-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-02-19
AI Technical Summary
The prior art is difficult to efficiently handle isolated points in point cloud compression, resulting in increased computing burden and memory requirements, and impact on compression performance.
Direct encoding and decoding mode (DCM) is used to encode and decode point cloud data, and improve compression performance by sorting and pairwise encoding and decoding the coordinates of two or more disordered points in the sub-volume.
The encoding and decoding time is reduced, the number of nodes processed is reduced, and the compression performance is improved, especially in multi-point situations, which significantly speeds up the encoding and decoding speed.
Smart Images

Figure CN115152148B_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to point cloud compression and, in particular, to methods and apparatuses for improved compression of directly encoding and decoding unordered points in point cloud compression. Background Art
[0002] Data compression is used in communication and computer networking to efficiently store, transmit, and reproduce information. There is increasing interest in the representation of three-dimensional objects or spaces, which may involve large data sets, and efficient and effective compression would be very useful and valuable. In some cases, a three-dimensional object or space can be represented using a point cloud, which is a set of points, each having three coordinate positions (X, Y, Z) and, in some cases, other attributes such as color data (e.g., luminance and chrominance), transparency, reflectivity, normal vectors, etc. Point clouds can be static (stationary objects, or snapshots of an environment / object at a single point in time) or dynamic (a chronological sequence of point clouds).
[0003] Example applications for point clouds include terrain and mapping applications. Autonomous vehicles and other machine visualization applications can rely on point cloud sensor data in the form of 3D scans of the environment, such as data from LiDAR scanners. Virtual reality simulations can rely on point clouds.
[0004] It should be understood that point clouds can involve large amounts of data, and it is of great significance to compress (encode and decode) this data quickly and accurately. Therefore, it would be advantageous to provide methods and apparatuses that can compress point cloud data more quickly, efficiently, and / or effectively. Brief Description of the Drawings
[0005] Example embodiments of the present application will now be referred to by way of example, and in the drawings:
[0006] Figure 1 An example method for encoding a point cloud using an inferred direct encoding and decoding mode is shown in flowchart form;
[0007] Figure 2 An example method for decoding encoded point cloud data using an inferred direct encoding and decoding mode is shown in flowchart form;
[0008] Figure 3 An example sub-volume containing two points is shown;
[0009] Figure 4 A simplified flowchart illustrating a paired encoding process for coordinate value pairs is shown;
[0010] Figure 5 A direct encoding and decoding mode encoding process is shown in flowchart form;
[0011] Figure 6 The direct codec mode decoding process is shown in the form of a flowchart;
[0012] Figure 7 A simplified flowchart showing an example process for encoding and decoding a triple of bits during a three-point DCM codec process is shown;
[0013] Figure 8 An example simplified block diagram of an encoder is shown; and
[0014] Figure 9 An example simplified block diagram of a decoder is shown.
[0015] Similar reference numerals may have been used in different drawings to denote similar components. Detailed Description
[0016] This application describes methods for encoding and decoding point clouds, as well as encoders and decoders for encoding and decoding point clouds. Generally, this application describes methods and apparatuses for encoding and decoding a point cloud by encoding and decoding the coordinates of two or more points within a sub-volume associated with a current node using a direct codec mode. Methods for encoding and decoding for compression of oriented codec points in a sub-volume are described. In some cases, the coordinates of two or more unordered points in a sub-volume are encoded by sorting the points based on their respective coordinate values and then using pairwise encoding of the bits in the corresponding positions in the respective coordinates.
[0017] In one aspect, this application describes a method for encoding a point cloud to generate a compressed point cloud data bitstream, where a current sub-volume contains a first point and a second point, the first point having a first position within the sub-volume defined by a first coordinate value and the second point having a second position within the sub-volume defined by a second coordinate value. The method may include sorting the first coordinate value and the second coordinate value, where the first coordinate value and the second coordinate value are binary; starting from the most significant bit position, pairwise encoding the current position of the first coordinate value and the current position of the second coordinate value by: encoding a same bit flag that indicates whether the bits in the current position in the first coordinate value and the second coordinate value are the same; when the bits in the current position are different, not encoding the bits in the current position and encoding any remaining bits of the first coordinate value and any remaining bits of the second coordinate value; when the bits in the current position are the same, then encoding a bit value flag that indicates whether both bits are 1 or both bits are 0; and recursively repeating the pairwise encoding for the next position in the first coordinate value and the next position in the second coordinate value until the first coordinate value and the second coordinate value are encoded.
[0018] In some implementations, the sorting is from low to high. In some implementations, the sorting is from high to low.
[0019] In some implementations, encoding the same bit flag can include entropy encoding the same bit flag. In some cases, entropy encoding the same bit flag uses a dedicated context for entropy encoding the same bit flag.
[0020] In some implementations, the first coordinate value and the second coordinate value correspond to directions in a Cartesian coordinate system. In some cases, the directions are the x direction, the y direction, or the z direction.
[0021] In some implementations, the first position is further defined by a third coordinate value and the second position is further defined by a fourth coordinate value, the third coordinate value and the fourth coordinate value corresponding to the same direction in the coordinate system. The method may further include: determining that the first coordinate value and the second coordinate value are the same, and as a result sorting the third coordinate value and the fourth coordinate value, and recursively performing pairwise encoding with respect to the third coordinate value and the fourth coordinate value until the third coordinate value and the fourth coordinate value are encoded.
[0022] In some implementations, the method may further include: first determining that the sub - volume includes at least two points and determining that direct encoding and decoding is to be applied to the at least two points. In some cases, determining that direct encoding and decoding is to be applied may include: determining that the number of points in the sub - volume is less than a threshold number. In some cases, the method may further include encoding a direct encoding and decoding mode flag that signals that direct encoding and decoding is used to encode the at least two points.
[0023] In another aspect, the present application describes a method of decoding a compressed point cloud data bitstream to produce a reconstructed point cloud, where a current sub - volume includes a first point and a second point, the first point having a first position defined by a first coordinate value within the sub - volume and the second point having a second position defined by a second coordinate value within the sub - volume. The method may include starting from the most significant bit position and pairwise - decoding the current positions of the first coordinate value and the second coordinate value by: decoding a same bit flag that indicates whether the bits in the current positions of the first coordinate value and the second coordinate value are the same; when the bits in the current positions are different, reconstructing the bits in the current positions and decoding any remaining bits of the first coordinate value and any remaining bits of the second coordinate value; when the bits in the current positions are the same, decoding a bit value flag that indicates whether both bits in the current position are 1 or both are 0; recursively repeating the pairwise - decoding for the next positions in the first coordinate value and the next positions in the second coordinate value until the first coordinate value and the second coordinate value are decoded; and outputting the reconstructed first coordinate value for the first point and the reconstructed second coordinate value for the second point.
[0024] In another aspect, the present application describes an encoder and a decoder configured to implement such encoding and decoding methods.
[0025] In yet another aspect, the present application describes a non-transitory computer-readable medium storing computer-executable program instructions that, when executed, cause one or more processors to perform the described encoding and / or decoding methods.
[0026] In yet another aspect, the present application describes a computer-readable signal comprising program instructions that, when executed by a computer, cause the computer to perform the described encoding and / or decoding methods.
[0027] Those of ordinary skill in the art will understand other aspects and features of the present application from a review of the following description of the examples in conjunction with the accompanying drawings.
[0028] In the following description, the terms "node" and "sub-volume" may sometimes be used interchangeably. It should be understood that a node is associated with a sub-volume. A node is a specific point on a tree, which can be an internal node or a leaf node. A sub-volume is the bounded physical space represented by a node. The term "volume" may be used to refer to the maximum bounded space defined to enclose a point cloud. The volume is recursively divided into sub-volumes for the purpose of constructing a tree structure of interconnected nodes for encoding and decoding point cloud data.
[0029] In the present application, the term "and / or" is intended to cover all possible combinations and sub-combinations of the listed elements, including any single element of the listed elements, any sub-combination, or all elements, without necessarily excluding additional elements.
[0030] In the present application, the phrase "at least one of... or..." is intended to cover any one or more of the listed elements, including any single element of the listed elements, any sub-combination, or all elements, without necessarily excluding any additional elements and without necessarily requiring all elements.
[0031] A point cloud is a set of points in a three-dimensional coordinate system. These points are typically used to represent the outer surface of one or more objects. Each point has a position (location) in the three-dimensional coordinate system. The position can be represented by three coordinates (X, Y, Z), which can be a Cartesian coordinate system or any other coordinate system. These points can have other associated attributes such as color, which can also be a three-component value in some cases, such as R, G, B or Y, Cb, Cr. Other associated attributes can include transparency, reflectivity, normal vectors, etc., depending on the desired application for the point cloud data.
[0032] Point clouds can be static or dynamic. For example, a detailed scan or mapping of an object or terrain can be static point cloud data. A LiDAR-based scan of an environment for machine visualization purposes can be dynamic since the point cloud (at least potentially) changes over time, e.g., each successive scan of a volume. Thus, a dynamic point cloud is a chronological sequence of point clouds.
[0033] Point cloud data can be used in many applications, including, for example, conservation (scanning of historical or cultural objects), mapping, machine visualization (such as autonomous or semi-autonomous vehicles), and virtual reality systems. Dynamic point cloud data for applications such as machine visualization can be quite different from static point cloud data for conservation purposes. For example, automotive visualization typically involves relatively small-resolution, non-color, high-dynamic point clouds obtained via LiDAR (or similar) sensors with a high capture frequency. The purpose of such point clouds is not for human consumption or viewing but for machine object detection / classification in a decision-making process. For example, a typical LiDAR frame contains tens of thousands of points, while high-quality virtual reality applications require millions of points. As computing speeds increase and new applications are discovered, there may be a need for higher-resolution data over time.
[0034] Although point cloud data is useful, the lack of effective and efficient compression (i.e., the encoding and decoding processes) can impede adoption and deployment.
[0035] One of the more common mechanisms for encoding and decoding point cloud data is by using a tree-based structure. In a tree-based structure, the bounding three-dimensional volume of the point cloud is recursively divided into sub-volumes. The nodes of the tree correspond to the sub-volumes. The decision of whether to further divide a sub-volume can be based on the resolution of the tree and / or the presence of any points contained within the sub-volume. Leaf nodes can have occupancy flags that indicate whether their associated sub-volume contains points. Split flags can signal whether a node has children (i.e., whether the current volume has been further divided into sub-volumes). In some cases, these flags can be entropy encoded, while in some cases, predictive encoding can be used.
[0036] A commonly used tree structure is the octree. In this structure, the volume / sub-volume is a cube, and each split of the sub-volume generates eight additional sub-volumes / sub-cubes. Another commonly used tree structure is the KD-tree, where the volume (cube or rectangular cuboid) is recursively divided into two parts by a plane orthogonal to one of the axes. The octree is a special case of the KD-tree, where the volume is divided by three planes, each plane orthogonal to one of the three axes. Both of these examples are related to cubes or cuboids; however, the present application is not limited to such tree structures, and in some applications, the volume and sub-volumes can have other shapes. The division of the volume does not necessarily have to be into two sub-volumes (KD-tree) or eight sub-volumes (octree), but can involve other divisions, including division into non-rectangular shapes or involving non-adjacent sub-volumes.
[0037] For the sake of explanation, the present application may refer to the octree, and since the octree is a popular candidate tree structure for automotive applications, it should be understood that the methods and devices described herein can be implemented using other tree structures or encoding / decoding structures other than trees.
[0038] The recursive encoding / decoding of a tree structure typically involves splitting an occupied sub-volume into additional sub-volumes and encoding the occupancy status of each of these additional sub-volumes. There are various techniques that can be used to encode / decide the occupancy bits and determine the context for encoding / deciding these occupancy bits, or collectively referred to as the occupancy "pattern", i.e., an eight-bit sequence corresponding to the eight sub-volumes of an octree-based split.
[0039] One problem in the problem of compressed point cloud data in a tree structure is that it may not handle isolated points well. The recursive splitting of sub-volumes and the position of points within the split sub-volumes involve computational burden and time, and in terms of bandwidth / memory storage as well as computational time and resources, signaling the recursive splitting of sub-volumes for precisely locating the position of one or a few isolated points can be expensive. In addition, the isolated points "contaminate" the distribution of the pattern, resulting in many patterns having only one occupied child node, thus changing the balance of the distribution and penalizing the encoding / decoding of other patterns.
[0040] Thus, in some cases, the encoder and decoder can directly encode the position information of the isolated points. The direct encoding and decoding of the position of a point (e.g., its spatial coordinates within a volume or sub-volume) is referred to as the Direct Coding Mode (DCM). Using DCM for all points would be very inefficient. One option is to use a dedicated flag for each occupied node to signal whether DCM will be used for any point within that node. Another option is to evaluate a set of criteria for whether a node is "eligible" to use DCM and only encode and decode the DCM flag when the node is eligible. This technique is described by the present applicant in PCT patent publication WO / 2019 / 140508 entitled "Methods and Devices Using Direct Coding in Point Cloud Compression", the content of which is incorporated herein by reference. In the art, these techniques may be referred to as Inferred Direct Coding Mode (IDCM).
[0041] The use of IDCM has proven to be valuable as it reduces the complexity of point cloud data encoding for sparse datasets. With IDCM, the total number of nodes to be encoded and decoded is less because branches are truncated early to directly encode and decode point positions. IDCM decoded points can be quickly used for further processing. The bypass encoding and decoding of point coordinate values is less complex than the entropy encoding of occupancy flags, thus speeding up the computational speed at both the encoder and decoder. This is particularly useful in applications such as automotive visualization, e.g., LiDAR-based scans.
[0042] To ensure that IDCM does not have a significant negative impact on the compression of point cloud data, the eligibility criteria need to be carefully selected. An explicit condition for using direct coding is that the current sub-volume contains fewer than a threshold number of points. The threshold can be set to two, three, or some other value. IDCM also employs an implicit condition that looks at the occupancy status and other factors related to adjacent sub-volumes at the same encoding depth, parent depth, or grandparent depth. The need to access parent or grandparent level occupancy information for eligibility assessment can impose a significant computational and memory burden on both the encoder and decoder. Additionally, due to the eligibility criteria, IDCM does not prune the tree as much as DCM without eligibility criteria, resulting in more nodes to be tracked than DCM. It has been observed that the speed of encoding and decoding point clouds is directly related to the number of nodes processed and may be involved in eligibility determination as they involve too much memory occupancy to rely on caching. Memory access operations can be a significant bottleneck in the speed of the encoding and decoding process. Within reasonable limits, the computational burden within the processing nodes affects the total processing time in a second order.
[0043] Thus, a way to reduce the encoding and decoding time is to reduce the number of nodes processed, for example, by improving tree pruning.
[0044] One option is to eliminate the implicit eligibility conditions for evaluating occupancy information for nearby volumes and rely only on explicit threshold criteria as the condition for enabling DCM. That is, the use of DCM can be conditioned on the number of points within a sub-volume being below a threshold, without considering additional implicit criteria. This simplification of the eligibility criteria may lead to a significant increase in the use of DCM and reduce the encoding and decoding complexity by reducing the total number of nodes to be encoded and decoded; however, the increased use of DCM will have a negative impact on the compression performance.
[0045] To counteract the negative impact of the increased use of DCM on compression, this application describes a method for encoding and decoding DCM coordinate values that can improve the compression performance when the sub-volume contains at least two points. In some examples, the improved compression performance not only counteracts the negative impact of the increased use of DCM, making the resulting encoding and decoding performance at least as good (measured in bits per point) and significantly faster in terms of compression. In some cases, the encoding and decoding speed can be increased by an order of magnitude. By taking advantage of the fact that they are unordered points, the encoding and decoding of DCM coordinate values in the multi-point case can be improved. That is, the order in which two (or more) points are directly encoded and decoded is not important to the encoder and decoder. Thus, as will be further described below, in a sub-volume containing two or more points, the corresponding coordinate values for the two or more points can be sorted from low to high, and pair-wise encoding is applied to the bit pairs from these two values to improve the encoding and decoding compression.
[0046] Before describing the encoding and decoding of the coordinate values in DCM, the IDCM process from PCT Publication WO 2019 / 140508 is described.
[0047] Now refer to Figure 1 , Figure 1An example method 100 for encoding point clouds using IDCM is shown in the form of a flowchart. The method 100 in this example involves recursive splitting of occupied nodes (sub-volumes) for encoding and decoding and breadth-first traversal of the tree. In operation 104, with respect to the current occupied node, e.g., the current sub-volume associated with a node of the tree occupied by at least one point, the encoder evaluates whether the sub-volume is eligible for DCM. If not, then in operation 106, the node is split and encoded according to the normal tree encoding and decoding process. That is, at least in this example, the sub-volume is split into sub-sub-volumes, as shown in operation 116, and in operation 118, the occupancy patterns of these sub-sub-volumes are entropy encoded and decoded. In operation 120, any of these sub-sub-volumes occupied by at least one point are buffered for further splitting / encoding (put into a FIFO buffer). Although not explicitly stated, it should be understood that method 100 incorporates a stopping condition such as a maximum tree depth after which method 100 does not further split the sub-volume.
[0048] If the node is evaluated in operation 104 and it is determined that the node is eligible for DCM, then in operation 108, the number of points contained in the sub-volume is evaluated according to a threshold. If the number of points in the sub-volume is less than the threshold, DCM is used. If the number of points is equal to or greater than the threshold, DCM is not used. The threshold is pre-set and can be hard-coded or user-determined. It can be transmitted from the encoder to the decoder in the header information. The threshold can be equal to or greater than 2. It should be understood that if the number is less than or equal to the threshold, the evaluation can be modified to enable DCM, and in this case the threshold can be set one point lower to achieve the same result. In any case, if DCM is not used, then in operation 110, the DCM flag is set to negative (in some implementations, it can be signaled as value 0) and output in the bitstream to notify the decoder that DCM is not used in this sub-volume. Method 100 then loops back to operation 106 to split and encode the sub-volume in the normal manner.
[0049] If DCM is to be used, then in operation 112, the DCM flag is set to positive (in some implementations, it can be value 1), and in operation 114, at least some of the points within the sub-volume can be encoded by encoding their coordinate positions within the sub-volume. This can include encoding the X, Y, and Z Cartesian coordinate positions relative to the corners of the sub-volume in some implementations. For example, the corner can be the vertex of the sub-volume closest to the origin of the coordinate system. Depending on the implementation, various techniques for encoding the coordinates can be applied, including prediction operations, differential encoding and decoding, etc.
[0050] Operation 114 was described above as encoding at least some points rather than all points because in some possible implementations, a rate-distortion optimization process may be applied to evaluate whether the rate cost of DCM encoding / decoding the points exceeds the distortion cost of not encoding / decoding the points. Note that if such an RD optimization evaluation affects whether a parent node will be "occupied", then the RD optimization may need to be performed earlier in the encoding / decoding process and / or the process may involve two-pass encoding / decoding.
[0051] Once a node has been encoded, either by using DCM or by conventional encoding of the mode, method 100 fetches the next occupied node / sub-volume from the FIFO buffer, as shown in operation 122, and loops back to evaluate whether the node / sub-volume is eligible for DCM. As described above, the stopping condition will eventually stop the further subdivision of the sub-volumes, and all the nodes in the FIFO will be processed.
[0052] The eligibility evaluation in operation 104 is based on the occupancy data for previously encoded nodes. This allows both the encoder and the decoder to independently make the same eligibility determination.
[0053] Now refer Figure 2 , Figure 2 FIG. shows an example method 200 for decoding a bitstream of encoded point cloud data using IDCM in the form of a flowchart.
[0054] In operation 202, the decoder evaluates whether the currently occupied node of the point cloud data tree is eligible for DCM. The decoder uses the same eligibility determination as used by the encoder. Typically, the eligibility determination is based on some occupancy data from sibling nodes or neighboring nodes, as exemplified above.
[0055] If the node is not eligible, the decoder splits the node in operation 204 and entropy decodes the occupancy pattern, and then pushes any occupied child nodes into the FIFO buffer in operation 206. However, if the node is eligible for DCM, the decoder decodes the DCM flag in operation 208. The decoded DCM flag indicates whether DCM was actually used to encode the points in the current node, as shown in operation 210. In this example, a DCM flag value of 1 corresponds to the use of DCM and a flag value of 0 corresponds to not using DCM. If the DCM flag indicates that DCM was not used, method 200 proceeds to operations 204 and 206 to decode the pattern as normal. If the DCM flag indicates that DCM was used, then in operation 214, the decoder decodes the coordinate point data of any points in the node.
[0056] If the encoder and decoder are configured to use DCM when there is more than one point per node, in operation 212, the decoder decodes the number of points. It should be understood that this value can be encoded as a number less than 1, as it is known that this value must be equal to or greater than 1. Once the decoder knows the number of encoded points, it decodes the coordinate data for each of the points in operation 214.
[0057] After the decoder has decoded the pattern or decoded the point coordinate data, in operation 216, the decoder fetches the next occupied node from the FIFO buffer and returns to operation 202 to evaluate its eligibility for DCM encoding and decoding.
[0058] It should be understood that the above IDCM process involves evaluating eligibility, then encoding and decoding the DCM flag to signal whether DCM is used, and then encoding the number of points (if the threshold allows DCM to be used in the case of two or more points), and encoding the coordinate values for each point. According to one aspect of the present application, the eligibility evaluation is removed, and if the sub - volume includes at least two points to be directly encoded and decoded, the encoding and decoding of at least one pair of coordinate values involves sorting the coordinate values and applying bit - position - based pairwise encoding of the bits of the coordinate values.
[0059] Now reference will be made to Figure 3 , Figure 3 which shows an example sub - volume 300 containing two points: a first point 302 and a second point 304. The first point 302 is located at a position that can be defined using Cartesian coordinates (x1, y1, z1) relative to the origin of the coordinate system. In this example, the origin is the vertex of the sub - volume 300. Similarly, the second point 304 is located at a position defined by (x2, y2, z2).
[0060] When DCM is applied to the sub - volume 300, the position of each of the points is encoded by directly encoding and decoding the coordinates of the position of the point. That is, the coordinate values x1, y1, z1 are each encoded. At the decoder, when the DCM flag indicating that one or more points are encoded and decoded by DCM is decoded, the decoder determines the number of points and then decodes the three coordinate values for each point, where the three coordinate values define the coordinate position of the point.
[0061] The coordinate values can be represented in binary. The length of the binary coordinate values can depend on the depth of the sub - volume and the resolution of the point cloud, i.e., the maximum depth of the encoding and decoding tree. Each of the coordinate values in the coordinate values is related to a direction in the coordinate system. For example, in this example involving Cartesian coordinates, the position of the first point is represented by the x - direction coordinate value x1, the y - direction coordinate value y1, and the z - direction coordinate value z1.
[0062] These points are "unordered" because the encoding and decoding processes do not depend on whether the first or second point is encoded or decoded first. As will be explained below, this property can be used to obtain improved compression when encoding and decoding coordinate values.
[0063] In the example case of two points, starting from one of the coordinate directions (such as the x - direction), the two points can be sorted based on the ascending order of the x - direction coordinates. That is, they are sorted to ensure that x1 is less than or equal to x2. The two x - direction coordinate values (binary) can be represented as:
[0064] x1: b 11 b 12 b 13 ... b 1n , where n is the number of bits in x1.
[0065] x2: b 21 b 22 b 23 ... b 2n , where n is the number of bits in x2.
[0066] The points are sorted such that x1 ≤ x2. The examples described herein are based on the ascending order of coordinate values. It should be understood that in some implementations, the sorting can be in descending order.
[0067] Starting from the most significant bit position, i.e., including bits b 11 and b 21 , these bits will have values (0, 0), (0, 1), or (1, 1). These two bits cannot produce the value (1, 0) because the points have been sorted to ensure x1 ≤ x2. Thus, the encoder can encode the same bit flag to signal whether the two bits b 11 and b 21 are the same. If not, the decoder knows they must be (0, 1). If they are the same, the encoder encodes the bit - value flag to signal whether they are both 0 or both 1. In fact, the bit - value flag may just encode the binary value of the bits, i.e., whether they are 0 or 1.
[0068] In this way, these two bits are encoded using one bit or two bits, depending on whether they are the same. If the two bits b 11 and b 21 are different, after encoding the same bit flag to signal the bits as (0, 1), the encoder uses bypass encoding and decoding to encode the remaining bits of each of the coordinate values x1 and x2. If the two bits b 11 and b 21are the same, then the bit pairs at the next position (e.g., b 12 and b 22 ) undergo the same pairwise encoding process.
[0069] Now referring to Figure 4 , Figure 4 FIG. 400 is a simplified flowchart illustrating a pairwise encoding process for coordinate value pairs. In the case of a Cartesian coordinate system, the coordinate values can be x-direction coordinates, y-direction coordinates, or z-direction coordinates. Starting from the most significant bit position for each of the coordinate values, the bits in that bit position are compared, and the encoder determines in operation 402 whether they are the same. If not, the encoder encodes the same bit flag in operation 404 to signal that they are different. This indicates that these bits are (0, 1) because due to the sorting of the coordinate values, they cannot be (1, 0). In some implementations, the fact that the bits are different from each other is signaled by setting the same bit flag to zero. The encoder then bypass-encodes the remaining bits of the two coordinate values. For example, as indicated in operation 406, bits b 12 to b 1n and bits b 22 to b 2n .
[0070] If the two bits from the current bit position in the coordinate values are the same, then in operation 408, the same bit flag is set to signal that they are the same and is encoded. In some cases, the same bit flag can be set to 1 to signal that the two bits are the same. The encoder then encodes and decodes the bit values in operation 410 to signal whether both of these two bits are 0 or both are 1.
[0071] The pairwise encoding process then returns to operation 402 to evaluate the two bits in the next (subsequent) bit position, e.g., bits b 12 and b 22 .
[0072] The direct coding and decoding mode is typically applied to the case of sparse sub-volumes. The condition for applying DCM is typically that the number of points within the sub-volume is less than or equal to a threshold, where the threshold can be two or three in some examples. However, it has been found that although these points are located in a sparse sub-volume, there is a non-negligible spatial correlation between two points within the sub-volume. This correlation affects the likelihood that the two corresponding coordinate values for the two points are similar, i.e., the probability that the corresponding bits of these coordinate values are the same. Therefore, further compression gain can be achieved by using entropy coding of the same bit flag. Context adaptive binary arithmetic coding and decoding can be used, e.g., using a CABAC (Context Adaptive Binary Arithmetic Coding) engine. A dedicated context can be assigned to encode and decode the same bit flag.
[0073] In some cases, all bits of the first coordinate value and the second coordinate value can be the same. That is, x1 can be equal to x2. After the paired encoding process, if the two coordinate values are equal, the paired encoding and decoding process can be used to encode and decode the coordinate values corresponding to the other coordinate direction. For example, if x1 = x2, the encoder can sort the coordinate values y1 and y2 and apply the above paired encoding and decoding process to these coordinate values.
[0074] Now refer to Figure 5 , Figure 5 which shows the direct coding and decoding mode encoding process 500 in the form of a flowchart. Process 500 is executed by an encoder, which can be implemented using software containing processor-executable instructions. In the following examples, when executed by one or more processing units, the processor-executable instructions cause the processing units to perform the described operations. It should be understood that process 500 is a sub-process in a larger overall point cloud encoding process. In some cases, the entire point cloud encoding process is a tree-based encoding and decoding of occupancy states using recursively partitioned volume spaces. As will be described, process 500 is selectively applied in cases where a sub-volume is eligible for the direct coding and decoding mode.
[0075] In operation 502, the encoder determines whether DCM is to be applied to the current sub-volume. As described above, DCM can be conditioned on the number of points in the sub-volume being less than a threshold number. In some cases, DCM can also be conditioned on the depth of the sub-volume in the encoding and decoding tree. That is, for some unsuitable tree depths, DCM may not be enabled. If the encoder determines that DCM will not be applied, for example, if there are more than a threshold number of points in the sub-volume, then in operation 504, the occupancy state of the sub-volume is encoded using the normal point cloud encoding process.
[0076] If DCM is applied, then in operation 506, the encoder evaluates whether the sub-volume contains more than one point. If the sub-volume contains only one point, the encoder encodes the position of that one point using the normal DCM encoding and decoding process, as shown in operation 508. However, if the sub-volume includes at least two points, the multi-point DCM encoding and decoding process of the present application can be applied.
[0077] As shown in operation 510, the two points are sorted based on a first coordinate value. In some examples, the first coordinate value can be an x-direction value, but in other examples, the first coordinate value can be a y-direction value or a z-direction value. Once sorted, in operation 512, the encoder determines whether the bits in the same bit position in the two sorted coordinate values are the same. Process 500 starts with the most significant bit position as the current bit position. If the two bits in the current bit position (e.g., the most significant bit position in the first iteration) are different, then due to the sorting operation, they must be (0, 1). Thus, in operation 514, the encoder entropy-encodes the same-bit flag that signals the two bits are different, and then in operation 516, bypass-codes the remaining bits of the two coordinate values, if any.
[0078] If it is determined in operation 512 that the two bits are the same, then in operation 518, the encoder entropy-encodes the same-bit flag that signals the two bits in the current bit position are the same. Then in operation 520, the encoder also encodes the bit values of these bits, i.e., whether they are both 0 or both 1. In some implementations, the encoding of the bit values can be bypass-coding.
[0079] In operation 522, the encoder evaluates whether there are more bits in the coordinate values. If so, it advances the current bit position to the next or subsequent bit position, such as by incrementing a pointer or index, and returns to operation 512 to determine whether the bits in the new current bit position are the same. This paired-encoding process is repeated until the two bits are different or until the entire coordinate values are encoded and the same, e.g., x1 = x2. If the two coordinate values are the same, then in operation 524, the encoder determines whether it can apply the paired-coding process to the next coordinate, such as a y coordinate or a z coordinate. If so, process 500 returns to operation 510 to sort the points based on the next coordinate.
[0080] In one variation of process 500 not shown above, operation 510 can include decoding an equal-value flag to signal whether the two coordinate values are the same. If so, the remainder of process 500 can be applied to the next coordinate value. In some cases, the equal-value flag can be entropy-encoded.
[0081] Figure 6 An example direct-coding mode decoding process 600 is shown in flowchart form. Decoding process 600 can be performed by a decoder, and in some implementations, the decoder can be implemented by software including processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the described operations.
[0082] In operation 602, the decoder receives a bitstream of encoded point cloud data. In some examples, the bitstream can be received via a communication channel or read from memory. It will be understood that portions of process 600 are implemented in the context of an overall point cloud decoding operation that results in the output of a decoded point cloud. In certain cases, the decoded point cloud can be rendered and / or displayed, or can be input into additional processes, such as object detection or collision warning processes in some applications. The encoded data in the bitstream can be based on the encoding of tree-based occupancy state information for a volume space.
[0083] In operation 604, the decoder determines that the current sub-volume is encoded and decoded using DCM. In some examples, this includes decoding a DCM flag that signals whether DCM is applied to the sub-volume. As described above, DCM can be applied to sub-volumes containing a number of points equal to or less than a threshold. In certain cases, additional eligibility criteria can be evaluated, such as the depth of the sub-volume in the encoding tree. If the sub-volume is not eligible based on this criterion, the DCM flag is not decoded because the sub-volume is not eligible.
[0084] Assuming the sub-volume is eligible and the decoded DCM flag signals that DCM is applied, in operation 606, the decoder determines the number of points in the sub-volume. This can include decoding multiple point values. If the threshold for DCM is two or fewer points, operation 606 involves decoding a flag that signals whether one or two points are present in the sub-volume. If the threshold allows for more than two points, operation 606 involves decoding the number or number-less-one, since the sub-volume cannot be empty.
[0085] If only one point is present in the sub-volume, as shown in operation 608, the decoder decodes the position data of the point using DCM. If more than one point is present in the sub-volume, the present application applies the multi-point encoding and decoding process described herein, which begins at operation 610.
[0086] The decoder can assume that the encoder has sorted two coordinate values. The selected coordinate values can be pre-determined, for example the x-direction coordinate value can be selected by default or signaled in the bitstream. The signaling of the coordinate values used in process 600 can be in the header of the point cloud data or elsewhere in the metadata associated with the point cloud. In operation 610, the decoder entropy decodes the same bit flag. As shown in operation 612, the decoded same bit flag indicates whether the bits in the current bit positions of the two coordinate values are the same. If not, the decoder knows the corresponding bits are (0, 1) based on the fact that the encoder has sorted these points. Thus, in operation 614, the decoder takes these bits (e.g., bit b 11 and b 21)Reconstructed as (0, 1), and the remaining bits of x1 and x2 (if any) are decoded bypass. The length of the coordinate value will be known to the decoder based on the resolution of the encoding and decoding tree and the depth of the sub - volume in the encoding and decoding tree.
[0087] If the decoded same bit flag signals that the two bits in the current bit position are the same, then in operation 616, the decoder decodes the bit value flag that signals whether both bits are 0 or both are 1. Based on this, the decoder reconstructs these bits (e.g., bits b 11 and b 21 ) as (0, 0) or (1, 1). Then, in operation 618, the decoder evaluates whether there are any other bits in the coordinate value. If so, it increments the current bit position to the subsequent or next bit position in the coordinate value and returns to operation 610 to decode another same bit flag corresponding to the corresponding bit in the current bit position now.
[0088] If there are no other bits to decode for the coordinate value in operation 618, the coordinate value is the same. Then the decoder can evaluate in operation 620 whether it can apply the paired decoding process to the next coordinate value. For example, if the x - direction coordinate value is decoded as the same, it can be configured to apply the process to the y - direction value based on the encoder sorting the points based on the y - direction value and encoding them using paired encoding. If so, it returns to operation 610 to decode the same bit flag related to the y - direction coordinate value.
[0089] Continue to recursively apply process 600 until the coordinate values are different, at which time the remaining bits of the current coordinate value and any additional coordinate values are decoded bypass.
[0090] For ease of illustration, the above example uses the case of two points, but other embodiments of the process can be applied to the case of three or more points. For example, joint encoding of three points can be implemented. If there are three points in the sub - volume and these three points have, for example, x - coordinate values, the encoder can sort them such that x1 ≤ x2 ≤ x3.
[0091] Starting from the most significant bit position, i.e., containing bits b 11 and b 21 and b 31, these bits will have values (0, 0, 0), (0, 0, 1), (0, 1, 1), or (1, 1, 1). A 3-tuple of three bits cannot produce the values (0, 1, 0), (1, 0, 0), (1, 0, 1), or (1, 1, 0) because these points have been sorted to ensure x1 ≤ x2 ≤ x3. To encode and decode the bits, in cases where the first two bits are the same and equal to zero, e.g., (0, 0, b3), the encoder can utilize the addition flag as before to apply the pairwise process to encode whether the third bit b3 is 0 or 1.
[0092] Figure 7 A simplified example process 700 for encoding and decoding a triple of bits in the case of a three-point DCM encoding and decoding process is shown. Process 700 includes evaluating in operation 702 whether b1 and b2 are the same. If they are different, the only possible triple is (0, 1, 1), so in operation 704 the same-bit flag for b1 and b2 is encoded and the remaining bits are bypass-encoded. If the first two bits are the same, in operation 706 the same-bit flag is encoded, and a bit-value flag is encoded to signal whether both of them are 0 or both of them are 1, as shown in operations 708, 710, and 712. If both of them are 1, the triple is (1, 1, 1). If the first two bits are both zero, in operations 714, 716, and 718, the value of b3 is encoded to signal whether the triple is (0, 0, 1) or (0, 0, 0).
[0093] The extension to 4-tuples is straightforward. If b1 and b2 are different, the 4-tuple is (0, 1, 1, 1). If b1 and b2 are the same and both are 1, the 4-tuple is (1, 1, 1, 1). If both of them are zero and b3 is 1, the 4-tuple is (0, 0, 1, 1). An additional flag (beyond the flags encoded in the 3-tuple example) is encoded in the case where b1, b2, and b3 are (0, 0, 0), in which case b4 is encoded to signal whether the 4-tuple is (0, 0, 0, 1) or (0, 0, 0, 0). It will be understood from the above description how the process can be further extended to n-tuples.
[0094] Now refer to Figure 8 , Figure 8A simplified block diagram illustrating an example embodiment of an encoder 800 is shown. The encoder 800 includes a processor 802, a memory 804, and an encoding application 806. The encoding application 806 may include a computer program or application stored in the memory 804 and containing instructions that, when executed, cause the processor 802 to perform operations such as those described herein. For example, the encoding application 806 may encode and output a bitstream encoded according to the processes described herein. It should be understood that the encoding application 806 may be stored on a non-transitory computer-readable medium, such as an optical disc, a flash device, a random access memory, a hard disk drive, etc. When the instructions are executed, the processor 802 performs the operations and functions specified in the instructions so as to operate as a dedicated processor for implementing the (one or more) described processes. In some examples, such a processor may be referred to as a "processor circuit" or "processor circuitry".
[0095] Now also referring to Figure 9 , Figure 9 A simplified block diagram illustrating an example embodiment of a decoder 900 is shown. The decoder 900 includes a processor 902, a memory 904, and a decoding application 906. The decoding application 906 may include a computer program or application stored in the memory 904 and containing instructions that, when executed, cause the processor 902 to perform operations such as those described herein. It should be understood that the decoding application 906 may be stored on a computer-readable medium, such as an optical disc, a flash device, a random access memory, a hard disk drive, etc. When the instructions are executed, the processor 902 performs the operations and functions specified in the instructions so as to operate as a dedicated processor for implementing the (one or more) described processes. In some examples, such a processor may be referred to as a "processor circuit" or "processor circuitry".
[0096] It should be understood that the decoder and / or encoder according to the present application may be implemented in a plurality of computing devices, including but not limited to servers, appropriately programmed general-purpose computers, machine visualization systems, and mobile devices. The decoder or encoder may be implemented by software containing instructions for configuring one or more processors to perform the functions described herein. The software instructions may be stored on any suitable non-transitory computer-readable memory, including CDs, RAMs, ROMs, flash memories, etc.
[0097] It should be understood that the decoders and / or encoders described herein, as well as the modules, routines, procedures, threads, or other software components that implement the described methods / procedures for configuring an encoder or decoder, can be implemented using standard computer programming techniques and languages. This application is not limited to specific processors, computer languages, computer programming conventions, data structures, or other such implementation details. Those skilled in the art will recognize that the described procedures can be implemented as part of computer-executable code stored in volatile or non-volatile memory, as part of a dedicated integrated chip (ASIC), etc.
[0098] This application also provides a computer-readable signal for encoding data generated by applying the encoding process according to this application.
[0099] Impact on compression performance
[0100] Tests using pairwise encoding and decoding in the case of two points have been performed on multiple example point clouds with different characteristics at different resolutions. The tests evaluated the encoding and decoding compression of bits per point ("bpp"), as well as the encoding and decoding complexity measured in terms of the number of tree nodes in the encoding and decoding. The tests involved the current implementation of the Moving Picture Experts Group (MPEG) test model using IDCM, a variant of the test model using "simple" DCM without any implicit eligibility assessment and only using a threshold point count test to determine whether to apply DCM, and the current implementation of the pairwise encoding and decoding process for DCM.
[0101] The use of "simple" DCM results in a significant reduction in the number of tree nodes processed, for example, an increase in speed, a reduction in memory requirements and memory access operations. In some tests, the number of nodes was reduced by a factor of 6. However, "simple" DCM also results in an increase in bpp of 0.1 to more than 1.
[0102] Using pairwise encoding and decoding for DCM as described herein results in the same reduction in the number of tree nodes (since the DCM application criteria are the same), but results in a bpp that is approximately the same as that of IDCM, or in some cases a reduced bpp. Thus, the methods and systems of the embodiments of this application not only result in faster processing of point cloud data, but in some cases, they also provide a compression gain.
[0103] Certain adaptations and modifications can be made to the described embodiments. Thus, the embodiments discussed above are considered illustrative rather than restrictive.
Claims
1. A method for encoding a point cloud to generate a compressed point cloud data bitstream, wherein a current sub-volume includes a first point and a second point, the first point having a first position within the sub-volume defined by a first coordinate value and the second point having a second position within the sub-volume defined by a second coordinate value, the method comprising: Sorting the first coordinate value and the second coordinate value, wherein the first coordinate value and the second coordinate value are binary; Starting from the most significant bit position, pairwise encoding the current position of the first coordinate value and the current position of the second coordinate value by: Encoding a same bit flag that indicates whether the bits in the current positions of the first coordinate value and the second coordinate value are the same; When the bits in the current positions are different, not encoding the bits in the current positions and bypass encoding any remaining bits of the first coordinate value and any remaining bits of the second coordinate value; When the bits in the current position are the same, encoding a bit value flag that indicates whether both bits are 1 or both bits are 0; and recursively repeating the pairwise encoding for the next position in the first coordinate value and the next position in the second coordinate value until the first coordinate value and the second coordinate value are encoded.
2. The method according to claim 1, wherein encoding the same bit flag comprises: Entropy encoding the same bit flag.
3. The method according to claim 2, wherein entropy encoding the same bit flag uses a dedicated context for entropy encoding of the same bit flag.
4. The method according to claim 1, wherein the first coordinate value and the second coordinate value correspond to directions in a Cartesian coordinate system.
5. The method according to claim 4, wherein the direction is the x-direction, the y-direction, or the z-direction.
6. The method according to claim 1, wherein the first position is further defined by a third coordinate value and the second position is further defined by a fourth coordinate value, the third coordinate value and the fourth coordinate value corresponding to the same direction in a coordinate system, wherein the method further comprises: Determining that the first coordinate value and the second coordinate value are the same, and as a result Sorting the third coordinate value and the fourth coordinate value, and Recursively performing the pairwise encoding with respect to the third coordinate value and the fourth coordinate value until the third coordinate value and the fourth coordinate value are encoded.
7. The method according to claim 1, further comprising: First determining that the sub-volume includes at least two points and determining that direct encoding and decoding is to be applied to the at least two points.
8. The method according to claim 7, wherein determining that direct encoding / decoding is to be applied comprises: Determining that the number of points in the sub-volume is less than a threshold number.
9. The method according to claim 7 further comprises: Encoding a direct encoding and decoding mode flag that signals that direct encoding and decoding is used to encode the at least two points.
10. An encoder for encoding a point cloud to generate a compressed point cloud data bitstream, wherein a current sub-volume includes a first point and a second point, the first point having a first position within the sub-volume defined by a first coordinate value and the second point having a second position within the sub-volume defined by a second coordinate value, the encoder comprising: A processor; Memory; And A coding application, comprising instructions executable by the processor, which when executed cause the processor to perform: Sort the first coordinate value and the second coordinate value, wherein the first coordinate value and the second coordinate value are binary; Starting from the most significant bit position, pair - code the current positions of the first coordinate value and the second coordinate value as follows: Code a same - bit flag, which indicates whether the bits in the current positions of the first coordinate value and the second coordinate value are the same; When the bits in the current positions are different, do not code the bits in the current positions, and bypass - code any remaining bits of the first coordinate value and any remaining bits of the second coordinate value; When the bits in the current position are the same, code a bit - value flag, which indicates whether both bits are 1 or both bits are 0; And Recursively repeat the pair - coding for the next positions in the first coordinate value and the next positions in the second coordinate value until the first coordinate value and the second coordinate value are coded.
11. A method for decoding a compressed point - cloud data bitstream to produce a reconstructed point cloud, wherein a current sub - volume includes a first point and a second point, the first point having a first position within the sub - volume defined by a first coordinate value and the second point having a second position within the sub - volume defined by a second coordinate value, the method comprising: Starting from the most significant bit position, pair - decode the current positions of the first coordinate value and the second coordinate value as follows: Decode a same - bit flag, which indicates whether the bits in the current positions of the first coordinate value and the second coordinate value are the same; When the bits in the current positions are different, reconstruct the bits in the current positions, and bypass - decode any remaining bits of the first coordinate value and any remaining bits of the second coordinate value; When the bits in the current positions are the same, decode a bit - value flag, which indicates whether both bits in the current position are 1 or both bits are 0; Recursively repeat the pair - decoding for the next positions in the first coordinate value and the next positions in the second coordinate value until the first coordinate value and the second coordinate value are decoded; And Output the reconstructed first coordinate value for the first point and the reconstructed second coordinate value for the second point.
12. The method according to claim 11, wherein decoding the same bit flag comprises: Entropy - decode the same - bit flag.
13. The method according to claim 12, wherein entropy - decoding the same - bit flag uses a dedicated context for entropy - decoding the same - bit flag.
14. The method according to claim 11, wherein the first coordinate value and the second coordinate value correspond to directions in a Cartesian coordinate system.
15. The method according to claim 14, wherein the direction is the x - direction, the y - direction, or the z - direction.
16. The method according to claim 11, wherein the first position is further defined by a third coordinate value and the second position is further defined by a fourth coordinate value, the third coordinate value and the fourth coordinate value corresponding to the same direction in a coordinate system, and wherein the method further comprises: determining that the first coordinate value and the second coordinate value are the same, and as a result recursively performing the paired decoding with respect to the third coordinate value and the fourth coordinate value until the third coordinate value and the fourth coordinate value are reconstructed.
17. The method according to claim 11 further comprises: decoding a direct encoding / decoding mode flag that signals that direct encoding / decoding is used to encode the points in the sub-volume.
18. The method according to claim 17 further comprises: decoding a value of the number of points to determine that the number of points in the sub-volume is greater than 1.
19. A decoder for decoding a compressed point cloud data bitstream to produce a reconstructed point cloud, wherein a current sub-volume includes a first point and a second point, the first point having a first position within the sub-volume defined by a first coordinate value and the second point having a second position within the sub-volume defined by a second coordinate value, the decoder comprising: a processor; a memory; and a decoding application including instructions executable by the processor, the instructions when executed causing the processor to perform: starting from a most significant bit position, performing paired decoding of a current position of the first coordinate value and a current position of the second coordinate value by: decoding a same bit flag that indicates whether bits in the current positions of the first coordinate value and the second coordinate value are the same; when the bits in the current positions are different, reconstructing the bits in the current positions and bypass decoding any remaining bits of the first coordinate value and any remaining bits of the second coordinate value; when the bits in the current positions are the same, decoding a bit value flag that indicates whether both bits in the current positions are 1 or both are 0; recursively repeating the paired decoding for a next position in the first coordinate value and a next position in the second coordinate value until the first coordinate value and the second coordinate value are decoded; and outputting a reconstructed first coordinate value for the first point and a reconstructed second coordinate value for the second point.
20. A non-transitory processor-readable medium storing processor-executable instructions for decoding a compressed point cloud data bitstream to produce a reconstructed point cloud, wherein a current sub-volume includes a first point and a second point, the first point having a first position within the sub-volume defined by a first coordinate value and the second point having a second position within the sub-volume defined by a second coordinate value, the processor-executable instructions when executed by a processor causing the processor to perform: starting from a most significant bit position, performing paired decoding of a current position of the first coordinate value and a current position of the second coordinate value by: Decode a same bit flag that indicates whether bits in the current positions in the first coordinate value and the second coordinate value are the same; When the bits in the current position are different, reconstruct the bits in the current position and bypass-decode any remaining bits of the first coordinate value and any remaining bits of the second coordinate value; When the bits in the current position are the same, decode a bit value flag that indicates whether both bits in the current position are 1 or both are 0; Recursively repeat the paired decoding for the next positions in the first coordinate value and the next positions in the second coordinate value until the first coordinate value and the second coordinate value are decoded; And Output the reconstructed first coordinate value for the first point and the reconstructed second coordinate value for the second point.
21. A non-transitory processor-readable medium storing processor-executable instructions for a method of encoding a point cloud to generate a compressed point cloud data bitstream, where a current subvolume includes a first point and a second point, the first point having a first position within the subvolume defined by a first coordinate value and the second point having a second position within the subvolume defined by a second coordinate value, the processor-executable instructions causing the processor to perform when executed by the processor: Sort the first coordinate value and the second coordinate value, where the first coordinate value and the second coordinate value are binary; Starting from the most significant bit position, perform paired encoding of the current position of the first coordinate value and the current position of the second coordinate value by: Encoding a same bit flag that indicates whether bits in the current positions in the first coordinate value and the second coordinate value are the same; When the bits in the current position are different, do not encode the bits in the current position and perform bypass encoding on any remaining bits of the first coordinate value and any remaining bits of the second coordinate value; When the bits in the current position are the same, then encode a bit value flag that indicates whether both bits are 1 or both are 0; and recursively repeat the paired encoding for the next positions in the first coordinate value and the next positions in the second coordinate value until the first coordinate value and the second coordinate value are encoded.
Citation Information
Patent Citations
Methods and devices using direct coding in point cloud compression
WO2019140508A1
Method and apparatus for efficient transform unit encoding
CN104012092A
Point cloud compression
US20190311500A1