Method, encoder and decoder for encoding and decoding 3D point clouds

CN119631408BActive Publication Date: 2026-09-29BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280098723.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2026-09-29
Estimated Expiration
2042-08-11

Smart Images

  • Figure CN119631408B_ABST
    Figure CN119631408B_ABST
Patent Text Reader

Abstract

An encoding method and a decoding method, and an encoder and a decoder, for encoding a 3D point cloud into a bitstream. The geometry of the point cloud is defined by an octree structure comprising a plurality of nodes having a parent-child relationship and representing three-dimensional positions of an object, the point cloud being located within a volume space of a three-dimensional coordinate system, the volume space being hierarchically partitioned into sub-volumes and containing points of the point cloud, wherein a volume is partitioned into a number of sub-volumes, each sub-level sub-volume being associated with one of the plurality of nodes of the octree structure, wherein occupancy information associated with each of the sub-level sub-volumes indicates whether the respective sub-level sub-volume contains at least one of the points of the point cloud, the method comprising: along each axis X, Y, Z of the coordinate system, obtaining, in a lexicographic order, nodes of a depth d containing at least a portion of the point cloud; entropy decoding occupancy information of child nodes of each node from the bitstream based on the lexicographic order.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a method for encoding 3D point clouds into a bitstream. Furthermore, the purpose of this application is to provide a method for decoding 3D point clouds from a bitstream. Additionally, the purpose of this application is to provide an encoder and decoder, a bitstream encoded according to this application, and software. In particular, the purpose of this application is to provide a method for improving encoding efficiency and reducing processing latency. Background Technology

[0002] Point clouds, as a format for representing 3D data, have attracted much attention due to their versatility in representing all types of 3D objects or scenes. Therefore, point clouds can handle a variety of use cases, including:

[0003] Post-production of the film

[0004] • Real-time 3D immersive remote presence or VR (virtual reality) / AR (augmented reality) applications.

[0005] • Free-viewpoint video (e.g., for watching sports events),

[0006] Geographic Information System (also known as mapping),

[0007] • Cultural heritage (scanning and storing rare objects in digital form),

[0008] • Autonomous driving, including 3D mapping of the environment and real-time LiDAR data acquisition.

[0009] A point cloud is a set of points in 3D space, each of which can have additional values. These additional values ​​are often called point attributes. Therefore, a point cloud is a combination of geometry (the 3D position of each point) and attributes.

[0010] These properties can be, for example, three-component color, material properties (such as reflectivity), and / or two-component normal vectors of the surface associated with the point.

[0011] Point clouds can be captured by various types of devices, such as camera arrays, depth sensors, LiDAR, and scanners, or they can be generated by computers (e.g., in film post-production). Depending on the use case, point clouds in mapping applications can contain anywhere from thousands to billions of points.

[0012] In the raw representation of point clouds, each point requires a very high number of bits; each spatial component X, Y, or Z has at least a dozen bits, and optionally, attributes require even more bits—for example, color requires three times the number of 10 bits. The practical deployment of point cloud-based applications necessitates compression techniques to enable the storage and distribution of point clouds under reasonable storage and transmission infrastructure conditions.

[0013] Compression may be lossy (similar to lossy video compression) during the distribution of point clouds to end users and their visualization, such as on AR / VR glasses or any other 3D-enabled device. Other use cases require lossless compression, such as medical applications or autonomous driving, to avoid compromising decision-making outcomes obtained through analysis of compressed and transmitted point clouds.

[0014] Until recently, point cloud compression (also known as PCC) had not been solved by the mass market, and there was no standardized point cloud codec available. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11 (also known as the Moving Picture Experts Group or MPEG) launched a working project on point cloud compression.

[0015] • MPEG-I Part 5 (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC); and

[0016] • MPEG-I Part 9 (ISO / IEC 23090-9) or Geometry-based Point Cloud Compression (G-PCC).

[0017] The V-PCC and G-PCC standards completed their first versions at the end of 2020 and will soon be available on the market.

[0018] The V-PCC encoding method compresses point clouds by projecting 3D objects multiple times to obtain 2D patches that are packed into images (or into videos, suitable for moving point clouds). Existing image / video codecs are then used to compress the resulting images or videos, enabling the use of already deployed image and video solutions. By its very nature, V-PCC is only effective on dense and continuous point clouds because image / video codecs cannot compress non-smooth patches obtained from projections of sparse geometry acquired from sources such as LiDAR.

[0019] The G-PCC encoding method has two schemes for compressed geometry.

[0020] The first approach is based on an occupancy tree (octree / quadtree / binary tree) representation of the point cloud geometry. Occupied nodes are branched until a certain size is reached, at which point the occupied leaf nodes provide the location of the points, typically the center of these nodes. High compression can be achieved for dense point clouds by using neighborhood-based prediction techniques. Sparse point clouds are also addressed by directly encoding the location of points within nodes with non-minimum sizes, stopping tree construction when only isolated points exist within a node; this technique is called Direct Coding Mode (DCM).

[0021] The second approach is based on a prediction tree, where each node represents the 3D location of a point, and the relationship between nodes is a spatial prediction from parent to child nodes. This method can only solve sparse point clouds and has the advantages of lower latency and simpler decoding than occupancy-based trees. However, compared to the first occupancy-based approach, its compression performance is only slightly better, and it has complex encoding, densely searching for the best predictor (from a long list of potential sub-volume predictors) when constructing the prediction tree.

[0022] In both schemes, attribute encoding is performed after geometric encoding (decoding) is complete, resulting in two-pass coding. Therefore, low latency is achieved by using slices that decompose the 3D space into independently coded sub-volumes without requiring prediction between sub-volumes. However, using many slices can severely impact compression performance.

[0023] One important use case is the transmission of dynamic AR / VR point clouds. "Dynamic" means that the point cloud evolves relative to time. Furthermore, AR / VR point clouds are typically locally 2D because they represent the surface of objects most of the time. Therefore, AR / VR point clouds are highly correlated (or dense), meaning that points are rarely isolated but rather have many neighboring points.

[0024] In GPCC, although neighboring nodes are reduced to the nearest neighboring nodes, the number of possible configurations is still too high to be directly used to select the context that will be used for the current node’s occupancy entropy encoding.

[0025] Several techniques have been proposed to further reduce the neighborhood configuration to a reduced number of feasible configurations.

[0026] Firstly, some "physical" arguments are introduced, such as another occupied neighboring node placed between the occupied neighboring node and the current node, which masks the occupied neighboring node; in this case, the information of the occupied neighboring node is discarded because the information should be weaker than the information of the other occupied neighboring node mentioned above.

[0027] Secondly, based on learning, a reduction method using lookup tables (LUTs) is proposed.

[0028] Reduced configuration = LUT[configuration]

[0029] The output range of the LUT is less than the number of entries. The construction of the LUT depends heavily on the type of point cloud being learned. The LUT is used during the encoding of child node occupancy. The LUT is a lookup table of encoder indices configured based on the occupancy information of the currently encoded child node, and it maps the occupancy information configuration to select the encoder index, with the selected binary encoder used to encode the current child node.

[0030] Thirdly, a more flexible technique called "Optimal Binary Coder with Update on the Fly (OBUF)" has been introduced to reduce the number of contexts in entropy encoders (such as CABAC (Context-based Adaptive Binary Arithmetic Coding)) by:

[0031] • Use a finite number of adaptive entropy encoders (e.g., 32).

[0032] • Associate the encoder index (e.g., 0 to 31) with each neighborhood configuration.

[0033] • After each encoding of the current node's occupancy status, update the encoder index associated with the current node's neighborhood configuration.

[0034] OBUF combines the two reduction techniques previously used by OBUF to limit the number of encoder indices that depend on the reduction configuration.

[0035] A practical implementation for entropy encoding of the occupancy bits of the current geometric element is also proposed, using pre-encoded occupancy information of neighboring nodes, as follows:

[0036] • Use easily computed, low-memory reduction functions to reduce neighborhood occupancy configurations;

[0037] • A reduced configuration is used as input to a LUT that points to a partial entropy encoder; this LUT is updated after the bit-occupying encoding. This is what is known in the prior art as the OBUF process;

[0038] • Use the entropy encoder pointed to by the LUT entry indicated by the reduction configuration to entropy encode the occupied bits.

[0039] Once a sufficient number of occupied bits have been encoded, and the LUT has been updated a sufficient number of times to converge and consistently point to an entropy encoder whose encoding probability is well correlated with the reduction configuration, the process is very efficient.

[0040] However, when the point cloud is small or at the beginning of the encoding process for a large point cloud, the OBUF LUT has not yet converged and cannot select a sufficient entropy encoder. This leads to poor compression of small point clouds or degraded compression of large point clouds.

[0041] In the geometric 3D octree representation of a point cloud, the number of possible occupancy configurations for a neighborhood is 2N, where N is the number of nodes / voxels involved in the neighborhood. Even with only a small number of N neighbors, this number becomes very large.

[0042] Therefore, it is proposed to add an additional reduction step after the (optional) fixed reduction function. This new reduction will be called "dynamic reduction," and it will be updated according to the encoding progress of the point cloud.

[0043] A typical example of the order of magnitude of the context information involved when processing adjacent nodes in an octree representing point cloud geometry is as follows.

[0044] There are 26 neighboring nodes that share faces, edges, or vertices with the current node's parent node. The occupied bits of the current node need to be encoded. When using a breadth-first search of an octree, all the occupied bits of these 26 neighboring nodes have already been encoded and can be used as context information.

[0045] Depending on the scan sort (Morton sort of raster scan sort), most of the time, the occupied bits of at least 7 of the child nodes out of the 26 neighboring nodes have already been encoded. This adds 8 * 7 = 56 occupied bits that can be used as context information.

[0046] Finally, some of the current node's sibling nodes may have already been encoded. There are at most 7 of them. Therefore, by considering only the nearest neighbor that touches the current node, a total of up to 26 + 56 + 7 = 89 encoded occupied bits can be used as context information. This results in 2 89 ≈1000 9 =10 27 A possible neighborhood configuration is used as context information.

[0047] When encoding point clouds, not all of these configurations are accessed. However, they must be preserved to cover a wide variety of point clouds. For example, encoding dense AR / VR point clouds will result in access to some reduced contextual configurations, but encoding sparse LiDAR point clouds will result in access to other reduced contextual configurations. This is the main goal of dynamic scaling: to further scale down configurations that are rarely or never accessed, and to give full potential (i.e., no scaling or slight scaling) to frequently accessed configurations. The dynamism comes from the on-the-fly determination of frequently accessed configurations during encoding and the updating of the function of dynamic scaling based on that determination.

[0048] Therefore, essentially, the dynamically reduced function DR is initialized with an initial function DR0 that has the maximum reduction for all configurations, and then updated to DR. 1 DR 2 , ...,DR n Each update depends on the configuration accessed during encoding and is determined by relaxing the reduction of frequently accessed configurations. A tower structure for the function DR is obtained, with an ever-increasing image size equal to the number of configurations that can be dynamically reduced for the possible output.

[0049] #Im DR 0 ≤#Im DR 1 ≤…≤#Im DR n ≤#Im DR n+1 ≤…

[0050] Furthermore, a dynamically reduced function DR is proposed to operate by preserving certain bits of contextual information (CI). It is assumed that the contextual information CI consists of a series of K bits.

[0051] CI=β1…βK

[0052] Then we can maintain the top k of CI. n (CI) bits are used to simply define the dynamic shrinkage function DR n This makes the dynamically reduced context information CI' the

[0053] CI' = β1…βkn(CI)

[0054] And the function DR n Completely composed of array k n (CI) definition. The value of k0 can be initialized to zero.

[0055] k n (CI)=0

[0056] And by using array k n(CI) is changed to the new array k. n+1 (CI) to execute from DR n To DR n+1 The update is performed by increasing the number of hold bits.

[0057] k n (CI)≤k n+1 (CI)

[0058] A series of dynamic reduction functions DR based on binary tree construction are proposed. n Iterative construction (exceeding n). Initialization function DR n =0 corresponds to a tree that has only a root node.

[0059] exist Figure 6 An example function DR is shown in the figure. n It is defined by a binary tree. For a node of depth d, each branch of the tree is selected based on the value (0 or 1) of the d-th bit βd of the context information CI. Leaf nodes correspond to a dynamically shrinking configuration.

[0060] like Figure 6 As shown in the example tree, there are three dynamically reduced configurations corresponding to context information CIs with binary forms 0xxxxxx, 10xxxxx, and 11xxxxx. These three configurations are appended with three numbers N(0xxxxxx), N(10xxxxx), and N(11xxxxx), which track the number of accesses so far during the point cloud encoding process. These numbers count the number of accesses to each reduced configuration CI' during encoding.

[0061] Because the function DR is used n The binary representation of k n (CI), so the function can be derived from array k. n (CI) is represented as follows:

[0062] kn(0xxxxxx) = 1 and kn(1xxxxxx) = 2

[0063] When the quantity N exceeds the threshold th, the function DR n Evolves into the updated function DR n+1 .

[0064] Figure 7 The update function DR is shown when N(10xxxxx) has exceeded the threshold th. n+1 .

[0065] The tree described above has been extended from the node corresponding to configuration 10xxxxx to two child nodes corresponding to configurations 100xxxx and 101xxxx. In general, when N(C) exceeds the threshold of configuration C, the node corresponding to configuration C is extended to two child nodes associated with the two new configurations that replace configuration C. These two new configurations are then increased by one bit after the reduction.

[0066] After the update, the function DR n+1 It can be determined by array k n (CI) is represented as follows:

[0067] k n (0xxxxxx) = 1, k n (11xxxxx) = 2, and k n (10xxxxx) = 3.

[0068] However, based on current technology, there is still room for improvement in encoding efficiency. In particular, the child occupancy of each node is encoded by Morton sort, and the encoding order of the nodes is also Morton sort.

[0069] Therefore, there are at least two drawbacks in the prior art, described as follows:

[0070] Firstly, in existing technologies, when considering the encoding of the child occupancy of the current node, not all adjacent child volume occupancy has been encoded. Therefore, some child volume occupancy information of adjacent nodes is lost. This will be explained below.

[0071] Using breadth-first Morton scan sorting, available adjacent child nodes are as follows: Figure 8 As shown. The occupancy of the child nodes of the 7 neighboring nodes at the lower xyz corner (a) is always encoded and therefore available. However, the occupancy of the child nodes of the 7 neighboring nodes at the higher xyz corner (b) is never encoded and therefore unavailable. In between, the occupancy of the child nodes of the remaining 12 neighboring nodes may or may not have been encoded, and therefore may or may not be available.

[0072] The availability of these 12 adjacent nodes depends on the current node's position in the octree and its corresponding Morton code. Specifically, according to the first order, it depends on the current node's position within its parent node.

[0073] Therefore, one of the main drawbacks of using Merton sort is that the neighboring patterns used to determine the entropy context or make predictions can vary considerably depending on the position of different subvolumes in the Merton sort. This complicates the design of efficient entropy encoding for the occupancy tree and may also lead to suboptimal compression.

[0074] Secondly, a known advantage of using Merton sort (or some other known similar space-filling curves) is that the sort tends to allow handling nodes that are close together in 3D, and close together (i.e., the distance between Merton codes / indices) within the 1D Merton sort. Of course, this is only a trend.

[0075] However, there are exceptions where two nodes at a certain depth might be close in 3D but very far in 1D Morton sorting. This can happen, for example, when these nodes are close but opposite the center of a larger node. In one example, the first voxel belongs to the root node. Figure 1 The first voxel corresponds to the volume of b0 and is close to the center of the root node; the second voxel belongs to the volume of b7 and is also close to the center of the parent node. In extreme cases, these two voxels may be so close that they will share an edge, but after several depths in the octree traversal, the Morton codes associated with the refined nodes containing these voxels become increasingly distant in the 1D Morton sort.

[0076] For example, given two points (x0, y0, z0) and (x1, y1, z1), where (x0, y0, z0) = (255, 255, 255) and (x1, y1, z1) = (256, 256, 255), they can be represented in binary as follows:

[0077] (x0, y0, z0) = (255, 255, 255) = (011111111b, 011111111b, 011111111b), (x1, y1, z1) = (256, 256, 255) = (100000000b, 100000000b, 011111111b).

[0078] By combining the first bit values ​​of their three coordinates, namely 000b for point (x0, y0, z0) and 110b for point (x1, y1, z1), we can determine that point (x0, y0, z0) belongs to node b0, and point (x1, y1, z1) belongs to node b6 in the root node hierarchy. For example... Figure 1 As shown, they share the same edges.

[0079] And its Morton code can be:

[0080] m0=000111111111111111111111111b=16 777 275

[0081] m1=110001001001001001001001001b=103 060 041

[0082] Then the Euclidean distance in Cartesian space can be

[0083]

[0084] Furthermore, the Euclidean distance in the Morton code space can be

[0085]

[0086] Therefore, the Morton code associated with the refined node containing 2 points becomes further away in the 1D Morton sort. Summary of the Invention

[0087] Therefore, the purpose of this application is to provide a method for encoding 3D point clouds into a bitstream, and a method for decoding the geometry of 3D point clouds from the bitstream, which has higher encoding efficiency.

[0088] The problem is solved by the decoding method of claim 1, the encoding method of claim 2, the encoder of claim 12, the decoder of claim 13, the bitstream of claim 14, and the software of claim 15.

[0089] A first aspect of this application provides a method for decoding a 3D point cloud from a bitstream, preferably applied to a decoder. The geometry of the point cloud is defined by an octree structure comprising multiple nodes having parent-child relationships and representing the three-dimensional position of objects. The point cloud lies within a volume space of a three-dimensional coordinate system, the volume space being progressively divided into sub-volumes containing points of the point cloud. Each sub-volume is associated with one of the multiple nodes of the octree structure, and occupancy information associated with each sub-volume indicates whether the corresponding sub-volume contains at least one point of the point cloud. The method includes:

[0090] Along each axis X, Y, Z of the coordinate system, nodes of depth d containing at least a portion of the point cloud are obtained in a lexicographical order.

[0091] Based on the dictionary sorting, the occupancy information of each node's child nodes is entropy decoded from the bitstream.

[0092] On the other hand, a method for encoding 3D point clouds into a bitstream is provided, preferably applied to an encoder. The geometry of the point cloud is defined by an octree structure comprising multiple nodes having parent-child relationships and representing the three-dimensional position of objects. The point cloud lies within a volume space of a three-dimensional coordinate system, the volume space being progressively divided into sub-volumes containing points of the point cloud. Each sub-volume is associated with one of the multiple nodes of the octree structure, and occupancy information associated with each sub-volume indicates whether the corresponding sub-volume contains at least one point of the point cloud. The method includes:

[0093] Along each axis X, Y, Z of the coordinate system, nodes of depth d containing at least a portion of the point cloud are obtained in a lexicographical order.

[0094] Based on the dictionary sorting, the occupancy information of each node's child nodes is entropy encoded into the bit stream.

[0095] A point cloud is a set of points in a three-dimensional coordinate system. These points are typically intended to represent the outer surface of one or more objects. Each point has a location (position) in the three-dimensional coordinate system. The position can be represented by three coordinates (X, Y, Z), which can be Cartesian or any other coordinate system. Therefore, according to this application, for a given depth d in an octree, nodes are obtained in a lexicographical order, and the occupancy of child nodes is encoded at full resolution in a lexicographical order. Preferably, the nodes are stored in a linked data structure (e.g., a linked list). The linked data structure has elements connected to each other, such that the elements are arranged in a sorted manner. Preferably, a list L can be used. d Nodes are stored in a list, and nodes can be retrieved iteratively from the list. Utilizing lexicographical sorting (regardless of the axis order in the lexicographical sort), the occupancy of any sub-volume with lower z-position, lower y-position, and lower x-position can be used for contextual analysis. Therefore, the prediction of different sub-volumes is more stable compared to existing techniques (e.g., Merton sort). This improves encoding efficiency and reduces processing latency.

[0096] Preferably, the method further includes storing occupied child nodes at depth d+1 in lexicographical sort. Therefore, if a node at depth d is not a leaf node, the octree can continue to be traversed in lexicographical sort. The child nodes at the next depth d+1 are already stored in the preferred sort, thereby further facilitating the next iteration.

[0097] Preferably, the lexicographical sorting of nodes at depth d and / or the lexicographical sorting of nodes at depth d+1 is a raster scan sort.

[0098] Preferably, the occupancy information of the child nodes of each node at depth d is entropy-encoded into a bitstream based on a lexicographical sort of the child nodes at depth d+1, or entropy-decoded from the bitstream. Therefore, for any child node whose position is lower than the already encoded child node z, y, x, its occupancy can be used as context. Then, a predetermined pattern of adjacent child nodes can be defined and used to select the encoding context for any encoded occupancy bits.

[0099] Preferably, the occupancy information of each child node is further entropy encoded into a bitstream based on the same adjacent child node occupancy pattern indicating the context of each child node, or entropy decoded from the bitstream. Therefore, the adjacent occupancy pattern used to construct the context for each child node can be the same and predefined, and by using the same adjacent child node occupancy pattern, the construction of context information can be simplified, thereby improving encoding efficiency.

[0100] Preferably, the context of each child node is selected based on the child node occupancy pattern and / or the occupancy pattern of adjacent parent nodes. If the adjacent node occupancy pattern is used for the selection of the entropy coding context, the context selection can also depend on the child occupancy index within the node, where the index indicates the position of the child node within the node. This is because the information provided by the occupancy bits of the parent node's six adjacent nodes will not have the same importance in selecting a suitable entropy coding context for encoding the sub-volume / node. Therefore, by considering the child occupancy index, i.e., the position of the child node within the parent node, the importance of each occupancy bit of adjacent nodes can be distinguished. Preferably, each index is associated with a weight indicating the importance of adjacent nodes.

[0101] Preferably, occupied child nodes at depth d+1 are stored in lexicographical order by traversing each node four times. Here, four pipelines can be defined for each node at a certain depth. Each pipeline is formed by the corresponding pipelines of all nodes with fixed coordinate positions in a plane formed by two of the three coordinate axes. For example, refer to... Figure 9 and Figure 10 The pipeline is described in detail. To encode all child nodes of a node, each node is processed 4 times, and each time one of the node's 4 pipelines is traversed.

[0102] Preferably, the entropy encoding of the occupancy information of the child nodes of each node into the bitstream, or the entropy decoding of the occupancy information of the child nodes of each node from the bitstream, is based on a lexicographical sort of nodes at depth d and a fixed sort of the child nodes of a given node. Therefore, occupied nodes are processed using a raster scan sort, but the occupancy bits of all sub-volumes of a node are encoded consecutively in a fixed order. Compared to existing techniques that encode nodes using Merton sort, the use of lexicographical sorting for scanning results in more reliable adjacency information and better encoding performance.

[0103] Preferably, storing the occupied child nodes at depth d+1 is performed by processing each of all obtained nodes at once. Preferably, the child nodes are sorted lexicographically using intermediate storage. More preferably, raster scan sorting is performed. For example, an intermediate list can be used to sort the child nodes within a node in a single traversal, thus avoiding the need to process each node four times to encode all child nodes. The complexity of encoding is reduced without affecting the performance of encoding.

[0104] Preferably, the entropy encoding or entropy decoding of each child node applies a dynamic, continuously updated optimal binary encoder (OBUF). Therefore, it is more efficient when a sufficient number of occupied bits have already been encoded, due to the additional beneficial effects of the dynamic OBUF.

[0105] Another aspect of this application provides an encoder for encoding 3D point clouds into a bitstream. The encoder includes a memory and a processor, wherein instructions are stored in the memory and, when executed by the processor, perform the steps of the aforementioned encoding method.

[0106] Another aspect of this application provides a decoder for decoding 3D point clouds from a bitstream. The decoder includes a memory and a processor, wherein instructions are stored in the memory and, when executed by the processor, perform the steps of the aforementioned decoding method.

[0107] Another aspect of this application provides a bitstream that is encoded by the steps of the aforementioned encoding method.

[0108] Another aspect of this application provides a computer-readable storage medium including instructions for performing the steps of the method described above for encoding a 3D point cloud into a bit stream.

[0109] Another aspect of this application provides a computer-readable storage medium including instructions for performing the steps of the method described above for decoding a 3D point cloud from a bitstream. Attached Figure Description

[0110] The present application is described in more detail below with reference to the accompanying drawings.

[0111] The attached diagram shows:

[0112] Figure 1 The Morton sort of child nodes and occupied bits is shown.

[0113] Figure 2 This shows the neighborhood formed by nodes at the same depth as the current node or at a greater depth.

[0114] Figure 3An example of an occupied node in the neighborhood is shown, involving a neighboring node with a depth greater than the current depth.

[0115] Figure 4 The diagram shows the occupancy map (white cube) known from the previously encoded occupancy bits when the occupancy bits of the child nodes (striped cubes) are encoded and decoded using Morton sort.

[0116] Figure 5 A raster scan of a voxel is shown.

[0117] Figure 6 The tree-based dynamic shrinkage function DR is shown. n .

[0118] Figure 7 This illustrates the tree-based dynamic reduction function DR. n Evolved into DR n+1 .

[0119] Figure 8 The diagram shows the neighboring child nodes of the current node (solid gray) when sorted using Morton's scan, where (a) is always available, (b) is never available, and (c) is sometimes available.

[0120] Figure 9 The raster scanning process of the parent node is shown to encode the occupancy of its child nodes using raster scanning.

[0121] Figure 10 The encoded order of slices 0 and 1 of the occupied nodes with the same fixed x at the current depth is shown.

[0122] Figure 11 An embodiment of a decoder employing the proposed method is shown.

[0123] Figure 12 An embodiment of an encoder employing the proposed method is shown.

[0124] Figure 13 The diagram shows the occupancy map (white cube) known from the previously encoded occupancy bits when the occupancy bits of the child nodes (striped cubes) are encoded using raster scan sorting.

[0125] Figure 14 An example of adjacent sub-occupancy patterns is shown, including 3D (left) and 2D layers (right).

[0126] Figure 15 An example of the neighboring node occupancy pattern is shown when the coder occupies bits (strip volume).

[0127] Figure 16The parent node of the raster scan sort is shown, along with the raster scan sort of the child nodes, which has 8 encoded child occupancy bits.

[0128] Figure 17 An example is shown (decoding method of parent nodes sorted by raster scan, where 8 child bits are encoded together).

[0129] Figure 18 An example of an encoder for the second embodiment (an encoding method employing raster scan sorting) is shown.

[0130] Figure 19 The diagram shows the occupancy map (white cube) known from the previously encoded occupancy bits when the occupancy bits of the child node (striped cube) are encoded and decoded using the raster scan sort of the node, but all occupancy bits of the child node are encoded using the Morton sort.

[0131] Figure 20 An encoder block diagram of one embodiment is shown, wherein child nodes are sorted into raster scan order by processing nodes in a single operation.

[0132] Figure 21 A decoder block diagram of one embodiment is shown, in which child nodes are sorted into a raster scan sort by processing nodes in one go.

[0133] Figure 22 This shows the sorting of node N by raster scan. k The occupied child node is placed in L d+1 The process.

[0134] Figure 23 The buffer L of depth d is shown. d+1 and buffer L d+1 The composition of the buffer L0.

[0135] Figure 24 A schematic flowchart of the method used for encoding is shown.

[0136] Figure 25 A schematic flowchart of the method used for decoding is shown.

[0137] Figure 26 An encoder according to this application is shown.

[0138] Figure 27 A decoder according to this application is shown. Detailed Implementation

[0139] In the occupancy tree of G-PCC, each node of the tree represents an occupied cube volume of a certain size / volume, which can be determined based on its depth in the tree. Occupancy in the tree is encoded depth-wise, starting from the root node representing the full volume of the point cloud. In other words, a breadth-first traversal of the tree is used to encode and decode the occupancy tree. For simplicity, our description is limited to an octree representation (without direct encoding mode enabled), where each of the three dimensions is doubled in precision for each level of depth in the octree. At a given depth (d) in the tree, the coordinates of all nodes in the tree are processed according to the Morton order. For each node, the occupancy of the 8 sub-volumes at the next depth level is encoded, representing 2×2×2 sub-volumes, and is indicated by providing child nodes in the next depth of the tree for the occupied sub-volumes.

[0140] When using a quadtree partition at a given depth level (d), the occupancy of the sub-volume at the next depth becomes a 2×2×1, 2×1×2, or 1×2×2 sub-volume (i.e., 2D); and the processing order of the child nodes at the next depth (d+1) then becomes an encoded order for their parent node (i.e., the node at depth d), which is refined in the 2D space by its Merton sort (i.e., Z-scan sort along the two partition directions). When using a binary tree partition, the occupancy becomes a 2x1x1, 1x2x1, or 1x1x2 sub-volume (i.e., 1D); and the processing order of the child nodes at the next depth then becomes an encoded order for their parent node refined in the 1D space by its Merton sort (i.e., lexicographical sort along the partition directions).

[0141] Essentially, to process nodes following Merton's order, a queue structure for nodes can be used. The process begins with a queue containing a single node (corresponding to the root node of the complete point cloud volume). Then, while the queue is not empty, nodes at the beginning (front) of the queue are removed, their sub-volume occupancy is encoded, and if these sub-volumes are not leaf nodes, their child nodes are appended to the end (back) of the queue, following their respective Merton's order. Thus, the queue structure acts as a First-In-First-Out (FIFO) buffer: when processing the next depth, the nodes pushed in first are encoded first.

[0142] Figure 1 The diagram shows eight possible child nodes of a parent node sorted by Morton's algorithm. Occupied bits b0 to b7 indicate the presence or absence of a point in each sub-volume within the corresponding sub-volume. These bits are encoded when processing the parent node, and the node corresponding to the occupied child node is appended to the processing queue.

[0143] In the octree representation of point clouds, processing nodes in a depth-first order, as in MPEG (Moving Picture Experts Group) GPCC (Geometry-based Point Cloud Compression), may benefit from knowing the occupancy of deeper nodes, such as... Figure 2 As shown.

[0144] Nodes with a depth greater than the current node are used to obtain geometric information in the region that has not yet been encoded at the current depth (where the y-value is greater than the y-value of the current node).

[0145] In MPEG G-PCC, the combination of the current depth and the neighboring nodes at the current depth plus 1 is used to define a neighborhood. However, to limit the possible number of neighborhood configurations, the neighborhood has been restricted to a subset of the set of nodes adjacent to the current node.

[0146] Figure 3 An example of a neighborhood is shown, in which only the occupied nodes are drawn, and also the neighboring nodes with a depth greater than the current depth are involved.

[0147] When using breadth-first search sorting, it can be understood that when encoding sub-occupancy, the occupancy of all adjacent nodes at the same depth as the current node is always known, because these nodes were encoded at previous depths.

[0148] Figure 4 This diagram illustrates the occupied volume (composed of white cubes) at a known depth when encoding the occupied bits of the current child node (striped cube) while nodes are processed using Morton sort and occupied bits are also encoded using Morton sort. Based on this diagram, the nature of the different sub-volumes that can be mapped within the occupied map can also be understood.

[0149] Raster scan sorting is typically used to encode the attributes of dense content (e.g., grayscale levels or 3 color channels) that does not require geometry: geometry covering the entire surface of the image (in the case of 2D) or the entire volume (in the case of 3D, biomedical images, or video content).

[0150] In some embodiments, a mask is provided. The mask defines a region of interest, restricting attribute encoding to that region of interest. The encoding of the mask can be assimilated into the encoding of geometric occupancy in the point cloud. Preferably, the mask is simply encoded by sequentially encoding bits for each possible voxel location in the volume under consideration using raster scan ordering and context entropy encoding, each bit indicating whether the corresponding voxel is occupied (and therefore whether an attribute is encoded for that voxel location).

[0151] In 2D, raster scan sorting of pixels is simply a lexicographical ordering of pixels based on their 2D coordinates x and y. Typically, lexicographical sorting in (y, x) is used on images: pixels are processed first by x and then by y; pixels are sorted by incrementing the x-coordinate for a given y-coordinate and then by incrementing the y-coordinate. 3D raster scan sorting is similar, but in three-dimensional space (i.e., 3D raster scan sorting is a lexicographical ordering over the arrangement of 3D coordinates). Voxel raster scan sorting follows only a lexicographical ordering, but lexicographical ordering in three dimensions x, y, and z, such as (z, y, x), can be used for volumetric images (or t, y, x for time series images, such as in video).

[0152] like Figure 5 As shown, in the voxel representation of the geometry of a point cloud, voxels (= the 3D volume associated with a single point) can be scanned in raster scan order, such as... Figure 5 As shown.

[0153] exist Figure 5 In the diagram, the sorting is shown first in y, then in x, and finally in z, following a raster scan. This is a lexicographical sort in xyz. And... Figure 5 The diagram illustrates the causal neighborhood (composed of white cubes of already encoded nodes / voxels) of the current node / voxel to be encoded. The occupancy of already encoded neighboring nodes is known and can be used to predict and encode the occupancy of the current node / voxel. Encoded nodes / voxels are those whose lexicographical order is lower than that of the current node / voxel.

[0154] The proposed method encodes / decodes sub-volume occupancy bits based on the raster scan sorting of nodes at a given depth in the occupancy tree. In one embodiment of the proposed method, the raster scan sorting of sub-volumes is used to encode / decode their corresponding occupancy bits; in another embodiment, the raster scan sorting of nodes is used to encode / decode (together) the occupancy bits of all their corresponding sub-volumes.

[0155] In one embodiment, the proposed method is Figure 9 The diagram illustrates this. For better understanding, the entire occupied volume is shown in the figure to better illustrate the raster scan order of the processing. However, it should be clarified that if an ancestor node has been marked as unoccupied, a portion of this volume will not be encoded / processed.

[0156] like Figure 9As shown, a lexicographical sort in (x, y, z) (i.e., a cooperative lexicographical sort in (z, y, x)) is used. For a given depth in the octagon, the occupancy of child nodes is encoded according to the raster scan sort of the child nodes. To implement the raster scan sort of child nodes, four tubes (tube 0, tube 1, tube 2, tube 3) are defined for each node at the current depth, each tube being made up of each node with the same x and y. The occupancy bits of the child nodes of tube 0 are encoded by encoding the occupancy bits b0 and b1 of all occupancy nodes with the same fixed x and y, and the nodes are sorted by incrementing the z of all occupancy nodes. The occupancy of tube 1 is then encoded by encoding the occupancy bits b2 and b3 of the same nodes and in the same sort as tube 0.

[0157] The encoding process for encoding occupied pipe 0 and occupied pipe 1 of all occupied nodes with the same fixed x and y is repeated for each y in ascending order. This provides raster scan encoding of slice 0, consisting of each bit b0, b1, b2, b3 of all occupied nodes with the same x. A diagram illustrating the raster scan order of these nodes in slice 0 is shown in [the diagram]. Figure 10 As shown in (a), where each child node is represented in 2D for better illustration. After encoding slice 0 of the node with a fixed x, slice 1 of these nodes is then encoded with the same x, and the encoding process of slice 0 is reproduced for slice 1 by encoding consecutively in the same order. Occupying pipes 2 and then 3 consist of occupied bits b4, b5, then b6, b7 of the same node. And the raster scan order of these nodes' slice 1 illustrates the encoding process in... Figure 10 As shown in (b), where each child node is represented in 2D for better illustration. The encoding process of slice 0 followed by slice 1 is repeated for every x nodes at the current depth of all occupied nodes in ascending order.

[0158] An embodiment of the decoder employing the proposed method is in Figure 11 The diagram is shown in the middle, where list L is... d This represents the number of nodes N in an octree of depth d. k (x k y k , z k A list of nodes sorted by raster scan from N0 to N1. last Sort the list L, and k is a list L. d The index of the node in the [database / database]. Figure 12 The following two are shown:

[0159] • The process of decoding the occupancy of child nodes (sIdx, tIdx, i) at depth d+1 from a bitstream that follows raster scan order;

[0160] And putting child nodes into list L d+1 The process of generating ordered nodes by raster scanning at the next depth.

[0161] The variable sIdx represents the index of the slice of all occupied nodes with the same x; the variable tIdx represents the index of the pipe of all occupied nodes with the same x and y; and the variable i represents the index of the child node in each pipe.

[0162] In other words, the integer value obtained from the binary word formed by (sIdx, tIdx, i) is the sub-index of the node (a decimal value from 0 to 7).

[0163] The variable kSStart stores the sorted index of the first node of the slice with all occupied nodes having the same x; and the variable kTStart stores the sorted index of the first node of the pipe with all occupied nodes having the same x and y.

[0164] The depth dMax is the maximum depth in the tree and can be determined from the bitstream.

[0165] The detailed process in the decoder can be described as follows:

[0166] • To initialize the algorithm, list L0 is set to contain only the root node of the tree, and

[0167] Set the initial depth d to 0;

[0168] • Tag 0:

[0169] • Obtain a list L of nodes at storage depth d (First In First Out, FIFO) sorted by raster scan. d ;

[0170] • Obtain an empty list L d+1 To generate child nodes for decoding the raster scan sorting of the next depth;

[0171] • Set the node index k to 0 (initialize the variable k with the value 0);

[0172] • Tag 1:

[0173] • Set the slice index sIdx to 0 (initialize the variable sIdx with the value 0);

[0174] • Set the index of the first node in the slice to k (initialize the variable kSStart with the value k);

[0175] • Tag 2:

[0176] • Set the pipe index tIdx to 0 (initialize the variable tIdx with the value 0);

[0177] • Set the index of the first node in the pipeline to k (initialize the variable kTStart with the value of k);

[0178] • Tag 3:

[0179] • From list L d Obtain node N k (x k y k , z k );

[0180] • Set the sub-index in the pipe to 0 (initialize variable i with the value 0);

[0181] Process 0: Decode and generate L d Middle node N k The ordered child nodes of the child nodes in the pipeline tIdx of the slice sIdx;

[0182] • The number of bits occupied by the i-th child stage of the pipe tIdx of the bitstream decoded slice sIdx (which is N) k The child nodes (sIdx, tIdx, i));

[0183] • If the decoded bit indicates that the child node is occupied, then in list L d+1 At the end of the pipeline tIdx, append the new node of the i-th child node of the slice sIdx (which is N). k The child node (sIdx, tIdx, i)). This new node has coordinates (2×x). k +sIdx, 2×y k +tIdx,2×z k +i);

[0184] • If i is not equal to 1, increment i and repeat process 0; otherwise, proceed to the next step.

[0185] · Determine N k Is it the last node of depth d (i.e., in list L)? d (the last node):

[0186] If N k If it is not the last node of depth d, then start from list L. d Get the next occupied node N k+1 (x k+1 y k+1 , z k+1Then determine if it is x. k ≠x k+1 Or y k ≠y k+1 ;

[0187] If x k ≠x k+1 or y k ≠y k+1 This means that the next node does not belong to the same pipeline as an occupied node with the same x and y values. (Then determine if...)

[0188] tIdx==1;

[0189] If tIdx == 1, it means that the processing of both pipelines (in the pipelines of all occupying nodes with the same x and y) of the child nodes of the current slice of the child node has been completed. Then determine whether x k ≠x k+1 ;

[0190] If x k ≠x k+1 If so, it means the next node does not belong to the same slice as the occupied node with the same x. Then check if sIdx == 1;

[0191] ● If sIdx == 1, it means that the processing of the two slices of the child node...

[0192] (In the slice of all occupied nodes with the same x) it has been completed.

[0193] Therefore, N k+1 It is the starting point of the next slice of the occupied node and the next pipeline of the occupied node to be processed;

[0194] ●Increment k by 1 and loop through Label1 to begin processing the first child node pipeline of the first child node slice of the next occupied node slice;

[0195] ● Otherwise, sIdx == 0. This means we can begin processing the second slice of the occupied node;

[0196] ●Tag 4:

[0197] ●Increment the slice index sIdx by 1, and reset k to the index of the first node of all slices with the same x:

[0198] k = kSStart;

[0199] • Loop back to label 2 to begin processing the third pipeline of child nodes;

[0200] ●Otherwise x k ==xk+1 This means the next node belongs to the same slice as the occupied node, but to the next pipe after the occupied node. Therefore,

[0201] N k+1 It is the starting point of the next pipeline to be processed;

[0202] ●Increment k by 1 and loop to Label2 to begin processing the first pipe of the child node of the next pipe of the occupied node (the current slice sIdx of the child node);

[0203] ● Otherwise, tIdx == 0. This means we can begin processing the second pipeline of the child nodes of the current slice of the child node;

[0204] ●Tag 5:

[0205] ●Increment the pipe index tIdx by 1 and reset k to the index of the first node of all pipes with the same x and y: k = kTStart;

[0206] ● Loop back to label 3 to begin processing the second pipeline of child nodes;

[0207] Otherwise x k ==x k+1 And y k ==y k+1 This means the next node belongs to the same pipeline as the occupied node. Increment k by 1 and loop to Label3 to continue processing the same pipeline for child nodes;

[0208] Otherwise N k ==N last This node is the last node at depth d. Then determine if...

[0209] tIdx==1;

[0210] If tIdx == 1, it means that the processing of the last pipeline of the child nodes in the current slice of the child node has been completed. And the next slice of the child node (if there is one) can be processed.

[0211] Then determine if tIdx == 1;

[0212] If sIdx == 1, it means that the processing of the last slice of the child nodes of the occupied node has been completed. And the next depth in the occupied tree can be processed (if there is more). Then check if d == dMax;

[0213] ●If d==dMax, it means that the final depth of the occupied tree has been decoded.

[0214] Then complete the tree occupation;

[0215] ●Otherwise, d<dMax. This means that the next depth of the occupancy tree can be processed.

[0216] Increase the depth d by 1 and loop back to Label0 to process the next depth in the occupancy tree;

[0217] ·Otherwise, sIdx == 0. This means processing of the second slice of the child nodes of the last slice of an occupied node can start. Go to label 4 to start;

[0218] ·Otherwise, tIdx == 0. This means processing of the last pipeline of the child nodes of the current slice of a child node can start. Go to label 5 to start.

[0219] In a preferred embodiment, the list L can be omitted d+1 , and when d == dMax, child nodes are not appended to this list, since this list will not be used because there are no more depths to process.

[0220] The coordinates of a child node are the coordinates of its parent node refined by one bit of precision. In other words, some physical spatial coordinates of the node center can be obtained by multiplying the coordinates of the node center by the physical dimension of the node volume scaled according to the depth: ((x k , y k , z k )+(0.5, 0.5, 0.5))×(width, length, height) / 2 深度 .

[0221] In a preferred embodiment, when d == dMax, if a child node is occupied, a point with coordinates ((2×x k +sIdx, 2×y k +tIdx, 2×z k +i)+(0.5, 0.5, 0.5))×(width, length, height) / 2 dMax is output (e.g., in a buffer). Point geometry in raster-scan order is generated.

[0222] According to the above encoding process for child nodes, it can be observed that each node N k needs to be accessed / processed 4 times: the first access / processing is for encoding occupancy pipeline 0, the second for encoding occupancy pipeline 1, the third for encoding occupancy pipeline 2, and the last for encoding occupancy pipeline 3.

[0223] This encoding process constructs the nodes required for raster-scan ordering of occupied child nodes The sorting mentioned is obtained from the raster scan sorting of the processed nodes. And the raster scan sorting of the nodes is obtained directly from the raster scan sorting of the child node occupancy encoding of the parent node (i.e., during the previous depth of the octree occupancy encoding).

[0224] Similarly, one embodiment of the encoder employing the proposed method is in Figure 12 The diagram is shown in the middle.

[0225] The effectiveness of the provided method is Figure 13 The diagram shows the occupancy map coverage of a known sub-volume when encoding occupancy bits for a striped sub-volume, as the occupancy bits are encoded according to the raster scan order of the node's sub-volumes. Figure 13 The diagram also illustrates adjacent positions that can be used to construct the context for bit entropy encoding. Occupancy of any sub-volume with lower z-position, lower y-position, and lower x-position can be used for context analysis using raster scan sorting (regardless of the coordinate axis order in lexicographical sorting). Therefore, a predetermined pattern of adjacent sub-volumes can be defined and used to select the encoding context for any encoded bit occupancy.

[0226] In some embodiments, when each child node is encoded using entropy coding with raster scan sorting, the adjacent occupancy pattern used to construct the context for each child node can be the same. As an example of an adjacent occupancy pattern, Figure 14 An example of a sub-neighbor occupancy pattern (occupancy of white sub-volumes) is provided. This sub-neighbor occupancy pattern can be used to determine the entropy encoding context for encoding the occupancy bits of striped sub-volumes when sub-occupancy is encoded according to raster scan order. On the left, this neighborhood is represented in 3D, while on the right, it is represented in 2D layers. In one example, each number on the right can be used to sort the occupancy bits from the most important / valid one (position 0) to the least important / valid one (position 15) for use with the previously described dynamic OBUF. This will produce 65,536 entropy contexts. In some embodiments that do not use dynamic OBUF, the number of neighboring nodes can be limited, for example, to positions 0 to 11, to reduce the number of contexts to 4,096.

[0227] Figure 15 This illustrates the adjacency pattern formed by the node's six directly adjacent nodes (nodes sharing a face) when encoding different occupancy bits of the current node's subvolume. As in the prior art, using Morton sort, the occupancy of adjacent parent nodes can provide useful information for improving compression efficiency.

[0228] Even for raster scan encoding of sub-volumes, the occupancy neighborhood of the same sub-occupancy pattern can be used to encode any one of the sub-volume occupancy bits. As can be understood, the information provided by the occupancy bits of the parent node's six neighboring volumes / nodes will not have the same importance in selecting a suitable entropy encoding context for encoding the sub-volume, depending on the sub-volume's location. Therefore, if neighboring node occupancy is used for selecting the entropy encoding context, the context selection can also depend on the sub-occupancy index. This also applies to modeling the entropy encoding context using other neighboring nodes consisting of any neighboring nodes / volumes that share a face, edge, or wedge with the current node, e.g., 36 neighboring nodes.

[0229] In some embodiments, the proposed method uses raster scan sorting of nodes to (together) encode / decode the occupancy bits of all their respective sub-volumes at a given depth in the occupancy tree. Preferably, as Figure 16 As shown, occupied nodes are processed using raster scan sorting, but the occupied bits of all sub-volumes of a node are sequentially encoded on the node by sorting in a single channel, for example, using Morton scan sorting of child node occupied bits (meaning that the sub-volumes of each node are encoded using Morton sorting in a single channel on the node); and in order to sort child nodes by raster scan sorting for the next depth, the same child node sorting method as in the embodiment is used, which uses 4 channels on each node at the current depth and Figure 7 As shown in the image.

[0230] Figure 17 Examples of child nodes of the decoder occupying the decoding process are shown in some embodiments. Figure 18 Examples of encoders in some embodiments are shown, and Figure 17 and Figure 18 The sorting process of child nodes is not shown in the text.

[0231] The detailed process in the encoder / decoder can be described as follows:

[0232] • Obtain the FIFO list L of the node at storage depth d by raster scanning and sorting. d ,

[0233] • From the FIFO list L d Node N is obtained at depth d. k (x k, y k , z k) And initially at each depth, from the first node N0(x) 0, y 0, z 0) start,

[0234] • Encode the bits occupied by child nodes into a bitstream according to a certain order (such as Morton sort), or decode from a bitstream.

[0235] If N k Not a FIFO list L d The last node in the list L is then used for the FIFO list. d The next node (N) k+1 (x k+1 y k+1 , z k+1 Encode / decode until L is reached. d The last node in,

[0236] • Encode the child node occupancy of the next depth until the final depth is reached.

[0237] Figure 19 It shows Figure 16 The described alternative encoding sorting is used to encode the occupancy bits of striped sub-volumes based on a known child node occupancy graph. Utilizing a raster scan sort of nodes (regardless of the axis sorting used in a lexicographical sort), the occupancy of any sub-volume of a node with a lower z, lower y, or lower x position than the current node (the parent node of the striped sub-volume) can be used to contextualize the entropy encoding of the striped sub-volume occupancy bits, as well as any sub-volume occupancy bits previously encoded for the current node. Using this method, it is clear that there is a distinct pattern among the available neighboring nodes for each position of the encoded sub-volume occupancy. Therefore, in the case of occupancy in an occupancy tree, there are 8 distinct patterns. Consequently, the entropy encoding of the occupancy bits needs to be tuned according to each pattern (i.e., for each occupancy bit index) for better optimization.

[0238] Compared to Figure 11 The embodiment shown complicates the encoded bits, but is still superior to patterns in octrees with nodes encoded in Merton ordering in the prior art. In the prior art, for Merton ordering of nodes, the patterns occupied by "unknown" (i.e., not yet encoded) neighboring nodes vary more with their position in the ancestor nodes (and therefore with the Merton index), so some neighbor information is not always reliable when constructing the entropy context, and it may introduce bias into the entropy coding probabilities constructed in the adaptive entropy coding context, thereby degrading coding performance.

[0239] [0116. It should be understood that in point cloud encoding, occupied nodes and child nodes will rarely (and preferably never) be associated with the example.] Figure 4 , Figure 9 , Figure 13 and Figure 16The graphs are densely represented to provide a better view of available causal occupancy information and the order of node / child node processing. In a typical point cloud, several volumes will be unoccupied, therefore, the nodes for these volumes will not exist and will be processed. The processing order for occupied nodes is the same as in the case of dense content, and for a given occupied node, the occupancy of all child volumes is always encoded, but unoccupied nodes / volumes are not processed at all because they do not exist in the tree structure. The processing of these unoccupied nodes / volumes is skipped compared to the processing order for dense content.

[0240] For example, in Figure 16 In the illustrated embodiment, the raster scan sort of the child nodes at depth d+1 is obtained by processing the parent node at depth d four times, which increases the runtime complexity.

[0241] To reduce runtime complexity, preferably, the method sorts child nodes using raster scan sorting instead of traversing each node four times. In some preferred embodiments, such as Figure 20 As shown. In the raster scan sorting, N with depth d are... k The occupied child node is placed into L d+1 In this step, an intermediate list is used to sort the child nodes on the node in one go. The occupancy pipeline of child nodes 0 to 3 is maintained in 4 separate lists, namely L t0 L t1 L t2 and L t3 These lists will be used to construct the occupied slices for child nodes 0 and 1, which are maintained in two separate lists, namely L s0 and L s1 Finally, list L d+1 From list L s0 and L s1 Build.

[0242] In detail, for N of depth d (k starts from 0) k (x k y k , z k In this step:

[0243] Add the new node of each occupied child node of pipe 0 to list L (appended to the end). t0 middle,

[0244] Add the new node of each occupied child node of pipe 1 to list L. t1 middle,

[0245] Add the new node of each occupied child node of pipe 2 to list L. t2 middle,

[0246] Add the new node of each occupied child node of pipe 3 to list L. t3 middle,

[0247] The occupied child nodes of each pipeline are processed from bottom to top.

[0248] After iterating over all child nodes of the pipeline at depth d with the same x and y (parent) node, the pipeline list can be appended to the slice list: L t0 Attached to L s0 At the end, L t1 Attached to L s0 The end; L t2 Attached to L s1 At the end, then L t3 Attached to L s1 At the end. This provides the value for each slice list L. s0 and L s1 The child nodes are sorted using raster scan sorting. Then, an empty pipeline list is used to process the next pipeline of the (parent) node.

[0249] After iterating over all child nodes of a slice at depth d with the same x-axis parent node, the slice list can be appended to the next depth list L. d+1 :L s0 Attached to L d+1 At the end, L s1 Attached to L d+1 The end. An empty slice list is used to process the next slice of the (parent) node. This provides the next slice at depth L. d+1 The list of child nodes sorted by raster scan. Figure 22 The diagram shows the sorting of node N by raster scan. k The occupied child node is placed into L d+1 The detailed process.

[0250] If lists are implemented using memory blocks, appending one list to another can be time-consuming: it requires copying memory and may require memory allocation. To avoid this problem, in a preferred embodiment, a (dynamically) linked list is used to implement list L. t0 L t1 L t2 L t3 L s0 L s1 L d and L d+1 Therefore, appending one list to another can have a time complexity of O(1). These intermediate buffers allow for one-time processing of nodes.

[0251] exist Figure 23 (a) and (b) show the obtained buffer L for storing child nodes in raster scan order for the next depth. d+1 , Figure 23 (a) shows the buffer L d+1 Contains a buffer L with the same x node x (from L) x0 To L x(2^d)-1 ), Figure 23 (a) shows the buffer L x0 The structure is composed of L t0 L t1 L t2 L t3 Composition. And buffer L d+1 Organized as:

[0252] ·Buffer L t0 and L t1 :L t0 Pipeline 0 for occupied nodes with the same x and y coordinates

[0253] (exist Figure 23 (b) marked), and L t1 Pipeline 1 (for occupied nodes with the same x and y) Figure 23 (b) marked), and sorted by increasing y, they are linked

[0254] Continuously appended to the buffer L of slice 0 with the same x-occupying node s0 ;

[0255] ·Buffer L t2 and L t3 :L t2 Pipeline 2 for occupied nodes with the same x and y coordinates

[0256] (exist Figure 23 (b) marked), and L t3 Pipeline 3 (for occupied nodes with the same x and y) Figure 23 (b) marked), and sorted by increasing y, they are linked

[0257] Continuously appended to the buffer L of slice 1 with the same x-axis. s1 ;

[0258] • Buffer L of occupied slice 0 of occupied nodes with the same x s0 The buffer L is attached to the occupied slice 1 of the occupied node with the same x. s1 To form a buffer L x, its in Figure 23 (a) is represented as, for example, L x0 (for x = x1) and L x1 (For x = x0);

[0259] ·Buffer L x Appended to buffer L x-1 At the end. At depth d, buffer L is appended one after another by increasing x. x To obtain buffer L d+1 The child nodes in the raster scan sorting (are built for the next depth occupancy encoding).

[0260] Now for reference Figure 24 It shows a schematic flowchart of a method for encoding 3D point clouds into a bitstream according to this application.

[0261] The method includes:

[0262] Step S10: Along each axis X, Y, Z of the coordinate system, obtain nodes with depth d that contain at least a portion of the point cloud in a dictionary order;

[0263] Step S11: Encode the occupancy information entropy of the child nodes of each node into a bit stream based on the dictionary sorting.

[0264] Now for reference Figure 25 The diagram illustrates a schematic flowchart of a method for decoding 3D point clouds from a bitstream according to this application.

[0265] The method includes:

[0266] Step S20: Along each axis X, Y, Z of the coordinate system, obtain nodes with depth d that contain at least a portion of the point cloud in a dictionary order;

[0267] Step S21: Based on the dictionary sorting, entropy decode the occupancy information of the child nodes of each node from the bit stream.

[0268] Now for reference Figure 26The diagram illustrates a simplified block diagram of an example embodiment of encoder 1100. Encoder 1100 includes a processor 1102 and a storage device 1104. Storage device 1104 may store a computer program or application containing instructions that, when executed, cause processor 1102 to perform certain operations such as those described herein. For example, instructions may be encoded and output as a bitstream encoded according to the methods described herein. It should be understood that instructions may be stored on a non-transitory computer-readable medium, such as an optical disk, flash memory, random access memory, hard disk, etc. When instructions are executed, processor 1102 performs the operations and functions specified in the instructions to operate as a dedicated processor implementing the described processes. In some examples, such a processor may be referred to as a "processor circuit" or a "processor circuit system".

[0269] refer to Figure 27 The diagram illustrates a simplified block diagram of an example embodiment of decoder 1200. Decoder 1200 includes a processor 1202 and a storage device 1204. Storage device 1204 may include a computer program or application containing instructions that, when executed, cause processor 1202 to perform certain operations such as those described herein. It should be understood that the instructions may be stored on a computer-readable medium, such as an optical disc, flash memory, random access memory, hard disk drive, etc. When the instructions are executed, processor 1202 performs the operations and functions specified in the instructions to operate as a dedicated processor implementing the described processes and methods. In some examples, such a processor may be referred to as a "processor circuit" or a "processor circuit system".

[0270] It should be understood that the decoder and / or encoder according to this application can be implemented in multiple computing devices, including but not limited to servers, appropriately programmed general-purpose computers, machine vision systems, and mobile devices. The decoder or encoder can be implemented by software containing instructions for configuring one or more processors to perform the functions described herein. The software instructions can be stored on any suitable non-transitory computer-readable storage medium, including CD, RAM, ROM, flash memory, etc.

[0271] It should be understood that the decoders and / or encoders described herein, as well as the modules, routines, procedures, threads, or other software components implementing the methods / processes for configuring the encoder or decoder, can be implemented using standard computer programming techniques and languages. This application is not limited to specific processors, computer languages, computer programming conventions, data structures, or other such implementation details. Those skilled in the art will recognize that the described processes can be implemented as part of computer-executable code stored in volatile or non-volatile memory, as part of an application-specific integrated chip (ASIC), etc.

[0272] This application also provides computer-readable signals encoded by applying the encoding process according to this application.

[0273] Certain adjustments and modifications can be made to the described embodiments. Therefore, the above embodiments are considered illustrative rather than restrictive. In particular, the embodiments can be freely combined with each other.

Claims

1. A method for decoding a three-dimensional 3D point cloud from a bitstream, the method being applied to a decoder, the geometry of the point cloud being defined by an octree structure comprising a plurality of nodes having parent-child relationships and representing the three-dimensional position of objects, the point cloud being located in a volume space of a three-dimensional coordinate system, the volume space being progressively divided into sub-volumes and containing points of the point cloud, wherein the volume is divided into several sub-volumes, each sub-volume being associated with one of the plurality of nodes of the octree structure, wherein occupancy information associated with each sub-volume indicates whether the sub-volume contains at least one point of the point cloud, the method comprising: Along each axis X, Y, Z of the coordinate system, nodes of depth d containing at least a portion of the point cloud are obtained in a lexicographical order. Store occupied child nodes at a depth of d+1 in dictionary order; Based on the dictionary sorting, the occupancy information of each node's child nodes is entropy decoded from the bitstream.

2. The method according to claim 1, wherein, Lexicographic sorting of nodes at depth d and / or lexicographic sorting of child nodes at depth d+1 is raster scan sorting.

3. The method according to claim 2, wherein, Lexicographic sorting based on child nodes at depth d+1 performs entropy decoding on the occupancy information of each node's child nodes from the bitstream.

4. The method according to any one of claims 1-3, further comprising entropy decoding of the occupancy information of the child nodes of each node from the bitstream based on the same adjacent child node occupancy pattern indicating the context of each child node.

5. The method according to claim 4, wherein, The context of each child node is selected based on the occupancy pattern of the child node and / or the occupancy pattern of the adjacent parent node.

6. The method according to claim 5, wherein if the occupancy pattern of the adjacent parent node is used for the selection of the entropy coding context, the context selection depends on the child occupancy index within the parent node, wherein, The child occupancy index indicates the position of the child node within the parent node.

7. The method according to any one of claims 2-3, wherein, By traversing each node four times, the occupied child nodes at a depth of d+1 are stored in lexicographical order.

8. The method according to any one of claims 1-2, wherein, Entropy decoding of the occupancy information of each node's child nodes from the bitstream is performed based on a lexicographical sort of nodes at depth d and a fixed sort of the child nodes of a given node.

9. The method according to claim 8, wherein, Storing occupied child nodes at depth d+1 is performed by processing each of all the nodes obtained at once.

10. The method according to claim 9, wherein the occupied child nodes at depth d+1 are sorted lexicographically using intermediate storage.

11. The method according to claim 10, wherein the occupied child nodes at depth d+1 are sorted by raster scan.

12. The method according to any one of claims 1-3, wherein, The entropy decoding of each child node applies a dynamic, up-to-date, optimal binary encoder, OBUF.

13. A method for encoding a three-dimensional 3D point cloud into a bitstream, the method being applied to an encoder, the geometry of the point cloud being defined by an octree structure comprising a plurality of nodes having parent-child relationships and representing the three-dimensional position of an object, the point cloud being located in a volume space of a three-dimensional coordinate system, the volume space being progressively divided into sub-volumes and containing points of the point cloud, wherein the volume is divided into several sub-volumes, each sub-volume being associated with one of the plurality of nodes of the octree structure, wherein occupancy information associated with each sub-volume indicates whether the sub-volume contains at least one point of the points of the point cloud, the method comprising: Along each axis X, Y, Z of the coordinate system, nodes of depth d containing at least a portion of the point cloud are obtained in a lexicographical order. Store occupied child nodes at a depth of d+1 in dictionary order; Based on the dictionary sorting, the occupancy information of each node's child nodes is entropy encoded into the bit stream.

14. The method according to claim 13, wherein, Lexicographic sorting of nodes at depth d and / or lexicographic sorting of child nodes at depth d+1 is raster scan sorting.

15. The method according to claim 14, wherein, Lexicographic sorting of child nodes at depth d+1 encodes the occupancy information of each node's child nodes into a bit stream.

16. The method according to any one of claims 13-15, further comprising entropy encoding the occupancy information of the child nodes of each node into a bit stream based on the same adjacent child node occupancy pattern indicating the context of each child node.

17. The method according to claim 16, wherein, The context of each child node is selected based on the occupancy pattern of the child node and / or the occupancy pattern of the adjacent parent node.

18. The method of claim 17, wherein if the occupancy pattern of the adjacent parent node is used for the selection of the entropy coding context, the context selection depends on the child occupancy index within the parent node, wherein, The child occupancy index indicates the position of the child node within the parent node.

19. The method according to any one of claims 13-15, wherein, By traversing each node four times, the occupied child nodes at a depth of d+1 are stored in lexicographical order.

20. The method according to any one of claims 13-14, wherein, The occupancy information of each node's child nodes is encoded into a bit stream based on a lexicographical sort of nodes at depth d and a fixed sort of the child nodes of a given node.

21. The method according to claim 20, wherein, Storing occupied child nodes at depth d+1 is performed by processing each of all the nodes obtained at once.

22. The method of claim 21, wherein the occupied child nodes at depth d+1 are sorted lexicographically using intermediate storage.

23. The method according to claim 22, wherein the occupied child nodes at depth d+1 are sorted by raster scan.

24. The method according to any one of claims 13-15, wherein, The entropy encoding of each child node applies a dynamic, up-to-date, optimal binary encoder, OBUF.

25. An encoder for encoding a three-dimensional 3D point cloud into a bitstream, the encoder comprising at least one processor and a memory, wherein, The memory stores instructions that, when executed by the processor, perform the steps of the method according to any one of claims 13-24.

26. A decoder for decoding a three-dimensional 3D point cloud from a bitstream, the decoder comprising at least one processor and a memory, wherein, The memory stores instructions that, when executed by the processor, perform the steps of the method according to any one of claims 1-12.

27. A computer-readable storage medium comprising instructions that, when executed by a processor, perform the steps of the method according to any one of claims 1-24.

Citation Information

Patent Citations

  • Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

    CN112041888A