Tree-based depth entropy model for point cloud compression

By introducing an octree-based depth entropy model in point cloud compression, the problem of not being able to effectively control the level of detail in the prior art is solved, and efficient point cloud compression and information retention of subsequent tasks is achieved.

CN120077646APending Publication Date: 2025-05-30INTERDIGITAL PATENT HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073777.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-18
Filing Date
2023-10-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art lacks a point cloud compression method using a deep entropy model learned based on octree, and cannot effectively control the level of detail to achieve efficient compression performance.

Method used

A method is proposed to achieve efficient compression of point clouds by decoding a tree-like structured point cloud from the bitstream, using a learning-based entropy model to predict the occupancy symbol distribution of nodes, and generating a bitstream through an adaptive entropy decoder.

Benefits of technology

Through this method, point cloud data can be stored and processed at affordable computing costs, achieving more efficient compression effects, and retaining task-specific information in subsequent machine tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077646A_ABST
    Figure CN120077646A_ABST
Patent Text Reader

Abstract

A method and apparatus for encoding and decoding a 3D point cloud. An octree-based learned depth entropy model is proposed for lossless compression of 3D point cloud data. The self-supervised compression is composed of an adaptive entropy encoder that operates based on a tree structure conditional entropy model. Information from local neighborhoods and global topologies is utilized from an octree structure. In an embodiment, features from a parent level are up-sampled to a resolution of the current level, and then further feature aggregation is performed. In order to process dense massive point clouds and facilitate parallel processing, a block-based compression scheme is proposed to reduce the required computing and time resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This principle generally relates to the field of point cloud processing. This field aims to develop tools for analyzing, interpolating, representing, and understanding point cloud signals. This document is also understood in the context of encoding, formatting, and decoding data representing point clouds for 3D rendering on end-user devices such as mobile devices or head-mounted displays (HMDs). Background Art

[0002] This section aims to introduce the reader to aspects of the prior art that may be relevant to the following description and / or the various aspects claimed of this principle. This discussion is considered to help provide background information for the reader to better understand the various aspects of this principle. Therefore, it should be understood that these statements should be understood in this light and not as an admission of prior art.

[0003] Point clouds are a data format across multiple business domains, covering autonomous driving, robotics, augmented reality / virtual reality (AR / VR), civil engineering, computer graphics, and the animation / movie industry. 3D light detection and ranging sensors (LIDAR) have been applied to autonomous vehicles, and several companies have also launched affordable LIDAR sensors. With the rapid development of sensing technology, 3D point cloud data has become more practical than ever and is expected to be the ultimate enabler for the above applications.

[0004] Point cloud data is also considered to consume a large amount of network traffic, such as in 5G network-connected vehicles and immersive communications (VR / AR). The understanding and communication of point clouds will essentially give rise to efficient representation formats. Specifically, the raw point cloud data needs to be properly organized and processed for world modeling and sensing. In addition, point clouds can represent continuous scans of the same scene, which contain multiple moving objects. Compared with static point clouds captured from static scenes or static objects, they are called dynamic point clouds. Dynamic point clouds are usually frame-based, and different frames are collected at different time points.

[0005] 3D point cloud data contains discrete samples of the surface of an object or a scene. To represent the real world completely with point samples, a large number of points are required. For example, a typical VR immersive scene contains millions of points, while point clouds usually contain hundreds of millions of points. Therefore, processing such a large-scale point cloud is computationally expensive, especially for consumer devices with limited computing power (such as smartphones, tablets, and car navigation systems). The first step for any processing or inference on point clouds is to have an efficient storage method. To store and process the input point cloud at an affordable computational cost, one solution is to downsample it first. The downsampled point cloud summarizes the geometric data of the input point cloud while containing fewer points. Then the downsampled point cloud is fed into subsequent machine tasks for further use. However, the original point cloud data (original or downsampled) can be converted into a bitstream through entropy coding techniques to achieve lossless compression, thus further reducing the storage space. A better entropy model produces a smaller bitstream, and thus achieves more efficient compression. In addition, the entropy model can also be paired with downstream tasks, enabling the entropy encoder to retain task-specific information during compression. Octree is a format for encoding any type of point cloud.

[0006] There is currently no point cloud compression (PCC) method using a deep entropy model based on octree learning that controls the level of detail by changing the depth of the tree or the quantization level of the input point cloud, thus providing efficient compression performance. Summary of the Invention

[0007] The following provides a brief overview of the present principle to provide a basic understanding of certain aspects of the present principle. This overview is not a comprehensive review of the present principle, nor is it intended to indicate the important or key elements of the present principle. The following overview only introduces certain aspects of the present principle in a simplified form as a prelude to the more detailed description below.

[0008] The present principle relates to a method for decoding a tree-structured point cloud from a bitstream. The method includes obtaining a node from the bitstream representing the compressed tree-structured point cloud and initializing the context information of the node. By using a learning-based entropy model, the occupancy symbol distribution for the node is predicted using the context information of the node and the feature information from adjacent nodes and the parent node previously obtained from the bitstream. An adaptive entropy decoder decodes the occupancy symbol for the node based on the occupancy symbol distribution. An extended tree is output based on the occupancy symbol.

[0009] In an embodiment, the entropy model uses the depth feature information of the sibling nodes of the parent node of the node to predict the occupancy symbol distribution through a learning-based module for the sibling nodes of the node. In another embodiment, the entropy model uses the depth feature information of all the nodes of the parent to predict the occupancy symbol distribution through a learning-based module for the sibling nodes of the node. In yet another embodiment, the entropy model uses the depth feature information of all the nodes of the parent to predict the occupancy symbol distribution through a separate learning-based module.

[0010] This application also relates to a device, which includes a memory associated with a processor configured to implement the above method.

[0011] The principle of this application relates to a method for encoding a tree-structured point cloud in a bitstream. The method includes obtaining a point cloud constructed as a tree. For each node in the tree, initializing the context information, predicting the occupancy symbol distribution by using a learning-based entropy model with the context information of the node and the feature information from adjacent nodes and the parent node, and encoding the occupancy symbol in the bitstream based on the occupancy symbol distribution by using an adaptive entropy encoder. Generating a combined bitstream for the tree.

[0012] This document also relates to a device, which includes a memory associated with a processor configured to implement the above method. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The present disclosure will be better understood and other specific features and advantages will be known by reading the following description with reference to the accompanying drawings, in which:

[0014] Figure 1 A 3D point cloud encoding method according to this principle is shown;

[0015] Figure 2 A decoding method according to this principle is shown;

[0016] Figure 3 An architecture of a depth entropy model according to this principle is shown;

[0017] Figure 4 An implementation of a depth entropy encoding module in an encoder according to a first embodiment of this principle is shown;

[0018] Figure 5 An implementation of a depth entropy encoding module according to a second embodiment of this principle is shown;

[0019] Figure 6 An implementation of a depth entropy encoding module according to a fourth embodiment of this principle is shown;

[0020] Figure 7A 、 Figure 7B and Figure 7CShows a sixth embodiment according to the present principle;

[0021] Figure 8 Schematically shows a transformer block used in a fourth variant of a sixth embodiment according to the present principle;

[0022] Figure 9 Shows a deep entropy encoding / decoding method in which features from a parent are upsampled to match the resolution of the current octree level;

[0023] Figure 10 and Figure 11 Shows a scenario according to the present principle in which an original point cloud is converted into blocks through a shallow octree;

[0024] Figure 12 Shows an example architecture of a device configured to implement the encoding and / or decoding method described with respect to Figures 1 to 11 ;

[0025] Figure 13 Shows an example of a syntax embodiment of a stream when transmitting data through a packet-based transport protocol. Detailed Description

[0026] The present principle will be described more fully hereinafter with reference to the accompanying drawings, in which examples of the present principle are shown. However, the present principle may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Thus, while the present principle is susceptible to various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and will be described in detail herein. However, it should be understood that the present principle is not intended to be limited to the particular forms disclosed, but on the contrary, the disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope defined by the claims of the present principle.

[0027] The terms used in this document are only for describing specific examples and are not intended to limit the present principle. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" used in this document also include the plural forms. It should also be understood that the terms "comprises", "comprising", "includes", and / or "including" used in this specification refer to the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their combinations. In addition, when an element is referred to as being "responsive" to or "connected" to another element, it can be directly responsive to or connected to the other element, or there can be intermediate elements. In contrast, when an element is referred to as being "directly responsive" to or "directly connected" to another element, there are no intermediate elements. The term "and / or" used in this document includes any and all combinations of one or more of the related listed items and can be abbreviated as " / ".

[0028] It should be understood that although the terms "first", "second", etc. may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of the present principle, the first element can be referred to as the second element, and similarly, the second element can also be referred to as the first element.

[0029] Although some diagrams contain arrows on the communication paths to indicate the main communication directions, it should be understood that the communication can occur in the direction opposite to the shown arrows.

[0030] Some examples are described in conjunction with block diagrams and operation flowcharts, where each block represents a circuit element, module, or section of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in other implementations, the functions noted in the blocks may not be executed in the noted order. For example, two consecutively shown blocks may actually be executed substantially simultaneously, or sometimes in the reverse order, depending on the functions involved.

[0031] The phrase "according to an example" or "in an example" mentioned in this document means that the specific features, structures, or characteristics described in conjunction with the example can be included in at least one implementation of the present principle. The phrase "according to an example" or "in an example" that appears multiple times in the specification does not necessarily refer to the same example, and individual examples or alternative examples are not necessarily mutually exclusive with other examples.

[0032] The reference numerals appearing in the claims are for illustration purposes only and shall not have any limiting effect on the scope of the claims. Although not explicitly stated, the examples and their variations may be used in any combination or sub-combination.

[0033] The automotive industry and self-driving cars are areas where point clouds can be used. Self-driving cars should be able to "see" their surroundings and make correct driving decisions based on the actual situation of their surroundings. Typical sensors like LiDAR generate (dynamic) point clouds that are used by the perception engine. These point clouds are not meant to be seen by the human eye, they are usually sparse, not necessarily in color, and are dynamic and captured very frequently. They may contain other properties, such as reflectivity provided by LiDAR, as this property indicates the material of the sensed object and may help in making decisions.

[0034] Virtual reality (VR) and immersive worlds have become a hot topic, with many seeing them as the future of 2D flat video. The basic idea is to immerse the viewer in all the surroundings, unlike traditional TV where the viewer can only see the virtual world in front of them. There are multiple levels of immersion depending on how much freedom the viewer has in the environment. Point clouds are an ideal format for distributing VR worlds. They can be static or dynamic and are usually of moderate size, with no more than millions of points at a time.

[0035] Point clouds can also be used for a variety of purposes, such as cultural heritage / architecture, where objects such as statues or buildings are 3D scanned to share their spatial configuration without having to send or visit them. It can also ensure that information about the object is preserved in the event that the object (such as a temple) is damaged due to an earthquake. Such point clouds are usually static, colorful, and huge.

[0036] Another use case is terrain and mapping, where using 3D representations, maps are not limited to flat surfaces but can also include landforms. Google Maps is now a good example of a 3D map, but it uses meshes instead of point clouds. Nevertheless, point clouds may be a suitable data format for 3D maps, and such point clouds are usually static, colorful, and huge.

[0037] World modeling and sensing through point clouds may be the key technology to allow machines to understand the 3D world around them, which is crucial for the above applications.

[0038] Point cloud compression refers to the problem of concisely representing the surface manifold of the objects contained in a point cloud. For this problem, several methods have been explored and can be classified into the following categories: PCC in the input domain, PCC in the raw domain, PCC in the transform domain, and PCC via entropy coding. PCC in the input domain refers to downsampling the original point cloud by selecting or generating points that represent the underlying surface manifold. Although there are various machine learning techniques (deep learning-based) and classical machine learning techniques in this field, PCC in the input domain is only applicable to low compression ratios because it is restricted to remain within the input domain and is mainly used to summarize the point cloud for subsequent downstream processing. PCC in the raw domain is also closely related to this field, where instead of key points, primitives (regular two-dimensional / three-dimensional geometric data) are generated, aiming to closely follow the underlying object manifold. PCC in the transform domain refers to first transforming the original point cloud data to another domain by a learning-based method or a classical method, and then compressing the representation in the new domain to obtain more efficient compression. Finally, there is the case of PCC via entropy coding, where the original point cloud data or another (easily obtainable) representation of the point cloud is entropy-coded by an adaptive learning-based method or a classical method.

[0039] This principle is related to this last case of PCC via entropy coding and further belongs to the category of learned hierarchical entropy models. Existing learned hierarchical entropy methods for point clouds only utilize the direct information in the hierarchical chain. According to this principle, not only the direct neighborhood information is integrated at the current hierarchical level, but also the available global information is integrated at the upper levels. This principle aims to perform point cloud compression as an independent method and be used for subsequent machine tasks (e.g., classification, segmentation, etc.). The goal of this principle is to perform lossless compression on the relevant point cloud geometric data and handle other features such as color or reflectivity.

[0040] Figure 1Illustrates a 3D point cloud encoding method 10 according to this principle. Regarding the point cloud encoding system, first, the input point cloud X containing N points is processed (i.e., transformed). For example, it can be quantized to a certain precision to obtain M points. Then, these M points are further transformed into a tree representation until a specified tree depth is reached. Possible tree representations include octree representation, quadtree plus binary tree (QTBT) representation, or prediction tree representation, etc. Octree representation is a direct method of partitioning and representing positions in 3D space, where the cube containing the entire point cloud is subdivided into 8 sub-cubes. Then, an 8-bit code, called occupancy code or occupancy symbol, is generated by associating a 1-bit value with each sub-cube. It is used to indicate whether the sub-cube contains points (i.e., value 1) or does not contain points (i.e., value 0). This partitioning process is performed recursively to form a tree, where only the sub-cubes containing multiple points are further partitioned. Similar to octree representation, QTBT also partitions the three-dimensional space recursively but allows for more flexible partitioning using a quadtree or a binary tree. It is particularly useful for representing sparsely distributed point clouds. Different from octree and QTBT that partition the three-dimensional space recursively, the prediction tree defines a prediction structure between the three-dimensional points in the three-dimensional point cloud. Using the prediction tree for geometric encoding is mainly beneficial for Class 3 content (LIDAR sequences) in PCC. Through this transformation step, the compression of the original point cloud geometric data becomes the compression of the tree representation.

[0041] In this paper, the octree representation is used as an example for illustration without loss of generality. After converting the original point cloud into an octree structure 11, a deep learning-based conditional tree-structured entropy model is used to predict the occupancy symbol distribution 12 for all nodes in the tree. This conditional entropy model operates in a node-by-node manner and provides the occupancy symbol distribution 13 of a node based on the context of the node and the features of adjacent nodes in the tree. The occupancy symbol of a node refers to the binary occupancy rate of each of its eight child nodes and is represented as an 8-bit integer by the 8-bit binary child node occupancy rate. The context of a given node contains the following information: the occupancy rate of the parent node (8-bit integer), the octree depth / hierarchy of the given node, the octant of the given node, and finally the spatial position of the current node. Then, the conditional symbol distribution is fed into a lossless adaptive entropy encoder, which compresses the occupancy rate of each node to generate a bitstream 14.

[0042] Figure 2Depicts a decoding method 20 according to this principle. Given a compressed bitstream 21 of a point cloud, the decoding method first starts by generating a default context for the root node. Then, the depth entropy model uses the default context of the root node to generate an occupancy symbol distribution. The adaptive entropy decoder uses this distribution and the portion of the bitstream corresponding to the root node to decode the root occupancy symbol. Now, the contexts of all the child nodes of this root node can be initialized, and the same process is repeated multiple times to expand and decode the entire tree structure. After the entire tree is decoded, it is transformed back to obtain the reconstructed point cloud 22.

[0043] The depth entropy model is used to predict the occupancy symbol distribution, such as in "OctSqueeze: Octree-Structured Entropy Models for LiDAR Compression" by Huang, Lila, et al. in the Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition 2020. However, different from methods that only use local information from the parent node to predict the distribution, according to this principle, more global information is available. Specifically, when predicting the occupancy symbol distribution of the current node, information from sibling nodes as well as all ancestor nodes is considered.

[0044] Figure 3 Depicts the architecture of a depth entropy model 30 according to this principle. Given an octree structure, a deep conditional tree-structured entropy model is trained. This conditional entropy model operates in a node-by-node manner (its operations can be parallelized) and predicts the probability distribution for the occupancy symbol of the node based on the context of the node and its neighboring nodes (including its sibling nodes and ancestor nodes). Then, the adaptive entropy encoder or decoder further uses this conditional occupancy symbol distribution to compress or decompress the tree structure, as Figure 1 and Figure 2 shown. To train the depth entropy model, an octree representation of a point cloud dataset containing a large number of independent point clouds is used. The model takes as input the context of each node and the features of its neighboring nodes and outputs the conditional occupancy symbol distribution. Then, the cross-entropy loss on the i-th node is calculated as where y ij is the one-hot encoded ground truth symbol of the i-th node, and q i represents the predicted distribution for the symbol of the i-th node. The network is trained in a self-supervised manner to minimize the cross-entropy loss for all nodes in all octrees.

[0045] Existing methods for point cloud compression are trained for only one input precision and use the trained model to compress point clouds with multiple different input precisions. This process is not optimal and results in poor compression performance for input precisions not encountered during the training phase. According to this principle, the model is trained with the highest set of geometric precisions of the input and multiple quantized versions (low-resolution versions of the point cloud). Quantization is achieved by multiplying the positions of the points in the point cloud by a factor less than 1, thereby obtaining a quantized version of the point cloud. This training enables the model to perform at multiple precision levels. This can improve the robustness of the training and achieve better generalization.

[0046] In addition to the precision of the point cloud, the distribution of the points in the point cloud also needs to be considered. Since different acquisition methods can lead to significant differences in the distribution of points in the original point cloud, when a high compression ratio is required, the model is trained based on one type of data (such as LIDAR or VR / AR data). However, if a more general model that performs well on average across several different types of datasets is needed, different types of datasets can be combined into one, and then the model can be trained on this combined dataset. Therefore, this training scheme can adapt to the target application and the target compression performance.

[0047] Figure 4 Shows the implementation of the depth entropy coding module in the encoder according to the first embodiment of this principle. In the first embodiment, the conditional entropy model combines the depth features of sibling nodes to simultaneously predict the occupancy symbols of all sibling nodes. This embodiment contrasts with the prior art, where the neighborhood information from sibling nodes is not fully utilized. Let the context of a given node i be represented as the vector c i . In addition, consider a deep neural network consisting of multiple back-to-back multi-layer perceptron (MLP) modules, where the k-th MLP module is denoted as MLP (k) . Then, the initial depth features of a specific node are obtained. Starting from this initial feature for each node (which can be obtained in parallel), subsequent depth features can be obtained (also in parallel), i.e., where is the depth feature of the parent node of the i-th node in MLP (k-1) . It should be noted that each MLP (k) is shared among all nodes. The final MLP is denoted as MLP (k) , which is special because it requires an additional depth feature as input, which is composed of the depth features of all sibling nodes, i.e., where MLP (sib)First, operations are performed on the depth features of each sibling node (including itself) of the $i$-th node respectively, and then the pooling function (e.g., Max(.)) operates on the resulting features in each dimension to generate an overall combined depth feature with the same length as the input feature. In this way, the final depth feature of the $i$-th node is obtained, that is Then, this final feature is passed through a linear layer and then processed by "Softmax", thereby generating a 256-dimensional probability vector for each of the 256 possible 8-bit occupancy symbols.

[0048] Figure 4 The depth entropy coding module 40 in the encoder is shown. In Figure 4 it, the feature extractor 41 refers to the MLP (0) and the sibling node feature extractor 42 refers to the MLP (sib) The remaining MLPs are bundled into the feature aggregator 43. The first MLP in our model has five layers deep and has 128-dimensional hidden features. All hidden MLPs contain three residual layers and the same 128-dimensional hidden features. The additional MLP (sib) is a single-layer MLP and also has 128-dimensional hidden features. The final linear layer produces a 256-dimensional output, which is then processed by Softmax, thereby generating a 256-dimensional probability vector. The layers in all MLPs are linear layers followed by ReLU, without normalization.

[0049] Although the proposed depth entropy model operates node by node and uses the depth features from ancestor nodes, during the encoding process, each MLP can be executed in parallel on all nodes. This is because the depth features used from ancestor nodes are the output of the previous MLP module. However, when using the depth entropy model during the decoding process, the ancestor nodes must be fully decoded before moving down the octree, so only parallel operations (i.e., decoding) can be performed on sibling nodes. Therefore, all embodiments in this disclosure have the same function: during the encoding process, the model can operate in parallel on all nodes; during the decoding process, the model can only operate in parallel on sibling nodes.

[0050] Figure 5 The implementation of the depth entropy coding module according to the second embodiment of this principle is shown. Compared with the first embodiment, in this architecture, the MLP (sib) 42 (which is used to collect the features of all sibling nodes in the first embodiment) collects the depth features of the parent node 51 and all sibling nodes under the same grandparent node. This generates richer depth features that contain more information from the parent, and since the MLP is shared, no additional parameters are required.

[0051] In a third embodiment of the present principles, richer features can be extracted from the parent level by utilizing features from all nodes of the parent level (rather than only utilizing sibling nodes of the parent level). Using features from all nodes of a higher level, a feature vector can be generated that represents a rough global topological structure of an object including a point cloud.

[0052] Figure 6 FIG. 6 shows an implementation of a deep entropy coding module according to a fourth embodiment of the present principles. In this implementation, the first MLP 62 collects deep features of all sibling nodes of a particular node, and the second MLP 61 collects features of all nodes of the parent level. The two MLPs operate at different scales. MLP 62 processes local information within the neighborhood, while MLP 61 extracts information from the global manifold shape. Due to the difference in scale, a separate MLP is introduced (pa) 61 to collect the characteristics of the parent.

[0053] The fifth embodiment of the present principles is based on the previous embodiment. However, in addition to converting the original point cloud into an octree structure, this embodiment also allows the use of other tree representations, such as QTBT or prediction tree.

[0054] Figure 7A , Figure 7B and Figure 7C A sixth embodiment of the present principles is shown. The sixth embodiment is based on embodiments 1 to 4. However, instead of using MLP to extract and propagate features for each node on each octree level, more advanced architectures such as convolution, sparse convolution, ResNet (residual network), Inception ResNets and Transformer (attention mechanism-based model) are used. They can also be used to extract and propagate features. This is intended to provide enhanced feature aggregation capabilities.

[0055] In a first variation of the sixth embodiment, the feature extractor is a series of sparse 3D convolutional layers, each followed by a ReLU activation function, such as Figure 7A CONV D 71 represents a sparse 3D convolutional layer with D output channels.

[0056] In the second variation of the sixth embodiment, the feature aggregation module adopts the ResNet architecture, such as Figure 7B This example shows the architecture of a ResNet module for aggregating features with D channels. Figure 7A Compared to the first variant, Figure 7B The second variant of introduces residual connections in the input and adds them to the output of the convolutional layer.

[0057] In the third variant of the sixth embodiment, the feature aggregation module adopts the Inception-ResNet (IRN) architecture, as Figure 7C shown. This example shows the architecture of the IRN module for aggregating features with D channels.

[0058] Figure 8 Schematically shows the transformer module used in the fourth variant of the sixth embodiment of this principle. In this variant, the feature propagation module adopts a transformer architecture similar to the voxel transformer, for example, as proposed by Mao, Jiageng et al. in "Voxel Transformers for 3D Object Detection" published in the Proceedings of the IEEE / CVF International Conference on Computer Vision in 2021. Figure 8 The transformer module in consists of a self-attention block with residual connections and an MLP block (composed of MLP layers) with residual connections. Given the current feature vector f A associated with the voxel position A, i and k adjacent features associated with the voxel position A i where A Ai (0 ≤ i ≤ k - 1) are the k nearest neighbors of A in the input sparse tensor, the self-attention block attempts to update the feature f A based on all adjacent features f Ai . First, based on the coordinates of A, the point A i is obtained through a k-nearest neighbor (kNN) search. Then, the query embedding Q A of A is calculated:

[0059] Q A = MLP Q (f A ),,

[0060] After that, the key embedding K Ai and value embedding V Ai of all nearest neighbors of A are calculated:

[0061]

[0062] where MLP Q (·), MLP K (·) and MLP V (·) are MLP layers for obtaining the query, key, and value respectively; E Ai is the positional encoding between voxel A and A i , and the calculation formula is:

[0063]

[0064] where MLP P (·) is the MLP layer for obtaining the positional encoding, PA and P Ai are three-dimensional coordinates, respectively representing the centers of voxels A and A i . The feature of position A output by the self-attention block is:

[0065]

[0066] where σ(·) is the Softmax normalization function, d is the length of the feature vector f A , and c is a predefined constant.

[0067] The transformer block updates the features of all occupied positions in the sparse tensor in the same way and then outputs the updated sparse tensor. As a simplified example, MLP Q (·), MLP K (·), MLP V (·), and MLP P (·) can contain only one fully connected layer, which corresponds to a linear projection.

[0068] In a variant, multiple feature aggregation blocks can be concatenated together to further enhance performance. The feature aggregation blocks can be of the same type. For example, they are all transformer blocks. In this case, the parameters of their neural network layers can be shared or not shared. The feature aggregation blocks can also be a mixture of different types of feature aggregation blocks, such as a mixture of IRN blocks and transformer blocks.

[0069] Figure 9 FIG. shows a depth entropy encoding / decoding method, in which the features from the parent are upsampled to match the resolution of the current octree level. These upsampled features are propagated to the child nodes for depth occupancy probability estimation. When encoding / decoding the nodes at the current level, the features from the parent are already available. The features from the parent are upsampled to obtain the unique features of all the child nodes at the current level. This upsampling can be performed by an MLP-based module that receives the feature vector and index corresponding to the child node to output the features of the corresponding child node; or by a module based on conventional or sparse convolution that receives the entire feature map of the parent and outputs the upsampled feature map containing the features of all the nodes at the current level. Then, this feature is paired (concatenated or added) with the feature of the current node obtained from its neighborhood occupancy information by an MLP or a module based on conventional / sparse convolution. After that, this combined feature can be propagated again through any feature aggregator architecture (as described in the sixth embodiment) to obtain the final depth feature. The probability generation module (PGM) uses this depth feature to output the predicted probability of each byte occupancy symbol. In addition, this depth feature is also sent to the next level and used as the feature of its parent.

[0070] In a variant, to reduce complexity, the upsampled features are directly combined with the context information (occupancy status, node position, etc.) of the current level (instead of being combined with the features obtained through the context), and the combined features are propagated to obtain the final depth features. Figure 9 The feature upsampler and feature aggregator modules are shown. The feature upsampler consists of a sparse convolutional layer and a sparse upsampling convolutional layer (as described in the sixth embodiment), and the feature aggregator consists of multiple Inception-ResNet layers (as described in the sixth embodiment).

[0071] Figure 10 and 11 A scenario of this principle is shown, where the original point cloud is converted into blocks through a shallow octree. Embodiments 1 to 7 relate to scenarios where the entire point cloud is converted into a single octree representation and the octree is compressed in a lossless manner. However, as the geometric data accuracy and the point density in the point cloud increase, this process becomes increasingly time-consuming and computationally intensive. In addition, the process of converting the original data into an octree representation also takes longer. To solve this problem, in this embodiment 8, the original point cloud is first converted into blocks through a shallow octree, the data points in each block are converted from the original coordinates to local block coordinates (by shifting the origin of each block), and finally the data of each block is converted into a separate octree. Through this process, each block contains a smaller part of the point cloud and can be converted into an octree more quickly in parallel. After lossless compression and decompression, the points recovered from each block are combined and restored to the original coordinates. The auxiliary information about the block division from the shallow octree is also compressed using uniform entropy coding and added to the bitstream.

[0072] Figure 12 An example architecture of device 120 is shown, which can be configured to implement Figures 1 to 11 the encoding and / or decoding methods described in. Alternatively, each circuit of the encoder and / or decoder according to this principle can be a device conforming to Figure 12 the architecture, for example, connected together through its bus 121 and / or I / O interface 126.

[0073] Device 120 includes the following elements connected together through data and address buses 31:

[0074] - A microprocessor 122 (or CPU), such as a DSP (Digital Signal Processor);

[0075] - A ROM (Read-Only Memory) 123;

[0076] - A RAM (Random Access Memory) 124;

[0077] - A storage interface 125;

[0078] - An I / O interface 126 for receiving data to be transmitted from an application; and

[0079] - A power source, such as a battery.

[0080] According to the example, the power source is located outside the device. In each memory, the term "register" used in this specification may refer to a small-capacity (a few bits) area or a very large area (e.g., the entire program or a large amount of received or decoded data). The ROM 123 contains at least one program and some parameters. The ROM 123 can store algorithms and instructions for implementing the related technologies of this principle. After the CPU 122 is turned on, the program is loaded into the RAM and the corresponding instructions are executed.

[0081] The RAM 124 contains in the register: a program executed and loaded by the CPU 122 after the device 120 is turned on, input data in the register, intermediate data in different states of the method in the register, and other variables in the register for executing the method.

[0082] The implementations described herein can be implemented, for example, in the form of a method or process, a device, a computer program product, a data stream, or a signal. Even if discussed only in the context of a single implementation form (e.g., discussed only as a method or a device), the implementation of the discussed features can also be implemented in other forms (e.g., a program). The device can be implemented, for example, in the form of appropriate hardware, software, and firmware. The method can be implemented, for example, in a device such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0083] According to the example, the device 120 is configured to implement the method described in combination with Figure 9 and Figure 10 and belongs to the following set:

[0084] - Mobile devices;

[0085] - Communication devices;

[0086] - Gaming devices;

[0087] - Tablet computers (or tablet PCs);

[0088] - Laptop computers;

[0089] - Still image cameras;

[0090] - Video cameras;

[0091] - Encoding chips;

[0092] - A server (e.g., a broadcast server, a video - on - demand server, or a web server).

[0093] Figure 13 An example embodiment of the stream syntax when data is transmitted via a packet - based transport protocol is shown. Figure 13 An example structure 130 of a volumetric video stream is shown. The video stream is constructed as a container that organizes the stream into independent syntax elements. This structure may include a header portion 131, which is a set of data common to each syntax element in the stream. For example, the header portion contains some metadata about the syntax elements, describing the nature and role of each syntax element. The structure includes a payload that contains a syntax element 132 and at least one syntax element 133. The syntax element 132 contains data representing a tree - structured point cloud of a point cloud sequence. These tree - structured point clouds have been compressed according to a depth - entropy compression method.

[0094] The syntax element 133 is part of the data stream payload and may contain metadata about how the frames of the syntax element 42 are encoded. Such metadata can be associated with each frame or a group of frames of the video (also known as a group of pictures (GoP) in video compression standards).

[0095] The implementations described herein can be implemented, for example, in a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when discussed in the context of a single implementation form (e.g., only as a method or a device), the implementation of the discussed features can be implemented in other forms (e.g., a program). The apparatus can be implemented, for example, in appropriate hardware, software, and firmware. A method can be implemented, for example, in a device such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as smartphones, tablets, computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end - users.

[0096] The implementation of the various processes and features described herein can be embodied in a variety of different devices or applications, particularly, for example, devices or applications related to data encoding, data decoding, view generation, texture processing, and other image - and related texture information and / or depth - information processing. Examples of such devices include encoders, decoders, post - processors that process decoder outputs, pre - processors that provide inputs to encoders, video encoders, video decoders, video codecs, web servers, set - top boxes, laptop computers, personal computers, mobile phones, PDAs, and other communication devices. It should be clear that these devices can be mobile devices and can even be installed in mobile vehicles.

[0097] In addition, these methods can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) can be stored on a processor-readable medium, such as an integrated circuit, a software carrier, or other storage devices, such as a hard disk, an optical disc ("CD"), an optical disc (e.g., DVD, commonly referred to as a digital versatile disc or digital video disc), a random access memory ("RAM"), or a read-only memory ("ROM"). These instructions can form an application program tangibly embodied on the processor-readable medium. The instructions can exist, for example, in the form of hardware, firmware, software, or a combination of the three. The instructions can exist, for example, in an operating system, a separate application program, or a combination of both. Thus, a processor can be described as, for example, either a device configured to execute a process or a device that includes a processor-readable medium (e.g., a storage device) having instructions for executing a process. In addition, the processor-readable medium can store data values generated by the implementation, as a supplement to or in place of the instructions.

[0098] Those skilled in the art will understand that a variety of implementations can generate a variety of formatted signals for carrying information that can be stored or transmitted, for example. This information can include, for example, instructions for executing a method or data generated by one of the implementations. For example, a signal can be formatted to carry, in data form, rules for the syntax for writing to or reading from the embodiments, or the actual syntax values written in the embodiments. Such signals can be formatted, for example, as electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. Signals can be stored on a processor-readable medium.

[0099] A variety of implementations have been described. However, it should be understood that various modifications can be made. For example, elements in different implementations can be combined, supplemented, modified, or removed to yield other implementations. In addition, those skilled in the art should understand that other structures and processes can be substituted for the disclosed implementations, and the resulting implementations will perform at least substantially the same functions in at least substantially the same manner to achieve at least substantially the same results as the disclosed implementations. Accordingly, this application covers these and other implementations.

Claims

1. A method, comprising: obtaining nodes of a tree-structured point cloud from a bitstream; initializing context information for the nodes; predicting an occupancy symbol distribution for the nodes by using a learning-based entropy model, utilizing the context information of the nodes and feature information from adjacent nodes and parent nodes previously obtained from the bitstream; decoding an occupancy symbol for the nodes based on the occupancy symbol distribution by using an adaptive entropy decoder; and outputting an extended tree based on the occupancy symbol.

2. The method according to claim 1, wherein the steps of the method are iteratively performed for each node of the tree-structured point cloud.

3. The method according to claim 1 or 2, wherein the entropy model predicts the occupancy symbol distribution by using, through a learning-based module for sibling nodes of the nodes, feature information of sibling nodes of the parent node of the nodes.

4. The method according to any one of claims 1 to 3, wherein the entropy model predicts the occupancy symbol distribution by using, through a learning-based module for sibling nodes of the nodes, feature information of all nodes of the parent level.

5. The method according to any one of claims 1 to 3, wherein the entropy model predicts the occupancy symbol distribution by using, through a separate learning-based module, feature information of all nodes of the parent level.

6. The method according to any one of claims 1 to 5, wherein the point cloud is configured as an octree.

7. An apparatus, comprising a memory associated with a processor, the processor configured to: obtain nodes of a tree-structured point cloud from a bitstream; initialize context information for the nodes; predict an occupancy symbol distribution for the nodes by using a learning-based entropy model, utilizing the context information of the nodes and feature information from adjacent nodes and parent nodes previously obtained from the bitstream; decode an occupancy symbol for the nodes based on the occupancy symbol distribution by using an adaptive entropy decoder; and output an extended tree based on the occupancy symbol.

8. The apparatus according to claim 7, wherein the processor is further configured to iteratively perform the process for each node of the tree-structured point cloud.

9. The apparatus according to claim 7 or 8, wherein the entropy model predicts the occupancy symbol distribution by using, through a learning-based module for sibling nodes of the nodes, feature information of sibling nodes of the parent node of the nodes.

10. The apparatus according to any one of claims 7 to 9, wherein the entropy model predicts the occupancy symbol distribution by using, through a learning-based module for sibling nodes of the nodes, feature information of all nodes of the parent level.

11. The apparatus according to any one of claims 7 to 9, wherein the entropy model predicts the occupancy symbol distribution by using, through a separate learning-based module, feature information of all nodes of the parent level.

12. The apparatus according to any one of claims 7 to 11, wherein the point cloud is configured as an octree.

13. A method, comprising: obtaining a point cloud configured as a tree; For each node of the tree, initialize the context information; predict the occupancy symbol distribution by using a learning-based entropy model with the context information of the node and the feature information from adjacent nodes and the parent node of the node; encode the occupancy symbols in a bitstream based on the occupancy symbol distribution by using an adaptive entropy encoder; and generate a combined bitstream for the tree.

14. The method according to claim 13, wherein the entropy model predicts the occupancy symbol distribution by using a learning-based module for sibling nodes of the node with the feature information of sibling nodes of the parent node of the node.

15. The method according to claim 14, wherein the entropy model predicts the occupancy symbol distribution by using a learning-based module for sibling nodes of the node with the feature information of all nodes of the parent level.

16. The method according to claim 14, wherein the entropy model predicts the occupancy symbol distribution by using a separate learning-based module with the feature information of all nodes of the parent level.

17. The method according to claim 14, wherein the entropy model first upsamples the feature information from the parent of each node to match the level of each node.

18. The method according to any one of claims 14 to 17, wherein the point cloud is constructed as an octree.

19. An apparatus, comprising a memory associated with a processor, the processor configured to: obtain a point cloud constructed as a tree; for each node in the tree, initialize the context information; predict the occupancy symbol distribution by using a learning-based entropy model with the context information of the node and the feature information from adjacent nodes and the parent node of the node; encode the occupancy symbols in a bitstream based on the occupancy symbol distribution by using an adaptive entropy encoder; and generate a combined bitstream for the tree.

20. The apparatus according to claim 19, wherein the entropy model predicts the occupancy symbol distribution by using a learning-based module for sibling nodes of the node with the feature information of sibling nodes of the parent node of the node.

21. The apparatus according to claim 19, wherein the entropy model predicts the occupancy symbol distribution by using a learning-based module for sibling nodes of the node with the feature information of all nodes of the parent level.

22. The apparatus according to claim 19, wherein the entropy model predicts the occupancy symbol distribution by using a separate learning-based module with the feature information of all nodes of the parent level.

23. The apparatus according to claim 19, wherein the entropy model first upsamples the feature information from the parent of each node to match the level of each node.

24. The apparatus according to any one of claims 19 to 23, wherein the point cloud is constructed as an octree.