Deep learning based point cloud lossless compression method and apparatus
Patent Information
- Application Number
- CN202211080672.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-05
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-09-05
AI Technical Summary
[0004]本发明提供基于深度学习的点云无损压缩方法和装置,用以解决现有技术中点云压缩效果不佳的缺陷,通过划分小段序列并利用自注意力神经网络模型对各个小段序列进行分组编码的方式,将更多地节点上下文信息应用到编码过程中,进而提升编码的压缩率
[0052]本发明提供的基于深度学习的点云无损压缩方法和装置,改进了利用八叉树实现点云压缩的技术,具体包括:将目标点云转换为八叉树;按广度优先顺序将八叉树每一层节点铺平为节点序列;然后利用上下文窗口将八叉树每一层节点铺平后的节点序列进行分段,得到所述八叉树每一层对应的小段序列;按照八叉树层级由上至下的顺序,对所述八叉树每一层对应的小段序列进行并行无损压缩,以得到所述目标点云的压缩结果。本发明通过对小段序列进行分层的并行压缩,大大提升点云压缩速度;同时对于任意一个小段序列,利用自注意力神经网络模型对小段序列进行分组编码,使得编码能够利用更多的节点上下文信息,进而提升编码的压缩率。
Smart Images

Figure CN115471576B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of point cloud compression technology, and more particularly to a lossless point cloud compression method and apparatus based on deep learning. Background Technology
[0002] With the rapid development of 3D scanning technology, people can now effectively capture complex geometric information using fine-grained point clouds. Point cloud compression is crucial for the storage and transmission of point cloud data, therefore, the development of point cloud compression technology has received widespread attention.
[0003] Currently, common point cloud compression methods include OctAttention and VoxelContext-Net. OctAttention, a self-attention-based compression method, transforms the point cloud into an octree. It then uses a breadth-first traversal to create a sequence of nodes, each containing its level, position within its parent node, and OctValue. Next, based on the context, a neural network with self-attention learns the OctValue probability distribution for each node. Finally, entropy encoding is applied to the OctValue probability distribution of each node to obtain the compressed point cloud file. However, this method only uses the node information of nodes preceding it as context when encoding each node, resulting in uneven distribution of context in the space. This affects the accuracy of OctValue probability distribution prediction, thus impacting the encoding performance. VoxelContext-Net differs from OctAttention in that it uses the node information of nodes spatially adjacent to each node as context when encoding, employing a convolutional neural network to predict the OctValue probability distribution. This approach uses the OctValue of nodes in the same layer less when encoding each node. Moreover, the computational power of convolutional neural networks is limited and cannot encompass a large space, resulting in very little contextual information available. The prediction accuracy of the OctValue probability distribution is poor, which in turn leads to poor encoding performance. Summary of the Invention
[0004] This invention provides a lossless point cloud compression method and apparatus based on deep learning to address the shortcomings of poor point cloud compression performance in existing technologies. By dividing the sequence into small segments and using a self-attention neural network model to encode each segment, more node context information is applied to the encoding process, thereby improving the compression ratio.
[0005] In a first aspect, the present invention provides a lossless point cloud compression method based on deep learning, comprising:
[0006] Representing the target point cloud using an octree;
[0007] Perform a breadth-first traversal on the octree to obtain the node sequence corresponding to each level of the octree;
[0008] By using a context window to separate the node sequence corresponding to each level of the octree, a small segment sequence corresponding to each level of the octree can be obtained;
[0009] Following the top-to-bottom order of the octree hierarchy, the small segments corresponding to each level of the octree are compressed in parallel without loss to obtain the compression result of the target point cloud.
[0010] Lossless compression of short sequences is achieved by using a self-attention neural network model to encode the short sequences in groups.
[0011] According to the point cloud lossless compression method based on deep learning provided by the present invention, the step of separating the node sequence corresponding to each layer of the octree using a context window to obtain the small segment sequence corresponding to each layer of the octree includes:
[0012] The node sequence corresponding to each level of the octree is divided with the context window length as the length of the segment sequence to obtain the segment sequence corresponding to each level of the octree.
[0013] The point cloud lossless compression method based on deep learning provided by the present invention performs lossless compression on small sequences, including:
[0014] Group the nodes in the short sequence;
[0015] Using a self-attention neural network model, a first mask tensor, a second mask tensor, and node information of each node and its ancestor nodes in the short sequence, the OctValue probability distribution of each node in the short sequence is determined.
[0016] Based on the OctValue probability distribution of each node in the short sequence, the nodes of the short sequence are grouped and encoded in ascending order of group number to obtain the encoding result of the short sequence.
[0017] The context node includes the level of the octree in which the node is located, the position of the node in the parent node, and the OctValue of the node; the OctValue of the node is an eight-bit binary number representing the existence of the node's eight subspace points.
[0018] The first mask tensor aims to mask the OctValue of all nodes in the short sequence; the second mask tensor aims to mask the OctValue of nodes in the short sequence whose group number is greater than the group number to which the corresponding node belongs during the learning process of the OctValue probability distribution of each node in the short sequence.
[0019] According to the point cloud lossless compression method based on deep learning provided by the present invention, the step of grouping the nodes in the small sequence includes:
[0020] Set the node indexes of nodes in the same group to be separated by LM;
[0021] Where M is the number of groups and L is an integer greater than or equal to 1.
[0022] According to the point cloud lossless compression method based on deep learning provided by the present invention, the step of determining the OctValue probability distribution of each node in the small segment sequence by utilizing a self-attention neural network model, a first mask tensor, a second mask tensor, and the context nodes of each node and its ancestor nodes in the small segment sequence includes:
[0023] The context node of each node and its ancestor node in the short sequence is represented in the form of a three-dimensional tensor.
[0024] Based on the first mask tensor and the second mask tensor, the three-dimensional tensor is processed using the self-attention neural network model to obtain the OctValue probability distribution of each node in the short sequence.
[0025] According to the point cloud lossless compression method based on deep learning provided by the present invention, the three-dimensional tensor has a size of (N,K,3), and the neural network module includes an embedding layer, a first reshape layer, a first transformer network structure, a second transformer network structure, a second reshape layer, a first linear layer, and a softmax layer.
[0026] The step of processing the three-dimensional tensor using the self-attention neural network model based on the first mask tensor and the second mask tensor to obtain the OctValue probability distribution of each node in the short sequence includes:
[0027] In the Embedding layer, the size of the three-dimensional tensor is converted to (N,K,S);
[0028] In the first Reshape layer, the three-dimensional tensor of size (N,K,S) is shaped into a two-dimensional tensor of size (N,KS).
[0029] In the first transformer network structure, the first mask tensor is used to learn the node OctValue probability distribution of the two-dimensional tensor to obtain a first learning tensor of size (N,KS).
[0030] In the second transformer network structure, the node OctValue probability distribution of the two-dimensional tensor is learned by using the second mask tensor to obtain a second learning tensor of size (N,KS).
[0031] In the second Reshape layer, the sum of the first learned tensor and the second learned tensor is shaped into a two-dimensional tensor of (N, KS).
[0032] Using the first Linear layer, the two-dimensional tensor output by the second Reshape layer is regularized to obtain a two-dimensional tensor of size (N, 255);
[0033] In the SoftMax layer, a two-dimensional tensor of size (N, 255) is normalized to obtain the OctValue probability distribution of each node in the small sequence from the normalization result; where N is the number of nodes in the small sequence, K is the number of ancestor nodes traced upwards, and S is any integer greater than 3.
[0034] According to the point cloud lossless compression method based on deep learning provided by the present invention, the first transformer network structure and the second transformer network structure have the same structure, both of which are composed of R encoders stacked in series.
[0035] The encoder includes a first sublayer connection structure and a second sublayer connection structure connected to it;
[0036] The first sub-layer connection structure includes a multi-head self-attention sub-layer, a normalization layer, and a residual connection;
[0037] The second sub-layer connection structure includes a second Linear layer, a normalization layer, and a residual connection;
[0038] In the first transformer network structure, the node OctValue probability distribution is learned by using the first mask tensor on the two-dimensional tensor, resulting in a first learning tensor of size (N, KS), including:
[0039] The two-dimensional tensor is used as the input of the multi-head self-attention sub-layer of the first encoder in the first transformer network structure. The first mask tensor is used as the input mask tensor of the multi-head self-attention sub-layer of each encoder in the first transformer network structure. The two-dimensional tensor is used to learn the node OctValue probability distribution using the first transformer network structure to obtain a first learning tensor of size (N, KS).
[0040] In the second transformer network structure, the node OctValue probability distribution is learned by using the second mask tensor on the two-dimensional tensor, resulting in a second learned tensor of size (N, KS), including:
[0041] The two-dimensional tensor is used as the input of the multi-head self-attention sublayer of the first encoder in the second transformer network structure, and the second mask tensor is used as the input mask tensor of the multi-head self-attention sublayer of each encoder in the second transformer network structure. The two-dimensional tensor is used to learn the node OctValue probability distribution using the second transformer network structure to obtain a second learning tensor of size (N, KS).
[0042] According to the point cloud lossless compression method based on deep learning provided by the present invention, the self-attention neural network model is constructed on the basis of a pre-trained neural network;
[0043] The pre-trained neural network is trained using small sequence samples of octree nodes with their OctValue nodes randomly set to zero.
[0044] Secondly, the present invention provides a point cloud lossless compression device based on deep learning, comprising:
[0045] The octree representation module is used to represent the target point cloud using an octree;
[0046] The breadth-first traversal module is used to perform a breadth-first traversal on the octree to obtain the node sequence corresponding to each level of the octree;
[0047] The separation module is used to separate the node sequence corresponding to each level of the octree using a context window to obtain the small segment sequence corresponding to each level of the octree;
[0048] The compression module is used to perform parallel lossless compression on the small segment sequence corresponding to each level of the octree in order from top to bottom, so as to obtain the compression result of the target point cloud.
[0049] Lossless compression of short sequences is achieved by using a self-attention neural network model to encode the short sequences in groups.
[0050] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the lossless point cloud compression method based on deep learning as described in the first aspect.
[0051] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the lossless point cloud compression method based on deep learning as described in the first aspect.
[0052] This invention provides a lossless point cloud compression method and apparatus based on deep learning, which improves the technique of point cloud compression using octrees. Specifically, it includes: converting the target point cloud into an octree; flattening each layer of the octree into a node sequence in breadth-first order; then segmenting the flattened node sequence using a context window to obtain a small segment sequence corresponding to each layer of the octree; and performing parallel lossless compression on the small segment sequences corresponding to each layer of the octree in top-to-bottom order to obtain the compressed result of the target point cloud. This invention significantly improves the point cloud compression speed by performing layered parallel compression on the small segment sequences; simultaneously, for any small segment sequence, a self-attention neural network model is used to group and encode the small segment sequence, enabling the encoding to utilize more node context information, thereby improving the compression ratio. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0054] Figure 1 This is a flowchart illustrating the lossless point cloud compression method based on deep learning provided by the present invention.
[0055] Figure 2 This is a schematic diagram of the transformer network structure provided by the present invention;
[0056] Figure 3 This is a schematic diagram of the point cloud lossless compression device based on deep learning provided by the present invention.
[0057] Figure 4This is a schematic diagram of the structure of an electronic device that implements a point cloud lossless compression method based on deep learning, as provided by the present invention. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0059] The following explains abbreviations specific to certain professional fields:
[0060] Voxel: A data structure that uses a fixed-size cube as the smallest unit to represent a three-dimensional object.
[0061] Octree: A tree-like data structure used to describe three-dimensional space. Each node in an octree represents a volume element of a cube, and each node has eight child nodes, corresponding to the eight equal-sized spaces of the cube. The sum of the volume elements represented by the eight child nodes equals the volume of the parent node.
[0062] OctValue: An eight-bit binary number consisting of the eight child nodes of an octree node.
[0063] Compression ratio (bpp): The average number of bits occupied per point after point cloud compression.
[0064] Context Window: The context window.
[0065] bptt: A sequence of small segments divided by a context window.
[0066] The following is combined with Figures 1-4 The present invention describes a deep learning-based lossless point cloud compression method and apparatus.
[0067] Firstly, this invention provides a lossless point cloud compression method based on deep learning, such as... Figure 1 As shown, the method includes:
[0068] S11. Represent the target point cloud using an octree;
[0069] Using an octree to represent the target point cloud can reduce spatial redundancy.
[0070] Specifically, an octree is a tree-like data structure for describing 3D space, where each internal node has eight child nodes. Each node in an octree can represent a space, and the corresponding eight child nodes subdivide this space into eight ternary models. Therefore, an 8-bit binary number consisting of the presence status of the eight child cube points (1 for presence, 0 for absence) is used to abstractly describe the point cloud information of this space.
[0071] S12. Perform a breadth-first traversal on the octree to obtain the node sequence corresponding to each level of the octree.
[0072] Breadth-first traversal, also known as level-order traversal, visits each level sequentially from top to bottom. Within each level, nodes are visited from left to right (or right to left). After visiting one level, the process moves to the next level until no more nodes can be visited.
[0073] S13. Use a context window to separate the node sequence corresponding to each level of the octree to obtain the small segment sequence corresponding to each level of the octree.
[0074] It is understandable that separating the node sequence representing any level of the octree can yield multiple smaller sequences.
[0075] S14. Following the top-to-bottom order of the octree hierarchy, perform parallel lossless compression on the small sequence corresponding to each level of the octree to obtain the compression result of the target point cloud.
[0076] In other words, small sequences belonging to the same octree level are encoded simultaneously, with sequences from octree levels closer to the root node being encoded first, compared to sequences from octree levels farther from the root node. Entropy coding is used when encoding each group of nodes.
[0077] For example, an octree has seven levels, from top to bottom: level 1, level 2, ..., level 7. The small segments separated from the node sequences of each level are compressed in parallel, with the segments separated from the first level node sequence compressed before those separated from the second level node sequence, and so on, until the compressed result of the target point cloud is obtained.
[0078] To accelerate decoding speed, this invention proposes an octree-based hierarchical encoding method, which specifies that small segments of the same level are encoded in parallel, and that information from all octree nodes in the upper K levels can be used when encoding each level.
[0079] Lossless compression of short sequences is achieved by using a self-attention neural network model to encode the short sequences in groups.
[0080] Existing research indicates that information about ancestor nodes and sibling nodes is equally important for encoding the current node.
[0081] Existing compression techniques either employ insufficient contextual information or introduce intolerable decoding complexity (for example, the OctAttention method can only decode one node at a time, requiring a call to the self-attention neural network for each node, resulting in significant time overhead and extremely slow decoding speed). Therefore, this invention proposes a grouping encoding strategy based on a self-attention neural network model. This strategy develops a sufficient and effective context for each node and decodes a group of nodes at a time, reducing decoding complexity and improving decoding speed.
[0082] This invention provides a lossless point cloud compression method based on deep learning, which improves upon the technique of using octrees for point cloud compression. Specifically, it includes: converting the target point cloud into an octree; flattening each level of the octree into a node sequence in breadth-first order; then segmenting the flattened node sequence using a context window to obtain a small segment sequence corresponding to each level of the octree; and performing parallel lossless compression on the small segment sequences corresponding to each level of the octree in top-to-bottom order to obtain the compressed result of the target point cloud. This invention significantly improves the point cloud compression speed by performing layered parallel compression on the small segment sequences; simultaneously, for any small segment sequence, a self-attention neural network model is used to group and encode the small segment sequence, allowing the encoding to utilize more node context information, thereby improving the compression ratio.
[0083] Based on the above embodiments, as an optional embodiment, the step of using a context window to separate the node sequence corresponding to each level of the octree to obtain the small segment sequence corresponding to each level of the octree includes:
[0084] The node sequence corresponding to each level of the octree is divided with the context window length as the length of the segment sequence to obtain the segment sequence corresponding to each level of the octree.
[0085] That is, for the node sequence representing each layer of nodes, zero nodes are first added to the end of the node sequence so that the length of the node sequence is an integer multiple of the length of the context window;
[0086] Then, using the context window length as the length of the small sequence, the node sequence is divided from left to right to obtain multiple small sequences.
[0087] For example, the node sequence is {x1, x2, ... x} 36 The context window can cover 8 nodes, which can be separated into {x1, x2…x8}, {x8, x9…x8}, and {x9…x8}. 16}、{x17 x 18 …x 24}、{x 25 x 26 …x 32}、{x 33 x 34 x 35 x 36 The sequence consists of five segments: 0, 0, 0, 0, 0.
[0088] This operation ensures that each node within a context window forms a small sequence, which is convenient for processing in self-attention neural network models.
[0089] Based on the above embodiments, as an optional embodiment, the lossless compression of small sequences includes:
[0090] Group the nodes in the short sequence;
[0091] This refers to the group number corresponding to each node in the short sequence. It is best to use a uniform distribution method for grouping to ensure that the context information is evenly distributed in space.
[0092] Using a self-attention neural network model, a first mask tensor, a second mask tensor, and node information of each node and its ancestor nodes in the short sequence, the OctValue probability distribution of each node in the short sequence is determined. The node information includes the level of the octree in which the node is located, the node's position within its parent node, and the node's OctValue. The OctValue of a node is an eight-bit binary number representing the presence of points in the node's eight subspaces. The first mask tensor aims to mask the OctValues of all nodes in the short sequence. The second mask tensor aims to mask the OctValues of nodes in the short sequence whose group number is greater than the group number to which the corresponding node belongs during the learning process of the OctValue probability distribution of each node in the short sequence.
[0093] The self-attention neural network model contains two Transformer structures. The first Transformer structure introduces a first mask tensor so that when learning the OctValue probability distribution of each node in the small sequence, the context information used includes the node information of the ancestor node traced by each node in the small sequence, the level of the octree in which each node in the small sequence is located, and the position of each node in the small sequence in the parent node. The OctValue of any node in the small sequence is not used.
[0094] The second Transformer structure introduces a second mask tensor so that when learning the OctValue probability distribution of each node in the small sequence, the context information used includes the node information of the ancestor node traced by each node in the small sequence, the level of the octree in which each node in the small sequence is located, the position of each node in the parent node, and the OctValue of the nodes of the encoded groups in the small sequence (for the p-th group, the encoded groups are the p-1th group to the 1st group). The OctValue of the nodes of the unencoded groups in the small sequence (for the p-th group, the unencoded groups are the p+1th group to the Mth group) is not used.
[0095] Since the block decoding has a specific order, the mask tensor to be added must ensure that the OctValue of later decoded nodes is not used during encoding. The above setting conforms to the decoding rules. The OctValue probability distribution can be learned by summing the inputs of the first Transformer structure and the inputs of the second Transformer structure and passing it through a linear layer.
[0096] Assuming there are 4 groups, the second mask tensor can be as shown in Table 1.
[0097] Table 1
[0098] 1 0 0 2 0 0 0 0 3 0 0 0 0 0 0 4 0 0 0 0 0 0 0 0 1 0 0 2 0 0 0 0 3 0 0 0 0 0 0 4 0 0 0 0 0 0 0 0
[0099] In this case, 0 indicates that the OctValue information can be used, meaning that the later group can use the OctValue information of the earlier group, but the earlier group cannot use the OctValue information of the later group.
[0100] Correspondingly, when represented graphically, 0 corresponds to black, and the rest to white.
[0101] This invention designs two Transformer channels and uses different mask operations to encode the octree in order to achieve efficient decoding and reduce decoding complexity.
[0102] Based on the OctValue probability distribution of each node in the short sequence, the nodes of the short sequence are grouped and encoded in ascending order of group number to obtain the encoding result of the short sequence.
[0103] In practical applications, since the OctValue of each node in a short sequence is known, it is possible to encode each group node simultaneously.
[0104] Decoding is the reverse process of encoding. Since the decoded node group needs to use the OctValue of the previous node group, and the OctValue of the previous node group can only be obtained after the previous node group is decoded, the decoding must be performed strictly in ascending order of group number. Here, the OctValue of the first node group is initialized to 0.
[0105] This invention utilizes a self-attention neural network model to develop sufficient and effective contextual information for each node, thereby improving the prediction performance of OctValue and thus increasing bpp compression.
[0106] Based on the above embodiments, as an optional embodiment, the grouping of nodes in the short sequence includes:
[0107] Set the node indexes of nodes in the same group to be separated by LM;
[0108] Where M is the number of groups selected according to the requirements, and L is an integer greater than or equal to 1.
[0109] For example: Let M be 4, and the existing short sequence is {y1, y2, ... y...} 16 After grouping, {y1, y5, y9, y} 13} belongs to the first group, {y2, y6, y 10 y 14} belongs to the second group, {y3, y7, y 11 y 15} belongs to the third group, {y4, y8, y 12 y 16} belongs to the fourth group,
[0110] This invention employs a uniform grouping method, which ensures that when encoding each node in a short sequence, the sibling nodes with available OctValue information are partly located before the current node and partly located after the current node, and are evenly distributed. This is beneficial for improving node context information and thus improving the compression ratio.
[0111] Based on the above embodiments, as an optional embodiment, determining the OctValue probability distribution of each node in the short sequence using the self-attention neural network model, the first mask tensor, the second mask tensor, and the node information of each node and its ancestor nodes in the short sequence includes:
[0112] The node information of each node and its ancestor nodes in the short sequence is represented in the form of a three-dimensional tensor.
[0113] That is, the size of the three-dimensional tensor can be (N,K,3), where N is the number of nodes contained in the short sequence, K is the number of ancestor nodes traced upwards, and 3 refers to 3 types of node information.
[0114] Based on the first mask tensor and the above, the three-dimensional tensor is processed using the self-attention neural network model to obtain the OctValue probability distribution of each node in the short sequence.
[0115] This invention employs self-attention technology to extract contextual information from octree nodes, enabling octree nodes to utilize more contextual information during encoding and improving compression results.
[0116] Based on the above embodiments, as an optional embodiment, the size of the three-dimensional tensor is (N,K,3), and the neural network module includes an Embedding layer, a first Reshape layer, a first transformer network structure, a second transformer network structure, a second Reshape layer, a first Linear layer, and a SoftMax layer;
[0117] The step of processing the three-dimensional tensor using the self-attention neural network model based on the first mask tensor and the second mask tensor to obtain the OctValue probability distribution of each node in the short sequence includes:
[0118] In the Embedding layer, the size of the three-dimensional tensor is converted to (N,K,S);
[0119] In the first Reshape layer, the three-dimensional tensor of size (N,K,S) is shaped into a two-dimensional tensor of size (N,KS).
[0120] In the first transformer network structure, the first mask tensor is used to learn the node OctValue probability distribution of the two-dimensional tensor to obtain a first learning tensor of size (N,KS).
[0121] In the second transformer network structure, the node OctValue probability distribution of the two-dimensional tensor is learned by using the second mask tensor to obtain a second learning tensor of size (N,KS).
[0122] In the second Reshape layer, the sum of the first learned tensor and the second learned tensor is shaped into a two-dimensional tensor of (N, KS).
[0123] Using the first Linear layer, the two-dimensional tensor output by the second Reshape layer is regularized to obtain a two-dimensional tensor of size (N, 255);
[0124] In the SoftMax layer, a two-dimensional tensor of size (N, 255) is normalized to obtain the OctValue probability distribution of each node in the small sequence from the normalization result; where N is the number of nodes in the small sequence, K is the number of ancestor nodes traced upwards, and S is any integer greater than 3.
[0125] It should be noted that a Positional Encoding layer can be included between the first Reshape layer and the transformer network structure to label the positions of sequence nodes. Between the first Linear layer and the SoftMax layer, a Truncate layer is also included to remove the zero-node portion of the (N, 255) two-dimensional tensor, resulting in a (N0, 255) two-dimensional tensor. This allows the SoftMax layer to normalize the (N0, 255) two-dimensional tensor to obtain the OctValue probability distribution of N0 nodes, where N0 is the number of non-zero nodes in the short sequence segment. The first Reshape layer can be a Multilayer Perceptron (MPL).
[0126] The self-attention neural network model of this invention uses a structure similar to the Transformer encoder. The mask tensor masks out information that cannot be used during the decoding process according to the group decoding order, taking into account context information and decoding complexity, thereby improving compression speed and compression ratio.
[0127] Based on the above embodiments, as an optional embodiment, the first transformer network structure and the second transformer network structure are identical. Figure 2 A schematic diagram of the second transformer network structure is shown below. Figure 2 As shown, it consists of R encoders stacked together;
[0128] The encoder includes a first sublayer connection structure and a second sublayer connection structure that takes the output of the first sublayer connection structure as input.
[0129] The first sub-layer connection structure includes a multi-head self-attention sub-layer, a normalization layer, and a residual connection;
[0130] The second sub-layer connection structure includes a second Linear layer, a normalization layer, and a residual connection;
[0131] In the first transformer network structure, the node OctValue probability distribution is learned by using the first mask tensor on the two-dimensional tensor, resulting in a first learning tensor of size (N, KS), including:
[0132] The two-dimensional tensor is used as the input of the multi-head self-attention sub-layer of the first encoder in the first transformer network structure. The first mask tensor is used as the input mask tensor of the multi-head self-attention sub-layer of each encoder in the first transformer network structure. The two-dimensional tensor is used to learn the node OctValue probability distribution using the first transformer network structure to obtain a first learning tensor of size (N, KS).
[0133] In the second transformer network structure, the node OctValue probability distribution is learned by using the second mask tensor on the two-dimensional tensor, resulting in a second learned tensor of size (N, KS), including:
[0134] The two-dimensional tensor is used as the input of the multi-head self-attention sublayer of the first encoder in the second transformer network structure, and the second mask tensor is used as the input mask tensor of the multi-head self-attention sublayer of each encoder in the second transformer network structure. The two-dimensional tensor is used to learn the node OctValue probability distribution using the second transformer network structure to obtain a second learning tensor of size (N, KS).
[0135] Specifically, the second Linear layer can be a Multilayer Perceptron (MLP). The first and second transformer network structures are identical except for the input mask tensor. The execution process of the second transformer network structure is explained below:
[0136] Because the encoders are stacked in series, the input of the first encoder in the second transformer network structure is the two-dimensional tensor, the output of the last encoder in the second transformer network structure is the output of the second transformer network structure, and the output of the a-th encoder is the input of the (a+1)-th encoder.
[0137] The process by which the first encoder in the second transformer network structure processes the two-dimensional tensor is as follows:
[0138] In the first sub-layer connection structure, the two-dimensional tensor is self-attentionally learned by a multi-head self-attention sub-layer to obtain a weighted two-dimensional tensor.
[0139] The result of the weighted two-dimensional tensor after normalization is added to the two-dimensional tensor through residual connection, and then input into the second sub-layer connection structure;
[0140] In the second sub-layer connection structure, the result of the input after being processed by the second Linear layer and the normalization layer is added to the input through the residual connection. The result of the addition is the output of the first encoder of the second transformer network structure.
[0141] Other encoders are similar and will not be described in detail.
[0142] Specifically, a multi-head self-attention sublayer is used to perform self-attention learning on the two-dimensional tensor to obtain a weighted two-dimensional tensor, including:
[0143] The two-dimensional tensor is split into K sheets, each of which is...
[0144] K two-dimensional tensors are treated as a whole and input into the fourth, fifth, and sixth linear layers. The fourth, fifth, and sixth linear layers can all use multilayer perceptrons (MLPs). MLPs allow feature vectors to fully cross between different dimensions, enabling the network to capture more nonlinear and combined features.
[0145] The transposes of the output matrices of the fourth Linear layer (K heads) and the fifth Linear layer (K heads) are substituted into the first MatMui layer for matrix multiplication, yielding the K heads, each corresponding to a size of... Matrix;
[0146] In the Scale layer, the second mask matrix of size (N,N) is associated with the K heads, each corresponding to a size of... The matrix is multiplied element by element, and the result is passed through the SoftMax layer. Then, in the second MatMui layer, it is multiplied with the output matrix of the sixth Linear layer. The concatenation results in a weighted two-dimensional tensor.
[0147] Here, the multi-head attention mechanism helps the network capture rich feature information, and the residual connection can solve the problem of gradient shrinkage and the problem of weight matrix weakening.
[0148] This invention employs a multi-head attention mechanism to deeply learn the features of the two-dimensional tensor and performs weight allocation based on the mask tensor, thereby improving the learning accuracy of the OctValue probability distribution.
[0149] Based on the above embodiments, as an optional embodiment, the self-attention neural network model is built on the basis of a pre-trained neural network;
[0150] The pre-trained neural network is trained using small sequence samples of octree nodes with their OctValue nodes randomly set to zero.
[0151] The pretraining strategy adopted in this invention is inspired by BERT. It sets the OctValue of the input sample to zero with a 50% probability, uses the original OctValue as the target to train the pretrained neural network to obtain the initial weights of the network, and then performs formal training on this basis. The results show that the self-attention neural network model after pretraining significantly outperforms the self-attention neural network model without pretraining.
[0152] The pre-training operation of this invention improves the learning effect and generalization ability of the self-attention neural network model.
[0153] In summary, the hierarchical grouping point cloud encoding and decoding strategy introduced in this scheme can effectively improve the compression ratio of lossless point cloud compression in deep learning, achieving improvements on both LiDAR and Object datasets compared to the state-of-the-art OctAttention. More importantly, due to the introduction of hierarchical and grouping methods, the point cloud compression speed is reduced by 97% compared to OctAttention, equivalent to 1 / 30th of the original speed. This advancement can promote the practical application of neural networks in point cloud compression.
[0154] Furthermore, building upon this approach, employing a larger context window, adjusting layer parameters, channel count, and attention modules could all potentially achieve even better results. Additionally, in practical applications, parallel implementation of the neural network computation could be considered, significantly improving compression / decompression speed.
[0155] Therefore, complete point cloud compression also includes a decoding part, which is the reverse process of the encoding part. Thus, the decoding process is as follows:
[0156] According to the top-to-bottom order of the octree hierarchy, the small segment sequences corresponding to each level of the octree are decompressed in parallel without loss to obtain the decompressed octree.
[0157] The octree is decompressed and converted into a point cloud structure.
[0158] The lossless decompression of small sequences is achieved by using a self-attention neural network model to decode the small sequences in groups.
[0159] Specifically, lossless decompression of small sequences includes:
[0160] Initialize OctValue to 0;
[0161] The OctValue of the first group of nodes in the short sequence is obtained by decoding the first group of nodes in the short sequence using a self-attention neural network model.
[0162] Based on the OctValue of the first group of nodes in the short sequence, the OctValue of the second group of nodes in the short sequence is obtained by decoding the second group of nodes in the short sequence using a self-attention neural network model.
[0163] Based on the OctValue of the second group of nodes in the short sequence, the OctValue of the third group of nodes in the short sequence is obtained by decoding the third group of nodes in the short sequence using a self-attention neural network model.
[0164] By analogy, the decompression results of the small sequence segments are obtained.
[0165] The lossless point cloud compression device based on deep learning provided by the present invention will be described below. The lossless point cloud compression device based on deep learning described below can be referred to in correspondence with the lossless point cloud compression method based on deep learning described above. Figure 3 A schematic diagram of a point cloud lossless compression device based on deep learning is shown, such as... Figure 3 As shown, it includes:
[0166] Octree representation module 21 is used to represent the target point cloud using an octree;
[0167] The breadth-first traversal module 22 is used to perform a breadth-first traversal on the octree to obtain the node sequence corresponding to each level of the octree;
[0168] Separation module 23 is used to separate the node sequence corresponding to each level of the octree using a context window to obtain the small segment sequence corresponding to each level of the octree;
[0169] Compression module 24 is used to perform parallel lossless compression on the small segment sequence corresponding to each level of the octree in the order from top to bottom of the octree hierarchy, so as to obtain the compression result of the target point cloud;
[0170] Lossless compression of short sequences is achieved by using a self-attention neural network model to encode the short sequences in groups.
[0171] Based on the above embodiments, as an optional embodiment, the separation module is specifically used for:
[0172] The node sequence corresponding to each level of the octree is divided with the context window length as the length of the segment sequence to obtain the segment sequence corresponding to each level of the octree.
[0173] Based on the above embodiments, as an optional embodiment, the compression module specifically includes:
[0174] A grouping unit is used to group the nodes in the short sequence.
[0175] The determining unit is used to determine the OctValue probability distribution of each node in the small segment sequence by using the self-attention neural network model, the first mask tensor, the second mask tensor, and the node information of each node and its ancestor nodes in the small segment sequence.
[0176] A group coding unit is used to perform node group coding on the small segment sequence based on the OctValue probability distribution of each node in the small segment sequence and in ascending order of group number to obtain the coding result of the small segment sequence;
[0177] The node information includes the level of the octree in which the node is located, the position of the node in the parent node, and the OctValue of the node; the OctValue of the node is an eight-bit binary number representing the existence of the node's eight subspace points.
[0178] The first mask tensor aims to mask the OctValue of all nodes in the short sequence; the second mask tensor aims to mask the OctValue of nodes in the short sequence whose group number is greater than the group number to which the corresponding node belongs during the learning process of the OctValue probability distribution of each node in the short sequence.
[0179] Based on the above embodiments, as an optional embodiment, the grouping unit is specifically used for:
[0180] Set the node indexes of nodes in the same group to be separated by LM;
[0181] Where M is the number of groups and L is an integer greater than or equal to 1.
[0182] Based on the above embodiments, as an optional embodiment, the determining unit includes...
[0183] The three-dimensional tensor representation submodule is used to represent the node information of each node and its ancestor nodes in the small sequence in the form of a three-dimensional tensor.
[0184] The processing submodule is used to process the three-dimensional tensor based on the first mask tensor and the second mask tensor using the self-attention neural network model to obtain the OctValue probability distribution of each node in the short sequence.
[0185] Based on the above embodiments, as an optional embodiment, the size of the three-dimensional tensor is (N,K,3), and the neural network module includes an Embedding layer, a first Reshape layer, a first transformer network structure, a second transformer network structure, a second Reshape layer, a first Linear layer, and a SoftMax layer;
[0186] The processing submodule specifically includes:
[0187] Tensor transformation subunit, used to convert the size of the three-dimensional tensor to (N,K,S) in the Embedding layer;
[0188] The first shaping subunit is used to shape a three-dimensional tensor of size (N,K,S) into a two-dimensional tensor of size (N,KS) in the first Reshape layer.
[0189] The first learning subunit is used to learn the node OctValue probability distribution of the two-dimensional tensor using the first mask tensor in the first transformer network structure, so as to obtain a first learning tensor of size (N,KS).
[0190] The second learning subunit is used to learn the node OctValue probability distribution of the two-dimensional tensor using the second mask tensor in the second transformer network structure, so as to obtain a second learning tensor of size (N,KS).
[0191] The second shaping subunit is used to shape the sum of the first learned tensor and the second learned tensor into a two-dimensional tensor of (N, KS) in the second Reshape layer;
[0192] The regularization subunit is used to regularize the two-dimensional tensor output by the second Reshape layer using the first Linear layer to obtain a two-dimensional tensor of size (N, 255).
[0193] The normalization subunit is used to normalize a two-dimensional tensor of size (N, 255) in the SoftMax layer to obtain the OctValue probability distribution of each node in the small sequence from the normalization result; where N is the number of nodes in the small sequence, K is the number of ancestor nodes traced upwards, and S is any integer greater than 3.
[0194] Based on the above embodiments, as an optional embodiment, the first transformer network structure and the second transformer network structure have the same structure, both consisting of R encoders stacked in series;
[0195] The encoder includes a first sublayer connection structure and a second sublayer connection structure connected to it;
[0196] The first sub-layer connection structure includes a multi-head self-attention sub-layer, a normalization layer, and a residual connection;
[0197] The second sub-layer connection structure includes a second Linear layer, a normalization layer, and a residual connection;
[0198] The first learning subunit is specifically used for:
[0199] The two-dimensional tensor is used as the input of the multi-head self-attention sub-layer of the first encoder in the first transformer network structure. The first mask tensor is used as the input mask tensor of the multi-head self-attention sub-layer of each encoder in the first transformer network structure. The two-dimensional tensor is used to learn the node OctValue probability distribution using the first transformer network structure to obtain a first learning tensor of size (N, KS).
[0200] The second learning subunit is specifically used for:
[0201] The two-dimensional tensor is used as the input of the multi-head self-attention sublayer of the first encoder in the second transformer network structure, and the second mask tensor is used as the input mask tensor of the multi-head self-attention sublayer of each encoder in the second transformer network structure. The two-dimensional tensor is used to learn the node OctValue probability distribution using the second transformer network structure to obtain a second learning tensor of size (N, KS).
[0202] Based on the above embodiments, as an optional embodiment, the self-attention neural network model is built on the basis of a pre-trained neural network;
[0203] The pre-trained neural network is trained using small sequence samples of octree nodes with their OctValue nodes randomly set to zero.
[0204] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include a processor 410, a communication interface 420, a memory 430, and a communication bus 440. The processor 410, communication interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a deep learning-based lossless point cloud compression method. This method includes: representing the target point cloud using an octree; performing a breadth-first traversal of the octree to obtain the node sequence corresponding to each level of the octree; separating the node sequence corresponding to each level of the octree using a context window to obtain a small segment sequence corresponding to each level of the octree; and performing parallel lossless compression of the small segment sequence corresponding to each level of the octree in descending order of the octree hierarchy to obtain the compressed result of the target point cloud. The lossless compression of the small segment sequence is achieved by using a self-attention neural network model to group and encode the small segment sequence.
[0205] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0206] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute a deep learning-based lossless point cloud compression method. The method includes: representing the target point cloud using an octree; performing a breadth-first traversal of the octree to obtain the node sequence corresponding to each level of the octree; separating the node sequence corresponding to each level of the octree using a context window to obtain a small segment sequence corresponding to each level of the octree; and performing parallel lossless compression on the small segment sequence corresponding to each level of the octree in descending order of the octree hierarchy to obtain the compression result of the target point cloud. The lossless compression of the small segment sequence is achieved by using a self-attention neural network model to group and encode the small segment sequence. In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a deep learning-based lossless point cloud compression method. The method includes: representing a target point cloud using an octree; performing a breadth-first traversal of the octree to obtain a node sequence corresponding to each level of the octree; separating the node sequence corresponding to each level of the octree using a context window to obtain a small segment sequence corresponding to each level of the octree; and performing parallel lossless compression on the small segment sequence corresponding to each level of the octree in a top-down order according to the octree hierarchy to obtain a compressed result of the target point cloud. The lossless compression of the small segment sequence is achieved by using a self-attention neural network model to group and encode the small segment sequence.
[0207] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0208] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0209] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A lossless point cloud compression method based on deep learning, characterized in that, include: Representing the target point cloud using an octree; Perform a breadth-first traversal on the octree to obtain the node sequence corresponding to each level of the octree; By using a context window to separate the node sequence corresponding to each level of the octree, a small segment sequence corresponding to each level of the octree can be obtained; Following the top-to-bottom order of the octree hierarchy, the small segments corresponding to each level of the octree are compressed in parallel without loss to obtain the compression result of the target point cloud. Lossless compression of short sequences is achieved by using a self-attention neural network model to encode short sequences in groups. The lossless compression of small sequences includes: Group the nodes in the short sequence; Using a self-attention neural network model, a first mask tensor, a second mask tensor, and node information of each node and its ancestor nodes in the short sequence, the OctValue probability distribution of each node in the short sequence is determined. Based on the OctValue probability distribution of each node in the short sequence, the nodes of the short sequence are grouped and encoded in ascending order of group number to obtain the encoding result of the short sequence.
2. The point cloud lossless compression method based on deep learning according to claim 1, characterized in that, The step of using a context window to separate the node sequence corresponding to each level of the octree to obtain a small segment sequence corresponding to each level of the octree includes: The node sequence corresponding to each level of the octree is divided with the context window length as the length of the segment sequence to obtain the segment sequence corresponding to each level of the octree.
3. The point cloud lossless compression method based on deep learning according to claim 1, characterized in that, The node information includes the level of the octree in which the node is located, the position of the node in the parent node, and the OctValue of the node; the OctValue of the node is an eight-bit binary number representing the existence of the node's eight subspace points. The first mask tensor aims to mask the OctValue of all nodes in the short sequence; the second mask tensor aims to mask the OctValue of nodes in the short sequence whose group number is greater than the group number to which the corresponding node belongs during the learning process of the OctValue probability distribution of each node in the short sequence.
4. The point cloud lossless compression method based on deep learning according to claim 3, characterized in that, The grouping of nodes in the short sequence includes: Set the node indexes of nodes in the same group to be separated by LM; Where M is the number of groups and L is an integer greater than or equal to 1.
5. The point cloud lossless compression method based on deep learning according to claim 3 or 4, characterized in that, The step of determining the OctValue probability distribution of each node in the short sequence using a self-attention neural network model, a first mask tensor, a second mask tensor, and node information of each node and its ancestor nodes in the short sequence includes: The node information of each node and its ancestor nodes in the short sequence is represented in the form of a three-dimensional tensor. Based on the first mask tensor and the second mask tensor, the three-dimensional tensor is processed using the self-attention neural network model to obtain the OctValue probability distribution of each node in the short sequence.
6. The point cloud lossless compression method based on deep learning according to claim 5, characterized in that, The size of the three-dimensional tensor is The neural network module includes an Embedding layer, a first Reshape layer, a first transformer network structure, a second transformer network structure, a second Reshape layer, a first Linear layer, and a SoftMax layer; The step of processing the three-dimensional tensor using the self-attention neural network model based on the first mask tensor and the second mask tensor to obtain the OctValue probability distribution of each node in the short sequence includes: The size of the three-dimensional tensor is converted in the Embedding layer. ; In the first Reshape layer The size of the three-dimensional tensor is shaped as A two-dimensional tensor; In the first transformer network structure, the first mask tensor is used to learn the node OctValue probability distribution of the two-dimensional tensor, resulting in a value of... The first learning tensor; In the second transformer network structure, the node OctValue probability distribution of the two-dimensional tensor is learned using the second mask tensor, resulting in a size of... The second learning tensor; In the second Reshape layer, the sum of the first and second learned tensors is shaped into... A two-dimensional tensor; Using the first Linear layer, the two-dimensional tensor output by the second Reshape layer is regularized to obtain a tensor of size [value missing]. A two-dimensional tensor; In the SoftMax layer, for a size of The two-dimensional tensor is normalized to obtain the OctValue probability distribution of each node in the small sequence from the normalization result. in, This represents the number of nodes contained in the short sequence. S represents the number of ancestor nodes traced upwards, where S is any integer greater than 3.
7. The point cloud lossless compression method based on deep learning according to claim 6, characterized in that, The first transformer network structure and the second transformer network structure have the same structure, both consisting of R encoders stacked in series; The encoder includes a first sublayer connection structure and a second sublayer connection structure connected to the first sublayer connection structure; The first sub-layer connection structure includes a multi-head self-attention sub-layer, a normalization layer, and a residual connection; The second sub-layer connection structure includes a second Linear layer, a normalization layer, and a residual connection; In the first transformer network structure, the first mask tensor is used to learn the node OctValue probability distribution of the two-dimensional tensor, resulting in a value of... The first learning tensor includes: The two-dimensional tensor is used as the input to the multi-head self-attention sublayer of the first encoder in the first transformer network structure. The first mask tensor is used as the input mask tensor of the multi-head self-attention sublayer of each encoder in the first transformer network structure. The node OctValue probability distribution is learned using the first transformer network structure to obtain a value of size [missing information]. The first learning tensor; In the second transformer network structure, the node OctValue probability distribution of the two-dimensional tensor is learned using the second mask tensor, resulting in a size of... The second learning tensor includes: The two-dimensional tensor is used as the input to the multi-head self-attention sublayer of the first encoder in the second transformer network structure. The second mask tensor is used as the input mask tensor of the multi-head self-attention sublayer of each encoder in the second transformer network structure. The node OctValue probability distribution is learned using the second transformer network structure to obtain a value of size [missing information]. The second learning tensor.
8. The point cloud lossless compression method based on deep learning according to claim 5, characterized in that, The self-attention neural network model is built on the basis of a pre-trained neural network; The pre-trained neural network is trained using small sequence samples of octree nodes with their OctValue nodes randomly set to zero.
9. A point cloud lossless compression device based on deep learning, characterized in that, To implement the deep learning-based lossless point cloud compression method according to any one of claims 1-8, the method comprises: The octree representation module is used to represent the target point cloud using an octree; The breadth-first traversal module is used to perform a breadth-first traversal on the octree to obtain the node sequence corresponding to each level of the octree; The separation module is used to separate the node sequence corresponding to each level of the octree using a context window to obtain the small segment sequence corresponding to each level of the octree; The compression module is used to perform parallel lossless compression on the small segment sequence corresponding to each level of the octree in order from top to bottom, so as to obtain the compression result of the target point cloud. Lossless compression of short sequences is achieved by using a self-attention neural network model to encode the short sequences in groups.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the point cloud lossless compression method based on deep learning as described in any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the deep learning-based lossless point cloud compression method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
A method and apparatus for processing a sequence model
CN109543824A
Multi-level image compression method using Transform
CN113709455A