Point Cloud Compression Method and Device Based on Multi-Scale Octree Attention Mechanism
Through the point cloud compression method of multi-scale octree attention mechanism, the multi-head attention mechanism and three-dimensional sparse convolution are used to solve the problem of low point cloud compression efficiency, and efficient point cloud data storage and transmission are achieved, while maintaining the quality and details of point clouds.
Patent Information
- Application Number
- CN202510541659.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The prior art is difficult to effectively utilize the geometric space information of point clouds, resulting in low compression efficiency and high storage and transmission costs.
The point cloud compression method based on the multi-scale octree attention mechanism is adopted. By constructing a multi-scale octree representation and multi-head attention mechanism, the spatial context information of point cloud data is captured, the important node features are dynamically paid attention to, and feature extraction and reconstruction are combined with three-dimensional sparse convolution.
Improve point cloud compression efficiency, reduce bit overhead, maintain point cloud quality and details, improve the quality of reconstructed point clouds, and reduce the amount of computing.
Smart Images

Figure CN120075476B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and point cloud compression, and particularly relates to a point cloud compression method and device based on a multi-scale octree attention mechanism. Background Art
[0002] As a data form that can efficiently represent 3D shapes or objects, point clouds are gradually becoming a hot topic in fields such as computer vision, computer graphics, and machine learning. With the gradual improvement of the performance of point cloud acquisition devices, point cloud models with multiple levels of detail can be obtained from three-dimensional scenes, and the scale of point cloud data is getting larger and larger. The large volume of point cloud data has increased the costs of transmission and storage, which has brought burdens and challenges to storage space capacity and network transmission bandwidth. Compression of point cloud data has become one of the key technologies to solve this problem.
[0003] Efficient point cloud compression technology can significantly reduce data storage requirements, reduce transmission bandwidth consumption, and at the same time maintain the quality and details of point cloud data. In recent years, with the introduction of point cloud compression standards and the application of advanced technologies such as deep learning, point cloud compression technology has developed rapidly. However, how to further utilize the geometric spatial information of point clouds to improve the efficiency of point cloud compression remains an important problem in point cloud compression research. Summary of the Invention
[0004] The purpose of the present invention is to provide a point cloud compression method and device based on a multi-scale octree attention mechanism, which can effectively improve the efficiency of point cloud compression and reduce bit overhead on the premise of ensuring the same point cloud quality.
[0005] The present invention adopts the following technical solutions:
[0006] In a first aspect, a point cloud compression method based on a multi-scale octree attention mechanism includes:
[0007] Constructing and training a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; the point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; the decoder network includes an octree decoding module and an upscaling feature reconstructor;
[0008] Input the point cloud data into the trained point cloud compression model. The downscaling feature extractor downsamples and extracts features from the input point cloud data to obtain a downscaled deep feature point cloud. Input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain the octree representation of the point cloud. Input the octree representation of the point cloud into the context construction module for context construction to obtain a context window. Input the context window into the multi-head attention module for calculation to obtain the occupancy probability distribution of the octree nodes. Input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream. Input the bitstream into the octree decoding module to obtain a downscaled reconstructed point cloud. Input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original one.
[0009] Preferably, the downscaling feature extractor includes three sequentially connected downsampling feature extraction modules, namely the first downsampling feature extraction module, the second downsampling feature extraction module, and the third downsampling feature extraction module. Each downsampling feature extraction module includes a first feature extraction layer, a downsampling layer, and a second feature extraction layer connected in sequence.
[0010] Preferably, inputting the point cloud data into the trained point cloud compression model, and the downscaling feature extractor downsamples and extracts features from the input point cloud data to obtain a downscaled deep feature point cloud, specifically including:
[0011] Input the point cloud data into the first downsampling feature extraction module, and perform preliminary feature extraction on the input point cloud through the first feature extraction layer to obtain a feature point cloud after preliminary feature extraction. The calculation process of the first feature extraction layer is as follows:
[0012] ;
[0013] Where, represents the input point cloud data, represents the feature point cloud passing through the first feature extraction layer; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a channel number of C, and a scaling factor of 1; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a channel number of , and a scaling factor of 1; represents the ReLU activation function; represents the concatenation operation;
[0014] Input the feature point cloud into the downsampling layer to obtain a feature point cloud after one downsampling. The calculation process of the downsampling layer is as follows:
[0015] ;
[0016] Among them, represents the feature point cloud after downsampling; represents a three-dimensional sparse convolution with a convolutional kernel of 2×2×2, a number of channels of C, and a scaling factor of ;
[0017] Input the feature point cloud into the second feature extraction layer to obtain the downsampled deep feature point cloud. The calculation process of the second feature extraction layer is as follows:
[0018] ;
[0019] Among them, represents the downsampled deep feature point cloud after passing through the second feature extraction layer;
[0020] Input the downsampled deep feature point cloud into the second downsampled feature extraction module to obtain the downsampled deep feature point cloud after passing through the second downsampled feature extraction module ;
[0021] Input the downsampled deep feature point cloud into the third downsampled feature extraction module to obtain the downsampled deep feature point cloud after passing through the third downsampled feature extraction module , as the downscaled deep feature point cloud.
[0022] Preferably, input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain the octree representation of the point cloud, specifically including:
[0023] According to translate the deep feature point cloud so that its minimum coordinate value is zero; among them, is the downscaled deep feature point cloud; is the deep feature point cloud after translation, is the translation offset; , and respectively represent the X-axis, Y-axis, and Z-axis coordinates of the downscaled deep feature point cloud, , and respectively represent the minimum values of its coordinates;
[0024] According to and the given quantization depth L, quantize the deep feature point cloud after translation; among them, represents the quantized deep feature point cloud; represents the floor operation; quantization step ;
[0025] The quantized deep feature point cloud is recursively divided into eight equal-sized sub-cubes, represented as the child nodes of an octree, and finally the octree representation of the point cloud is obtained .
[0026] Preferably, the octree representation of the point cloud is input into the context construction module for context construction to obtain a context window, specifically including:
[0027] Traverse the octree representation of the point cloud in breadth-first order to obtain a sequence . For each node in the sequence , a context window of length is constructed , which includes the current node and its previous N - 1 sibling nodes, and embeds the K - 1 ancestor nodes of each of these sibling nodes, finally obtaining a context window containing nodes ; where .
[0028] Preferably, the context window is input into the multi-head attention module for calculation to obtain the occupancy probability distribution of the octree nodes, specifically including:
[0029] Input the context window into the multi-head attention module for calculation as follows:
[0030] ;
[0031] where represents the occupancy probability distribution of the octree nodes; represents a multi-layer perceptron; represents a normalization operation; represents the function corresponding to the multi-head attention layer; represents a feature embedding operation; represents the ReLU activation function;
[0032] The calculation process of the multi-head attention layer is as follows:
[0033] ;
[0034] where represents the weighted context window; represents the SoftMax activation function; represents the context window after embedding features; represents a matrix multiplication operation.
[0035] Preferably, the upscaling feature reconstructor includes three upsampling feature reconstruction modules connected in sequence, namely the first upsampling feature reconstruction module, the second upsampling feature reconstruction module, and the third upsampling feature reconstruction module; each upsampling feature reconstruction module includes a first feature reconstruction layer, an upsampling layer, and a second feature reconstruction layer connected in sequence.
[0036] Preferably, the downscaled reconstructed point cloud is input into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original, specifically including:
[0037] Input the downscaled reconstructed point cloud into the first upsampling feature reconstruction module, and perform preliminary feature reconstruction on the downscaled reconstructed point cloud through the first feature reconstruction layer to obtain a feature point cloud after preliminary feature reconstruction. The calculation process of the first feature reconstruction layer is as follows:
[0038] ;
[0039] Among them, represents the downscaled reconstructed point cloud; represents the feature point cloud passing through the first feature reconstruction layer; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a channel number of C, and a scaling factor of 1; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a channel number of , and a scaling factor of 1; represents the ReLU activation function; represents the concatenation operation;
[0040] Input the feature point cloud into the upsampling layer to obtain a feature point cloud after one upsampling. The calculation process of the upsampling layer is as follows:
[0041] ;
[0042] Among them, represents the feature point cloud after upsampling, represents a three-dimensional transposed sparse convolution with a convolution kernel of 2×2×2, a channel number of C, and a scaling factor of ;
[0043] Input the feature point cloud into the second feature reconstruction layer to obtain a deep feature point cloud after upsampling. The calculation process of the second feature reconstruction layer is as follows:
[0044] ;
[0045] Among them, Denote the upsampled deep feature point cloud after passing through the first upsampling feature reconstruction module;
[0046] Input the feature point cloud into the second upsampling feature reconstruction module to obtain the upsampled deep feature point cloud after passing through the second upsampling feature reconstruction module ;
[0047] Input the upsampled deep feature point cloud into the third upsampling feature reconstruction module to obtain the upsampled deep feature point cloud after passing through the third upsampling feature reconstruction module , which serves as the reconstructed point cloud consistent with the original resolution .
[0048] In a second aspect, a point cloud compression device based on a multi-scale octree attention mechanism includes:
[0049] A model construction module configured to construct and train a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; the point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; the decoder network includes an octree decoding module and an upscaling feature reconstructor;
[0050] A point cloud compression and reconstruction module configured to input point cloud data into the trained point cloud compression model, the downscaling feature extractor performs downsampling and feature extraction on the input point cloud data to obtain a downscaled deep feature point cloud; input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain an octree representation of the point cloud; input the octree representation of the point cloud into the context construction module for context construction to obtain a context window; input the context window into the multi-head attention module for calculation to obtain the occupancy probability distribution of octree nodes; input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream; input the bitstream into the octree decoding module to obtain a downscaled reconstructed point cloud; input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud consistent with the original resolution.
[0051] In a third aspect, an electronic device includes:
[0052] One or more processors;
[0053] A storage device for storing one or more programs;
[0054] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the point cloud compression methods based on the multi-scale octree attention mechanism.
[0055] In a fourth aspect, a computer-readable storage medium stores a computer program which, when executed by a processor, implements any of the point cloud compression methods based on the multi-scale octree attention mechanism.
[0056] In a fifth aspect, a computer program product includes a computer program which, when executed by a processor, implements any of the point cloud compression methods based on the multi-scale octree attention mechanism.
[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0058] (1) The point cloud compression method based on the multi-scale octree attention mechanism proposed by the present invention adopts a strategy based on the octree attention mechanism, which can effectively capture the spatial context information in the point cloud data. By constructing an octree representation and using the multi-head attention mechanism to fuse the octree node features, the model can more accurately predict the occupancy probability of the nodes, thereby improving the compression efficiency. The attention mechanism can also dynamically focus on the node features that have a greater impact on the compression performance, further enhancing the compression ability of the model;
[0059] (2) The downscaling feature extractor and upscaling feature reconstructor in the point cloud compression method based on the multi-scale octree attention mechanism proposed by the present invention gradually extract or restore the multi-scale features of the point cloud through multi-level downsampling feature extraction or upsampling feature reconstruction operations. This multi-scale processing method helps the model better capture the detailed information and global structure of the point cloud, thereby retaining more useful information during the compression process and improving the quality of the reconstructed point cloud;
[0060] (3) The three-dimensional sparse convolution adopted by the point cloud compression method based on the multi-scale octree attention mechanism proposed by the present invention can effectively handle the irregularity and sparsity of the point cloud data while maintaining high computational efficiency. Compared with traditional dense convolution operations, three-dimensional sparse convolution can greatly reduce the computational amount and improve the running efficiency of the model. Sparse convolution can also better adapt to the spatial distribution of the point cloud data and extract more representative features, further enhancing the compression performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a schematic flow chart of the point cloud compression method based on the multi-scale octree attention mechanism according to an embodiment of the present invention;
[0062] Figure 2Schematic diagram of the point cloud compression model of the point cloud compression method based on the multi-scale octree attention mechanism according to the embodiment of the present invention;
[0063] Figure 3 Schematic diagram of the feature extraction layer of the point cloud compression method based on the multi-scale octree attention mechanism according to the embodiment of the present invention;
[0064] Figure 4 Schematic diagram of the multi-head attention module of the point cloud compression method based on the multi-scale octree attention mechanism according to the embodiment of the present invention;
[0065] Figure 5 Schematic diagram of the multi-head attention layer of the point cloud compression method based on the multi-scale octree attention mechanism according to the embodiment of the present invention;
[0066] Figure 6 Block diagram of the structure of the point cloud compression device based on the multi-scale octree attention mechanism according to the embodiment of the present invention;
[0067] Figure 7 Schematic diagram of the hardware structure of the electronic device according to the embodiment of the present invention. Detailed implementation manners
[0068] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0069] Refer to Figure 1 As shown, a point cloud compression method based on the multi-scale octree attention mechanism in this embodiment includes the following steps.
[0070] S101. Construct a point cloud compression model based on the multi-scale octree attention mechanism and train it to obtain a trained point cloud compression model; the point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; the decoder network includes an octree decoding module and an upscaling feature reconstructor.
[0071] Specifically, refer to Figure 2As shown in the figure, the point cloud compression model proposed in the embodiment of the present invention has two parts: an encoder network and a decoder network, which are respectively used to encode point cloud data into a bitstream and decode the encoded bitstream into a reconstructed point cloud. First, the input point cloud is downsampled and feature-extracted by a downscaling feature extractor to obtain a downscaled deep feature point cloud; an octree is constructed recursively to obtain an octree representation of the point cloud; a context construction module is designed to obtain a context window for octree nodes; by combining the multi-head attention mechanism and a multi-layer perceptron, the occupancy probability of octree nodes is obtained; an octree encoding module is designed to encode the node occupancy probability into a bitstream; an octree decoding module is designed to decode the bitstream into a reconstructed point cloud; and the reconstructed point cloud is upsampled and feature-reconstructed by an upscaling feature reconstructor to obtain a reconstructed point cloud with the same resolution as the original point cloud, which can effectively achieve point cloud compression.
[0072] S102, input the point cloud data into the trained point cloud compression model. The downscaling feature extractor downsamples and feature extracts the input point cloud data to obtain a downscaled deep feature point cloud; input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain an octree representation of the point cloud; input the octree representation of the point cloud into the context construction module for context construction to obtain a context window; input the context window into the multi-head attention module for calculation to obtain the occupancy probability distribution of octree nodes; input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream; input the bitstream into the octree decoding module to obtain a downscaled reconstructed point cloud; input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original.
[0073] In a specific embodiment, the downscaling feature extractor includes three sequentially connected downsampling feature extraction modules, namely a first downsampling feature extraction module, a second downsampling feature extraction module, and a third downsampling feature extraction module; each downsampling feature extraction module includes a first feature extraction layer, a downsampling layer, and a second feature extraction layer connected in sequence.
[0074] Input the point cloud data into the trained point cloud compression model. The downscaling feature extractor downsamples and feature extracts the input point cloud data to obtain a downscaled deep feature point cloud, as follows.
[0075] Input the point cloud data into the first downsampling feature extraction module. The first feature extraction layer preliminarily extracts features from the input point cloud data to obtain a feature point cloud after preliminary feature extraction, as Figure 3 shown. The calculation process of the first feature extraction layer is as follows:
[0076] ;
[0077] Among them, represents the input point cloud data, represents the feature point cloud passing through the first feature extraction layer, represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of 1, represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2 and a number of channels of , and a scaling factor of 1, represents the ReLU activation function, represents the concatenation operation.
[0078] Input the feature point cloud after the preliminary feature extraction into the downsampling layer to obtain the feature point cloud after one downsampling. The calculation process of the downsampling layer is as follows:
[0079] ;
[0080] Among them, represents the feature point cloud after downsampling, represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of .
[0081] Input the downsampled feature point cloud into the second feature extraction layer to obtain the downsampled deep feature point cloud. The calculation process of the second feature extraction layer is as follows:
[0082] ;
[0083] Among them, represents the downsampled deep feature point cloud after passing through the second feature extraction layer.
[0084] Input the downsampled deep feature point cloud after the first downsampling feature extraction module into the second downsampling feature extraction module to obtain the downsampled deep feature point cloud after passing through the second downsampling feature extraction module .
[0085] Input the downsampled deep feature point cloud after passing through the second downsampling feature extraction module into the third downsampling feature extraction module to obtain the downsampled deep feature point cloud after passing through the third downsampling feature extraction module .
[0086] In a specific embodiment, input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain the octree representation of the point cloud, as follows.
[0087] According to Translate the deep feature point cloud so that its minimum coordinate value is zero; where, is the downscaled deep feature point cloud; is the translated deep feature point cloud, is the translation offset; , and respectively represent the X-axis, Y-axis, and Z-axis coordinates of the downscaled deep feature point cloud, , and respectively represent the minimum values of their coordinates.
[0088] According to and the given quantization depth L, quantize the translated deep feature point cloud, represents the quantized deep feature point cloud, represents the floor operation, and the quantization step .
[0089] Divide the quantized deep feature point cloud recursively into eight equal-sized sub-cubes, represented as the child nodes of the octree, and finally obtain the octree representation of the point cloud . Among them, the occupancy status of each sub-cube constitutes an eight-bit binary occupancy code. An empty sub-cube is marked as 0, and a non-empty sub-cube is marked as 1 and further subdivided until the quantization depth L is reached and the division terminates; at the leaf node, an eight-bit occupancy code represents eight small cubes with side length Q, and the points in the point cloud are merged into the nearest corresponding small cubes.
[0090] In a specific embodiment, input the octree representation of the point cloud into the context construction module for context construction to obtain a context window, specifically as follows.
[0091] Traverse the octree representation of the point cloud in breadth-first order to obtain a sequence , for each node of the sequence , construct a context window of length N, which includes the current node and its previous N - 1 sibling nodes, and embed the K - 1 ancestor nodes of each of these sibling nodes, and finally obtain a context window containing nodes. Among them, , the length N of the context window can be set according to different point cloud inputs and is default set to 1024.
[0092] In a specific embodiment, the context window is input into the multi-head attention module for calculation to obtain the occupancy probability distribution of the octree nodes, as follows.
[0093] Input the context window into the multi-head attention module. As Figure 4 shown, first, perform feature embedding on the input context window to map it into a high-dimensional feature vector, obtaining the context window after embedding features; then, through the multi-head attention layer, perform global interaction on the embedded features to obtain a weighted context window that fuses global information; perform a normalization operation on it to obtain a normalized weighted context window; then, through a multi-layer perceptron, obtain the weighted context window after linear transformation; finally, introduce non-linearity through an activation function to obtain the occupancy probability of the output node.
[0094] The specific calculation process is as follows:
[0095] ;
[0096] Among them, represents the occupancy probability distribution of the octree nodes, represents the multi-layer perceptron, represents the normalization operation, represents the function corresponding to the multi-head attention layer, represents the feature embedding operation.
[0097] The multi-head attention layer is as Figure 5 shown. First, perform a normalization operation on the context window after embedding features to obtain a normalized context window after feature embedding; it performs a linear transformation through 3 parallel multi-layer perceptrons to generate a query vector, a key vector, and a value vector respectively; perform a matrix multiplication operation on the query vector and the key vector, and then apply an activation function for non-linearity to obtain the attention weights; finally, perform a matrix multiplication operation on the attention weights and the value vector to finally obtain the weighted context window.
[0098] The specific calculation process is as follows:
[0099] ;
[0100] Among them, represents the weighted context window, represents the SoftMax activation function, represents the context window after embedding features, represents the matrix multiplication operation.
[0101] In a specific embodiment, the upscaling feature reconstructor includes three upsampling feature reconstruction modules connected in sequence, namely a first upsampling feature reconstruction module, a second upsampling feature reconstruction module, and a third upsampling feature reconstruction module; each upsampling feature reconstruction module includes a first feature reconstruction layer, an upsampling layer, and a second feature reconstruction layer connected in sequence.
[0102] Input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original, as follows.
[0103] Input the downscaled reconstructed point cloud into the first upsampling feature reconstruction module, and perform preliminary feature reconstruction on the downscaled reconstructed point cloud through the first feature reconstruction layer to obtain a feature point cloud after preliminary feature reconstruction. The calculation process of the first feature reconstruction layer is as follows:
[0104] ;
[0105] Wherein, represents the downscaled reconstructed point cloud, represents the feature point cloud passing through the first feature reconstruction layer.
[0106] Input the feature point cloud after preliminary feature reconstruction into the upsampling layer to obtain a feature point cloud after one upsampling. The calculation process of the upsampling layer is as follows:
[0107] ;
[0108] Wherein, represents the feature point cloud after upsampling, represents a 3D transposed sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of .
[0109] Input the feature point cloud after upsampling into the second feature reconstruction layer to obtain a deep feature point cloud after upsampling. The calculation process of the second feature reconstruction layer is as follows:
[0110] ;
[0111] Wherein, represents the deep feature point cloud after upsampling through the first upsampling feature reconstruction module.
[0112] Input the deep feature point cloud after upsampling through the first upsampling feature reconstruction module into the second upsampling feature reconstruction module to obtain a deep feature point cloud after upsampling through the second upsampling feature reconstruction module .
[0113] The upsampled deep feature point cloud after the second upsampling feature reconstruction module is input into the third upsampling feature reconstruction module to obtain the upsampled deep feature point cloud after the third upsampling feature reconstruction module .
[0114] The upsampled deep feature point cloud after the third upsampling feature reconstruction module is the reconstructed point cloud consistent with the original resolution .
[0115] In a specific embodiment, the experimental environment used includes a workstation equipped with an Intel(R) Xeon(R) Gold 6226R processor (2.90 GHz), equipped with an NVIDIA RTX 3090 graphics card (24 GB video memory) and 128 GB of DDR4 memory, and the operating system is Ubuntu 20.04 LTS; the PyTorch deep learning framework is used in the experiment, and CUDA 11.8 acceleration is enabled.
[0116] In a specific embodiment, the selected dataset is SemanticKITTI, which is a large-scale LiDAR point cloud dataset for autonomous driving. It is obtained by scanning with a Velodyne HDL-64E sensor and contains a total of 45.49 million points; sequences 00 to 10 are selected as the training set, and sequences 11 to 21 are selected as the test set.
[0117] In a specific embodiment, the Adam optimizer is used during the training process. The initial learning rate is set to 1e-4, and the StepLR strategy is adopted to decay the learning rate to 0.1 of the original value every 10 epochs; the batch size is set to 32, and the number of training epochs is 100. All training is repeated under the same random seed to ensure the stability and reproducibility of the results; the loss function for training uses cross-entropy loss, as shown below:
[0118] ;
[0119] wherein represents the loss function, represents the predicted value of the node occupancy probability, represents the true value of the node occupancy probability.
[0120] Specifically, as shown in Figure 6 As an implementation of the methods shown in the above figures, an embodiment of a point cloud compression device based on a multi-scale octree attention mechanism is provided in this application. This device embodiment is related to Figure 1The method embodiments shown correspond to a device that can be specifically applied to various electronic devices.
[0121] A point cloud compression device based on a multi-scale octree attention mechanism, comprising:
[0122] A model construction module 601, configured to construct and train a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; the point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; the decoder network includes an octree decoding module and an upscaling feature reconstructor;
[0123] A point cloud compression and reconstruction module 602, configured to input point cloud data into the trained point cloud compression model, the downscaling feature extractor performs downsampling and feature extraction on the input point cloud data to obtain a downscaled deep feature point cloud; input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain an octree representation of the point cloud; input the octree representation of the point cloud into the context construction module for context construction to obtain a context window; input the context window into the multi-head attention module for calculation to obtain an occupancy probability distribution of octree nodes; input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream; input the bitstream into the octree decoding module to obtain a downscaled reconstructed point cloud; input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original.
[0124] The specific implementation of each module of a point cloud compression device based on a multi-scale octree attention mechanism is the same as that of a point cloud compression method based on a multi-scale octree attention mechanism, and this embodiment will not be repeated here.
[0125] See Figure 7 The following shows a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Figure 7 In this embodiment, the electronic device includes: a processor 701 and a memory 702; wherein the memory 702 is used to store computer execution instructions; the processor 701 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0126] Optionally, the memory 702 can be either independent or integrated with the processor 701.
[0127] When the memory 702 is independently provided, the electronic device further includes a bus 703 for connecting the memory 702 and the processor 701.
[0128] An embodiment of the present invention further provides a computer storage medium, in which computer-executable instructions are stored. When the processor 701 executes the computer-executable instructions, the above method is implemented.
[0129] An embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 701, the above method is implemented.
[0130] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.
[0131] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0132] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a hardware plus software functional unit.
[0133] The above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or the processor 701 to execute some steps of the methods in various embodiments of the present application.
[0134] It should be understood that the above-mentioned processor 701 can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor can be a microprocessor or the processor 701 can also be any conventional processor 701, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by the hardware processor 701, or executed and completed by a combination of the hardware and software modules in the processor 701.
[0135] The memory 702 may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0136] The bus 703 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus 703 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus 703 in the attached drawings of this application is not limited to only one bus 703 or one type of bus 703.
[0137] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0138] An exemplary storage medium is coupled to a processor 701, enabling the processor 701 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 701. The processor 701 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 701 and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0139] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A point cloud compression method based on a multi-scale octree attention mechanism, characterized in that Including: Construct and train a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; The point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; The decoder network includes an octree decoding module and an upscaling feature reconstructor; Input the point cloud data into the trained point cloud compression model. The downscaling feature extractor performs downsampling and feature extraction on the input point cloud data to obtain a downscaled deep feature point cloud; Input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain an octree representation of the point cloud; input the octree representation of the point cloud into the context construction module for context construction to obtain a context window; Input the context window into the multi-head attention module for calculation to obtain an occupancy probability distribution of the octree nodes; input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream; input the bitstream into the octree decoding module to obtain a downscaled reconstructed point cloud; input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original; Input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain an octree representation of the point cloud, specifically including: According to translate the deep feature point cloud so that its minimum coordinate value is zero; where is the downscaled deep feature point cloud; F t is the translated deep feature point cloud, is the translation offset; and respectively represent the X-axis, Y-axis, and Z-axis coordinates of the downscaled deep feature point cloud, and respectively represent the minimum values of its coordinates; According to and the given quantization depth L, quantize the translated deep feature point cloud; where F Q represents the quantized deep feature point cloud; round represents the floor operation; the quantization step Quantize the deep feature point cloud F Q Recursively divide it into eight equal-sized sub-cubes, which are represented as the child nodes of the octree, and finally obtain the octree representation F of the point cloud O ; Input the context window into the multi-head attention module for calculation to obtain an occupancy probability distribution of the octree nodes, specifically including: Input the context window into the multi-head attention module for calculation as follows: Among them, P represents the occupancy probability distribution of the octree node; MLP represents the multi-layer perceptron; Norm represents the normalization operation; MA(·) represents the function corresponding to the multi-head attention layer; Emb represents the feature embedding operation; represents the ReLU activation function; The calculation process of the multi-head attention layer is as follows: Among them, represents the weighted context window; represents the SoftMax activation function; represents the context window after embedding features; represents the matrix multiplication operation.
2. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 1, characterized in that The downscaling feature extractor includes three sequentially connected downsampling feature extraction modules, namely the first downsampling feature extraction module, the second downsampling feature extraction module, and the third downsampling feature extraction module; each downsampling feature extraction module includes a first feature extraction layer, a downsampling layer, and a second feature extraction layer connected in sequence.
3. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 2, wherein Input the point cloud data into the trained point cloud compression model. The downscaling feature extractor performs downsampling and feature extraction on the input point cloud data to obtain a downscaled deep feature point cloud, specifically including: Input the point cloud data into the first downsampling feature extraction module. The first feature extraction layer performs preliminary feature extraction on the input point cloud to obtain a feature point cloud after preliminary feature extraction. The calculation process of the first feature extraction layer is as follows: Among them, X represents the input point cloud data, represents the feature point cloud passing through the first feature extraction layer; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of 1; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2 and a number of channels of and a scaling factor of 1; represents the ReLU activation function; concat represents the concatenation operation; Input the feature point cloud into the downsampling layer to obtain the feature point cloud after one downsampling. The calculation process of the downsampling layer is as follows: Among them, represents the feature point cloud after downsampling; represents a three-dimensional sparse convolution with a convolutional kernel of 2×2×2, a number of channels of C, and a scaling factor of 2 3 ↓. Input the feature point cloud into the second feature extraction layer to obtain the downsampled deep feature point cloud. The calculation process of the second feature extraction layer is as follows: Among them, represents the downsampled deep feature point cloud after passing through the second feature extraction layer; Input the downsampled deep feature point cloud into the second downsampling feature extraction module to obtain the downsampled deep feature point cloud after passing through the second downsampling feature extraction module Input the downsampled deep feature point cloud into the third downsampling feature extraction module, and obtain the downsampled deep feature point cloud after passing through the third downsampling feature extraction module as the deep feature point cloud with reduced scale.
4. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 1, wherein Input the octree representation of the point cloud into the context construction module for context construction to obtain a context window, specifically including: Traverse the octree representation F of the point cloud in breadth-first order O to obtain a sequence For each node n in the sequence i , construct a context window of length N {n i-N+1 ,..., n i-1 , n i}, which includes the current node n i and its previous N - 1 sibling nodes, and embed the K - 1 ancestor nodes of each of these sibling nodes, finally obtaining a context window containing N × K nodes where i ∈ [0, N].
5. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 1, wherein The upscaling feature reconstructor includes three sequentially connected upsampling feature reconstruction modules, namely the first upsampling feature reconstruction module, the second upsampling feature reconstruction module, and the third upsampling feature reconstruction module; each upsampling feature reconstruction module includes a first feature reconstruction layer, an upsampling layer, and a second feature reconstruction layer connected in sequence.
6. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 5, characterized in that, Input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original, specifically including: Input the downscaled reconstructed point cloud into the first upsampling feature reconstruction module. Perform preliminary feature reconstruction on the downscaled reconstructed point cloud through the first feature reconstruction layer to obtain the feature point cloud after preliminary feature reconstruction. The calculation process of the first feature reconstruction layer is as follows: Among them, X0 represents the downscaled reconstructed point cloud; represents the feature point cloud passing through the first feature reconstruction layer; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of 1; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2 and a number of channels of a scaling factor of 1; represents the ReLU activation function; concat represents the concatenation operation; Input the feature point cloud into the upsampling layer to obtain the feature point cloud after one upsampling. The calculation process of the upsampling layer is as follows: Among them, represents the feature point cloud after upsampling, represents a three-dimensional transposed sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of 2 3 ↑. Input the feature point cloud into the second feature reconstruction layer to obtain the upsampled deep feature point cloud. The calculation process of the second feature reconstruction layer is as follows: Among them, represents the upsampled deep feature point cloud after passing through the first upsampling feature reconstruction module; Input the feature point cloud into the second upsampling feature reconstruction module to obtain the upsampled deep feature point cloud after passing through the second upsampling feature reconstruction module Input the upsampled deep feature point cloud into the third upsampling feature reconstruction module, and obtain the upsampled deep feature point cloud after passing through the third upsampling feature reconstruction module as the reconstructed point cloud consistent with the original resolution 7. A point cloud compression device based on a multi-scale octree attention mechanism, characterized in that A point cloud compression method based on a multi-scale octree attention mechanism according to any one of claims 1 to 6, comprising: A model construction module configured to construct and train a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; the point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; the decoder network includes an octree decoding module and an upscaling feature reconstructor; A point cloud compression and reconstruction module configured to input point cloud data into the trained point cloud compression model. The downscaling feature extractor performs downsampling and feature extraction on the input point cloud data to obtain a downscaled deep feature point cloud; input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain an octree representation of the point cloud; input the octree representation of the point cloud into the context construction module for context construction to obtain a context window; input the context window into the multi-head attention module for calculation to obtain an occupancy probability distribution of octree nodes; input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream; input the bitstream into the octree decoding module to obtain a downscaled reconstructed point cloud; input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original.
8. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Point cloud compression and decompression method based on octree coding and voxel context
CN113284203A
Point cloud coding and decoding method based on double octree structure
CN117692662A