Point cloud compression method and device based on multi-scale octree attention mechanism

Through the point cloud compression method based on the multi-scale octree attention mechanism, the multi-head attention mechanism is used to integrate node characteristics, and the problem of low point cloud compression efficiency in the existing technology is solved, and efficient point cloud data compression and reconstruction is achieved.

CN120075476AActive Publication Date: 2025-05-30HUAQIAO UNIVERSITY

Patent Information

Application Number
CN202510541659.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize the geometric spatial information of point clouds, resulting in low compression efficiency of point clouds and unable to effectively reduce the cost of data storage and transmission.

Method used

The point cloud compression method based on the multi-scale octree attention mechanism is adopted. By constructing a point cloud compression model, the multi-head attention mechanism is used to fusion the octree node features, and dynamically pay attention to the node features that have a greater impact on compression performance, thereby improving compression efficiency.

Benefits of technology

On the premise of ensuring the quality of point cloud, the efficiency of point cloud compression is significantly improved, bit overhead is reduced, and data storage and transmission efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075476A_ABST
    Figure CN120075476A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud compression method and device based on a multi-scale octree attention mechanism, and relates to the field of image processing, and the method comprises the steps: an encoder network receives point cloud data, carries out the down-sampling and feature extraction of a point cloud through a downscaling feature extractor, obtains a downscaled deep feature point cloud, and carries out the feature extraction of the point cloud; the method comprises the following steps of: firstly, encoding the data into an octree in a recursive manner, constructing a context window according to a relationship among octree nodes, introducing a multi-head attention mechanism to perform feature fusion on the octree nodes to obtain an occupancy probability of the octree nodes, and compressing the occupancy probability into a bit stream by using arithmetic encoding; and the decoder network decompresses the bit stream to obtain a reconstructed point cloud, upsampling and feature reconstruction are performed on the reconstructed point cloud by using an upscaling feature reconstruction device, and finally a reconstructed point cloud with the same resolution as the initial point cloud is obtained. According to the invention, on the premise of ensuring the quality of the same point cloud, the point cloud compression efficiency is effectively improved, and the bit overhead is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and point cloud compression, and particularly relates to a point cloud compression method and device based on a multi-scale octree attention mechanism. Background Art

[0002] As a data form that can efficiently represent 3D shapes or objects, point clouds are gradually becoming a hot topic in fields such as computer vision, computer graphics, and machine learning. With the gradual improvement of the performance of point cloud acquisition devices, point cloud models with multiple levels of detail can be obtained from three-dimensional scenes, and the scale of point cloud data is getting larger and larger. The large volume of point cloud data has increased the costs of transmission and storage, which has brought burdens and challenges to storage space capacity and network transmission bandwidth. Compression of point cloud data has become one of the key technologies to solve this problem.

[0003] Efficient point cloud compression technology can significantly reduce data storage requirements, reduce transmission bandwidth consumption, and at the same time maintain the quality and details of point cloud data. In recent years, with the introduction of point cloud compression standards and the application of advanced technologies such as deep learning, point cloud compression technology has developed rapidly. However, how to further utilize the geometric space information of point clouds to improve the efficiency of point cloud compression remains an important problem in point cloud compression research. Summary of the Invention

[0004] The purpose of the present invention is to provide a point cloud compression method and device based on a multi-scale octree attention mechanism, which can effectively improve the efficiency of point cloud compression and reduce bit overhead on the premise of ensuring the same point cloud quality.

[0005] The present invention adopts the following technical solutions:

[0006] In a first aspect, a point cloud compression method based on a multi-scale octree attention mechanism includes:

[0007] Constructing and training a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; the point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; the decoder network includes an octree decoding module and an upscaling feature reconstructor;

[0008] Input the point cloud data into the trained point cloud compression model. The downscaling feature extractor performs downsampling and feature extraction on the input point cloud data to obtain a downscaled deep feature point cloud. Input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain the octree representation of the point cloud. Input the octree representation of the point cloud into the context construction module for context construction to obtain a context window. Input the context window into the multi-head attention module for calculation to obtain the occupancy probability distribution of the octree nodes. Input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream. Input the bitstream into the octree decoding module to obtain a downscaled reconstructed point cloud. Input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original one.

[0009] Preferably, the downscaling feature extractor includes three sequentially connected downsampling feature extraction modules, namely the first downsampling feature extraction module, the second downsampling feature extraction module, and the third downsampling feature extraction module. Each downsampling feature extraction module includes a first feature extraction layer, a downsampling layer, and a second feature extraction layer connected in sequence.

[0010] Preferably, inputting the point cloud data into the trained point cloud compression model, and the downscaling feature extractor performing downsampling and feature extraction on the input point cloud data to obtain a downscaled deep feature point cloud specifically includes:

[0011] Input the point cloud data into the first downsampling feature extraction module, and perform preliminary feature extraction on the input point cloud through the first feature extraction layer to obtain a feature point cloud after preliminary feature extraction. The calculation process of the first feature extraction layer is as follows:

[0012] ;

[0013] Wherein, represents the input point cloud data, represents the feature point cloud passing through the first feature extraction layer; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a channel number of C, and a scaling factor of 1; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a channel number of , and a scaling factor of 1; represents the ReLU activation function; represents the concatenation operation;

[0014] Input the feature point cloud into the downsampling layer to obtain a feature point cloud after one downsampling. The calculation process of the downsampling layer is as follows:

[0015] ;

[0016] Among them, represents the feature point cloud after downsampling; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of ;

[0017] Input the feature point cloud into the second feature extraction layer to obtain the downsampled deep feature point cloud. The calculation process of the second feature extraction layer is as follows:

[0018] ;

[0019] Among them, represents the downsampled deep feature point cloud after passing through the second feature extraction layer;

[0020] Input the downsampled deep feature point cloud into the second downsampling feature extraction module to obtain the downsampled deep feature point cloud after passing through the second downsampling feature extraction module ;

[0021] Input the downsampled deep feature point cloud into the third downsampling feature extraction module to obtain the downsampled deep feature point cloud after passing through the third downsampling feature extraction module , as the deep feature point cloud with reduced scale.

[0022] Preferably, input the deep feature point cloud with reduced scale into the octree construction module for quantization processing to obtain the octree representation of the point cloud, specifically including:

[0023] According to translate the deep feature point cloud so that its minimum coordinate value is zero; among them, is the deep feature point cloud with reduced scale; is the deep feature point cloud after translation, is the translation offset; , and respectively represent the X-axis, Y-axis, and Z-axis coordinates of the deep feature point cloud with reduced scale, , and respectively represent the minimum values of its coordinates;

[0024] According to and the given quantization depth L, quantize the deep feature point cloud after translation; among them, represents the quantized deep feature point cloud; represents the floor operation; quantization step ;

[0025] The quantized deep feature point cloud is recursively divided into eight equal-sized sub-cubes, represented as the child nodes of an octree, and finally the octree representation of the point cloud is obtained .

[0026] Preferably, the octree representation of the point cloud is input into the context construction module for context construction to obtain a context window, specifically including:

[0027] Traverse the octree representation of the point cloud in breadth-first order to obtain a sequence , and for each node in the sequence , construct a context window with a length of , which includes the current node and its first N - 1 sibling nodes, and embeds the K - 1 ancestor nodes of each of these sibling nodes, finally obtaining a context window containing nodes ; where . .

[0028] Preferably, the context window is input into the multi-head attention module for calculation to obtain the occupancy probability distribution of the octree nodes, specifically including:

[0029] Input the context window into the multi-head attention module for calculation as follows:

[0030] ;

[0031] where represents the occupancy probability distribution of the octree nodes; represents a multi-layer perceptron; represents a normalization operation; represents the function corresponding to the multi-head attention layer; represents a feature embedding operation; represents the ReLU activation function;

[0032] The calculation process of the multi-head attention layer is as follows:

[0033] ;

[0034] where represents the weighted context window; represents the SoftMax activation function; represents the context window after embedding features; represents a matrix multiplication operation.

[0035] Preferably, the upscaling feature reconstructor includes three upsampling feature reconstruction modules connected in sequence, namely a first upsampling feature reconstruction module, a second upsampling feature reconstruction module, and a third upsampling feature reconstruction module; each upsampling feature reconstruction module includes a first feature reconstruction layer, an upsampling layer, and a second feature reconstruction layer connected in sequence.

[0036] Preferably, the downscaled reconstructed point cloud is input into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original, specifically including:

[0037] Input the downscaled reconstructed point cloud into the first upsampling feature reconstruction module, and perform preliminary feature reconstruction on the downscaled reconstructed point cloud through the first feature reconstruction layer to obtain a feature point cloud after preliminary feature reconstruction. The calculation process of the first feature reconstruction layer is as follows:

[0038] ;

[0039] Among them, represents the downscaled reconstructed point cloud; represents the feature point cloud passing through the first feature reconstruction layer; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a channel number of C, and a scaling factor of 1; represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a channel number of and a scaling factor of 1; represents the ReLU activation function; represents the concatenation operation;

[0040] Input the feature point cloud into the upsampling layer to obtain a feature point cloud after one upsampling. The calculation process of the upsampling layer is as follows:

[0041] ;

[0042] Among them, represents the feature point cloud after upsampling, represents a three-dimensional transposed sparse convolution with a convolution kernel of 2×2×2, a channel number of C, and a scaling factor of ;

[0043] Input the feature point cloud into the second feature reconstruction layer to obtain a deep feature point cloud after upsampling. The calculation process of the second feature reconstruction layer is as follows:

[0044] ;

[0045] Among them, Denote the upsampled deep feature point cloud after passing through the first upsampling feature reconstruction module;

[0046] Input the feature point cloud into the second upsampling feature reconstruction module to obtain the upsampled deep feature point cloud after passing through the second upsampling feature reconstruction module ;

[0047] Input the upsampled deep feature point cloud into the third upsampling feature reconstruction module to obtain the upsampled deep feature point cloud after passing through the third upsampling feature reconstruction module , which serves as the reconstructed point cloud consistent with the original resolution .

[0048] In a second aspect, a point cloud compression device based on a multi-scale octree attention mechanism includes:

[0049] A model construction module configured to construct and train a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; the point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; the decoder network includes an octree decoding module and an upscaling feature reconstructor;

[0050] A point cloud compression and reconstruction module configured to input point cloud data into the trained point cloud compression model. The downscaling feature extractor performs downsampling and feature extraction on the input point cloud data to obtain a downscaled deep feature point cloud; input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain an octree representation of the point cloud; input the octree representation of the point cloud into the context construction module for context construction to obtain a context window; input the context window into the multi-head attention module for calculation to obtain an occupancy probability distribution of octree nodes; input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream; input the bitstream into the octree decoding module to obtain a downscaled reconstructed point cloud; input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud consistent with the original resolution.

[0051] In a third aspect, an electronic device includes:

[0052] One or more processors;

[0053] A storage device for storing one or more programs;

[0054] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the point cloud compression methods based on the multi-scale octree attention mechanism.

[0055] In a fourth aspect, a computer-readable storage medium stores a computer program which, when executed by a processor, implements any of the point cloud compression methods based on the multi-scale octree attention mechanism.

[0056] In a fifth aspect, a computer program product includes a computer program which, when executed by a processor, implements any of the point cloud compression methods based on the multi-scale octree attention mechanism.

[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0058] (1) The point cloud compression method based on the multi-scale octree attention mechanism proposed by the present invention adopts a strategy based on the octree attention mechanism, which can effectively capture the spatial context information in the point cloud data. By constructing an octree representation and using the multi-head attention mechanism to fuse the octree node features, the model can more accurately predict the occupancy probability of the nodes, thereby improving the compression efficiency. The attention mechanism can also dynamically focus on the node features that have a greater impact on the compression performance, further enhancing the compression ability of the model;

[0059] (2) The downscaling feature extractor and upscaling feature reconstructor in the point cloud compression method based on the multi-scale octree attention mechanism proposed by the present invention gradually extract or restore the multi-scale features of the point cloud through multi-level downsampling feature extraction or upsampling feature reconstruction operations. This multi-scale processing method helps the model better capture the detailed information and global structure of the point cloud, thereby retaining more useful information during the compression process and improving the quality of the reconstructed point cloud;

[0060] (3) The three-dimensional sparse convolution adopted by the point cloud compression method based on the multi-scale octree attention mechanism proposed by the present invention can effectively process the irregularity and sparsity of the point cloud data while maintaining high computational efficiency. Compared with traditional dense convolution operations, three-dimensional sparse convolution can greatly reduce the computational amount and improve the running efficiency of the model. Sparse convolution can also better adapt to the spatial distribution of the point cloud data, extract more representative features, and further improve the compression performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a schematic flowchart of the point cloud compression method based on the multi-scale octree attention mechanism according to an embodiment of the present invention;

[0062] Figure 2Schematic diagram of the point cloud compression model of the point cloud compression method based on the multi-scale octree attention mechanism according to the embodiment of the present invention;

[0063] Figure 3 Schematic diagram of the feature extraction layer of the point cloud compression method based on the multi-scale octree attention mechanism according to the embodiment of the present invention;

[0064] Figure 4 Schematic diagram of the multi-head attention module of the point cloud compression method based on the multi-scale octree attention mechanism according to the embodiment of the present invention;

[0065] Figure 5 Schematic diagram of the multi-head attention layer of the point cloud compression method based on the multi-scale octree attention mechanism according to the embodiment of the present invention;

[0066] Figure 6 Block diagram of the structure of the point cloud compression device based on the multi-scale octree attention mechanism according to the embodiment of the present invention;

[0067] Figure 7 Schematic diagram of the hardware structure of the electronic device according to the embodiment of the present invention. Detailed implementation manners

[0068] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0069] Refer to Figure 1 As shown, a point cloud compression method based on the multi-scale octree attention mechanism in this embodiment includes the following steps.

[0070] S101. Construct a point cloud compression model based on the multi-scale octree attention mechanism and train it to obtain a trained point cloud compression model; the point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; the decoder network includes an octree decoding module and an upscaling feature reconstructor.

[0071] Specifically, refer to Figure 2As shown in the figure, the point cloud compression model proposed in the embodiment of the present invention has two parts: an encoder network and a decoder network, which are respectively used to encode point cloud data into a bitstream and decode the encoded bitstream into a reconstructed point cloud. First, a downscaling feature extractor downsamples and extracts features from the input point cloud to obtain a downscaled deep feature point cloud; an octree is constructed recursively to obtain an octree representation of the point cloud; a context construction module is designed to obtain a context window for octree nodes; by combining the multi-head attention mechanism and a multi-layer perceptron, the occupancy probability of octree nodes is obtained; an octree encoding module is designed to encode the node occupancy probability into a bitstream; an octree decoding module is designed to decode the bitstream into a reconstructed point cloud; the reconstructed point cloud is upsampled and feature-reconstructed by an upscaling feature reconstructor to obtain a reconstructed point cloud with the same resolution as the original point cloud, which can effectively achieve point cloud compression.

[0072] S102. Input the point cloud data into the trained point cloud compression model. The downscaling feature extractor downsamples and extracts features from the input point cloud data to obtain a downscaled deep feature point cloud; input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain an octree representation of the point cloud; input the octree representation of the point cloud into the context construction module for context construction to obtain a context window; input the context window into the multi-head attention module for calculation to obtain the occupancy probability distribution of octree nodes; input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream; input the bitstream into the octree decoding module to obtain a downscaled reconstructed point cloud; input the downscaled reconstructed point cloud into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original.

[0073] In a specific embodiment, the downscaling feature extractor includes three sequentially connected downsampling feature extraction modules, namely a first downsampling feature extraction module, a second downsampling feature extraction module, and a third downsampling feature extraction module; each downsampling feature extraction module includes a first feature extraction layer, a downsampling layer, and a second feature extraction layer connected in sequence.

[0074] Input the point cloud data into the trained point cloud compression model. The downscaling feature extractor downsamples and extracts features from the input point cloud data to obtain a downscaled deep feature point cloud, as follows.

[0075] Input the point cloud data into the first downsampling feature extraction module. The first feature extraction layer preliminarily extracts features from the input point cloud data to obtain a feature point cloud after preliminary feature extraction, as Figure 3 shown. The calculation process of the first feature extraction layer is as follows:

[0076] ;

[0077] Among them, represents the input point cloud data, represents the feature point cloud passing through the first feature extraction layer, represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of 1, represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2 and a number of channels of , and a scaling factor of 1, represents the ReLU activation function, represents the concatenation operation.

[0078] Input the feature point cloud after the preliminary feature extraction into the downsampling layer to obtain the feature point cloud after one downsampling. The calculation process of the downsampling layer is as follows:

[0079] ;

[0080] Among them, represents the feature point cloud after downsampling, represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of .

[0081] Input the downsampled feature point cloud into the second feature extraction layer to obtain the downsampled deep feature point cloud. The calculation process of the second feature extraction layer is as follows:

[0082] ;

[0083] Among them, represents the downsampled deep feature point cloud after passing through the second feature extraction layer.

[0084] Input the downsampled deep feature point cloud after the first downsampling feature extraction module into the second downsampling feature extraction module to obtain the downsampled deep feature point cloud after passing through the second downsampling feature extraction module .

[0085] Input the downsampled deep feature point cloud after passing through the second downsampling feature extraction module into the third downsampling feature extraction module to obtain the downsampled deep feature point cloud after passing through the third downsampling feature extraction module .

[0086] In a specific embodiment, input the downscaled deep feature point cloud into the octree construction module for quantization processing to obtain the octree representation of the point cloud, as follows.

[0087] According to Translate the deep feature point cloud so that its minimum coordinate value is zero; where, is the downscaled deep feature point cloud; is the translated deep feature point cloud, is the translation offset; , and respectively represent the X-axis, Y-axis, and Z-axis coordinates of the downscaled deep feature point cloud, , and respectively represent the minimum values of their coordinates.

[0088] According to and the given quantization depth L, quantize the translated deep feature point cloud, represents the quantized deep feature point cloud, represents the floor operation, quantization step .

[0089] Divide the quantized deep feature point cloud recursively into eight equal-sized sub-cubes, represented as the child nodes of the octree, and finally obtain the octree representation of the point cloud . Among them, the occupancy status of each sub-cube constitutes an eight-bit binary occupancy code, an empty sub-cube is marked as 0, a non-empty sub-cube is marked as 1, and further subdivision is performed until the quantization depth L is reached and the division is terminated; at the leaf node, an eight-bit occupancy code represents eight small cubes with side length Q, and the points in the point cloud are merged into the nearest corresponding small cubes.

[0090] In a specific embodiment, input the octree representation of the point cloud into the context construction module for context construction to obtain a context window, as follows.

[0091] Traverse the octree representation of the point cloud in breadth-first order to obtain the sequence , for each node of the sequence , construct a context window with length N, which includes the current node and its previous N - 1 sibling nodes, and embed the K - 1 ancestor nodes of each of these sibling nodes, and finally obtain a context window containing nodes. Among them, , the length N of the context window can be set according to different point cloud inputs, and the default is set to 1024.

[0092] In a specific embodiment, the context window is input into the multi-head attention module for calculation to obtain the occupancy probability distribution of the octree nodes, as follows.

[0093] Input the said context window into the said multi-head attention module. As Figure 4 shown, first, perform feature embedding on the input context window to map it into a high-dimensional feature vector, obtaining the context window after feature embedding; then, through the multi-head attention layer, perform global interaction on the embedded features to obtain a weighted context window that fuses global information; perform a normalization operation on it to obtain a normalized weighted context window; then, through a multi-layer perceptron, obtain a weighted context window after linear transformation; finally, introduce non-linearity through an activation function to obtain the occupancy probability of the output node.

[0094] Its calculation process is as follows:

[0095] ;

[0096] Among them, represents the occupancy probability distribution of the octree nodes, represents the multi-layer perceptron, represents the normalization operation, represents the function corresponding to the multi-head attention layer, represents the feature embedding operation.

[0097] The said multi-head attention layer is as Figure 5 shown. First, perform a normalization operation on the context window after feature embedding to obtain a normalized context window after feature embedding; it performs a linear transformation through 3 parallel multi-layer perceptrons to generate a query vector, a key vector, and a value vector respectively; perform a matrix multiplication operation on the query vector and the key vector, and then apply an activation function for non-linearity to obtain the attention weights; finally, perform a matrix multiplication operation on the attention weights and the value vector to finally obtain a weighted context window.

[0098] Its calculation process is as follows:

[0099] ;

[0100] Among them, represents the weighted context window, represents the SoftMax activation function, represents the context window after feature embedding, represents the matrix multiplication operation.

[0101] In a specific embodiment, the upscaling feature reconstructor includes three upsampling feature reconstruction modules connected in sequence, namely a first upsampling feature reconstruction module, a second upsampling feature reconstruction module, and a third upsampling feature reconstruction module; each upsampling feature reconstruction module includes a first feature reconstruction layer, an upsampling layer, and a second feature reconstruction layer connected in sequence.

[0102] The downscaled reconstructed point cloud is input into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud with the same resolution as the original, as follows.

[0103] The downscaled reconstructed point cloud is input into the first upsampling feature reconstruction module, and the downscaled reconstructed point cloud is preliminarily feature - reconstructed through the first feature reconstruction layer to obtain a feature point cloud after preliminary feature reconstruction. The calculation process of the first feature reconstruction layer is as follows:

[0104] ;

[0105] where, represents the downscaled reconstructed point cloud, represents the feature point cloud passing through the first feature reconstruction layer.

[0106] The feature point cloud after preliminary feature reconstruction is input into the upsampling layer to obtain a feature point cloud after one - time upsampling. The calculation process of the upsampling layer is as follows:

[0107] ;

[0108] where, represents the feature point cloud after upsampling, represents a 3D transposed sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling factor of .

[0109] The upsampled feature point cloud is input into the second feature reconstruction layer to obtain a deep feature point cloud after upsampling. The calculation process of the second feature reconstruction layer is as follows:

[0110] ;

[0111] where, represents the deep feature point cloud after upsampling through the first upsampling feature reconstruction module.

[0112] The deep feature point cloud after upsampling through the first upsampling feature reconstruction module is input into the second upsampling feature reconstruction module to obtain a deep feature point cloud after upsampling through the second upsampling feature reconstruction module .

[0113] Input the upsampled deep feature point cloud after the second upsampling feature reconstruction module into the third upsampling feature reconstruction module to obtain the upsampled deep feature point cloud after the third upsampling feature reconstruction module .

[0114] The upsampled deep feature point cloud after the third upsampling feature reconstruction module is the reconstructed point cloud consistent with the original resolution .

[0115] In a specific embodiment, the experimental environment used includes a workstation equipped with an Intel(R) Xeon(R) Gold 6226R processor (2.90 GHz), equipped with an NVIDIA RTX 3090 graphics card (24 GB video memory) and 128 GB of DDR4 memory, and the operating system is Ubuntu 20.04 LTS; the PyTorch deep learning framework is used in the experiment, and CUDA 11.8 acceleration is enabled.

[0116] In a specific embodiment, the selected dataset is SemanticKITTI, which is a large-scale LiDAR point cloud dataset for autonomous driving. It is obtained by scanning with a Velodyne HDL-64E sensor and contains a total of 45.49 million points; sequences 00 to 10 are selected as the training set, and sequences 11 to 21 are selected as the test set.

[0117] In a specific embodiment, the Adam optimizer is used during the training process. The initial learning rate is set to 1e-4, and the StepLR strategy is adopted to decay the learning rate to 0.1 of the original value every 10 epochs; the batch size is set to 32, and the number of training epochs is 100. All training is repeated under the same random seed to ensure the stability and reproducibility of the results; the loss function for training is the cross-entropy loss, which is specifically as follows:

[0118] ;

[0119] wherein represents the loss function, represents the predicted value of the node occupancy probability, represents the true value of the node occupancy probability.

[0120] Specifically, as shown in Figure 6 As an implementation of the methods shown in the above figures, an embodiment of a point cloud compression device based on a multi-scale octree attention mechanism is provided in this application. This device embodiment is related to Figure 1The method embodiments shown correspond to this, and this device can be specifically applied to various electronic devices.

[0121] A point cloud compression device based on a multi-scale octree attention mechanism, comprising:

[0122] A model construction module 601, configured to construct and train a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; the point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree construction module, a context construction module, a multi-head attention module, and an octree encoding module; the decoder network includes an octree decoding module and an upscaling feature reconstructor;

[0123] A point cloud compression and reconstruction module 602, configured to input point cloud data into the trained point cloud compression model, the downscaling feature extractor performs downsampling and feature extraction on the input point cloud data to obtain downscaled deep feature point clouds; input the downscaled deep feature point clouds into the octree construction module for quantization processing to obtain an octree representation of the point cloud; input the octree representation of the point cloud into the context construction module for context construction to obtain a context window; input the context window into the multi-head attention module for calculation to obtain an occupancy probability distribution of octree nodes; input the occupancy probability distribution into the octree encoding module for compression to obtain a bitstream; input the bitstream into the octree decoding module to obtain downscaled reconstructed point clouds; input the downscaled reconstructed point clouds into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain reconstructed point clouds with the same resolution as the original.

[0124] The specific implementation of each module of a point cloud compression device based on a multi-scale octree attention mechanism is the same as that of a point cloud compression method based on a multi-scale octree attention mechanism, and this embodiment will not be repeated here.

[0125] See Figure 7 The following is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Figure 7 In this, the electronic device of this embodiment includes: a processor 701 and a memory 702; wherein the memory 702 is used to store computer execution instructions; the processor 701 is used to execute the computer execution instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiments. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.

[0126] Optionally, the memory 702 can be either independent or integrated with the processor 701.

[0127] When the memory 702 is independently provided, the electronic device further includes a bus 703 for connecting the memory 702 and the processor 701.

[0128] An embodiment of the present invention also provides a computer storage medium, in which computer-executable instructions are stored. When the processor 701 executes the computer-executable instructions, the above method is implemented.

[0129] An embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 701, the above method is implemented.

[0130] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical or other forms.

[0131] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0132] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The units formed by the above modules can be implemented in the form of hardware or in the form of a hardware plus software functional unit.

[0133] The above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or the processor 701 to execute some steps of the methods in various embodiments of the present application.

[0134] It should be understood that the above-mentioned processor 701 can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor can be a microprocessor or the processor 701 can also be any conventional processor 701, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by the hardware processor 701, or by a combination of hardware and software modules in the processor 701.

[0135] The memory 702 may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0136] The bus 703 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus 703 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the bus 703 in the attached drawings of this application is not limited to only one bus 703 or one type of bus 703.

[0137] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0138] An exemplary storage medium is coupled to the processor 701, enabling the processor 701 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 701. The processor 701 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 701 and the storage medium can also exist as discrete components in an electronic device or a master control device.

[0139] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A point cloud compression method based on a multi-scale octree attention mechanism, characterized in that: include: Construct and train a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; The point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree building module, a context building module, a multi-head attention module, and an octree encoding module; The decoder network includes an octree decoding module and an upscaling feature reconstructor; The point cloud data is input into the trained point cloud compression model, and the downscaling feature extractor downsamples and extracts features from the input point cloud data to obtain a downscaled deep feature point cloud; The downscaled deep feature point cloud is input into the octree construction module for quantization processing to obtain the octree representation of the point cloud; the octree representation of the point cloud is input into the context construction module for context construction to obtain the context window; The context window is input into the multi-head attention module for calculation to obtain the occupancy probability distribution of the octree nodes; the occupancy probability distribution is input into the octree encoding module for compression to obtain a bit stream; the bit stream is input into the octree decoding module to obtain a downscaled reconstructed point cloud; the downscaled reconstructed point cloud is input into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud consistent with the original resolution.

2. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 1, characterized in that: The downscaling feature extractor includes three downsampling feature extraction modules connected in sequence, namely a first downsampling feature extraction module, a second downsampling feature extraction module and a third downsampling feature extraction module; each downsampling feature extraction module includes a first feature extraction layer, a downsampling layer and a second feature extraction layer connected in sequence.

3. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 2 is characterized in that: The point cloud data is input into the trained point cloud compression model. The downscaling feature extractor downsamples and extracts features from the input point cloud data to obtain a downscaled deep feature point cloud, which includes: The point cloud data is input into the first downsampling feature extraction module, and the input point cloud is subjected to preliminary feature extraction through the first feature extraction layer to obtain a feature point cloud after preliminary feature extraction. The calculation process of the first feature extraction layer is as follows: ; in, Represents the input point cloud data, Represents the feature point cloud that passes through the first feature extraction layer; Represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling of 1; Indicates that the convolution kernel is 2×2×2 and the number of channels is , three-dimensional sparse convolution with a scaling factor of 1; ReLU activation function. Represents a splicing operation; Feature point cloud Input to the downsampling layer to obtain the feature point cloud after one downsampling. The calculation process of the downsampling layer is as follows: ; in, Represents the feature point cloud after downsampling; Indicates that the convolution kernel is 2×2×2, the number of channels is C, and the scaling is Three-dimensional sparse convolution; Feature point cloud Input to the second feature extraction layer to obtain the downsampled deep feature point cloud. The calculation process of the second feature extraction layer is as follows: ; in, represents the downsampled deep feature point cloud after passing through the second feature extraction layer; Downsample the deep feature point cloud Input to the second downsampling feature extraction module to obtain the downsampled deep feature point cloud after passing through the second downsampling feature extraction module ; Downsample the deep feature point cloud Input to the third downsampling feature extraction module to obtain the downsampled deep feature point cloud after passing through the third downsampling feature extraction module , as the downscaled deep feature point cloud.

4. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 1, characterized in that: The downscaled deep feature point cloud is input into the octree construction module for quantization processing to obtain the octree representation of the point cloud, which includes: according to The deep feature point cloud is translated so that its minimum coordinate value is zero; It is the downscaled deep feature point cloud; is the deep feature point cloud after translation, is the translation offset; , and Respectively represent the X-axis, Y-axis, and Z-axis coordinates of the downscaled deep feature point cloud, , and Respectively represent the minimum value of their coordinates; according to And the given quantization depth L is used to quantize the translated deep feature point cloud; where, Represents the quantized deep feature point cloud; Indicates rounding down operation; quantization step size ; The quantized deep feature point cloud Recursively divide it into eight equal-sized sub-cubes, represented as child nodes of the octree, and finally obtain the octree representation of the point cloud .

5. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 1, characterized in that: The octree representation of the point cloud is input into the context building module for context building to obtain the context window, which includes: Traverse the octree representation of the point cloud in breadth-first order Get sequence , for the sequence Each node , construct a length of The context window , which includes the current node and its previous N-1 sibling nodes, and embed the K-1 ancestor nodes of each of these sibling nodes, and finally get Context window for nodes ;in, .

6. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 1, characterized in that: The context window is input into the multi-head attention module to calculate the occupancy probability distribution of the octree nodes, including: The context window Input into the multi-head attention module for calculation, as follows: ; in, Represents the occupancy probability distribution of octree nodes; represents a multi-layer perceptron; Represents a normalization operation; Represents the function corresponding to the multi-head attention layer; represents the feature embedding operation; ReLU activation function. The calculation process of the multi-head attention layer is as follows: ; in, represents the weighted context window; Represents the SoftMax activation function; Represents the context window after embedding features; Represents a matrix multiplication operation.

7. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 1, characterized in that: The upscaling feature reconstructor includes three upsampling feature reconstruction modules connected in sequence, namely a first upsampling feature reconstruction module, a second upsampling feature reconstruction module and a third upsampling feature reconstruction module; each upsampling feature reconstruction module includes a first feature reconstruction layer, an upsampling layer and a second feature reconstruction layer connected in sequence.

8. The point cloud compression method based on the multi-scale octree attention mechanism according to claim 7, characterized in that: The downscaled reconstructed point cloud is input into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud consistent with the original resolution, including: The downscaled reconstructed point cloud is input into the first upsampling feature reconstruction module, and the downscaled reconstructed point cloud is preliminarily reconstructed through the first feature reconstruction layer to obtain a feature point cloud after preliminary feature reconstruction. The calculation process of the first feature reconstruction layer is as follows: ; in, represents the downscaled reconstructed point cloud; Represents the feature point cloud through the first feature reconstruction layer; Represents a three-dimensional sparse convolution with a convolution kernel of 2×2×2, a number of channels of C, and a scaling of 1; Indicates that the convolution kernel is 2×2×2 and the number of channels is , three-dimensional sparse convolution with a scaling factor of 1; ReLU activation function. Represents a splicing operation; Feature point cloud Input to the upsampling layer to obtain the feature point cloud after one upsampling. The calculation process of the upsampling layer is as follows: ; in, represents the feature point cloud after upsampling, Indicates that the convolution kernel is 2×2×2, the number of channels is C, and the scaling is 3D transposed sparse convolution; Feature point cloud Input to the second feature reconstruction layer to obtain the upsampled deep feature point cloud. The calculation process of the second feature reconstruction layer is as follows: ; in, represents the upsampled deep feature point cloud after passing through the first upsampled feature reconstruction module; Feature point cloud Input to the second upsampling feature reconstruction module to obtain the upsampled deep feature point cloud after passing through the second upsampling feature reconstruction module ; Upsample the deep feature point cloud Input to the third upsampling feature reconstruction module to obtain the upsampled deep feature point cloud after passing through the third upsampling feature reconstruction module , as a reconstructed point cloud consistent with the original resolution .

9. A point cloud compression device based on a multi-scale octree attention mechanism, characterized in that: include: A model building module is configured to build and train a point cloud compression model based on a multi-scale octree attention mechanism to obtain a trained point cloud compression model; The point cloud compression model includes an encoder network and a decoder network; the encoder network includes a downscaling feature extractor, an octree building module, a context building module, a multi-head attention module, and an octree encoding module; The decoder network includes an octree decoding module and an upscaling feature reconstructor; The point cloud compression and reconstruction module is configured to input the point cloud data into the trained point cloud compression model, and the downscaling feature extractor downsamples and extracts features from the input point cloud data to obtain a downscaled deep feature point cloud; The downscaled deep feature point cloud is input into the octree construction module for quantization processing to obtain the octree representation of the point cloud; the octree representation of the point cloud is input into the context construction module for context construction to obtain the context window; The context window is input into the multi-head attention module for calculation to obtain the occupancy probability distribution of the octree nodes; the occupancy probability distribution is input into the octree encoding module for compression to obtain a bit stream; the bit stream is input into the octree decoding module to obtain a downscaled reconstructed point cloud; the downscaled reconstructed point cloud is input into the upscaling feature reconstructor for point cloud upsampling and point cloud feature reconstruction to obtain a reconstructed point cloud consistent with the original resolution.

10. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Point cloud compression and decompression method based on octree coding and voxel context

    CN113284203A

  • Point cloud coding and decoding method based on double octree structure

    CN117692662A

  • Point cloud coding method based on multilevel ball octree and graph-driven attention entropy model

    CN119006620A

  • Point Cloud Compression Using Octrees with Slicing

    US20210407147A1

  • Attention-Based Method for Deep Point Cloud Compression

    US20240282014A1

Cited By

  • Semantic scene graph-based double-flow point cloud compression method and system

    CN121937547A

  • Dual-stream point cloud compression method and system based on semantic scene graph

    CN121937547B