A Point Cloud Geometry Compression Method and System Based on Interlayer Residual and IRN Connection Residual
Through the point cloud geometric compression method of inter-layer residual and IRN connection residual, the problems of low efficiency and information loss in point cloud compression are solved, efficient point cloud reconstruction and resource conservation are achieved, and compression efficiency and reconstruction quality are improved.
Patent Information
- Application Number
- CN202510220228.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing point cloud compression methods are inefficient when dealing with irregular point clouds, and there are problems of information loss and resource waste. In particular, voxel-based methods are not efficient in processing sparse point clouds, and the encoder cannot effectively process floating-point type data.
The point cloud geometric compression method based on inter-layer residual and IRN connection residual is adopted, and features are extracted through sparse convolution, combined with the D-U residual mechanism and the IRN connection residual mechanism, the convolution neural network is optimized, and the point cloud information is processed using the octree codec and arithmetic encoder, and the scale super-priority modeling point cloud feature probability distribution is introduced, and the first k largest elements are selected for prediction.
It improves the efficiency and reconstruction quality of point cloud compression, reduces computational complexity and resource consumption, improves reconstruction accuracy and data set profile clarity at low bit rate, and shows compression efficiency and reconstruction effects that are better than existing methods.
Smart Images

Figure CN119722825B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point cloud compression, and in particular to a point cloud geometric compression method and system based on inter-layer residuals and IRN connection residuals. Background Technique
[0002] Geometry-based point cloud compression methods are applicable to static point clouds and dynamically acquired point clouds. By constructing and encoding the geometric relationships between points, they reduce the amount of data while maintaining the geometric structure. However, this method has limited ability to process irregular point clouds. In addition, the most commonly used partitioning method for geometry-based point cloud compression is the octree method, which recursively divides the three-dimensional space into eight subspaces to efficiently represent and compress the spatial distribution of the point cloud. This results in a sharp increase in the number of bits required by this method as the octree depth increases, and a significant increase in the storage space required and the bandwidth during transmission.
[0003] The point cloud geometric compression method based on plane approximation utilizes the nodes divided in a larger space by the pruned octree to capture the continuous point distribution within the nodes. This method is applicable to densely sampled and surface-smooth point clouds, but due to the inherent error of plane approximation, this method cannot achieve lossless compression.
[0004] Video-based point cloud compression projects the three-dimensional point cloud onto a two-dimensional plane, and then uses existing efficient video codecs (such as HEVC) to process the point cloud. This method is applicable to dense and dynamic point clouds, but during the projection process, this method inevitably loses a large number of occluded continuous points, resulting in obvious artifacts in the compressed point cloud.
[0005] With the powerful capabilities of deep learning in feature extraction and data encoding, learning-based point cloud compression methods have achieved a good balance between compression ratio and reconstruction quality. Therefore, learning-based point cloud compression methods have emerged as the times require. It can well solve the problems existing in the above-mentioned traditional point clouds by training to adapt to different types of point cloud data.
[0006] The voxel-based point cloud geometry compression method is a type of learning-based point cloud compression. It divides point cloud data into the form of three-dimensional grids and achieves efficient representation and compression through the processing of voxels. However, for sparse point clouds, most voxels are empty, but 3D convolution is still uniformly applied to each voxel. Therefore, this method does not effectively utilize the sparsity between point clouds, resulting in the consumption of a large amount of space and computing resources. In addition, in the end-to-end compression framework, due to the downsampling operation, while extracting the key features of the point cloud, it inevitably causes the loss of a part of the point cloud information, thereby reducing the accuracy of downsampling and the quality of reconstruction. When the encoder processes the point cloud attribute information, the entropy model itself cannot directly process continuous floating-point type data. Therefore, there is an urgent need for a point cloud geometry compression method based on inter-layer residuals and IRN connection residuals. Summary of the Invention
[0007] To solve the above-mentioned problems, the present invention provides a point cloud geometry compression method and system based on inter-layer residuals and IRN connection residuals.
[0008] In the first aspect, a point cloud geometry compression method based on inter-layer residuals and IRN connection residuals provided by the present invention adopts the following technical solutions:
[0009] A point cloud geometry compression method based on inter-layer residuals and IRN connection residuals includes:
[0010] Obtain point cloud data;
[0011] Use a convolutional neural network based on sparse convolution for feature extraction. Among them, perform upsampling after downsampling based on the D-U residual mechanism, and perform geometric reduction with the feature information before upsampling to obtain the up-down sampling residual; calculate the loss function using the up-down sampling residual, and optimize the convolutional neural network by minimizing the loss function;
[0012] Capture the extracted features based on the IRN connection residual mechanism to obtain global features;
[0013] Encode and decode the global features based on an encoder-decoder;
[0014] Predict the decoded information.
[0015] Further, the use of a convolutional neural network based on sparse convolution for feature extraction includes, for the input 3D raw point cloud data, first performing a sparse convolution operation to expand the input 1-channel data to 16 channels, then activating through the Relu function, and performing a downsampling operation using sparse convolution to expand the feature data from 16 channels to 32 channels while halving the size.
[0016] Further, the upsampling after downsampling based on the D-U residual mechanism is performed, and geometric reduction is carried out with the feature information before upsampling to obtain the up-downsampling residual, including the feature after downsampling is subjected to upsampling operation , and geometric reduction is carried out with the feature information before upsampling to obtain the residual , which is expressed as:
[0017]
[0018]
[0019]
[0020] Among them, represents the coordinate residual result value, and respectively represent the geometric coordinates of the i-th point before upsampling and downsampling, that is, the coordinate differences of each point on the three-dimensional components, which are expressed as ([[]] ), represents the upsampling operation, represents the downsampling operation.
[0021] Further, the loss function is calculated by using the up-downsampling residual, and the convolutional neural network is optimized by minimizing the loss function, including using the up-downsampling residual result value to calculate the loss function. During the training process, the model is optimized by minimizing the loss function, which is expressed as:
[0022]
[0023] Among them, represents the calculated residual loss value of up-downsampling at the encoder end, is the activation function, M is the number of points, is calculated by the formula in claim 3 and represents the result value of the coordinate residual.
[0024] Further, the global feature is captured by the IRN connection residual mechanism for the extracted feature, including stacking three IRN residuals by using the method of multi-layer residual connection. Among them, first, the feature information is processed by an IRN residual to obtain multi-scale feature information, and then the input feature information is connected with the multi-scale feature information and sent to the next IRN residual for processing, and the original feature information of the point cloud data is retained through residual connection, which is expressed as:
[0025]
[0026]
[0027] where is the output of the IRN connection residual module, is the input of the IRN connection residual module, is the input of the IRN module, is the connection operation.
[0028] Furthermore, the global feature is encoded and decoded based on the codec, including performing quantization processing on the global feature information, then using an octree codec to process the geometric information in the global feature, and using an arithmetic codec to process the attribute information of the global feature. At the same time, a scale hyperprior c is introduced to capture the scale changes of spatially adjacent points to model the probability distribution of the point cloud features.
[0029] Furthermore, the prediction of the decoded information includes selecting the top k largest elements for processing to reduce data redundancy and improve computational efficiency. The top k elements are calculated according to the binary cross-entropy loss, and the prediction probability is improved by minimizing the binary cross-entropy loss to reduce the reconstruction error, expressed as:
[0030]
[0031] where N is the number of point clouds to be predicted, is the occupancy situation in the voxel, where 1 represents occupancy and 0 represents non-occupancy, is the predicted probability value.
[0032] In a second aspect, a point cloud geometry compression system based on inter-layer residuals and IRN connection residuals includes:
[0033] A data acquisition module configured to acquire point cloud data;
[0034] A D-U residual module configured to use a convolutional neural network based on sparse convolution for feature extraction, where downsampling and then upsampling are performed based on the D-U residual mechanism, and geometric reduction is performed with the feature information before upsampling to obtain the up and down sampling residuals; the loss function is calculated using the up and down sampling residuals, and the convolutional neural network is optimized by minimizing the loss function;
[0035] An IRN connection residual module configured to capture the extracted features based on the IRN connection residual mechanism to obtain global features;
[0036] An encoding and decoding module configured to encode and decode the global features based on the codec;
[0037] A prediction module configured to predict the decoded information.
[0038] In a third aspect, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted to be loaded and executed by a processor of a terminal device to implement the method for geometric compression of point clouds based on inter-layer residuals and IRN connection residuals.
[0039] In a fourth aspect, the present invention provides a terminal device including a processor and a computer-readable storage medium. The processor is configured to implement each instruction; the computer-readable storage medium is configured to store a plurality of instructions adapted to be loaded and executed by the processor to implement the method for geometric compression of point clouds based on inter-layer residuals and IRN connection residuals.
[0040] In summary, the present invention has the following beneficial technical effects:
[0041] 1. By calculating the upsampling and downsampling residuals based on the D-U residual mechanism and using it to optimize the convolutional neural network, the present invention can reduce the information loss and quantization error during the downsampling process, effectively constraining the downsampling accuracy. At the same time, the IRN connection residual mechanism captures global features through multi-layer residual connections, retains the original feature information of the point cloud data, and avoids the problem of gradient disappearance. In the experiment, compared with G-PCC (octree), the present invention has an average BD-PSNR improvement of 11.70dB and 11.08dB in D1 and D2 respectively. At low bitrates, the reconstructed point cloud boundary regions perform better, and the dataset contours are clearer and more accurate. The reconstruction effect is better at the feet of the longdress dataset and the gun edges of the soldier dataset.
[0042] 2. In the encoding process, the octree codec is used to process geometric information, only encoding the subspaces containing points, combined with the arithmetic codec to process attribute information, and introducing the scale hyperprior c to model the probability distribution of point cloud features. This method improves the data compression efficiency. In the comparative experiments with G-PCC (octree), G-PCC (trisoup), and PCGC v2, the average BD-Rate gains of the present invention compared with G-PCC (octree) in D1 and D2 are 96% and 93% respectively, compared with G-PCC (trisoup) are 74% and 81% respectively, and compared with PCGC v2 are 6% and 9% respectively, demonstrating the advantages of the present invention in terms of compression efficiency.
[0043] 3. The present invention uses sparse convolution for feature extraction, only extracting features of occupied coordinates, avoiding invalid calculations for a large number of empty voxels, reducing computational complexity, and reducing the consumption of spatial and computational resources. At the same time, in the prediction stage, the top k largest elements are selected for processing, reducing data redundancy, improving computational efficiency, further reducing the overall resource consumption, and enabling the model to operate efficiently in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a schematic diagram of a point cloud geometric compression method based on inter-layer residuals and IRN connection residuals according to Embodiment 1 of the present invention;
[0045] Figure 2 is a schematic diagram of the IRN connection residual module according to Embodiment 1 of the present invention; wherein, Figure 2 A in is the IRN module, Figure 2 B in is the overall IRN connection residual module;
[0046] Figure 3 is a schematic diagram of the attention module according to Embodiment 1 of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The present invention will be further described in detail below with reference to the accompanying drawings.
[0048] Embodiment 1
[0049] Referring to Figure 1 , a point cloud geometric compression method based on inter-layer residuals and IRN connection residuals in this embodiment includes:
[0050] Obtain point cloud data;
[0051] Use a convolutional neural network based on sparse convolution for feature extraction, wherein, perform upsampling after downsampling based on the D-U residual mechanism, and perform geometric reduction with the feature information before upsampling to obtain up and down sampling residuals; calculate the loss function using the up and down sampling residuals, and optimize the convolutional neural network by minimizing the loss function;
[0052] Capture the extracted features based on the IRN connection residual mechanism to obtain global features;
[0053] Encode and decode the global features based on an encoder-decoder;
[0054] Predict the decoded information.
[0055] Specifically:
[0056] First, the present invention randomly selects 26,342 datasets from the ShapeNet dataset for training, and selects two different types of datasets, 8iVFB and Owlii, for testing. For the point cloud data in the above datasets, the present invention inputs this data into a convolutional neural network based on sparse convolution. In this neural network, the point cloud data is represented as a set of sparse tensors , including geometric coordinates (x, y, z positive occupancy) and feature attributes (such as color, opacity, reflectivity, etc.), which are represented by a set of coordinates and related attribute features . Among them, sparse convolution only extracts the features of the occupied coordinates. For each voxel , its feature vector is obtained by performing convolution calculations on the input voxels within its neighborhood. The specific calculation process is as follows:
[0057] (1)
[0058] where and are the coordinates of the input and output point clouds respectively, and are the feature vectors of the input and output point clouds at coordinate c respectively, represents the definition of a 3D convolution kernel centered on the offset i. is the value of the convolution kernel at offset i.
[0059] The present invention uses sparse convolution to reduce complexity and better extract the feature information of the point cloud. Each convolution operation in the present invention (i.e., the feature aggregation in downsampling and the point cloud reconstruction part in upsampling) is processed according to the above process.
[0060] In other words, for the input 3D raw point cloud data, the present invention first performs a sparse convolution operation to represent the point cloud data in the form of sparse tensors , and records the coordinates and related eigenvalue of the non-empty voxels. During the convolution process, the sparse convolution kernel acts on the neighborhood of the input voxels, calculates the feature information of each voxel within the neighborhood, and sums them up with weights to generate the feature information of the output voxel. Initially, the input data is 1 channel. After the sparse convolution operation, the 1-channel input data is expanded to 16 channels through the convolution kernel, and then activated by the Relu function to further increase the non-linear features, ensuring that the model can learn complex feature representations. Subsequently, we use sparse convolution for downsampling operations to reduce the spatial resolution of the features, expand the feature data from 16 channels to 32 channels, and halve the size at the same time. However, during the downsampling process, due to the halving of the data, it is inevitable that some point cloud information will be lost.
[0061] In the second step, in order to constrain the accuracy of point cloud downsampling as much as possible and reduce quantization errors, the present invention designs a D-U residual module. It upsamples the downsampled to obtain , and geometrically reduces it with the feature information before upsampling to obtain the residual . In this way, the module can clearly know the point cloud feature information lost during the downsampling process. Subsequently, the module constrains the accuracy of the downsampling process by minimizing this residual to reduce the loss of feature information and the residual introduced by the quantization process, thereby improving the reconstruction effect. The specific process is as follows:
[0062]
[0063]
[0064]
[0065] Among them, represents the coordinate residual result value, and respectively represent the geometric coordinates of the i-th point before upsampling and downsampling. That is, the coordinate difference of each point in the three-dimensional components can be expressed as ( ), represents the upsampling operation, and represents the downsampling operation.
[0066] The present invention records the above residual result value without transmission. It is mainly used for the calculation of the loss function. During the training process, the above model is optimized by minimizing this loss. The loss calculation process of this residual is as follows:
[0067]
[0068] Among them, represents the calculated residual loss value of upsampling and downsampling at the encoder end, is the activation function, M is the number of points, is calculated by the formula in claim 3 and represents the result value of the coordinate residual.
[0069] In the third step, the feature information after downsampling in the present invention is transmitted to the IRN connection residual module after being activated by Relu. First, this module introduces the IRN module, and then uses convolutional kernels of different sizes in the IRN module to extract feature information of different scales in parallel. Subsequently, this module stacks three IRN modules by using multi-layer residual connections. Specifically, this module first processes the input feature information through an IRN module to obtain multi-scale feature information, and then this module connects the multi-scale feature information with the input to ensure the correlation between them. The operations of the subsequent two IRN modules are the same, but after the processing of the last IRN module, this module retains the original feature information of the point cloud through residual connection. This operation not only ensures the correlation between layers, but also makes good use of the spatial correlation between blocks, and at the same time avoids the problem of gradient disappearance between blocks. The specific operation steps are as follows:
[0070]
[0071]
[0072] where is the output of the IRN connection residual module, is the input of the IRN connection residual module. is the input of the IRN module, is the connection operation.
[0073] In the fourth step, the subsequent two downsampling operations are the same. For the last convolutional operation at the encoder end, the present invention converts the number of channels to 8 channels for encoding operation.
[0074] In the fifth step, the present invention uses an octree codec to process geometric information, and recursively divides the 3D space into multiple eight subspaces for processing, and only encodes the subspaces containing points. During the encoding process, the octree codec uses a binary bit string to represent the octree structure, that is, each space is represented by 8 bits to indicate the occupancy of the subspace. If the bit is 1, it means that there is point cloud data in the subspace. If it is 0, it means that the subspace is empty and storage is skipped. Finally, a compact bitstream is obtained after encoding to represent the geometric information of the stored and transmitted point cloud. Subsequently, during the decoding process, the inverse process of the encoding operation is used to restore the geometric structure of the point cloud, and the point cloud is accurately restored by reading 8-bit bit information and combining the position information of the points.
[0075] The present invention uses arithmetic coding to process attribute information. Among them, before arithmetic coding, the present invention introduces a quantization operation, and adds uniform noise to the input tensor to achieve more efficient data storage and representation. The specific representation is as follows:
[0076]
[0077] Among them uniform noise x and y respectively represent the quantized feature information and the original feature information follows a uniform distribution centered on .
[0078] Subsequently, the present invention inputs the quantized feature information into an arithmetic codec for processing. The present invention uses a non-parametric fully factorized probability density model to encode the quantized feature information. At the same time, the present invention introduces a scale hyperprior c to capture the scale changes of spatially adjacent points, so as to more accurately model the probability distribution of point cloud features. In short, the present invention uses the Laplace distribution L to approximate the probability density function , and the calculation formula is as follows
[0079]
[0080] where and are respectively the mean and mean square deviation of each under the scale hyperprior c model
[0081] Specifically, the core idea of arithmetic coding is to use the probability distribution of data to allocate shorter bitstreams to high-probability events and longer bitstreams to low-probability events. For the input quantized feature data , its probability model is approximately estimated by the Laplace distribution. In the first step, set the initial interval , and read the quantized feature value of the first symbol . In the second step, read the cumulative distribution function (CDF) of the quantized feature from the probability model , and then determine the sub-interval of this quantized feature within : ; , so the quantized feature occupies this part of the interval . In the third step, for each , repeat the above second step, gradually narrowing the interval. The final interval after processing represents the complete point cloud feature data. In the fourth step, select any real number in the interval and convert it into the form of a binary bitstream. In the arithmetic decoding process, in the first step, convert the binary bitstream into a real number , and map it to the initial interval . In the second step, reverse-lookup the symbol (feature value) through the cumulative distribution function (CDF) to find the one that satisfies; of , that is, to determine the currently decoded symbol. In the third step, according to the interval calculation method in the second step of the encoding, the interval range is updated and set as the initial interval for the next symbol, and the next symbol is decoded. Iterate in turn until all point cloud features are restored.
[0082] Sixth step: After the above encoding and decoding operations, the present invention restores the decoded feature information through an upsampling operation, activates it using the Rule function, and expands the feature information. Similarly, in order to extract multi-scale feature information, the present invention also introduces an IRN residual connection module at the decoder end. In order to enhance the attention degree of the decoder end to significant features, the present invention also introduces an attention module. This module adopts two weighted operations of point-by-point multiplication, captures the global dependence range by directly calculating the relationship between the positions of scattered points, thereby further enhancing the interaction between features. At the end, while retaining the original point cloud information through residual connection, this module increases the fluidity of the gradient. The specific operation of this module is expressed as:
[0083]
[0084] where is the output of this module, is the input of this module.
[0085] Seventh step: Perform a sparse convolution operation on the feature information restored after the attention module, convert its output channels to 1, that is, each voxel only outputs a numerical value, indicating the probability that the voxel belongs to a certain category. The present invention prunes the voxels according to the occupancy of the voxels. In the voxel pruning stage of hierarchical reconstruction, the present invention sorts the result probabilities and considers the top k voxels as the most likely occupied voxels, and performs binary classification through sparse convolution with output channels of 1. The subsequent two upsampling operations are the same as the above operations, and finally the output is obtained to get the compressed and reconstructed point cloud data.
[0086] However, in order to continuously optimize the model during the training process and reduce the gap between the original point cloud and the reconstructed point cloud, the present invention adopts a binary cross-entropy loss function, and improves the prediction probability of the occupancy information of the reconstructed point cloud by minimizing this loss function. And this prediction probability is to map the feature vector finally output at the decoder end to a probability value through the Relu activation function, and this value indicates the occupancy of the voxels in the space. Then, adaptive thresholding classification is performed according to the number of occupied points in the original cube, and it is classified as binary 0 or 1.
[0087] During the prediction process, the present invention selects the top k largest elements for processing to reduce data redundancy and improve computational efficiency. These top k elements are calculated according to the binary cross-entropy loss in the following formula. By minimizing the binary cross-entropy loss, the prediction probability is increased and the reconstruction error is reduced:
[0088]
[0089] where N is the number of point clouds to be predicted, is the occupancy situation in the voxel, where 1 represents occupancy and 0 represents non-occupancy, is the predicted probability value.
[0090] Experimental verification:
[0091] The present invention randomly selected 26,342 datasets from the ShapeNet dataset for training. In the experiment, the Adam optimizer was used, and the initial learning rate was set to 0.0008. The learning rate was adjusted by a halving strategy in each round of training until it was reduced to 0.00001. The batch size was set to 16. According to experimental experience, the value of is selected between 0.25 and 5, and the value of is set to 10. This experiment was run on a computer with an NVIDIA RTX 4080 GPU (16G) and an i7 13790 (32G).
[0092] Objective comparison: The present invention selected two different types of datasets, 8iVFB and Owlii, for testing, and used the point-to-point geometric error (D1 PSNR) and the point-to-plane geometric error (D2 PSNR) to evaluate the reconstruction quality of the point cloud, and used Bits Per Point (bpp) to measure the storage efficiency after compression. To better verify the performance of the model proposed by the present invention, the proposed method was compared with G-PCC (octree), G-PCC (trisoup), and PCGC v2.
[0093] For fair comparison, the present invention follows the recommendations of MPEG Common Test Condition (CTC) and applies similar bit ranges to G-PCC (octree), G-PCC (trisoup), PCGC v2, and the method proposed by the present invention. For G-PCC, the present invention uses the latest TMC13-v23.0-rc2, and the parameter settings follow CTC. For G-PCC (octree), the present invention sets the positionQuantizationScale between 0.15 and 0.5. For G-PCC (trisoup), the present invention sets the trisoupNodeSizeLog2 to 2, 3, and 4. Other parameters remain unchanged with their default values.
[0094] The experimental results show that the method of the present invention has average BD-Rate gains of 96% and 93% on D1 and D2 respectively compared to G-PCC (octree), average BD-Rate gains of 74% and 81% respectively compared to G-PCC (trisoup), and average BD-Rate gains of 6% and 9% respectively compared to the state-of-the-art PCGC v2. Moreover, the method of the present invention exhibits better reconstruction quality in the low bitrate range. Compared with G-PCC (octree), the average BD-PSNR is increased by 11.70 dB and 11.08 dB on D1 and D2 respectively. Compared with G-PCC (trisoup), the average BD-PSNR is increased by 3.15 dB and 4.66 dB on D1 and D2 respectively. Compared with the state-of-the-art PCGC v2, the average BD-PSNR is increased by 0.22 dB and 0.37 dB on D1 and D2 respectively.
[0095] The present invention makes a visual comparison of the point clouds reconstructed by the proposed method with those reconstructed by the G-PCC (octree) and PCGC v2 methods. At a lower bitrate, for the feet of the longdress dataset, the method proposed by the present invention reconstructs the contour of the left foot more clearly and the details of the back of the right foot more completely. For the soldier dataset, the method proposed by the present invention can better focus on the edge part of the gun. In short, although at a lower bitrate, the method proposed by the present invention performs better in the boundary region, making the dataset contour clearer and more accurate.
[0096] Example 2
[0097] This embodiment provides a point cloud geometry compression system based on inter-layer residual and IRN connection residual, including:
[0098] A data acquisition module, configured to:
[0099] A computer-readable storage medium stores multiple instructions, and the instructions are adapted to be loaded and executed by a processor of a terminal device for a point cloud geometric compression method based on inter-layer residuals and IRN connection residuals as described above.
[0100] A terminal device includes a processor and a computer-readable storage medium. The processor is configured to implement each instruction, and the computer-readable storage medium is configured to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for a point cloud geometric compression method based on inter-layer residuals and IRN connection residuals as described above.
[0101] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.
Claims
1. A point cloud geometry compression method based on inter-layer residuals and IRN connection residuals, characterized in that: include: Get point cloud data; A convolutional neural network based on sparse convolution is used for feature extraction, in which upsampling after downsampling is performed based on the DU residual mechanism, and geometric reduction is performed with the feature information before upsampling to obtain up- and downsampling residuals; the up- and downsampling residuals are used to calculate the loss function, and the convolutional neural network is optimized by minimizing the loss function; The extracted features are captured based on the IRN connection residual mechanism to obtain global features; Encode and decode the global features based on the codec; Make predictions on the decoded information; The above method uses the up- and down-sampling residuals to calculate the loss function, and optimizes the convolutional neural network by minimizing the loss function, including using the up- and down-sampling residual result values to calculate the loss function. During the training process, the model is optimized by minimizing the loss function, which is expressed as: in, represents the residual loss value of upsampling and downsampling calculated on the encoder side, is the activation function, M is the number of points, The result value representing the coordinate residual; The upsampling after downsampling based on the DU residual mechanism is geometrically reduced with the feature information before upsampling to obtain the upsampling and downsampling residuals, including the downsampling feature information. Perform upsampling to obtain , and the feature information before upsampling Perform geometric reduction to obtain the residual , expressed as: in, Represents the coordinate residual result value, and They are respectively represented as the geometric coordinates of the i-th point before upsampling and downsampling, that is, the coordinate difference of each point in the three-dimensional component, expressed as ( ), represents the upsampling operation, represents the downsampling operation; The encoding and decoding of global features based on the encoder and decoder also includes introducing an attention module at the decoder end, using two point-by-point multiplication weighted operations based on the attention mechanism, capturing the global dependency range by directly calculating the relationship between the positions of scattered points, thereby enhancing the interactivity between features, and finally retaining the original point cloud information through residual connections while increasing the fluidity of the gradient, which is expressed as: in is the output of this module, is the input of this module.
2. A point cloud geometry compression method based on inter-layer residuals and IRN connection residuals according to claim 1, characterized in that: The feature extraction is performed using a convolutional neural network based on sparse convolution, including first performing a sparse convolution operation on the input 3D original point cloud data to expand the input 1-channel data to 16 channels, then activating it with a Relu function, and performing a downsampling operation using sparse convolution to expand the feature data from 16 channels to 32 channels while halving the size.
3. The point cloud geometry compression method based on inter-layer residuals and IRN connection residuals according to claim 2, characterized in that: The extracted features are captured based on the IRN connection residual mechanism to obtain global features, including stacking three IRN residuals using a multi-layer residual connection method, wherein the feature information is first processed by an IRN residual to obtain multi-scale feature information, and then the input feature information is connected with the multi-scale feature information and sent to the next IRN residual processing, and the original feature information of the point cloud data is retained through the residual connection, which is expressed as: in The output of the residual module is connected to IRN, Connect the input of the residual module to IRN, is the input of the IRN module, For connection operation.
4. The point cloud geometry compression method based on inter-layer residual and IRN connection residual according to claim 3, characterized in that: The encoding and decoding of global features based on the codec includes quantizing the global feature information, using the octree codec to process the geometric information in the global features, using the arithmetic codec to process the attribute information of the global features, and introducing the scale hyper-prior c to capture the scale changes of adjacent points in space to model the probability distribution of point cloud features.
5. The point cloud geometry compression method based on inter-layer residual and IRN connection residual according to claim 4, characterized in that: The predicting of the decoded information includes performing a sparse convolution operation on the decoded feature information. The predicted value of the voxel category is output by convolution based on the feature map of the previous layer, including the occupied and unoccupied conditions of the voxel; then the top k voxels with the highest scores are retained to ensure the performance of the model.
6. The point cloud geometry compression method based on inter-layer residuals and IRN connection residuals according to claim 5, characterized in that: The prediction of the decoded information also includes selecting the first k largest elements for processing to reduce data redundancy and improve calculation efficiency. The first k elements are calculated according to the binary cross entropy loss. The prediction probability is improved and the reconstruction error is reduced by minimizing the binary cross entropy loss, which is expressed as: Among them, N is the number of point clouds that need to be predicted, is the occupancy of the voxel, where 1 indicates occupancy and 0 indicates unoccupied. is the predicted probability value.
7. A point cloud geometry compression system based on inter-layer residuals and IRN connection residuals, executing the point cloud geometry compression method based on inter-layer residuals and IRN connection residuals as claimed in claim 1, characterized in that: include: The data acquisition module is configured to acquire point cloud data; The DU residual module is configured to perform feature extraction using a convolutional neural network based on sparse convolution, wherein upsampling after downsampling is performed based on the DU residual mechanism, and geometric reduction is performed with feature information before upsampling to obtain up- and downsampling residuals; a loss function is calculated using the up- and downsampling residuals, and the convolutional neural network is optimized by minimizing the loss function; The IRN connection residual module is configured to capture the extracted features based on the IRN connection residual mechanism to obtain global features; A codec module is configured to encode and decode the global feature based on the codec; The prediction module is configured to predict the decoded information.
Citation Information
Patent Citations
Image compression method and device based on deep learning
CN113259676A
Point cloud compression method, encoder, decoder and storage medium
CN113766228A