Point cloud geometric compression method based on sparse convolutional neural network
Through the point cloud geometric compression method based on sparse convolution neural network, large-scale point cloud data are processed, sparse and dense features are extracted, residual features are calculated and feature fusion is performed, and the problem of poor point cloud image quality in the existing technology is solved, and a more efficient encoding and decoding process is achieved.
Patent Information
- Application Number
- CN202510084233.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-20
AI Technical Summary
When existing point cloud geometric compression technology processes large-scale point cloud data, it is difficult to effectively preserve the detail richness, integrity and geometric accuracy of point clouds, resulting in poor image quality.
The point cloud geometric compression method based on sparse convolution neural network is adopted to process the sparse tensor of point cloud through sparse convolution, extract potential variables and sparse features, and calculate sparse dense residual features based on dense features, and perform feature fusion and reconstruction of point clouds.
Significantly improves the detail richness, integrity and geometric accuracy of the reconstructed point cloud, improves image quality, and reduces storage and transmission costs.
Smart Images

Figure CN119991834A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of point cloud compression coding, and in particular relates to a point cloud geometry compression method based on a sparse convolutional neural network. Background Art
[0002] Point cloud is a data set composed of multiple three-dimensional points, which is often used to represent the geometric structure of three-dimensional objects. Since point cloud data usually has high dimensionality and non-uniform properties, directly processing this data will consume a lot of computing resources. Therefore, it is necessary to compress the point cloud. Point cloud geometry compression (PCG) is an important technology in the field of computer vision and graphics processing. Its purpose is to reduce the redundancy of point cloud data through algorithms, thereby reducing its storage and transmission costs, while ensuring that the decompressed point cloud data can restore the original geometric structure as much as possible.
[0003] Point cloud compression technology mainly includes geometric information compression and attribute information compression. Geometric information compression usually uses a tree structure or block structure to divide and encode point clouds, such as the octree-based method. Attribute information compression reduces the redundancy between point cloud attributes through prediction, transform coding and other means.
[0004] With the advancement of deep learning technology, learning-based PCG compression methods have emerged. Most of them inherit the learned network architecture of 2D image compression, but use 3D convolution. Compared with traditional convolutional neural networks (CNN), sparse convolutional neural networks (Sparse Convolutional Neural Networks, Sparse CNN) take into account the sparsity of data and only calculate the convolution operation of non-zero elements. It can process sparse data more efficiently and is particularly suitable for processing data forms such as point clouds and sparse images.
[0005] The convolution kernel of sparse convolution is the same as that of traditional convolution, but its output is very different. There are two output definitions for sparse convolution. One is the regular output definition, just like normal convolution, the output point is calculated as long as the kernel covers an input point. The other is called submanifold output definition. The convolution output is calculated only when the center of the kernel covers the input point.
[0006] In terms of computational implementation of Sparse CNN: First, by establishing a serial number-coordinate hash table of the input and output tensors, the pixel coordinates of the input and output correspond to the serial number. Then, a Rulebook is constructed to establish a mapping relationship from the input serial number to the output serial number, which is the key step to implement sparse convolution. The implementation of sparse convolution is to query the Rulebook, match the convolution kernel weights and the input pixel values, and place the results in the corresponding position of the output tensor. It is implemented in parallel on the GPU to improve computational efficiency.
[0007] The data tensor is represented by a set of coordinates C = {(x i ,y i ,z i )} and related features F = {f i} to indicate that the convolution only gathers features at the occupied coordinates.
[0008] The computation of sparse convolution is defined as:
[0009] For u∈C out , C out Represents the output coordinates, C in Represents the input coordinates; represents the output feature vector at coordinate u, W represents the input feature vector at coordinate u. i Represents the kernel value of the 3D convolution kernel.
[0010] Define a three-dimensional convolution kernel: N 3 (u,C in )={i|u+i∈C in ,i∈N 3};
[0011] It covers a set of u-centered in This sparse convolution exploits the sparsity of point clouds to reduce complexity and only applies computations on occupied voxels.
[0012] Recent studies have shown that sparse convolutional neural networks (CNNs) are capable of processing large-scale point clouds and have very good performance in semantic segmentation and target detection, but there is still a lack of research on PCG in point cloud geometry compression. Summary of the invention
[0013] In order to solve the above technical problems, the present invention proposes a point cloud geometry compression method based on a sparse convolutional neural network, which includes:
[0014] S1: Get scene point cloud;
[0015] S2: Use sparse convolutional neural network to process the point cloud sparse tensor to obtain latent variables and sparse features respectively, and decode the sparse features to obtain sparse prediction features;
[0016] S3: Extract features from the dense tensor of the point cloud to obtain dense features;
[0017] S4: Calculate sparse and dense residual features according to the sparse prediction features and the dense features;
[0018] S5: Obtain a reconstructed feature by performing sparse convolution upsampling on the latent variable;
[0019] S6: Fusing the sparse and dense residual features and the reconstructed features to obtain fused features;
[0020] S7: Perform sparse convolution upsampling on the fused features to obtain the reconstructed point cloud.
[0021] Beneficial effects of the present invention: The present invention utilizes a sparse convolutional neural network to extract sparse features and latent variables of point clouds, and combines operations such as dense feature extraction and sparse-dense residual feature fusion to effectively improve the detail richness, completeness and geometric accuracy of the reconstructed point cloud, thereby significantly improving the image quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of a flow chart of an embodiment of the present invention;
[0023] Figure 2 is a flowchart of steps of an embodiment of the present invention;
[0024] Figure 3 Schematic diagram of the structure of a sparse convolution downsampling module in an embodiment of the present invention;
[0025] Figure 4 Schematic diagram of the structure of the initial residual network module IRN in an embodiment of the present invention;
[0026] Figure 5 Schematic diagram of the structure of a sparse convolution upsampling module in an embodiment of the present invention;
[0027] Figure 6 PSNR curve diagram of PCGC coding and the soldier test sequence tested by the present invention;
[0028] Figure 7 PSNR curve diagram of PCGC coding and the present invention testing the longdress test sequence. DETAILED DESCRIPTION
[0029] The terms "first", "second", "third", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is merely a way of distinguishing objects with the same attributes when describing the embodiments of the present application.
[0030] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0031] This paper studies how to combine point cloud coding with sparse convolutional neural network Sparse CNN, and utilizes the feature extraction and compression capabilities of neural networks to achieve more efficient encoding and decoding processes. The present invention uses sparse convolution for low-complexity tensor processing. At the same time, the initial residual network (Inception-Residual Network, IRN) unit is used for effective feature extraction. For downscaling on the encoder side, convolution is applied with a step size of 2, and each step halves the scale of each geometric dimension.
[0032] The embodiment of the present invention proposes a point cloud geometry compression method based on a sparse convolutional neural network. Figure 1 , 2 As shown, the method includes:
[0033] S1: Get scene point cloud.
[0034] In an optional embodiment, the point cloud data used is obtained from the open source dataset 8i Voxelized FullBodies (8iVFB v2).
[0035] For example, obtain the scene point cloud (Point Cloud), whose sparse tensor representation is: {C X , F X}, where C X Represents the coordinate information of each point in the point cloud X, usually a three-dimensional vector containing the position of the point in space, such as (x, y, z) coordinates. X Indicates the feature information of each point in the point cloud X. The feature information can be various attributes related to the point, such as color, intensity, normal, etc. The feature information can be a multidimensional vector, and its dimension depends on the type and number of features included. The embodiment of the present invention mainly focuses on the coordinate information part (C X ).
[0036] S2: Use a sparse convolutional neural network to process the point cloud sparse tensor to obtain latent variables and sparse features respectively, and then decode the sparse features to obtain sparse prediction features.
[0037] In a preferred embodiment, a sparse neural network is used to perform sparse convolution downsampling and residual connection on the sparse tensor of the point cloud to obtain latent variables and sparse features, and the sparse features are decoded to obtain sparse prediction features.
[0038] In a preferred embodiment, a sparse convolutional neural network is used to process the point cloud sparse tensor, and the specific processing process includes:
[0039] S201: Using the sparse convolution downsampling module in the sparse neural network Sparse CNN, downsample the sparse tensor of the point cloud to obtain a first sparse feature map.
[0040] Figure 3 Schematic diagram of the structure of the sparse convolution downsampling module in an embodiment of the present invention.
[0041] Reference Figure 3 As shown, the network structure of the sparse convolution downsampling module includes: a 3D sparse convolution layer with 16 channels and a convolution kernel of 3*3*3, a first ReLU activation function layer, a 3D sparse convolution layer with 32 channels and a convolution kernel of 3*3*3 / 2↓, and a second ReLU activation function layer.
[0042] Specifically, refer to Figure 3 As shown, the point cloud sparse tensor is input into the sparse neural network Sparse CNN, and downsampled through the sparse convolution downsampling module to obtain the first sparse feature map.
[0043] S202: Use an initial residual network IRN to perform a reversible bijective transformation on the first sparse feature map to obtain a latent variable Y, and obtain a sparse feature p according to the latent variable Y.
[0044] Figure 4 Schematic diagram of the structure of the initial residual network IRN in an embodiment of the present invention.
[0045] Reference Figure 4As shown, the network structure of the initial residual network IRN includes: N input channels, a first branch and a second branch, the output result of the first branch is added to the output result of the second branch, and then added to the N input channels to obtain the output result of the initial residual network IRN; the first branch includes: a 3d sparse convolution layer with N / 4 channels and a convolution kernel of 3*3*3, a third ReLU activation function layer, a 3d sparse convolution layer with N / 2 channels and a convolution kernel of 3*3*3, and a fourth ReLU activation function layer; the second branch includes: a 3d sparse convolution layer with N / 4 channels and a convolution kernel of 1*1*1, a fifth ReLU activation function layer, a 3d sparse convolution layer with N / 4 channels and a convolution kernel of 3*3*3, a sixth ReLU activation function layer, and a 3d sparse convolution layer with N / 2 channels and a convolution kernel of 1*1*1.
[0046] IRN implements sampling and obtains sparse features and latent variables in sparse feature maps through IRN.
[0047] For example, refer to Figure 4 As shown, x represents the input feature map, whose size is H×W×C; the scaling ratio is s and the model f θ,s , the model output is a low-resolution image p, whose size is By model f θ,s The input image x is calculated to generate two results (p, Y), where p is a low-resolution image, i.e., a sparse feature, and Y is a latent variable (i.e., latent representation); finally, the sparse feature p is returned.
[0048] S203: Decode the sparse feature p using a sparse neural network Sparse CNN to obtain a sparse prediction feature P'.
[0049] S3: Obtain reconstructed features by performing sparse convolution upsampling on the latent variable Y.
[0050] Figure 4 Schematic diagram of the structure of the sparse convolution upsampling module in an embodiment of the present invention.
[0051] Reference Figure 4 As shown, the network structure of the sparse convolution upsampling module includes: a 3D sparse convolution layer with 32 channels and a convolution kernel of 3*3*3 / 2↑, a seventh ReLU activation function layer, a 3D sparse convolution layer with 32 channels and a convolution kernel of 3*3*3, and an eighth ReLU activation function layer.
[0052] The reconstruction feature F r The expression is:
[0053] in, Represents the reconstruction feature F r The coordinates of .
[0054] The specific steps of upsampling are: the input is a sparse feature p, whose size is H×W×C, and also includes a scaling ratio s and a model f θ,s , the model output is a high-resolution image x, whose size is sH×sW×C.
[0055] Specifically: First, obtain the prior distribution p(z) through the latent variable Y. Randomly sample z from the prior distribution p(z), whose dimension is H×W×(s 2 -1). Then, through the inverse transformation f of the model θ,s -1 The high-resolution image x is calculated. Finally, the generated high-resolution image x is returned. x is based on the reconstructed feature F r The obtained feature map.
[0056] S4: Extract features from the dense tensor of the point cloud to obtain dense features.
[0057] In one embodiment, dense features are obtained by performing coordinate transformation, quantization processing, and removing duplicate points on the dense tensor of the point cloud.
[0058] Specifically: In the point cloud dense tensor, the position information of the point cloud Xn|n=1,…,N is usually represented by a floating point number and is located in the world coordinate system.
[0059] First, all points in the point cloud X are translated to the coordinate origin and converted to the object coordinate system, and then quantization is performed. The expression is:
[0060] X n '=(X n -T) / q;
[0061] T=(min(x n ),min(y n ),min(z n ))|n=1,…,N;
[0062] Where, X n ' represents the nth point in the point cloud X after quantization, X n represents the nth point in the point cloud X, T represents the minimum value of the point cloud X in each coordinate axis direction, q represents the quantization step size, q is set by the user, and the parameters T and q make all coordinates in the range [0,2^d), where d is a non-negative integer. The calculation formula is:
[0063] d=Ceil(Log2(max(x n ,y n ,z n )|n=1,…,N)+1);
[0064] In the formula, Ceil(.) represents the upward rounding function, specifically taking the smallest integer greater than or equal to (.), max(.) represents the maximum value function, (x n ,y n ,z n ) represents the coordinates of the nth point in the point cloud X, and N represents the number of points in the point cloud.
[0065] In order to facilitate division, the quantized position information also needs to be rounded to convert the floating point number into an integer, as shown in the following formula:
[0066] Int(X n ')=Round(X n '),
[0067] In the formula, Int(X n ') indicates rounding the position information of the midpoint of Xn, and Round(.) indicates the rounding function.
[0068] After quantization and rounding, there may be multiple points with the same geometric position. It is usually necessary to delete the duplicate points so that there is only one point with the same geometric position for easy division. After removing the duplicate points, the dense feature P" is obtained.
[0069] S5: Calculate sparse and dense residual features according to the sparse prediction features P' and the dense features P".
[0070] The sparse dense residual feature F c The expression is:
[0071]
[0072] In the formula, Represents the coordinates of sparse and dense residual features.
[0073] Residual calculation of sparse prediction features P' and dense features P" can effectively retain feature information, and the residual is smaller, and the features are easier to learn, especially when the input and output features are similar.
[0074] S6: Fusing the sparse and dense residual features and the reconstructed features to obtain fused features.
[0075] Specifically, the sparse and dense residual features F c and reconstruction feature F r They are concatenated and then processed through a multi-layer perceptron (MLP) to obtain fused features.
[0076] The fused features are quantized, duplicate points are removed, and the reconstructed point cloud is obtained through sparse convolution upsampling.
[0077] Hierarchical reconstruction based on binary classification is used to reconstruct the point cloud. Binary classification is used to classify whether the generated voxels are occupied. Sparse convolutional layers are used to generate the probability of voxels being occupied after continuous convolution.
[0078] The loss function is set during training, and its expression is:
[0079]
[0080] Among them, x i represents the voxel label that is actually occupied (1) or empty (0), p i Represents the probability of an occupied voxel, activated by a sigmoid function.
[0081] Experimental verification:
[0082] Simulation platform or software: Linux server with two 3090 graphics cards, software: vscode, cmake, git.
[0083] Data source used in simulation: The point cloud data used in this invention are all obtained from the open source dataset 8i Voxelized Full Bodies (8iVFB v2).
[0084] Evaluation indicators description:
[0085] BD Rate is a metric used to evaluate video coding efficiency, especially when comparing different video encoders or encoding parameters. BD Rate measures the change in bit rate required for one encoding method relative to another encoding method while maintaining the same quality level (such as PSNR or SSIM). The lower the BD Rate value, the more efficient the encoding method.
[0086] PSNR (Peak Signal-to-Noise Ratio) is a metric used to measure the quality of an image or video by comparing the difference between the original image and the compressed or processed image. The higher the PSNR value, the better the quality of the image or video, that is, the less distortion or noise. PSNR is usually expressed in decibels (dB).
[0087] Simulation results:
[0088] According to Table 1, the present invention significantly reduces the BD Rate and improves the coding efficiency.
[0089] Table 1 is a comparison of octree encoding, triangle soup encoding, PCGC encoding and the BD-Rate of the present invention.
[0090]
[0091] In Table 1, PointCloud represents point cloud, G-PCC (octree) represents point cloud geometry octree encoding, G-PCC (trisoup) represents point cloud geometry triangle soup encoding, PCGC represents point cloud geometry encoding, soldier represents point cloud sample soldier, longdress represents point cloud sample longdress, loot represents point cloud sample loot, and Average represents average bdrate.
[0092] Figure 6 PSNR curve diagram of PCGC coding and the present invention testing the soldier test sequence. Figure 6 Ourmethod is the PSNR curve obtained by the present invention when simulating the soldier test sequence. Figure 6 It can be seen that for the soldier test sequence, the PSNR of the present invention is higher and the image quality is better at various resolutions.
[0093] Figure 7 PSNR curve diagram of PCGC coding and the present invention testing the longdress test sequence. Figure 7 Ourmethod is the PSNR curve obtained by the present invention when simulating the longdress test sequence. Figure 7 It can be seen from the figure that for the longdress test sequence, the PSNR of the present invention is higher and the image quality is better at various resolutions.
[0094] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium, which can include: ROM, RAM, disk or CD, etc.
[0095] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A point cloud geometry compression method based on sparse convolutional neural network, characterized in that: This method includes: obtaining scene point cloud; The sparse convolutional neural network is used to process the point cloud sparse tensor to obtain latent variables and sparse features respectively, and the sparse features are decoded to obtain sparse prediction features; Obtaining a reconstructed feature by performing sparse convolution upsampling on the latent variable; Perform feature extraction on the dense tensor of the point cloud to obtain dense features; Calculating sparse and dense residual features according to the sparse prediction features and the dense features; Performing feature fusion on the sparse and dense residual features and the reconstructed features to obtain fused features; Sparse convolution upsampling is performed on the fused features to obtain the reconstructed point cloud.
2. The point cloud geometry compression method based on sparse convolutional neural network according to claim 1, characterized in that: Using sparse neural networks, the sparse tensors of point clouds are subjected to sparse convolution downsampling and residual connection to obtain latent variables and sparse features. After decoding the sparse features, sparse prediction features are obtained.
3. The point cloud geometry compression method based on sparse convolutional neural network according to claim 1 or 2, characterized in that: The sparse convolutional neural network is used to process the point cloud sparse tensor. The specific processing process includes: The sparse convolution downsampling module in the sparse neural network Sparse CNN is used to downsample the sparse tensor of the point cloud to obtain the first sparse feature map; Using an initial residual network IRN to perform a reversible bijective transformation on the first sparse feature map to obtain a latent variable Y, and obtaining a sparse feature p according to the latent variable Y; The sparse neural network Sparse CNN is used to decode the sparse feature p to obtain the sparse prediction feature P'.
4. The point cloud geometry compression method based on sparse convolutional neural network according to claim 1, characterized in that: The network structure of the sparse convolution downsampling module includes: a 3D sparse convolution layer with 16 channels and a convolution kernel of 3*3*3, a first ReLU activation function layer, a 3D sparse convolution layer with 32 channels and a convolution kernel of 3*3*3 / 2, and a second ReLU activation function layer.
5. The point cloud geometry compression method based on sparse convolutional neural network according to claim 1, characterized in that: The network structure of the initial residual network IRN includes: N input channels, a first branch and a second branch, the output result of the first branch is added to the output result of the second branch, and then added to the N input channels to obtain the output result of the initial residual network IRN; the first branch includes: a 3d sparse convolution layer with N / 4 channels and a convolution kernel of 3*3*3, a third ReLU activation function layer, a 3d sparse convolution layer with N / 2 channels and a convolution kernel of 3*3*3, and a fourth ReLU activation function layer; the second branch includes: a 3d sparse convolution layer with N / 4 channels and a convolution kernel of 1*1*1, a fifth ReLU activation function layer, a 3d sparse convolution layer with N / 4 channels and a convolution kernel of 3*3*3, a sixth ReLU activation function layer, and a 3d sparse convolution layer with N / 2 channels and a convolution kernel of 1*1*1.
6. The point cloud geometry compression method based on sparse convolutional neural network according to claim 1, characterized in that: The network structure of the sparse convolution upsampling module includes: a 3D sparse convolution layer with 32 channels and a convolution kernel of 3*3*3 / 2, a seventh ReLU activation function layer, a 3D sparse convolution layer with 32 channels and a convolution kernel of 3*3*3, and an eighth ReLU activation function layer.
7. The point cloud geometry compression method based on sparse convolutional neural network according to claim 1, characterized in that: After coordinate transformation, quantization and removal of duplicate points on the dense tensor of the point cloud, dense features are obtained.
Citation Information
Patent Citations
Point cloud geometric lossless compression method based on sparse convolutional neural network
CN113613010A
Large-scale point cloud geometric compression method based on double-branch neural network
CN116128985A
Point cloud geometric compression method based on space channel hybrid context model
CN117974818A
Hybrid framework for point cloud compression
CN118402234A