An end-to-end compression method for point clouds based on efficient sampling and optimized entropy model

By introducing efficient sampling and optimized entropy models into point cloud compression, the problem of point cloud data redundancy and low compression efficiency in the existing technology is solved, and efficient and low redundancy point cloud data compression is achieved.

CN119135925BActive Publication Date: 2025-05-23CHINA JILIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411621137.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-05-23
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The existing learning-based point cloud geometric compression method ignores the efficient sampling of point cloud data and the use of multiple context information, resulting in the inability to effectively handle geometric redundancy during the compression process, affecting compression efficiency and quality.

Method used

The point cloud end-to-end compression method based on efficient sampling and optimization entropy model is adopted, and effective data is extracted to optimize sampling efficiency and distortion rate of feature reconstruction through technical means such as grid regression, neural graph sampling, feature extraction and asymmetric spatial channel entropy module.

Benefits of technology

It realizes low redundancy and high efficiency compression of point cloud data, improves the sampling efficiency of the encoding process and the accuracy of feature reconstruction, and significantly improves the compression performance of point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119135925B_ABST
    Figure CN119135925B_ABST
Patent Text Reader

Abstract

The present invention discloses an end-to-end compression method for point cloud based on efficient sampling and optimized entropy model. In the prior art, a neural graph sampling module is designed by utilizing the construction of local graphs, graph feature embedding and sampling based on attention, and is embedded into the feature extraction network in the encoding process. A multi-module stacking method is adopted to extract the potential key points of the original point cloud, which effectively improves the sampling efficiency. The asymmetric spatial channel entropy module uses a deformable dynamic kernel to expand the spatial aggregation capability, and groups the potential variables along the channel dimension, which reduces statistical redundancy while maintaining the encoding efficiency. Based on the addition of the neural graph sampling module and the asymmetric spatial channel entropy module, the present invention realizes the acquisition of the most valuable sampling points and the effective optimization of the distribution parameter estimation, thereby realizing the end-to-end compression of massive point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and point cloud geometry compression, and specifically relates to an end-to-end compression method for point cloud based on efficient sampling and optimized entropy model. Background Art

[0002] Whether it is a dense point cloud or a sparse point cloud, it is composed of geometric information and attribute information of millions or even tens of millions of points. In addition, the point structure in the point cloud does not have topological connection and is often disordered and sparsely distributed in three-dimensional space, which makes it difficult to directly compress three-dimensional point clouds using traditional image or video encoding technology. Therefore, in the case of being unable to expand storage resources and network bandwidth in large quantities, how to reduce the volume of point cloud data, accelerate the retrieval and query of point cloud data, and achieve efficient transmission under limited resources is a huge challenge.

[0003] The existing methods all use learning-based point cloud geometry compression technology, which can more accurately reconstruct the shape and topological structure of point clouds compared to traditional compression methods. This technology uses an end-to-end training method and can directly use the original point cloud data for encoding and decoding. However, most of the current learning-based point cloud geometry compression methods ignore the efficient sampling of point cloud data and the use of multiple contextual information, resulting in the inability to effectively handle geometric redundancy during the compression process. The introduction of the efficient sampling module and the optimized entropy module can extract valid data from the original point cloud, improve the sampling efficiency of the encoding process, accurately estimate the residual features after quantization, optimize the distortion rate of potential feature reconstruction, and achieve low-redundancy and high-efficiency compression of point cloud data. Summary of the invention

[0004] In order to solve the shortcomings of the prior art, the present invention adopts the following technical solutions to achieve the purpose of efficient and fast compression encoding of point cloud data by further extracting key points and improving rate distortion performance:

[0005] A point cloud end-to-end compression method based on efficient sampling and optimized entropy model includes the following steps:

[0006] Step 1: Perform grid regression on the original point cloud, then obtain a bit stream with grid information through grid quantization and arithmetic coding, obtain the reconstructed grid through arithmetic decoding and grid restoration, and perform grid-to-point cloud conversion to generate a predicted point cloud;

[0007] Step 2: Perform neural graph sampling on the input raw point cloud; construct a local graph for each point, whose center point passes through a point dynamic filter to aggregate feature attributes, and then perform attention-based sampling to select a subset of points to well represent the input point;

[0008] Step 3: Extract features from the predicted point cloud and the original point cloud after neural graph sampling to obtain a multi-scale sparse tensor, and use coordinate encoding to obtain a coordinate bit stream, which is then subjected to feature mapping to obtain mapping features of the predicted point cloud. Based on the mapping features and the features extracted from the features, residual features are obtained by feature subtraction;

[0009] Step 4: Input the residual features into the asymmetric spatial channel entropy, and perform different offsets and vector projections on each group of the residual features by grouping the deformable dynamic kernel to enhance the expression ability of the module and obtain the latent variables. Then, the latent variables are grouped along the channel dimension. Each group except the first one captures the helpful spatial context from the previous groups, effectively using spatial information to reduce redundancy and optimize the residual features.

[0010] Step 5: Restore the grid from the grid bit stream through arithmetic decoding, then convert the grid to point cloud to generate a predicted point cloud, and then perform feature mapping on the predicted point cloud after feature extraction. Finally, combine the mapped features with the decoded coordinates and the residual features restored after the asymmetric spatial channel entropy module to obtain the features of the original point cloud. These features are sent to the feature propagation module to reconstruct the original point cloud.

[0011] Furthermore, the step 1 comprises the following steps:

[0012] Step 1.1: A mesh regression module that combines geometric changes and surface deformations to obtain prediction parameters using a pre-trained PointNet++ network and a probabilistic correspondence association module. The parameters are then mesh-quantized and encoded into a bitstream. The parameters include but are not limited to pose, shape, rotation, translation, and gender.

[0013] Step 1.2: To ensure that the encoder and decoder are fully synchronized, the aligned grid is restored from the quantization parameters of the grid operation block, and then the reconstructed grid is converted to a point cloud through the Poisson disk sampling algorithm to generate a predicted point cloud with a similar shape.

[0014] Furthermore, the step 2 comprises the following steps:

[0015] Step 2.1: Local graph construction: Use the K nearest neighbor algorithm to construct the K nearest neighbor graph of each point, and obtain the spatial K nearest neighbor graph features. For the original point cloud with attributes, based on the K nearest neighbor index in the spatial domain, the K nearest neighbor graph attributes of each point are obtained in the feature domain. The K nearest neighbor graph attributes and the spatial K nearest neighbor graph features are connected for feature embedding to obtain aggregated local geometric features.

[0016] Step 2.2: Graph feature embedding: construct dynamic filtering based on the K nearest neighbor graph to focus on local spatial relationship modeling, correspond points to convolution, and then embed features to obtain point-based adaptive convolution;

[0017] Step 2.3: Attention-based sampling: For the output of the graph feature embedding, select a representative subset with differentiable operations, and use an attention-based method to select a subset of points in the original point cloud to obtain the most valuable points for sampling.

[0018] Furthermore, in step 2.1, a set of three-dimensional point cloud geometric figures X'= {X i '∈R 3 , i = 1,···,N},X i '=(x i ,y i , z i ) represents the three-dimensional coordinates of each point. For the input point cloud of N×3 matrix, a directed K-nearest neighbor graph G = (V, E) is constructed to represent the local point cloud structure, V = {1, ···, N}, Used to represent vertices and edges, v i Represents X i ', by constructing a K nearest neighbor graph, we get the N×K×6 local structure of the point cloud, which contains the three-dimensional coordinates and edge features of each point in the spatial domain; for the point with attribute N×F i The original point cloud, F i Represents the input feature channel. Based on the K nearest neighbor index in the spatial domain, the K nearest neighbor graph attribute N×K×F of each point is obtained in the feature domain. i Finally, the K nearest neighbor graph attributes N×K×F i And the spatial K nearest neighbor graph feature N×K×6 is connected for feature embedding.

[0019] Furthermore, in the step 2.2, in order to adaptively extract the potential representation in the feature domain, a point dynamic filter is first generated based on the input K nearest neighbor graph to aggregate the neighbor information to the center point; then the point dynamic filter is applied to the input K nearest neighbor graph to output the dynamic filtering result, which focuses on local spatial relationship modeling, with one convolution for each point; finally, a fully connected layer is used for feature embedding, taking local structure and geometric changes as potential features, and the operation is performed twice to improve the feature embedding capability.

[0020] Furthermore, in step 2.3, for the graph feature embedding output , N represents the number of points in a specific point cloud, F o Represents the number of output feature channels, and the goal is to select a representative subset with differentiable operations , where the output points are sampled to N / 4, and the attention-based multi-instance learning pooling method (MIL) is used to obtain the most valuable points by selecting a subset of points in the original point cloud. The pooling output is expressed as:

[0021] Q = softmax(ωP T )

[0022] Among them, Q represents a weighted average of the standardized scores generated by the elements, ω represents a learnable weight parameter, softmax represents the normalization function, and the superscript T represents the transposed symbol.

[0023] Furthermore, the step 3 comprises the following steps:

[0024] Step 3.1: Feature extraction; based on sparse convolution to retain key point features, the intermediate results between modules are represented by sparse tensors; the sparse tensor X only uses the coordinate feature pairs Save the non-zero elements, each non-zero 3D coordinate (x i ,y i , z i )∈C, corresponding to the relevant feature f i ∈F, so that there is a rich context accurate representation. By applying convolution filtering only to non-zero elements, sparse convolution significantly reduces the computational complexity and memory usage. The sparse convolution calculation formula in three-dimensional space is as follows:

[0025]

[0026] Among them, C in and C out represents the input and output coordinates, W i Indicates i The weight of the kernel of points, set N 3 represents a three-dimensional kernel with arbitrary shape, whose subset includes C in The offset from the current three-dimensional coordinate u, C represents the coordinate set, and F represents the feature set;

[0027] Each downsampling block in feature extraction consists of a serial convolution, a VRN unit, and another serial convolution layer. The serial convolution is used to reduce the number of points, while the subsequent convolution layer is used to improve the extracted features to obtain the best performance. The VRN unit is located between two convolution layers, and skip connections are used to alleviate information loss during training. Parallel convolution layers with different kernel sizes, i.e., 1×1×1 or 3×3×3, are used to capture features in different ranges. By superimposing these downsampling blocks, the feature extraction module outputs multi-scale sparse tensors, including coordinates and features, forming a potential representation of the entire point cloud. Therefore, the extracted features can be used with a small number of output coordinates as necessary information to reduce the required bits, and these coordinates are losslessly encoded using a coordinate encoder.

[0028] Step 3.2: Feature Mapping; Convert the extracted features from the predicted point cloud to the coordinates of the original point cloud. This process is done in a coarse-to-fine manner. Each scale contains a hierarchical connection of a main sparse tensor from the feature extractor and an auxiliary sparse tensor from the previous block. To this end, the previous block utilizes a hierarchical convolution layer similar to the feature extraction module; Finally, the generalized sparse transposed convolution layer is used to convolve the connected features of the predicted point cloud on the downscaled coordinates of the original point cloud. The process is expressed as:

[0029]

[0030] in, S Represents the original point cloud, subscript T represents the predicted point cloud, L Indicates scale, Indicates the predicted point T In scale L The main sparse tensor of Indicates the predicted point T In scale L The auxiliary sparse tensor of , represents the connection operation, and the function G(·) represents the operation performed by the convolution layer on the target coordinates; after the convolution operation on the coordinates, the output sparse tensor includes the coordinates of the original point cloud And the mapping features of the predicted point cloud at the last scale .

[0031] Step 3.3: Feature subtraction: After obtaining the mapping features of the predicted point cloud, feature level subtraction is performed directly to calculate the extracted features and as a feature of motion mapping The feature-level residual between reduces the prediction error caused by motion compensation based on optical flow in the pixel domain. The process is expressed as:

[0032]

[0033] in, Represents the residual features in the last scale.

[0034] Furthermore, step 4 comprises the following steps:

[0035] Step 4.1: During encoding, the input features x are sent to the Encoder g a , then obtain the latent representation y, and encode the quantized residual between the latent representation y and the variance µ into the bitstream. In order to encode y and µ using arithmetic coding, use Q (·) performs quantization operation; then in the decoding process, the decoder gs Convert the decoded latency into reconstructed features, and at the same time utilize the following overall framework of dynamic spatial aggregation and asymmetric spatial channel entropy enhanced residual feature compression to further improve the rate-distortion performance;

[0036] Step 4.2: Most existing methods use stacked convolutions or window-based self-attention for transform coding, which can only aggregate features within a fixed spatial range. Therefore, an adaptive spatial aggregation is proposed to be achieved through deformable kernels. However, blindly introducing deformable convolutions may consume more computing resources. First, design an efficient and stable dynamic residual block group, and use depthwise convolutions to generate content-adaptive kernel offsets and modulation weights respectively; then divide the input features into several groups, and share the modulated kernel weights within each group, so that not only the advantages of deformable convolutions can be maintained, but also the computational burden will not be increased;

[0037] Step 4.3: By encoding the latent representation y as side information in the hyperprior part and introducing dynamic kernels to aggregate the global prior distribution, the global context expression of the generalized entropy model is strengthened; in addition, design an asymmetric spatial channel entropy transformation, group the latent variable y along the channel dimension, extract the channel context, and extract the spatial context from each group. Each group except the first group captures helpful spatial context from the previous groups. The first two groups use four-stage spatial context, and the subsequent groups use two-stage spatial context to effectively utilize spatial information to reduce redundancy. Combine the global context information in the hyperprior, and finally optimize the residual features through distribution parameter estimation.

[0038] Furthermore, in the above step 4.1, the overall framework of residual feature compression is as follows:

[0039]

[0040]

[0041]

[0042] where x and represent the input and decompressed features, g a and g s represent the analysis and synthesis transforms respectively, and represent the learned parameters of the analysis transform network and the synthesis transform network respectively, Q (·) represents the quantization operation.

[0043] Furthermore, in the above step 4.2, the method of using deformable kernels to achieve adaptive spatial aggregation is described as:

[0044]

[0045] in, G represents the number of split groups, , Indicates g The modulated kernel weights and additional learning offsets are shared within each subgroup. F (·) represents the feature map, Indicates that the feature map corresponds to the center of the convolution kernel, Indicates the relative coordinates of each position of the corresponding convolution kernel;

[0046] Finally, a network is used to project the feature, which is expressed as:

[0047]

[0048] Among them, Conv 1×1 (·) represents a 1×1 convolution operation, and GELU(·) represents the activation function.

[0049] The advantages and beneficial effects of the present invention are:

[0050] The present invention discloses an end-to-end compression method for point clouds based on efficient sampling and optimized entropy models. A neural graph sampling module is designed by utilizing the construction of local graphs, graph feature embedding, and sampling based on attention, and is embedded into a feature extraction network in the encoding process. A multi-module stacking method is used to extract potential key points of the original point cloud, which effectively improves the sampling efficiency. The asymmetric spatial channel entropy module uses a deformable dynamic kernel to expand the spatial aggregation capability and groups potential variables along the channel dimension, reducing statistical redundancy while maintaining encoding efficiency. The addition of the above two modules realizes the acquisition of the most valuable sampling points and the effective optimization of distribution parameter estimation, thereby achieving end-to-end compression of massive point cloud data. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flow chart of a method in an embodiment of the present invention.

[0052] Figure 2 Schematic diagram of the structure of the neural graph sampling module in an embodiment of the present invention.

[0053] Figure 3 Schematic diagram of the structure of an asymmetric spatial channel entropy module in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.

[0055] like Figure 1 As shown, a point cloud end-to-end compression method based on efficient sampling and optimized entropy model of the present invention comprises the following steps:

[0056] Step 1: During the encoding process, the original point cloud is used as input for grid regression, and then quantized and arithmetic encoded to obtain a bit stream with grid information, and then the grid is converted to point cloud to generate a predicted point cloud, which specifically includes the following steps:

[0057] Step 1.1: The mesh regression module combines geometric changes and surface deformations to model and utilizes the pre-trained PointNet++ network and the probabilistic correspondence association module. The resulting prediction parameters are quantized and encoded into a bitstream, which can be expressed as:

[0058]

[0059] in They represent posture, shape, rotation, translation and gender respectively (when the point cloud object is a human, gender is represented by male and female, and when it is an animal or plant, gender is represented by male and female, or male and female).

[0060] Step 1.2: To ensure that the encoder and decoder are fully synchronized, an aligned grid is restored from the quantization parameters in the grid operation block, and then the reconstructed grid is input into the grid-to-point cloud conversion through the Poisson disk sampling algorithm to generate a predicted point cloud with similar shape.

[0061] Step 2: Input the original point cloud into the neural graph sampling module, first construct a local graph for each point, and for each local graph, the center point of the graph aggregates the neighbor weights through the point dynamic filter to expand the relevant feature attributes. Then, perform attention-based sampling to select a subset of points to well represent the input point, such as Figure 2 As shown in Figure 1, the execution process of the neural graph sampling module includes the following steps:

[0062] Step 2.1: Local graph construction: Use the K nearest neighbor algorithm to construct a local graph of each point to aggregate local geometric features. Let X'= { X i '∈R 3 , i = 1,···,N} represents a set of 3D point cloud geometries with N points, and X i '=(x i ,y i , z i ) is the three-dimensional coordinate of each point. For the input point cloud represented as an N×3 matrix, a directed K-nearest neighbor graph G = (V, E) is constructed to represent the local point cloud structure, where V = {1,···,N}, are vertices and edges, v i Yes X i The neighbor set of .

[0063] By constructing a K-nearest neighbor graph, we can obtain an N×K×6 local structure representation of the point cloud, which contains the 3D coordinates and edge features of each point in the spatial domain. i , where F i As the input feature channel, based on the K nearest neighbor index in the spatial domain, the K nearest neighbor graph attribute N×K×F of each point is obtained in the feature domain i Finally, the K nearest neighbor graph attributes N×K×F i And the spatial K nearest neighbor graph features N×K×6 are connected together for feature embedding.

[0064] Step 2.2: Graph feature embedding: In order to adaptively extract potential representations in the feature domain, first generate a point dynamic filter based on the input K nearest neighbor graph to aggregate neighbor information to the center point; then apply the point dynamic filter to the input K nearest neighbor graph and output the dynamic filtering result. The dynamic filtering focuses on modeling local spatial relationships, with 1 convolution per point; finally, use a fully connected layer for feature embedding. This step uses local structure and geometric changes as potential features and is performed twice to improve feature embedding capabilities.

[0065] Step 2.3: Attention-based sampling: For graph feature embedding output , N is the number of points in a particular point cloud, F o is the number of output feature channels, and the goal is to select a representative subset with differentiable operations , where the output points are sampled to N / 4. And using the attention-based multi-instance learning pooling method (MIL), its pooling output is expressed as:

[0066] Q = softmax(ωP T )

[0067] Where Q is a weighted average of the standardized scores generated by the elements, and ω is a learnable weight parameter. By selecting a subset of points in the original point cloud, the most valuable points can be obtained. Softmax represents the normalization function, and the superscript T represents the transposed symbol.

[0068] Step 3: Input the predicted point cloud and the original point cloud after neural graph sampling into the feature extraction module, output the multi-scale sparse tensor, and use the coordinate encoder to obtain the coordinate bit stream, then pass it through the feature mapping module to obtain the mapping features, and obtain the residual features through feature subtraction, which specifically includes the following steps:

[0069] Step 3.1: Feature extraction: Based on sparse convolution to retain key point features, the intermediate results between modules are represented by sparse tensors. Specifically, the sparse tensor X only uses the coordinate feature pairs Save the non-zero elements, each non-zero coordinate (x i ,y i , z i )∈C, corresponding to the relevant feature f i ∈F, so that there is a rich context accurate representation. By applying convolution filters only to non-zero elements, sparse convolution significantly reduces computational complexity and memory usage. The sparse convolution calculation formula in three-dimensional space is as follows:

[0070]

[0071] Among them C in and C out are the input and output coordinates, W i For the i The weight of the kernel of points, set N 3 represents a 3D kernel with arbitrary shape, whose subset includes C in The offset from the current 3D coordinate u in .

[0072] Each downsampling block in the feature extractor consists of a serial convolution, a VRN unit, and another serial convolution layer. Specifically, the serial convolution reduces the number of points, while the subsequent convolution layer improves the extracted features to obtain the best performance; the VRN unit is located between the two convolution layers, using skip connections to mitigate information loss during training, and adopts parallel convolution layers with different kernel sizes, i.e., 1×1×1 or 3×3×3, to capture features in different ranges; by superimposing these downsampling blocks, the feature extraction module outputs multi-scale sparse tensors, including coordinates and features, forming a potential representation of the entire point cloud. Therefore, the extracted features can be used with a small number of output coordinates as necessary information to reduce the required bits, and these coordinates are losslessly encoded using a coordinate encoder.

[0073] Step 3.2: Feature Mapping: Aims to transform the extracted features from the predicted point cloud into the coordinates of the original point cloud. This process is done in a coarse-to-fine manner, with each scale containing a hierarchical connection of a main sparse tensor from the feature extractor and an auxiliary sparse tensor from the previous block, for which the previous block utilizes a hierarchical convolutional layer similar to the feature extraction module.

[0074] In the last block of the module, a generalized sparse transposed convolution layer is used to convolve the connection features of the predicted point cloud on the downscaled coordinates of the original point cloud. This process can be expressed as:

[0075]

[0076] in S is the original point cloud, subscript T To predict the point cloud, L For scale, Indicates the predicted point T In scale L The main sparse tensor of Indicates the predicted point T In scale L The auxiliary sparse tensor of , For the connection operation, the function G(·) represents the operation performed by the convolution layer on the target coordinates. After the convolution operation on the coordinates, the output sparse tensor includes the original point cloud The coordinates and the predicted point cloud of the last scale The mapping features.

[0077] Step 3.3: Feature subtraction: After obtaining the mapping features of the predicted point cloud, feature-level subtraction can be performed directly. By calculating the feature-level residual between the extracted features and the motion mapping features, the prediction error caused by motion compensation based on optical flow in the pixel domain is reduced. This process can be expressed as:

[0078]

[0079] in, Represents the residual features in the last scale.

[0080] Step 4: Input the residual features into the asymmetric spatial channel entropy module, and perform different offsets and vector projections on each group of residual features by grouping deformable dynamic kernels to enhance the expression ability of the module and obtain latent variables. Then, the latent variables are grouped along the channel dimension. Each group except the first one captures helpful spatial context from the previous groups, effectively using spatial information to reduce redundancy, such as Figure 3 As shown, the execution process of the asymmetric spatial channel entropy module includes the following steps:

[0081] Step 4.1: The overall framework of residual feature compression is as follows:

[0082]

[0083]

[0084]

[0085] Among them, x and represents the characteristics of input and decompression, g a and g s are the analysis and synthesis transformations, and are the learning parameters of the analysis transformation network and the synthesis transformation network, respectively. Q (·) indicates a quantization operation.

[0086] During encoding, the input features x are sent to a learning parameter Encoder g a , then obtain the latent representation y, and encode the quantized residual between y and the variance µ into the bitstream. In order to encode y and µ using arithmetic coding, quantize them with Q; then in the decoding process, the decoder g s The decoded delay The above framework is enhanced by using the following dynamic spatial aggregation and asymmetric spatial channel entropy model to further improve the rate-distortion performance.

[0087] Step 4.2: Most existing methods use stacked convolution or window-based self-attention for transform coding, which can only aggregate features within a fixed spatial range. Therefore, a deformable kernel is proposed to achieve adaptive spatial aggregation, but blindly introducing deformable convolution may consume more computing resources. First, an efficient and stable dynamic residual block group is designed, and deep convolution is used to generate content-adaptive kernel offsets and modulation weights respectively; then the input features are divided into several groups, and the modulated kernel weights are shared in each group. This not only maintains the advantages of deformable convolution, but also does not increase the computational burden. The method can be expressed as:

[0088]

[0089] in, G represents the number of split groups, , Indicates g The modulated kernel weights and additional learning offsets are shared within each subgroup. F (·) represents the feature map, Indicates that the feature map corresponds to the center of the convolution kernel, Indicates the relative coordinates of each position of the corresponding convolution kernel.

[0090] Finally, a network is used to project the feature, and its function can be expressed as:

[0091]

[0092] Among them, Conv 1×1 (·) represents a 1×1 convolution operation, and GELU(·) represents the activation function.

[0093] Step 4.3: By encoding the latent representation y as side information in the super-prior part and introducing a dynamic kernel to aggregate the global prior distribution, the global context expression of the generalized entropy model is strengthened; in addition, an asymmetric spatial channel entropy transform is designed to group the latent variable y along the channel dimension, extract the channel context, and extract the spatial context from each group. Each group except the first group captures helpful spatial context from the previous groups. The first two groups use four-stage spatial context, and the latter group uses two-stage spatial context to effectively utilize spatial information to reduce redundancy, combined with the global context information in the super-prior, and finally optimize the residual features through distribution parameter estimation.

[0094] Step 5: During the decoding process, the grid is restored from the grid bit stream through arithmetic decoding, and then the grid is converted to a point cloud to generate a predicted point cloud. The predicted point cloud is then input into the feature extraction module for feature mapping. Finally, the decoded coordinates and the residual features restored after the asymmetric spatial channel entropy module are combined to obtain the features of the original point cloud. These features are sent to the feature propagation module to reconstruct the original point cloud.

[0095] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A point cloud end-to-end compression method based on efficient sampling and optimized entropy model, characterized in that The steps include: Step 1: Perform grid regression on the original point cloud, then obtain a bit stream with grid information through grid quantization and arithmetic coding, obtain the reconstructed grid through arithmetic decoding and grid restoration, and perform grid-to-point cloud conversion to generate a predicted point cloud; Step 2: Perform neural graph sampling on the input raw point cloud; construct a local graph of each point to aggregate feature attributes, and then perform attention-based sampling to select a subset of points to represent the input points; specifically, the following steps are included: Step 2.1: Local graph construction: Use the K nearest neighbor algorithm to construct the K nearest neighbor graph of each point, and obtain the spatial K nearest neighbor graph features. For the original point cloud with attributes, based on the K nearest neighbor index in the spatial domain, obtain the K nearest neighbor graph attributes of each point in the feature domain, and connect the K nearest neighbor graph attributes and the spatial K nearest neighbor graph features for feature embedding; Step 2.2: Graph feature embedding; Dynamic filtering is constructed based on the K nearest neighbor graph to focus on local spatial relationship modeling, matching points with convolutions, and then embedding features to obtain point-based adaptive convolution; Step 2.3: Attention-based sampling; For the output of graph feature embedding, select a representative subset with differentiable operations, and use an attention-based method to select a subset of points in the original point cloud to obtain the points with the most sampling value; Step 3: Extract features from the predicted point cloud and the original point cloud after neural graph sampling to obtain a multi-scale sparse tensor, and use coordinate encoding to obtain a coordinate bit stream, which is then subjected to feature mapping to obtain mapping features of the predicted point cloud. Based on the mapping features and the features extracted from the features, residual features are obtained by feature subtraction; Step 4: Input the residual features into the asymmetric spatial channel entropy, perform different offsets and vector projections on each group of the residual features by grouping the deformable dynamic kernel, and then group the latent variables along the channel dimension. Each group except the first one captures helpful spatial context from the previous groups to optimize the residual features; The specific steps include: Step 4.1: During encoding, the input features x are sent to the encoder g with learned parameters φ a , then the potential representation y is obtained, and the quantized residual between the potential representation y and the variance μ is encoded into the bit stream for quantization; then in the decoding process, the decoder g s The decoded delay Convert to reconstruct features, and use dynamic spatial aggregation and asymmetric spatial channel entropy to enhance the overall framework of residual feature compression; Step 4.2: To achieve adaptive spatial aggregation through deformable kernels, first design a dynamic residual block group and use deep convolution to generate content-adaptive kernel offsets and modulation weights respectively; then divide the input features into several groups and share the modulated kernel weights within each group; Step 4.3: Aggregate the global prior distribution by encoding the latent representation y as side information in the super-prior part and introducing a dynamic kernel; in addition, design an asymmetric spatial channel entropy transform, group the latent variable y along the channel dimension, extract the channel context, and extract the spatial context from each group. Each group except the first group captures helpful spatial context from the previous groups. The first two groups use four-stage spatial context, and the latter group uses two-stage spatial context. Combined with the global context information in the super-prior, the residual features are finally optimized by estimating the distribution parameters; Step 5: Restore the grid from the grid bit stream through arithmetic decoding, then convert the grid to point cloud to generate a predicted point cloud, then perform feature mapping on the predicted point cloud after feature extraction, and finally combine the mapped features with the decoded coordinates and the residual features after the asymmetric spatial channel entropy module to obtain the features of the original point cloud for reconstructing the original point cloud.

2. The method for end-to-end compression of point cloud based on efficient sampling and optimized entropy model according to claim 1, characterized in that: The step 1 comprises the following steps: Step 1.1: A mesh regression module that combines geometric changes and surface deformations to obtain prediction parameters using a pre-trained network and a probabilistic correspondence association module, which are then mesh-quantized and encoded into a bitstream. The parameters include but are not limited to posture, shape, rotation, translation, and gender. Step 1.2: Restore the aligned grid from the quantized parameters of the grid operation block, and then convert the reconstructed grid to a point cloud through the Poisson disk sampling algorithm to generate a predicted point cloud with a similar shape.

3. The point cloud end-to-end compression method based on efficient sampling and optimized entropy model according to claim 1, characterized in that: In step 2.1, a set of three-dimensional point cloud geometric figures X'={X i '∈R 3 , i=1,…,N},X i '=(x i ,y i , z i ) represents the three-dimensional coordinates of each point. For the input point cloud of N×3 matrix, a directed K-nearest neighbor graph G=(V,E) is constructed to represent the local point cloud structure, V={1,…,N}, Used to represent vertices and edges, v i Represents X i 'neighborhood set, by constructing a K nearest neighbor graph, we get the N×K×6 local structure of the point cloud, which contains the three-dimensional coordinates and edge features of each point in the spatial domain; With the property N×F i The original point cloud, F i Represents the input feature channel. Based on the K nearest neighbor index in the spatial domain, the K nearest neighbor graph attribute N×K×F of each point is obtained in the feature domain. i Finally, the K nearest neighbor graph attributes N×K×F i And the spatial K nearest neighbor graph feature N×K×6 is connected for feature embedding.

4. The method for end-to-end compression of point cloud based on efficient sampling and optimized entropy model according to claim 1, characterized in that: In the step 2.2, a point dynamic filter is first generated based on the input K nearest neighbor graph to aggregate the neighbor information to the center point; then the point dynamic filter is applied to the input K nearest neighbor graph to output the dynamic filtering result, which focuses on local spatial relationship modeling, with one convolution for each point; finally, a fully connected layer is used for feature embedding, with local structure and geometric changes as potential features.

5. The method for end-to-end compression of point cloud based on efficient sampling and optimized entropy model according to claim 1, characterized in that: In step 2.3, for the graph feature embedding output N represents the number of points in a specific point cloud, and F o Represents the number of output feature channels, the goal is to select a representative subset with differentiable operations The output points are sampled to N / 4, and the attention-based multi-instance learning pooling method is used to select a subset of points in the original point cloud to obtain the most valuable points for sampling. The pooling output is expressed as: Q=softmax(ωP T ), Among them, Q represents a weighted average of the standardized scores generated by the elements, ω represents a learnable weight parameter, softmax represents the normalization function, and the superscript T represents the transposed symbol.

6. The method for end-to-end compression of point cloud based on efficient sampling and optimized entropy model according to claim 1, characterized in that: The step 3 comprises the following steps: Step 3.1: Feature extraction: Based on sparse convolution to retain key point features, the intermediate results between modules are represented by sparse tensors; The sparse tensor X uses only coordinate feature pairs Save the non-zero elements, each non-zero 3D coordinate (x i ,y i , z i )∈C, corresponding to the relevant feature f i ∈F, by applying convolution filtering only to non-zero elements, the sparse convolution in three-dimensional space is calculated as follows: Among them, C in and C out represents the input and output coordinates, W i Represents the weight of the kernel of the i-th point, set N 3 represents a three-dimensional kernel with arbitrary shape, whose subset includes C in The offset from the current three-dimensional coordinate u, C represents the coordinate set, and F represents the feature set; Each downsampling block in feature extraction consists of a serial convolution, a VRN unit, and another serial convolution layer. The serial convolution is used to reduce the number of points, while the subsequent convolution layer is used to improve the extracted features for optimal performance. The VRN unit is located between two convolution layers, using skip connections to mitigate information loss during training, and parallel convolution layers with different kernel sizes to capture features in different ranges. By superimposing the downsampling blocks, the feature extraction module outputs multi-scale sparse tensors, including coordinates and features, forming a potential representation of the entire point cloud. Step 3.2: Feature Mapping; Convert the extracted features from the predicted point cloud to the coordinates of the original point cloud. This process is done in a coarse-to-fine manner. Each scale contains a hierarchical connection of the main sparse tensor from the feature extractor and the auxiliary sparse tensor from the previous block. For this purpose, the previous block utilizes a hierarchical convolution layer; Finally, the generalized sparse transposed convolution layer is used to convolve the connected features of the predicted point cloud on the downscaled coordinates of the original point cloud. The process is expressed as: Among them, S represents the original point cloud, subscript T represents the predicted point cloud, and L represents the scale. Represents the main sparse tensor of the predicted point T at scale L, Represents the auxiliary sparse tensor of the predicted point T at scale L, represents the connection operation, and the function G(·) represents the operation performed by the convolution layer on the target coordinates; after the convolution operation on the coordinates, the output sparse tensor includes the coordinates of the original point cloud And the mapping features of the predicted point cloud at the last scale Step 3.3: Feature subtraction: After obtaining the mapping features of the predicted point cloud, feature level subtraction is performed directly to calculate the extracted features and as a feature of motion mapping The feature-level residual between , the process is expressed as: Where, ΔF L Represents the residual features in the last scale.

7. The method for end-to-end compression of point cloud based on efficient sampling and optimized entropy model according to claim 1, characterized in that: In step 4.1, the overall framework of residual feature compression is as follows: y=g a (x;φ), Among them, x and Represents the features of input and decompression, g a and g s denote the analysis and synthesis transformations, φ and θ denote the learning parameters of the analysis and synthesis transformation networks, respectively, and Q(·) denotes a quantization operation.

8. The method for end-to-end compression of point cloud based on efficient sampling and optimized entropy model according to claim 1, characterized in that: In step 4.2, the method of implementing adaptive spatial aggregation using a deformable kernel is expressed as: Among them, G represents the number of split groups, K represents the total number of sampling points, and w gk , Δp gk represents the kernel weights of the shared modulation and the additional learning offset within the g-th subgroup, F(·) represents the feature map, p represents the convolution kernel center corresponding to the feature map, and p k Indicates the relative coordinates of each position of the corresponding convolution kernel; Finally, the feature is projected and expressed as: MLP(·)=Conv 1×1 (GELU(Conv 1×1 (·))), Among them, Conv 1×1 (·) represents a 1×1 convolution operation, and GELU(·) represents the activation function.

Citation Information

Patent Citations

  • Point cloud geometric compression grid hole distortion repairing method based on projection

    CN117808682A

  • Tree point cloud classification method based on deep point feature aggregation network

    CN118747823A