Point cloud compression method based on cross-dimension feature collaboration
By projecting point clouds into distance images and combining deep convolutional neural networks and multi-head attention modules, the high-frequency details loss and computational redundancy problems in complex surface processing in the prior art are solved, and efficient point cloud compression and reconstruction are achieved.
Patent Information
- Application Number
- CN202510412800.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively model large-scale spatial continuity and unstructured point clouds, resulting in computing redundancy and lack of prior embedding of sensor physical characteristics. The existing compression methods have the problem of high-frequency details loss in complex surface processing.
A point cloud compression method based on cross-dimensional feature coordination is adopted. By projecting the input point cloud into a distance image and extracting features, combining a deep convolutional neural network and a multi-head attention module, point cloud and distance image features are fused, and compressed through a dimensionality reduction module and an entropy encoder, the point cloud is reconstructed using an adaptive decoder.
Significantly reduce encoding redundancy, improve geometric reconstruction accuracy, maintain geometric consistency of point clouds, and improve reconstruction quality.
Smart Images

Figure CN120339419A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of point cloud compression methods, and particularly relates to a point cloud compression method based on cross-dimensional feature collaboration. Background Art
[0002] As a geometric representation form of discrete sampling in three-dimensional space, point clouds have become the core data carriers in fields such as autonomous driving and digital twins because they can accurately describe the geometric structure and attribute information (such as color and reflection intensity) of the object surface. However, the amount of original point cloud data is huge (hundreds of thousands of points can be in a single-frame lidar point cloud), and its efficient compression and encoding technology has become a key challenge in practical applications. Current point cloud compression is mainly divided into two major directions: geometric compression focuses on the compact representation of three-dimensional coordinates, while attribute compression encodes accessory information such as reflectivity and color. Traditional compression methods (such as MPEG G-PCC, Draco) are based on octree decomposition and manually designed prediction models. Although they perform stably in specific scenarios, the octree encoding based on spatial partitioning is difficult to capture the local correlation of complex surfaces, resulting in the loss of high-frequency details.
[0003] With the rise of learning-based point cloud compression technology, the research direction has gradually shifted to deep learning methods. The point cloud compression method based on deep learning shows significant advantages in rate-distortion performance through end-to-end feature learning. Such methods usually adopt an autoencoder architecture and use a downsampling-upsampling mechanism to achieve joint modeling of geometry and attributes. However, existing learning-based methods rely on three-dimensional convolution or graph convolution of the point cloud itself to extract features, which are difficult to effectively model large-scale spatial continuity and directly process unstructured point clouds, resulting in computational redundancy, and lack prior embedding of sensor physical characteristics. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a point cloud compression method based on cross-dimensional feature collaboration, and the present invention proposes a lidar point cloud compression framework that combines distance image information and point cloud information. This framework realizes the combination of point cloud features and distance image features, and can significantly reduce encoding redundancy while improving the geometric reconstruction accuracy.
[0005] The technical solution protected by the present invention is: a point cloud compression method based on cross-dimensional feature collaboration, which is specifically carried out according to the following steps:
[0006] Step S1: Project the input point cloud into a distance image, and after preprocessing, use a deep convolutional neural network to extract the distance image feature F I ;
[0007] Step S2: Perform farthest point downsampling on the input point cloud data to extract features and coordinates. The point cloud feature is represented as F p , and the coordinate is represented as P;
[0008] Step S3: Project the distance image feature F obtained in Step S1 I and the point cloud feature F obtained in Step S2 p onto the same dimension through a learnable linear mapping, and then extract relevant features and remove redundancy through a multi-head attention module to form a fused feature F C ;
[0009] Step S4: The fused feature F C is dimensionally reduced through a dimensionality reduction module to obtain a feature F, and then F is further compressed into F' by an entropy encoder; at the same time, the bit rate of the point cloud coordinates P obtained in Step S2 is reduced to obtain P';
[0010] Step S5: Feed the F' and P' obtained in Step S4 into a decoder, and the decoder adaptively upsamples and decompresses the downsampled point cloud coordinates using the fused feature to reconstruct the point cloud;
[0011] Step S6: Use a standard rate-distortion loss function to constrain the entire network structure.
[0012] Furthermore, the specific process of obtaining the distance image feature F in Step S1 I is as follows:
[0013] Project the 3D lidar point cloud onto a 2D plane to generate a distance image, where each pixel value in the distance image represents the distance from the corresponding point to the sensor;
[0014] A rotating scanning lidar with m lines divides its vertical field of view FOV into two parts: FOV_up and FOV_down. The value of FOV_up is positive and the value of FOV_down is negative. Therefore, the vertical field of view FOV is the sum of the absolute values of FOV_up and FOV_down.
[0015] The point cloud obtained by the lidar rotating and scanning one week is equivalent to a hollow cylinder centered on itself. Unfolding this cylinder projects the point cloud onto an image plane, which is the distance image. The vertical resolution of the lidar is directly determined by the number of lasers H. Each laser corresponds to a fixed pitch angle θ and is arranged vertically. For example, for a 64-line lidar, H = 64 and the image height is 64 rows. The horizontal angular resolution of the lidar is ρ, so the shape of the distance image collected by the lidar is:
[0016] [H, W] = [H, <360 / ρ>] (1)
[0017] Among them, <> is the rounding operation. The height H of the distance image is directly determined by the number of lasers, and the width W of the distance image is obtained through the horizontal angular resolution ρ and the rounding operation. To project the 3D point cloud onto a 2D plane, spherical coordinates are required. For a point P = (x, y, z) in 3D space, the point cloud is projected onto a 2D pixel I = (w, h, r), where w and h are the vertical and horizontal indices respectively, and r represents the Euclidean distance between the point and the LiDAR origin. The specific calculation method is as follows:
[0018]
[0019] Among them, θ and are the azimuth angle and elevation angle respectively, is the minimum vertical angle, are respectively:
[0020]
[0021] The calculated distance image is preprocessed, interpolated, scaled, and extended to three channels so that the distance image better conforms to the input shape of ResNet50 to extract the feature F I 。
[0022] Furthermore, the deep convolutional neural network in step S1 adopts ResNet50.
[0023] Furthermore, the specific process of obtaining the fused feature F C in step S3 is as follows:
[0024] S31. First, project the distance image feature F I and the point cloud feature F p onto the same intermediate dimension through learnable linear projections, which are expressed as follows:
[0025]
[0026] Among them and are the adjusted features, and W I and W p are learnable linear mappings;
[0027] S32. Adopt the multi-head attention mechanism to process the adjusted features and . The feature after being processed by the multi-head attention module is denoted as F Q . Then, connect the feature F Q obtained by the multi-head attention module, the adjusted feature and to form the final fused feature F C 。
[0028] Further, the specific process of the dimensionality reduction module in S4 is as follows:
[0029] Fused feature F C First, it enters the reshaping layer to reshape the one-dimensional data into a tensor form with spatial dimensions. Subsequently, the data passes through a 1x1 convolutional layer, a ReLU activation function, a Dropout layer, and a 1x1 convolutional layer in sequence. Then, the ReLU activation function is used again to enhance the non-linearity. Subsequently, global average pooling operation is performed on the feature map. Finally, the data dimension is reduced to the required dimension through a fully connected operation, and a low-dimensional feature representation F that meets the requirements of subsequent tasks is output.
[0030] Further, the specific calculation formula of the loss function in step S5 is as follows:
[0031] Total training loss is expressed as:
[0032]
[0033] where D is the overall distortion loss, R is the bitrate, and λ is the trade-off parameter;
[0034] For distortion, the symmetric point-to-point chamfer distance is used to measure the difference between the reconstructed point cloud P' and the ground truth P. Since the decoder has S stages, the distortion loss of each stage is calculated and aggregated into D cha , the density loss D den is defined as:
[0035]
[0036] where |C(p)| is the set of points that collapse to the downsampled point p to form a collapsed point set. Similarly, |C′(P′)| is the set of points formed by upsampling. The first term of the numerator calculates the cardinality difference between the two sets, and the second term calculates the difference in the average distance of all points in the set to the center point P or P'. Therefore and are the average distances of all points in the corresponding sets at each level to the center point respectively, γ is the weight, |p′ s+1 | is the upsampled point at the s+1 level; for each stage s, another loss D card is used to measure the cardinality difference between the ground truth P and the reconstructed point cloud P':
[0037]
[0038] where |P s | and |P s '| are the cardinality values of the ground truth P and the reconstructed point cloud P' respectively.
[0039] In addition, the mean squared error (MSE) loss function is added, and the mean squared error loss is calculated based on the distance image of the reconstructed point cloud and the distance image of the original point cloud:
[0040]
[0041] where n is the total number of points of the calculated data, and y i is the distance value at the i-th position of the distance image projected from the original point cloud, and y i ' is the distance value at the i-th position of the distance image projected from the reconstructed point cloud.
[0042] Finally, the overall distortion loss is as follows:
[0043] D = D cha + αD den + βD card + MSE(9)
[0044] where α and β are the weights of their respective terms.
[0045] The present invention has the following advantages compared with the prior art:
[0046] 1. In view of the sparsity of lidar point clouds and the lack of rich spatial information features, the present invention designs a distance image feature extraction branch to enhance the spatial features of point clouds. At the same time, in view of the problem that the increase in feature dimensions will lead to an increase in the bit rate, the present invention proposes a feature dimensionality reduction module to reduce the computational amount while retaining key features.
[0047] 2. In order to effectively improve the quality of the reconstructed point cloud and maintain geometric consistency, the point cloud compression method of the present invention proposes a distance image reconstruction loss to calculate the difference between the distance image of the reconstructed point cloud and the distance image of the original point cloud.
[0048] 3. The present invention combines two-dimensional data features to enhance the point cloud feature strategy. Compared with other two-dimensional projection methods (such as multi-view orthogonal projection), the distance image avoids geometric distortion caused by view stitching and does not require complex calibration. Experiments show that the method of the present invention significantly reduces coding redundancy while improving the geometric reconstruction accuracy, indicating the effectiveness of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The present invention will be further described in detail below with reference to the accompanying drawings.
[0050] Figure 1 It is a schematic diagram of a point cloud compression framework based on cross-dimensional feature collaboration.
[0051] Figure 2 It is a diagram of the dimensionality reduction module of the present invention.
[0052] Figure 3This is a comparison data graph of the present invention with different solutions.
[0053] Figure 4 This is an ablation experiment data graph for the loss function design of the present invention.
[0054] Figure 5 This is an ablation experiment data graph for the intermediate dimensions mapped by two features of the present invention.
[0055] Figure 6 This is an ablation experiment data graph for the feature fusion module design of the present invention. Detailed implementation manners
[0056] To make the objectives, features, and advantages of the present invention obvious and understandable, the following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings.
[0057] The point cloud compression method based on cross-dimensional feature collaboration of the present invention adopts a network framework as Figure 1 shown. This framework is divided into two branches. One branch projects the input point cloud into a distance image, and after preprocessing, uses ResNet50 to extract features; the other branch downsamples the input point cloud to extract features. Then, the distance image features and the point cloud features are projected to the same dimension through a learnable linear projection, and then relevant features are extracted through a multi-head attention module to remove redundancy to form fused features. After that, a dimensionality reduction module is used to reduce the data volume while retaining important relevant features. Finally, the fused features are encoded by an entropy encoder, and the decoder can use the fused features to adaptively upsample and decompress the coordinates of the downsampled point cloud for reconstruction.
[0058] The above briefly describes the network framework adopted by the present invention. Based on the above network framework, the point cloud compression method of the present invention will be introduced in detail below.
[0059] The point cloud compression method based on cross-dimensional feature collaboration is specifically carried out according to the following steps:
[0060] Step S1: Project the input point cloud into a distance image, and after preprocessing, use a deep convolutional neural network to extract the distance image feature F I , and the specific process is as follows:
[0061] Project the three-dimensional lidar point cloud onto a two-dimensional plane to generate a distance image. Each pixel value in the distance image represents the distance from the corresponding point to the sensor, which can effectively capture the geometric information of the point cloud.
[0062] Assuming an m-line rotating scanning laser radar, its vertical field of view FOV is divided into two parts: FOV_up and FOV_down. The value of FOV_up is positive and the value of FOV_down is negative. Therefore, the vertical field of view FOV is the sum of the absolute values of FOV_up and FOV_down.
[0063] The point cloud obtained by the laser radar after a rotation scan is equivalent to a hollow cylinder with itself as the center. If the cylinder is unfolded, the point cloud is projected into an image plane. This image plane is the distance image. The vertical resolution of the laser radar is directly determined by the number of lasers H. Each laser corresponds to a fixed pitch angle θ and is arranged in the vertical direction. For example, the H of a 64-line laser radar is 64, and the image height is 64 lines. The horizontal angular resolution of the laser radar is ρ, so the shape of the distance image collected by the laser radar is:
[0064] [H,W]=[H,<360 / ρ>] (1)
[0065] Among them, <> is a rounding operation, the height H of the range image is directly determined by the number of lasers, and the width W of the range image is obtained by the horizontal angular resolution ρ and the rounding operation. Spherical coordinates are required to project a 3D point cloud onto a 2D plane. For a point P = (x, y, z) in a 3D space, the point cloud is projected onto a 2D pixel I = (w, h, r), where w and h are the vertical index and horizontal index respectively, and r represents the Euclidean distance between the point and the LiDAR origin. The specific calculation method is as follows:
[0066]
[0067] Among them, θ and are the azimuth and elevation angles, is the minimum vertical angle, They are:
[0068]
[0069] The calculated distance image is preprocessed, scaled and expanded to three channels using interpolation to make the distance image better conform to the input shape of ResNet50 to extract feature F I .
[0070] ResNet50 is a deep convolutional neural network with strong feature learning ability. It can extract the features of distance images well. With its deep residual structure, it can effectively alleviate the gradient vanishing problem and extract image features from low-level to high-level layer by layer. By fusing and screening feature maps at different levels, it provides a distance image feature description rich in semantic information for subsequent fusion with point cloud features. The extracted features are represented as FI For subsequent processing.
[0071] Step S2: Perform farthest point downsampling on the input point cloud data to extract features and coordinates. The point cloud features are represented as F p and the coordinates are represented as P.
[0072] Step S3: After the distance image features F I obtained in step S1 and the point cloud features F p are projected to the same dimension through a learnable linear mapping, then relevant features are extracted and redundancy is removed through a multi-head attention module to form fused features F C The specific process is as follows:
[0073] S31: First, project the distance image features F I and the point cloud features F p to the same intermediate dimension through a learnable linear projection, which is expressed as follows:
[0074]
[0075] Where and are the adjusted features, and W I and W p are learnable linear mappings;
[0076] S32: To more effectively capture the relationships between features, fuse the distance image features and the point cloud features, and extract rich feature representations. The present invention uses a multi-head attention mechanism to process the adjusted features and . The feature representation after being processed by the multi-head attention module is F Q . To further enrich the features, the final fused features are the features F Q obtained by the multi-head attention module and the adjusted features and concatenated to form features, which are represented as F C . However, after feature concatenation, the dimension increases and the data volume will increase. Therefore, we will perform dimensionality reduction processing on F C next.
[0077] Step S4: After the fused features F C are processed by a dimensionality reduction module to obtain features F, then F is further compressed into F' by an entropy encoder; at the same time, the bit rate of the point cloud coordinates P obtained in step S2 is reduced to obtain P'.
[0078] The dimensionality reduction module is as Figure 2 shown. First, F CEnter the reshaping layer. This operation reshapes the one-dimensional data into a tensor form with spatial dimensions to prepare for subsequent convolutional operations. Subsequently, the data passes through two 1x1 convolutional layers in sequence. The first convolutional layer performs preliminary feature extraction and channel transformation on the data. Immediately afterwards, the ReLU activation function is applied to introduce non-linearity into the model to enhance its expressive power. Then, the Dropout layer is introduced, randomly setting the outputs of some neurons to zero, aiming to prevent the model from overfitting and improve its generalization ability. The second convolutional layer further processes the data, and the ReLU activation function is used again to enhance non-linearity. Subsequently, global average pooling is performed on the feature map to reduce the dimensions, which effectively aggregates the information in the spatial dimensions and reduces the computational amount. Finally, through a fully connected operation, the data dimensions are reduced to the required dimensions, and a low-dimensional feature representation F that meets the needs of subsequent tasks is output.
[0079] This module aims to significantly reduce the data volume while retaining important relevant features. In this process, dimensionality reduction techniques are adopted to reduce the feature dimensions, making the calculation more efficient. For P, we use half-floating-point representation to reduce the bit rate and obtain P'. And F is further compressed into F' through an entropy encoder.
[0080] Step S5: Send the F' and P' obtained in step S4 into the decoder. The decoder uses the fused features to adaptively upsample and decompress the downsampled point cloud coordinates to reconstruct the point cloud.
[0081] Step S6: Use a standard rate-distortion loss function to constrain the entire network structure.
[0082] Total training loss Is expressed as:
[0083]
[0084] Where D is the overall distortion loss, R is the bit rate, and λ is the trade-off parameter;
[0085] For distortion, the symmetric point-to-point chamfer distance is used to measure the difference between the reconstructed point cloud P' and the ground truth P. Since the decoder has S stages, the distortion loss at each stage is calculated and aggregated into D cha , density loss D den Is defined as:
[0086]
[0087] Where |C(p)| is the set of all points collapsed into the downsampled point p, and similarly |C′(P′)| is the set of points formed by upsampling. The first term in the numerator calculates the cardinality difference between the two sets, and the second term calculates the difference in the average distance of all points in the set to the center point P or P'. Therefore and are the average distances from all points in each corresponding set at each level to the center point, γ is the weight, |p′ s+1 | is the upsampling point at the s+1 level; for each stage s, another loss D card is used to measure the cardinality difference between the ground truth P and the reconstructed point cloud P':
[0088]
[0089] where |P s | and |P s '| are the cardinality values of the measured ground truth P and the reconstructed point cloud P' respectively.
[0090] In addition, the mean square error MSE loss function is added, and the mean square error loss is calculated based on the distance image of the reconstructed point cloud and the distance image of the original point cloud:
[0091]
[0092] where n is the total number of points of the calculated data, y i is the distance value at the i-th position of the distance image projected by the original point cloud, y i ' is the distance value at the i-th position of the distance image projected by the reconstructed point cloud.
[0093] Finally, the overall distortion loss is as follows:
[0094] D = D cha + αD den + βD card + MSE(9)
[0095] where α and β are the weights of their respective terms.
[0096] Next, a simulation experiment is carried out on the point cloud compression method based on cross-dimensional feature collaboration of the present invention.
[0097] Experimental setup
[0098] We evaluated the method of the present invention on the SemanticKITTI dataset. First, all point clouds were normalized to a cube of 100, and SemanticKITTI was divided into 12 non-overlapping blocks, and each block was further normalized. Then, the point cloud was obtained by non-uniform sampling.
[0099] The present invention compares two state-of-the-art rule-based methods: Google Draco, MPEG Anchor, and G-PCC; learning-based: Depeco, PCGC, and DPCC. The symmetric point-to-point chamfer distance (CD) and point-to-plane PSNR are used to calculate the geometric accuracy, and bits per point (Bpp) are used to calculate the compression ratio.
[0100] First, the method of the present invention is compared with the state-of-the-art rate-distortion trade-off method. In Figure 3 , the CD and PSNR of all methods are shown with respect to bits per point (Bpp). Since the present invention is for lidar point cloud compression, the results of the present invention are referred to as LPCC. As Figure 3 shown, the decompression results of the model of the present invention show that under the same bitrate constraint, the sparse point cloud of SemanticKITTI has lower CD loss and higher PSNR scores. This indicates that the method of the present invention can produce better reconstruction performance. In particular, when Bpp < 3, the effect is better, showing a significant advantage in the case of low Bpp.
[0101] Ablation experiment
[0102] The present invention evaluates the effectiveness of the designs of three parts: the loss function, the intermediate dimension of the mapping, and the feature fusion module.
[0103] Effectiveness of the loss function design
[0104] As Figure 4 shown, the present invention conducts a detailed experimental analysis of the design of the loss function. These two figures are based on the SemanticKITTI dataset and show the performance of different loss calculation strategies with respect to the change of BPP (bits per pixel) in terms of CD and PSNR metrics.
[0105] Among them, LPCC_0 calculates the distortion and point number difference loss in stages, and the density loss is not calculated in stages. LPCC_1 calculates all losses in stages. LPCC_2 is the case where the distortion, density, and point number difference losses are not calculated in stages and are only calculated in the last stage. LPCC_3 adds the MSE loss on the basis of LPCC_0.
[0106] Overall, D cha and D card are calculated in stages, and D den is not calculated in stages and combines the MSE loss to show the best performance in controlling the distance error and improving the image quality. This is because the present invention enhances the features of the point cloud by combining the features of the distance image, but the enhancement process is not carried out in stages, so the density loss is selected not to be calculated in stages.
[0107] Effectiveness of the intermediate dimension design
[0108] As Figure 5As shown, the present invention conducts an ablation study on the design of mapping point cloud features and distance image features to different dimensions to evaluate the impact of different dimensional designs on the results. Since the feature dimension extracted from the distance image through ResNet50 is 2048 and the point cloud feature dimension is 8, the present invention designs three intermediate dimensions for the ablation experiment, namely 1024, 512, and 256. It can be seen from the results in the figure that a higher intermediate dimension (1024) performs better than lower intermediate dimensions (256 and 512) in reducing the point cloud distance and improving the reconstruction quality, indicating the effectiveness of the intermediate dimension design of the present invention.
[0109] Effectiveness of the feature fusion module design
[0110] The present invention conducts an ablation experiment on whether the features processed by the multi-head attention module are concatenated with the point cloud features and distance image features after linear projection, as Figure 6 shown. Among them, the St broken line is the feature obtained by concatenating the features processed by the multi-head attention module with the point cloud features and distance image features after linear projection as the enhanced feature, and Unst is the feature only processed by the multi-head attention module as the enhanced feature. It can be seen that concatenating the features processed by the multi-head attention module with the point cloud features and distance image features after linear projection can achieve better results, indicating the effectiveness of the feature fusion module design of the present invention.
[0111] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A point cloud compression method based on cross-dimensional feature collaboration, characterized in that: The specific steps are as follows: Step S1: Project the input point cloud into a distance image, and after preprocessing, use a deep convolutional neural network to extract the distance image feature F I ; Step S2: Perform farthest point downsampling on the input point cloud data to extract features and coordinates. The point cloud features are represented as F p , and the coordinates are represented as P; Step S3: Project the distance image feature F obtained in Step S1 I and the point cloud feature F obtained in Step S2 p , after projecting them to the same dimension through a learnable linear mapping, then extract relevant features and remove redundancies through a multi-head attention module to form a fused feature F C ; Step S4: Fused feature F C After dimensionality reduction by the dimensionality reduction module, the feature F is obtained, and then F is further compressed to F' by the entropy encoder; at the same time, the bit rate of the point cloud coordinates P obtained in step S2 is reduced to obtain P'; Step S5: Send the obtained F' and P' in step S4 into the decoder, and the decoder uses the fusion features to adaptively upsample and decompress the downsampled point cloud coordinates to reconstruct the point cloud; Step S6: Use a standard rate-distortion loss function to constrain the entire network structure.
2. The point cloud compression method based on cross-dimensional feature collaboration according to claim 1, wherein: The specific acquisition process of the distance image feature F in the step S1 I is as follows: Project the 3D lidar point cloud onto a 2D plane to generate a distance image, where each pixel value in the distance image represents the distance from the corresponding point to the sensor; A rotating scanning lidar with m lines, whose vertical field of view FOV is divided into two parts: FOV_up and FOV_down. The value of FOV_up is positive while the value of FOV_down is negative. Therefore, the vertical field of view FOV is the sum of the absolute values of FOV_up and FOV_down; The point cloud obtained by the lidar rotating and scanning one week is equivalent to a hollow cylinder centered on itself. Unfolding this cylinder projects the point cloud onto an image plane, which is the distance image. The vertical resolution of the lidar is directly determined by the number of lasers H. Each laser corresponds to a fixed pitch angle θ and is arranged vertically. So the shape of the distance image collected by the lidar is: [H, W] = [H, <360 / ρ>] (1) Among them, <> is the rounding operation. The height H of the distance image is directly determined by the number of lasers, and the width W of the distance image is obtained through the horizontal angular resolution ρ and the rounding operation. Projecting the 3D point cloud onto the 2D plane requires the use of spherical coordinates. For a point P = (x, y, z) in 3D space, project the point cloud onto a 2D pixel I = (w, h, r). w and h are the vertical index and horizontal index respectively, and r represents the Euclidean distance between this point and the LiDAR origin. The specific calculation method is as follows: where θ and are the azimuth angle and the elevation angle respectively, is the minimum vertical angle, θ, and σ are respectively: The calculated distance image is preprocessed, interpolated, scaled, and extended to three channels so that the distance image better conforms to the input shape of ResNet50 to extract feature F I .
3. The point cloud compression method based on cross-dimensional feature collaboration according to claim 2, wherein: The deep convolutional neural network in step S1 uses ResNet50.
4. The point cloud compression method based on cross-dimensional feature collaboration according to claim 3, wherein: The fused feature F in the step S3 C is obtained through the following specific process: S31. First, project the distance image feature F I and the point cloud feature F p onto the same intermediate dimension through a learnable linear projection, as shown below: Among them and are the adjusted features, and W I and W p are learnable linear mappings; S32. Use the multi-head attention mechanism to process the adjusted features and . The feature representation after being processed by the multi-head attention module is F Q . Then, connect the feature F Q obtained by the multi-head attention module, the adjusted features and to form the final fused feature F C .
5. The point cloud compression method based on cross-dimensional feature collaboration according to claim 4, wherein: The specific process of the dimensionality reduction module in S4 is as follows: Fused feature F C First, it enters the reshaping layer to reshape the one-dimensional data into a tensor form with spatial dimensions. Subsequently, the data passes through a 1x1 convolutional layer, a ReLU activation function, a Dropout layer, and a 1x1 convolutional layer in sequence. Then, the ReLU activation function is used again to enhance the non-linearity. Subsequently, global average pooling operation is performed on the feature map. Finally, the data dimension is reduced to the required dimension through a fully connected operation, and a low-dimensional feature representation F that meets the requirements of subsequent tasks is output.
6. The point cloud compression method based on cross-dimensional feature collaboration according to claim 5, wherein: The specific calculation formula of the loss function in step S5 is as follows: Total training loss Expressed as: Among them, D is the overall distortion loss, R is the bit rate, and λ is the trade-off parameter; For distortion, the symmetric point-to-point chamfer distance is used to measure the difference between the reconstructed point cloud P' and the ground truth P. Since the decoder has S stages, the distortion loss at each stage is calculated and aggregated into D cha , the density loss D den is defined as: where |C(p)| is the set of points that collapse to the downsampled point p, and similarly |C′(P′)| is the set of points formed by upsampling. The first term in the numerator calculates the difference in cardinality between the two sets, and the second term calculates the difference in the average distance of all points in the set to the center point P or P'. Therefore and are the average distances of all points in the corresponding sets at each level to the center point, γ is the weight, |p′ s+1 | is the upsampled point at the s+1-th level; for each stage s, another loss D card is used to measure the difference in cardinality between the ground truth P and the reconstructed point cloud P': where |P s | and |P s '| are respectively the base values for measuring the ground truth value P and the reconstructed point cloud P'; In addition, add the mean square error MSE loss function, and calculate the mean square error loss based on the distance image of the reconstructed point cloud and the distance image of the original point cloud: where n is the total number of points of the computed data, y i is the distance value at the i-th position of the distance image of the original point cloud projection, y i ' is the distance value at the i-th position of the reconstructed point cloud projected into the distance image; Finally, the overall distortion loss is as follows: D = D cha + αD den + βD card + MSE(9) Among them, α and β are the weights of their respective terms.