Machine vision-oriented point cloud real-time compression method

Through quantization, frequency perception and sparse convolution technology, point clouds are divided into high and low frequency components, abandoning high frequency components, and using low frequency adapters and downsampling network to compress point clouds, solving the problems of data redundancy and long encoding and decoding time in existing methods, and achieving efficient real-time compression and reconstruction in machine vision tasks.

CN120235964APending Publication Date: 2025-07-01BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411853897.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing point cloud compression methods are mainly aimed at human eye vision rather than machine vision, resulting in data redundancy. The deep learning-based method is too long to code and difficult to apply to real-time tasks.

Method used

Using quantization, frequency perception, low-frequency adapter and sparse convolution technology, point clouds are divided into high-frequency and low-frequency components, abandoning high-frequency components, compressing point clouds through low-frequency adapter and downsampling network, and reconstructing point clouds through upsampling and dequantization, combining BCE and Lagrangian loss functions to optimize the compression model.

Benefits of technology

It realizes efficient real-time compression of point clouds in machine vision tasks, maintains key feature information, reduces storage space, improves encoding and decoding efficiency, and is suitable for other machine vision tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235964A_ABST
    Figure CN120235964A_ABST
Patent Text Reader

Abstract

The invention provides a point cloud real-time compression method and system for machine vision. According to the method, radar point cloud data are converted into non-negative integers suitable for compression processing through quantization and de-quantization operation, and the data precision is kept. The frequency sensing module based on graph filtering carries out frequency sorting on the point cloud, the point cloud is divided into low-frequency and high-frequency components, and redundant high-frequency points are removed to reduce coding overhead. And the low-frequency adapter module dynamically adjusts the characteristics of the low-frequency points according to the quantization precision, adaptively optimizes the characteristic representation of the low-frequency points through sparse convolution and a full-connection layer, and ensures the retention of key space information. By integrating the modules, the real-time radar point cloud compression framework can maintain the accuracy of a machine vision task to the greatest extent while ensuring a high compression ratio and real-time performance. Through the combination of deep learning and frequency perception, an efficient and real-time point cloud compression scheme is provided, and the method is suitable for real-time point cloud processing application and has important technical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of point cloud compression, and particularly to a real-time point cloud compression method for machine vision. Background Art

[0002] As a key three-dimensional data representation, point cloud data is widely used in fields such as autonomous driving, 3D reconstruction, virtual reality, and augmented reality. However, with the continuous increase of large-scale and high-precision point cloud data, its huge data volume faces great challenges in transmission and storage.

[0003] The point cloud compression method is a strategy to reduce the amount of point cloud data. It can reduce the redundancy of point cloud data by extracting point cloud features, thereby reducing the amount of point cloud data. At present, the existing point cloud compression methods are mainly oriented to human vision rather than machine vision, and the differences between human vision and machine vision also cause a large amount of data redundancy. In addition, the existing deep learning-based point cloud compression methods have too long encoding and decoding times due to the complexity of the network structure, making it difficult to apply to real-time tasks. Therefore, how to compress point clouds in real time to the greatest extent while basically ensuring machine vision is still a challenge. Summary of the Invention

[0004] In view of this, the purpose of the present application is to provide a real-time point cloud compression method for machine vision. This method is divided into three stages. The first stage is the preprocessing stage of the point cloud, in which the point cloud is read and quantized, and then divided into high and low frequencies. The second stage is the encoding stage of the point cloud, in which the point cloud is compressed into a bitstream through downsampling and sparse convolution. The third stage is the decoding stage of the point cloud, in which the bitstream is reconstructed into a point cloud through upsampling and sparse convolution, and dequantization operation is performed to restore the coordinates of the point cloud to the original data range.

[0005] The embodiments of the present application provide a real-time point cloud compression method for machine vision. The point cloud compression network includes quantization, frequency perception, low-frequency adapter, point cloud downsampling, encoding, decoding, point cloud upsampling, and dequantization; the determination method includes:

[0006] Obtain the initial point cloud data of the target object;

[0007] Perform a quantization operation on the initial point cloud data to convert its coordinates into non-negative integers;

[0008] Input the quantized point cloud into the frequency perception module to calculate the score for each point therein for sorting from low frequency to high frequency;

[0009] Adaptive adjust the ratio of the high-frequency component and the low-frequency component of the point cloud according to the quantization accuracy of the point cloud, and discard the high-frequency component therein;

[0010] The remaining low-frequency point cloud is passed through the downsampling part of the network, which contains a low-frequency adapter to adapt the network to the low-frequency point cloud;

[0011] The downsampled point cloud is encoded to obtain the code stream, and the compression part is completed;

[0012] In practical applications, point clouds need to be reconstructed, so the method includes a point cloud reconstruction part, which first decodes the bitstream;

[0013] Pass the decoded point cloud through the upsampling part of the network, which also contains the low-frequency adapter;

[0014] Dequantize the obtained reconstructed point cloud to restore the point cloud data to the original data range;

[0015] The initial point cloud compression network is optimized based on a target value to obtain the point cloud compression network; wherein the target value includes at least one or more of the following items: the original point cloud data, the reconstructed point cloud data.

[0016] Furthermore, the point cloud quantization operation includes: finding the minimum value of the three dimensions x, y, and z of the original point cloud, that is, the quantized offset; setting the target data range of the point cloud; calculating the quantization step required for quantizing the point cloud; quantizing the point cloud data through the above parameters;

[0017] Inputting the point cloud obtained after the complete quantization operation into a frequency perception module to obtain a point cloud sorted from low to high according to the frequency of each point;

[0018] Further, find out the K neighboring points of each point in the point cloud, calculate the Euclidean distance between the K neighboring points, and obtain the distance matrix D, which is a symmetric matrix. Then, obtain the Gaussian matrix A through Gaussian transformation, which is also a symmetric matrix; then use the Haar graph filter to calculate the feature score ((IA)P) of each point, and calculate the L2-norm of the feature score as the final score of each point; sort each point in the point cloud according to the final score to obtain a frequency-sorted point cloud;

[0019] The ratio of high-frequency points to low-frequency points is calculated according to the quantization target range, and then the low-frequency points are retained according to the ratio, and the high-frequency points are discarded to obtain a low-frequency point cloud;

[0020] The low-frequency point cloud and the ratio of the high-frequency points to the low-frequency points are input into a low-frequency adapter, which adaptively transforms the feature representation of the low-frequency points to better capture and retain necessary spatial details; at the same time, the network is enabled to selectively focus on and retain key information, while reducing the possibility of loss of spatial details that often occurs when removing high-frequency points, maintaining a high fidelity to the original point cloud structure, and obtaining an adapted point cloud;

[0021] Furthermore, the low-frequency adapter includes two branches, the input of the first branch is the low-frequency point cloud, which is sequentially subjected to sparse convolution, nonlinear activation function, fully connected layer, and nonlinear activation function; the input of the second branch is the ratio of high and low frequency points, which is sequentially subjected to fully connected layer and nonlinear activation function; the results obtained from the two branches are multiplied and fused with the low-frequency point cloud to obtain an adapted point cloud 1;

[0022] Input the adapted point cloud 1 into the downsampling network to obtain the downsampling point cloud 1;

[0023] Furthermore, the downsampling network includes sparse convolution, nonlinear activation function, downsampling sparse convolution, nonlinear activation function, and three initial-residual network blocks;

[0024] Similar to the process, the downsampled point cloud 1 and the ratio of high and low frequency points are sequentially passed through the low frequency adapter, the downsampled network, the low frequency adapter, and the downsampled network to obtain the adapted point cloud 2, the downsampled point cloud 2, the adapted point cloud 3, and the downsampled point cloud 3;

[0025] Through the above operations, the features of the point cloud are continuously extracted while compressing the point cloud to ensure the reconstruction quality of the point cloud;

[0026] The downsampled point cloud is input into a geometric encoder for geometric encoding to obtain a bit stream, and the point cloud compression work is completed;

[0027] Inputting the bit rate into a geometry decoder for geometry decoding to obtain a rough point cloud;

[0028] Inputting the rough point cloud and the ratio of the high and low frequency points into the low frequency adapter to obtain an adapted point cloud 4;

[0029] Input the adapted point cloud 4 into the point cloud upsampling network to obtain an upsampled point cloud 1;

[0030] Furthermore, the point cloud upsampling network includes upsampling sparse convolution, nonlinear activation function, sparse convolution, nonlinear activation function, three initial-residual blocks, and sparse convolution;

[0031] Similar to the above process, the upsampled point cloud one and the high-low frequency point ratio are successively passed through the low-frequency adapter, the upsampling network, the low-frequency adapter, and the upsampling network to successively obtain the adapted point cloud five, the upsampled point cloud two, the adapted point cloud six, and the upsampled point cloud three;

[0032] Through the above operations, the point cloud is continuously refined during the point cloud reconstruction process to improve the reconstruction quality of the point cloud;

[0033] The upsampled point cloud three is input into the dequantization module to convert the coordinate range of the upsampled point cloud three to the initial range, that is, the range before the quantization operation, to obtain the reconstructed point cloud;

[0034] Based on the reconstructed point cloud and the original point cloud data, the loss function in point cloud compression is calculated;

[0035] Preferably, the BCE loss function and the Lagrangian loss function are used to calculate the error between the reconstructed point cloud and the original point cloud, so that the point cloud generated by point cloud compression and then reconstruction can be closer to the real uncompressed point cloud;

[0036] Furthermore, the calculation formula of the BCE loss function is as follows:

[0037]

[0038] where x i is the label of the voxel, which is 1 when occupied and 0 when empty, and p i is the probability that the voxel is occupied;

[0039] The calculation formula of the Lagrangian loss function is as follows:

[0040] J loss = R + λD

[0041] where R is the bit rate, D is the distortion, which is calculated by the BCE loss function, and the parameter λ is a hyperparameter for controlling the bit rate;

[0042] A computer system provided by the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the real-time point cloud compression method for machine vision described above is implemented.

[0043] Beneficial effects: Compared with the existing point cloud compression methods, the present method has the following advantages:

[0044] 1. The present invention represents point cloud data through sparse convolution and extracts features through sparse tensors, avoiding directly applying 3D convolution to operate on empty positions, which is beneficial to saving time consumption and better completing real-time tasks.

[0045] 2. The present invention divides the point cloud into high-frequency components and low-frequency components through a frequency perception module. Considering that for machine vision tasks, particularly perfect reconstruction quality is not required, but only the main features of the point cloud need to be retained. The high-frequency components, which mainly describe details and are difficult to compress, are discarded, and only the low-frequency components are retained. Through the above operations, the present invention can retain more low-frequency points in the same storage space, thereby achieving the retention of more feature information and ensuring the performance of machine vision tasks.

[0046] 3. The present invention adaptively adjusts the proportion of points in the high-frequency and low-frequency components according to the quantization scale of the point cloud. After discarding the high-frequency components, through the low-frequency adapter module, the points in the low-frequency components are re-adapted, and the features of the remaining low-frequency points are adjusted and refined. By applying a convolutional layer and a fully connected layer with a non-linear activation function, this module adaptively transforms the feature representation of the low-frequency points to better capture and retain the necessary spatial details. This enables the network to selectively focus on and retain key information while reducing the possibility of loss of spatial details that often occurs when high-frequency points are removed. Finally, through this mechanism, the low-frequency adapter minimizes the loss of important spatial details and maintains a high fidelity to the original point cloud structure. At the same time, it optimizes the memory usage and computational efficiency of the compression model, achieving faster encoding and decoding without affecting the quality of the compressed point cloud data.

[0047] 4. The present invention is a real-time point cloud compression method specifically for machine vision tasks, but it is not jointly trained with downstream tasks. This enables the method to be smoothly integrated into other machine vision task methods. Even if the downstream task network undergoes updates and upgrades, the present invention still applies.

[0048] 5. The experimental results in the following specific embodiments confirm the effectiveness and superiority of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic structural diagram of a real-time point cloud compression method for machine vision according to the present invention.

[0050] Figure 2 It is a schematic structural diagram of the low-frequency adapter according to the present invention.

[0051] Figure 3 It is a schematic structural diagram of the downsampling part according to the present invention.

[0052] Figure 4 It is a schematic structural diagram of the downsampling part according to the present invention.

[0053] Figure 5 It is a comparison chart of the time consumption of the present invention and other mainstream point cloud compression methods.

[0054] Figure 6 This is a performance comparison chart of the present invention and other mainstream point cloud compression methods in machine vision tasks (object detection). Detailed implementation manners

[0055] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0056] A real-time point cloud compression method for machine vision provided by an embodiment of the present invention improves the existing network architecture based on deep learning. Before compression, the point cloud data is quantized, and then through a frequency perception module, each point in the point cloud is sorted from low frequency to high frequency, and the division ratio is adaptively adjusted according to the quantization scale, and the high-frequency part of the point cloud is discarded, and only the low-frequency part of the point cloud is retained. For the low-frequency component of the point cloud, through a low-frequency adapter and a downsampling network, the features of the point cloud are extracted and compressed into a bitstream by an encoder. The bitstream is decoded into a point cloud by a decoder, and then through a low-frequency adapter and an upsampling network, and finally the point cloud is restored to the original data range through a dequantization operation to achieve the reconstruction of the point cloud. The network structure based on the embodiment of the present invention is as Figure 1 shown.

[0057] Specifically, a real-time point cloud compression method for machine vision provided by an embodiment of the present invention first quantizes the input initial point cloud to obtain a voxelized point cloud; then inputs the voxelized point cloud into a frequency perception module, calculates a score for each point in the point cloud, and arranges them from low to high according to the score, that is, the frequency of the point cloud is arranged from low to high; calculates the ratio of the high-frequency and low-frequency components according to the quantization scale, and discards the high-frequency components according to the ratio; inputs the remaining low-frequency component point cloud into a low-frequency adapter, and then performs another downsampling to obtain a first downsampled point cloud; then inputs the first downsampled point cloud into a low-frequency adapter and a downsampling network to obtain a second downsampled point cloud; then inputs the second downsampled point cloud into a low-frequency adapter and a downsampling network to obtain a third downsampled point cloud; encodes the coordinates and features of the third downsampled point cloud by an encoder to obtain a bitstream;

[0058] For the point cloud bitstream, a low-precision point cloud is obtained by decoding through a decoder; the low-precision point cloud is input into a low-frequency adapter and an upsampling network to obtain a first upsampled point cloud; the first upsampled point cloud is input into a low-frequency adapter and an upsampling network to obtain a second upsampled point cloud; the second upsampled point cloud is input into a low-frequency adapter and an upsampling network to obtain a third upsampled point cloud; the third upsampled point cloud undergoes a dequantization operation to obtain a reconstructed point cloud; the reconstructed point cloud can be used to complete machine vision tasks.

[0059] Among them, the frequency perception module first uses the K-nearest neighbor algorithm to construct the local point cloud of each point. For the local point cloud, calculate its Gaussian matrix, use a graph filter to implement high-pass filtering, and calculate the L2-norm as the score. For the scores of each point in the point cloud, sort them from low to high, and then the point cloud can be sorted according to the frequency.

[0060] For the low-frequency adapter module, by applying sparse convolutional layers and fully connected layers with non-linear activation functions, it adaptively transforms the feature representation of low-frequency points to better capture and retain necessary spatial details. This enables the network to selectively focus on and retain key information, while reducing the possibility of spatial detail loss that often occurs when high-frequency points are removed, maintaining a high degree of fidelity to the original point cloud structure.

[0061] The following combines Figures 1 to 4 to detail the specific steps of the above method:

[0062] Step 1: Obtain the initial point cloud data of the radar scene point cloud.

[0063] Step 2: Quantize the initial point cloud data;

[0064] Step 3: Input the quantized point cloud into the frequency perception module, sort each point in the point cloud from high to low according to the frequency, and adaptively adjust the high-low frequency ratio according to the quantization scale of the point cloud, divide the point cloud according to the ratio, and discard the high-frequency part;

[0065] Step 4: Input the remaining low-frequency point cloud into the low-frequency adapter and the downsampling network, a total of 3 times;

[0066] Step 5: Encode the point cloud after 3 times of downsampling into a bitstream through an encoder;

[0067] Step 6: Decode the bitstream into a rough point cloud through a decoder;

[0068] Step 7: Input the rough point cloud into the low-frequency adapter and the upsampling network, a total of 3 times;

[0069] Step 8: Perform a dequantization operation on the point cloud after 3 times of upsampling to obtain the reconstructed point cloud.

[0070] In this example, in Step 1, the initial point cloud data of the radar scene point cloud is obtained. It should be noted that the point cloud data refers to a set of discrete point data in a three-dimensional coordinate system. The data in the data set should at least include the position data of the point, such as the three-dimensional coordinate vector (x, y, z). Among them, in a real environment, the point cloud data of the target object can be obtained through three-dimensional sensors such as lidar.

[0071] In this example, the quantization operation of the point cloud in Step 2 is as follows:

[0072] 2-1. Denote the original point cloud data as P and the offset as offset. The calculation method of the offset is as follows:

[0073] offset = min(min(P x ), min(P y ), min(P z ))

[0074] where P x , P y , P z represent the coordinates of the point cloud in the three dimensions of x, y, and z respectively. That is, the offset is equal to the minimum value of all the dimensional coordinates of all the points in the point cloud.

[0075] 2-2. The target data range ql of the point cloud can be set arbitrarily according to the requirements of the actual application. After setting, the coordinates of the points will be quantized into the interval [0, ql - 1]; if the actual storage space is small, the data range can be set in a smaller range to achieve coarse quantization; while if the actual storage space is relatively sufficient, the data range can be appropriately expanded to achieve fine quantization and obtain a better reconstruction effect.

[0076] 2-3. The quantization step qs is jointly determined by the coordinate range of the original point cloud and the target data range of the point cloud. Its calculation formula is:

[0077]

[0078] 2-4. Among them, max(P) and min(P) represent the maximum value and the minimum value of the coordinates of the point cloud in the three dimensions of x, y, and z respectively.

[0079] 2-5. The method of quantizing the point cloud data can be expressed by the following formula:

[0080]

[0081] where P Q represents the quantized point cloud data, and round represents the rounding operation of the data.

[0082] In step 3, the quantized point cloud is input into the frequency perception module. The specific steps are as follows:

[0083] 3-1. The calculation of the frequency perception module includes three parts, namely the Gaussian matrix calculated from the Euclidean distance and the Gaussian function, the output features calculated from the Haar wavelet filter, and the score of each point calculated by calculating the L2-norm according to the output features.

[0084] 3-2. Input the initial point cloud data of N×3 into the frequency perception module, where the output can be simply summarized as the Euclidean distance matrix of K×K between each point and its K nearest neighbors, and the Gaussian matrix of K×K.

[0085] The calculation formula of the Euclidean distance matrix is:

[0086]

[0087] where i and j are positive integers not greater than K, representing the points within the K nearest neighbor point cloud, D i,j and D j,i are the Euclidean distances between point i and point j. Obviously, the obtained matrix is a symmetric matrix;

[0088] The calculation formula of the Gaussian matrix is:

[0089]

[0090] where i and j are positive integers not greater than K, representing the points within the K nearest neighbor point cloud, D i,j is the Euclidean distance between point i and point j, and A i,j is the Gaussian value between point i and point j. Obviously, the obtained matrix is a symmetric matrix;

[0091] 3-3. Calculate the features using the Haar wavelet filter, and the calculation formula is:

[0092] h(A)P = (I - A)P

[0093] 3-4. For each point in the point cloud, calculate the L2-norm using the above results, and use this L2-norm as the score. The score of low-frequency points is low, and the score of high-frequency points is high. Sort them from low to high to obtain the point cloud after frequency sorting.

[0094] 3-5. Adaptively adjust the ratio of high-frequency points, which can be obtained by the following formula:

[0095]

[0096] where, represents the ratio of high-frequency points to the total number of points in the quantized point cloud, round represents the rounding function, and ql represents the range of quantized data.

[0097] 3-6. Create a new sparse tensor, which consists of only a low-frequency number of points, where N represents the number of points in the quantized point cloud. By the above method, the high-frequency points are discarded, and the remaining point cloud are all low-frequency points that are convenient for encoding and compression.

[0098] Step 4 performs adaptation and downsampling on the low-frequency point cloud, and the specific steps are as follows:

[0099] 4-1. Input the low-frequency point cloud and the high-frequency ratio fl into the low-frequency adapter, where the ratio fl is broadcast into a sparse tensor, passed through a sparse fully connected layer and the SiLU activation function to obtain the scaling ratio tensor Scale; the low-frequency point cloud passes through a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with the same number of output channels as the input channels, and then successively through the ReLU activation function, a sparse fully connected layer, and the ReLU activation function. The obtained sparse tensor and the scaling ratio tensor Scale are added to the original low-frequency point cloud for feature fusion to obtain the first adapted point cloud; this result is passed through a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with 16 output channels, the ReLU activation function, a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with 32 output channels and 1 / 2 downsampling, the ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU functions to obtain the first downsampled point cloud.

[0100] 4-2. Similar to step 4-1, input the first downsampled point cloud and the high-frequency ratio fl into the low-frequency adapter to obtain the second adapted point cloud, and then pass the second adapted point cloud through a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with 32 output channels, the ReLU activation function, a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with 64 output channels and 1 / 2 downsampling, the ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU functions to obtain the second downsampled point cloud.

[0101] 4-3. Similar to steps 4-1 and 4-2, input the second downsampled point cloud and the high-frequency ratio fl into the low-frequency adapter to obtain the third adapted point cloud, and then pass the third adapted point cloud through a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with 64 output channels, the ReLU activation function, a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with 32 output channels and 1 / 2 downsampling, the ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU functions to obtain the third downsampled point cloud.

[0102] 4-4. Pass the third downsampled point cloud through a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with 8 output channels to obtain the downsampled point cloud.

[0103] Step 5 encodes the downsampled point cloud. This step uses the octree geometry encoder in G-PCC for lossless encoding to obtain the compressed bitstream.

[0104] Step 6 is the reverse operation of Step 5. The geometric decoder is used to decode the bitstream to obtain the rough point cloud.

[0105] In Step 7, the rough point cloud is adapted and upsampled. The specific steps are as follows:

[0106] 7-1. Input the rough point cloud and the high-frequency ratio fl into the low-frequency adapter to obtain the adapted point cloud IV. Then, pass the adapted point cloud IV through a sparse convolution with a convolution kernel size of 3 3 , an output channel number of 32, 2-fold upsampling, ReLU activation function, a convolution kernel size of 3 3 , a sparse convolution with an output channel number of 32, ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU function to obtain the upsampled point cloud component. Pass this component through a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with an output channel number of 1, and combine it with the aforementioned upsampled point cloud component to obtain the upsampled point cloud I.

[0107] 7-2. Similar to Step 7-1, input the upsampled point cloud I and the high-frequency ratio fl into the low-frequency adapter to obtain the adapted point cloud V. Then, pass the adapted point cloud V through a sparse convolution with a convolution kernel size of 3 3 , an output channel number of 64, 2-fold upsampling, ReLU activation function, a convolution kernel size of 3 3 , a sparse convolution with an output channel number of 64, ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU function to obtain the upsampled point cloud component. Pass this component through a sparse convolution with a convolution kernel size of 3 3 , a sparse convolution with an output channel number of 1, and combine it with the aforementioned upsampled point cloud component to obtain the upsampled point cloud II.

[0108] 7-3. Similar to Step 7-1 and Step 7-2, input the upsampled point cloud II and the high-frequency ratio fl into the low-frequency adapter to obtain the adapted point cloud VI. Then, pass the adapted point cloud VI through a sparse convolution with a convolution kernel size of 3 3 , an output channel number of 32, 2-fold upsampling, ReLU activation function, a convolution kernel size of 3 3 , a sparse convolution with an output channel number of 16, ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU function, and then directly pass it through a sparse convolution with a convolution kernel size of 3 3 , an output channel number of 1 to obtain the upsampled point cloud III.

[0109] In Step 8, a dequantization operation is performed on the upsampled point cloud to restore the original data range of the point cloud so as not to affect downstream tasks. The formula is as follows:

[0110]

[0111] Among them is the reconstructed point cloud, P Q is the point cloud before the dequantization operation, qs is the quantization step size calculated in step 2-3, and offset is the offset calculated in step 2-1.

[0112] The following is an example to illustrate the implementation effect of the present invention in the actual application process:

[0113] In an experiment, the voxel occupancy and rate distortion of the point cloud were used to optimize the point cloud compression network using the BCE loss function and the Lagrangian loss function respectively. The point cloud compression network was trained in two stages. First, the downsampling network and the upsampling network were trained in the first stage for 50 rounds; the initial learning rate was set to 0.0008 and gradually decayed to 0.00002; the batchsize was set to 8; subsequently, the low-frequency adapter module was trained in the second stage for 50 rounds, and the learning rate, batchsize, etc. were set the same as in the first stage. When training, the upsampling network and the downsampling network were frozen without parameter updates; here, this application uses PyTorch for code implementation and the Adam optimizer to optimize the model.

[0114] This application conducts experiments on the classic KITTI dataset, which is a common publicly available radar point cloud dataset, and tests the bits per point (bpp) of the compressed point cloud and the reconstructed point cloud in a machine vision task (here, the object detection task) on the corresponding test set.

[0115] The test results on the KITTI dataset are shown in Table 1 below. Since the present invention is a point cloud compression method for machine vision, in order to test the performance of this method in machine vision tasks and eliminate the differences brought by downstream tasks, for fairness, the same object detection method is used when comparing with other methods. The specific process is to send the point cloud into this method and the comparison method to obtain the bitstream; then use this method and the comparison method to decode respectively to obtain the reconstructed point cloud; then send the reconstructed point cloud into the same object detection network, and use the accuracy of object detection at the same bit rate to evaluate the advantages and disadvantages of the method in machine vision. The larger all indicators are, the higher the detection accuracy, that is, the better the effect.

[0116] Table 1 Object detection results of this method after compression and reconstruction on the KITTI dataset

[0117]

[0118] In Table 1, AP11 and AP40 are two interpolation evaluation methods used to evaluate the average accuracy of detection. For the same difficulty level of each category, different intersection over union (IoU) thresholds are used in the upper and lower rows, that is, only when the threshold is exceeded is the target considered to be successfully detected. For the results shown in Table 1, the same quantization scale is adopted. Since the present invention discards low-frequency points, the bit rate of the present invention is lower, but the detection accuracy is higher, highlighting the superiority of the present invention. It can be seen that the point cloud compression network in the present invention can complete the compression and reconstruction of point clouds in real time. For the specific comparison of compression time consumption, see Figure 5 In addition, the point cloud compression network in the present invention has achieved an advanced point cloud compression effect, and better maintains the performance of point clouds in machine vision tasks. It is superior to existing point cloud compression methods under the same bit rate. For the specific effect, see Figure 6 .

[0119] Based on the same inventive concept, an embodiment of the present invention provides a point cloud real-time compression system for machine vision, including an input module, a quantization module, a frequency perception module, an adapter module, a point cloud downsampling network, a point cloud geometry encoder, a point cloud geometry decoder, a point cloud upsampling network, and a dequantization module; the input module is used to obtain a three-dimensional point cloud containing coordinate information and input it into the quantization module; the quantization module is used to quantize the point cloud data and store it in the form of a sparse tensor; the frequency perception module is used to divide the point cloud into high-frequency and low-frequency parts; the adapter module is used to re-adapt the point cloud after discarding high-frequency points; the downsampling network is used to extract features from the point cloud; the point cloud geometry encoder is used to encode the point cloud into a bitstream; the point cloud geometry decoder is used to decode the bitstream into a rough point cloud; the point cloud upsampling network is used to finely reconstruct the point cloud; the dequantization module is used to restore the reconstructed point cloud data to the coordinate range of the original point cloud. For the specific implementation of each module of the network model, refer to the above method embodiment and will not be elaborated here. Those not detailed in the present invention are all prior arts.

Claims

1. A point cloud real-time compression method for machine vision, characterized in that: First, the input initial point cloud is quantized to obtain a voxelized point cloud; then the voxelized point cloud is input into the frequency perception module, and the score is calculated for each point in the point cloud, and the points are arranged from low to high according to the score, that is, the frequency of the point cloud is arranged from low to high; the ratio of high and low frequency components is calculated according to the quantization scale, and the high frequency components are discarded according to the ratio; the remaining low frequency component point cloud is input into the low frequency adapter, and then downsampled once to obtain the first downsampled point cloud; Then, the first down-sampled point cloud is input into the low-frequency adapter and the down-sampling network to obtain the second down-sampled point cloud; the second down-sampled point cloud is input into the low-frequency adapter and the down-sampling network to obtain the third down-sampled point cloud; the coordinates and features of the third down-sampled point cloud are encoded by the encoder to obtain a bitstream; for the point cloud bitstream, the low-precision point cloud is decoded by the decoder; the low-precision point cloud is input into the low-frequency adapter and the up-sampling network to obtain the first up-sampled point cloud; the first up-sampled point cloud is input into the low-frequency adapter and the up-sampling network to obtain the second up-sampled point cloud; the second up-sampled point cloud is input into the low-frequency adapter and the up-sampling network to obtain the third up-sampled point cloud; The third upsampled point cloud is dequantized to obtain the reconstructed point cloud.

2. The method according to claim 1, characterized in that: The frequency perception module first uses the K-nearest neighbor algorithm to construct a local point cloud for each point, calculates its Gaussian matrix for the local point cloud, uses a graph filter to implement high-pass filtering, and calculates the L2-norm as the score; the score of each point in the point cloud is sorted from low to high, and the point cloud can be sorted by frequency.

3. The method according to claim 1 is characterized in that: The low-frequency adapter adaptively transforms the feature representation of low-frequency points by applying sparse convolutional layers and fully connected layers with non-linear activation functions.

4. The method according to claim 1, characterized in that The following steps are involved: Step 1: Obtain the initial point cloud data of the radar scene point cloud; Step 2: Quantify the initial point cloud data; Step 3: Input the quantized point cloud into the frequency perception module, sort each point in the point cloud from high to low according to the frequency, and adaptively adjust the high-frequency and low-frequency ratio according to the quantization scale of the point cloud, divide the point cloud according to the ratio, and discard the high-frequency part; Step 4: Input the remaining low-frequency point cloud into the low-frequency adapter and downsampling network, a total of 3 times; Step 5: Encode the point cloud after 3 times of downsampling into a bit stream through the encoder; Step 6: Decode the bitstream into a rough point cloud through a decoder; Step 7: Input the rough point cloud into the low-frequency adapter and upsampling network, a total of 3 times; Step 8: Dequantize the point cloud after 3 times upsampling to obtain the reconstructed point cloud; Step 1: Obtain initial point cloud data of the radar scene point cloud. Point cloud data refers to a data set of a group of discrete points in a three-dimensional coordinate system. The data in the data set at least includes the position data of the point, such as a three-dimensional coordinate vector (x, y, z). Step 2: Quantization operation of point cloud. The specific steps are as follows: 2-1. Let the original point cloud data be P and the offset be offset. Then the offset is calculated as: offset=min(min(P x ),min(P y ),min(P z ) Where P x , P y , P z They represent the coordinates of the three dimensions of the point cloud, namely, the offset is equal to the minimum value of the coordinates of all dimensions of all points in the point cloud; 2-2. The target data range ql of the point cloud is set according to the actual application requirements. After setting, the coordinates of the points will be quantized to the interval [0,ql-1]; 2-3. The quantization step size qs is determined by the coordinate range of the original point cloud and the target data range of the point cloud. The calculation formula is: 2-4. Where max(P) and min(P) represent the maximum and minimum values ​​of the coordinates of the point cloud in the three dimensions of x, y, and z, respectively; 2-5. The method of quantifying point cloud data is expressed by the following formula: Among them, P Q Represents the quantized point cloud data, and round represents the rounding operation of the data; In step 3, the quantized point cloud is input into the frequency perception module. The specific steps are as follows: 3-1. The calculation of the frequency perception module includes three parts: the Gaussian matrix calculated by the Euclidean distance and the Gaussian function, the output features calculated by the Haar graph filter, and the score of each point calculated by the L2-norm based on the output features; 3-2. Input the initial point cloud data N×3 into the frequency perception module, where the output is the Euclidean distance matrix K×K and Gaussian matrix K×K of each point and its K neighboring points; The calculation formula of the Euclidean distance matrix is: Where i and j are positive integers not greater than K, representing the points in the K neighboring point cloud, D i,j and D j,i is the Euclidean distance between point i and point j, and the matrix is ​​a symmetric matrix; The calculation formula of Gaussian matrix is: Where i and j are positive integers not greater than K, representing the points in the K neighboring point cloud, D i,j is the Euclidean distance between point i and point j, A i,j is the Gaussian value between point i and point j. Obviously, the matrix is ​​a symmetric matrix. 3-3. Use Haar graph filter to calculate features. The calculation formula is: h(A)P=(IA)P 3-4. For each point in the point cloud, use the above results to calculate the L2-norm, and use this L2-norm as the score. The low-frequency points have low scores, and the high-frequency points have high scores. Sort them from low to high to obtain the frequency-sorted point cloud; 3-5. Adaptively adjust the ratio of high frequency points, obtained by the following formula: in, It represents the ratio of high-frequency points to the total number of points in the quantized point cloud, round represents the rounding function, and ql represents the range of the quantized data; 3-6. Create a new sparse tensor consisting only of low-frequency The number of points is composed of N, where N represents the number of points in the point cloud after quantization. The high-frequency points are discarded by the above method, and the retained point clouds are all low-frequency points that are convenient for encoding and compression. Step 4 adapts and downsamples the low-frequency point cloud. The specific steps are as follows: 4-1. Input the low-frequency point cloud and high-frequency ratio fl into the low-frequency adapter, where the ratio fl is broadcasted into a sparse tensor and passed through a sparse fully connected layer and SiLU activation function to obtain a scaled ratio tensor Scale; the low-frequency point cloud is convolutional with a kernel size of 3 3 , the sparse convolution with the same number of output channels as the input channels, and then passes through the ReLU activation function, the sparse fully connected layer, and the ReLU activation function in sequence. The sparse tensor and the scaling ratio tensor Scale are added to the original low-frequency point cloud for feature fusion to obtain the adapted point cloud 1; this result is passed through the convolution kernel size of 3 3 , sparse convolution with 16 output channels, ReLU activation function, and convolution kernel size of 3 3 , the number of output channels is 32, sparse convolution with 1 / 2 times downsampling, ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU function, and the downsampled point cloud 1 is obtained; 4-2. Input the downsampled point cloud 1 and the high frequency ratio fl into the low frequency adapter to obtain the adapted point cloud 2, and then pass the adapted point cloud 2 through the convolution kernel size of 3 3 , sparse convolution with 32 output channels, ReLU activation function, and convolution kernel size of 3 3 , the number of output channels is 64, sparse convolution with 1 / 2 times downsampling, ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU function, and the downsampled point cloud 2 is obtained; 4-3. Input the downsampled point cloud 2 and the high frequency ratio fl into the low frequency adapter to obtain the adapted point cloud 3, and then pass the adapted point cloud 3 through the convolution kernel size of 3 3 , sparse convolution with 64 output channels, ReLU activation function, and convolution kernel size of 3 3 , the number of output channels is 32, sparse convolution with 1 / 2 times downsampling, ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU function, and the downsampled point cloud three is obtained; 4-4. The downsampled point cloud 3 is convolved with a kernel size of 3 3 , the sparse convolution with an output channel number of 8 is used to obtain the downsampled point cloud; Step 5 encodes the downsampled point cloud. This step uses the octree geometry encoder in G-PCC for lossless encoding to obtain a compressed bitstream. Step 6 is the reverse operation of step 5, using the geometric decoder to decode the bitstream to obtain a rough point cloud; Step 7: adapt and upsample the rough point cloud. The specific steps are as follows: 7-1. Input the rough point cloud and high-frequency ratio fl into the low-frequency adapter to obtain the adapted point cloud 4, and then pass the adapted point cloud 4 through the convolution kernel size of 3 3 , the number of output channels is 32, sparse convolution with 2x upsampling, ReLU activation function, and convolution kernel size is 3 3 , sparse convolution with 32 output channels, ReLU activation function, and 3 initial-residual network blocks consisting of sparse convolution and ReLU function to obtain the upsampled point cloud component, which is then convolved with a convolution kernel size of 3 3 , output a sparse convolution with a channel number of 1, and combine it with the aforementioned upsampled point cloud component to obtain upsampled point cloud 1; 7-2. Input the upsampled point cloud 1 and the high frequency ratio fl into the low frequency adapter to obtain the adapted point cloud 5, and then pass the adapted point cloud 5 through the convolution kernel size of 3 3 , the number of output channels is 64, sparse convolution with 2x upsampling, ReLU activation function, and convolution kernel size is 3 3 , sparse convolution with 64 output channels, ReLU activation function, and 3 initial-residual network blocks composed of sparse convolution and ReLU function to obtain the upsampled point cloud component, which is convolved with a convolution kernel size of 3 3 , output a sparse convolution with a channel number of 1, and combine it with the aforementioned upsampled point cloud component to obtain upsampled point cloud 2; 7-3. Input the upsampled point cloud 2 and the high frequency ratio fl into the low frequency adapter to obtain the adapted point cloud 6, and then pass the adapted point cloud 6 through the convolution kernel size of 3 3 , the number of output channels is 32, sparse convolution with 2x upsampling, ReLU activation function, and convolution kernel size is 3 3 , sparse convolution with 16 output channels, ReLU activation function, and 3 initial residual network blocks consisting of sparse convolution and ReLU function, followed by a convolution kernel size of 3 3 , sparse convolution with an output channel number of 1, and upsampled point cloud three is obtained; Step 8: Dequantize the upsampled point cloud. The formula is as follows: in To reconstruct the point cloud, P Q is the point cloud before the dequantization operation, qs is the quantization step calculated in step 2-3, and offset is the offset calculated in step 2-1.