Compression method and apparatus for point cloud data, and electronic device and storage medium
Patent Information
- Application Number
- PCT/CN2025/104581
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-18
- Filing Date
- 2025-06-27
- Publication Date
- 2026-09-24
Smart Images

Figure CN2025104581_24092026_PF_FP_ABST
Abstract
Description
Methods, devices, electronic equipment, and storage media for compressing point cloud data
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202510321806.9, filed on March 18, 2025, entitled “Method, Apparatus, Electronic Device and Storage Medium for Compression of Point Cloud Data”, which is incorporated herein by reference in its entirety. Technical Field
[0003] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and storage medium for compressing point cloud data. Background Technology
[0004] Point cloud data, as an important representation of three-dimensional spatial information, has wide applications in fields such as autonomous driving, remote sensing and mapping, and virtual reality. However, point cloud data is typically characterized by massive volume and high dimensionality, posing significant challenges to storage and transmission. Therefore, point cloud compression technology has become a hot research topic.
[0005] Various point cloud compression techniques exist. For example, octree-based point cloud compression methods transform point cloud data into an octree structure, utilizing the hierarchical and spatial partitioning properties of octrees to organize and encode the data. However, existing octree-based deep learning models mostly employ Transformer architectures and large-scale context for prediction, resulting in high computational complexity and long inference times. Other point cloud compression methods also present a trade-off between compression efficiency and inference speed, making it difficult to meet the requirements of low-latency point cloud data transmission scenarios.
[0006] Therefore, how to efficiently achieve lossless compression of point clouds to meet the needs of low-latency lossless transmission of point cloud data is an urgent technical problem to be solved. Summary of the Invention
[0007] To address the aforementioned problems in the prior art, this application provides a method, apparatus, electronic device, and storage medium for compressing point cloud data, so as to efficiently achieve lossless compression of point clouds and meet the needs of low-latency lossless transmission scenarios of point cloud data.
[0008] This application provides a method for compressing point cloud data, including the following steps.
[0009] Acquire a depth image of the target point cloud; quantize the depth values in the depth image to obtain the quantization result of the depth image; generate a compressed target point cloud based on the quantization result of the depth image.
[0010] According to the point cloud data compression method provided in this application, the step of quantizing the depth values in the depth image to obtain the quantization result of the depth image includes: quantizing the depth values in the depth image using a first quantization step size to obtain a first quantization result of the depth image; predicting a second quantization step size based on the first quantization result using a quantization step size prediction model; wherein the quantization step size prediction model is a trained machine learning model; quantizing the difference between each element in the depth image and its corresponding element in the first quantization result based on the second quantization step size to obtain a second quantization result; and obtaining the quantization result of the depth image based on the first quantization result and the second quantization result.
[0011] According to a point cloud data compression method provided in this application, the step of generating a compressed target point cloud based on the quantization result of the depth image includes: extracting latent vectors of the quantization result of the depth image using a neural network; predicting a first probability distribution of depth values in the quantization result of the depth image using a depth value probability prediction model based on the quantization result of the latent vectors; wherein the depth value probability prediction model is a trained machine learning model used to predict the probability distribution of depth values in the quantization result of the corresponding depth image based on the quantization result of the latent vectors; entropy encoding the quantization result of the depth image according to the first probability distribution to obtain the compressed target point cloud; and entropy encoding the quantization result of the latent vectors to obtain the entropy encoding of the latent vectors, so that during the reconstruction process of the compressed target point cloud, the compressed target point cloud is decoded based on a second probability distribution of depth values in the quantization result of the depth image predicted using the entropy encoding of the latent vectors.
[0012] According to the point cloud data compression method provided in this application, the latent vector includes a first latent vector and a second latent vector, the neural network includes a first feature extraction network and a second feature extraction network, and the entropy encoding includes a first entropy encoding and a second entropy encoding. The method of extracting the latent vector of the quantization result of the depth image using the neural network includes: extracting the first latent vector from the quantization result of the depth image using the first feature extraction network; and extracting the second latent vector from the first latent vector using the second feature extraction network. The method of predicting a first probability distribution of depth values in the quantization result of the depth image using a depth value probability prediction model based on the quantization result of the latent vector includes: predicting the first probability distribution based on the quantization result of the first latent vector using the depth value probability prediction model. The method of entropy encoding the quantization result of the latent vector to obtain the entropy encoding of the latent vector includes: predicting the probability distribution of each element value in the first latent vector based on the quantization result of the second latent vector using a latent vector element value distribution probability prediction model; entropy encoding the quantization result of the first latent vector based on the probability distribution of each element value in the first latent vector to obtain the first entropy encoding; and entropy encoding the quantization result of the second latent vector to obtain the second entropy encoding.
[0013] According to the point cloud data compression method provided in this application, the depth value probability prediction model includes a first depth value probability prediction model, a second depth value probability prediction model, a third depth value probability prediction model, and a fourth depth value probability prediction model; the method further includes: dividing the first quantization result of the depth image into two parts based on a chessboard context structure to obtain a first context group and a second context group; dividing the second quantization result of the depth image into two parts based on a chessboard context structure to obtain a third context group and a fourth context group; and predicting the first probability distribution of depth values in the quantization result of the depth image using the depth value probability prediction model based on the quantization result of the latent vector, including: predicting the first context group using the first depth value probability prediction model based on the quantization result of the first latent vector. The third probability distribution of depth values is obtained; using the second depth value probability prediction model, based on the quantization result of the first latent vector and the first context group, the fourth probability distribution of depth values in the second context group is predicted; using the third depth value probability prediction model, based on the quantization result of the first latent vector, the first context group, and the second context group, the fifth probability distribution of depth values in the third context group is predicted; using the fourth depth value probability prediction model, based on the quantization result of the first latent vector, the first context group, the second context group, and the third context group, the sixth probability distribution of depth values in the fourth context group is predicted; based on the third probability distribution, the fourth probability distribution, the fifth probability distribution, and the sixth probability distribution, the first probability distribution is obtained.
[0014] This application provides a method for reconstructing point cloud data, including the following steps.
[0015] A compressed point cloud is obtained; wherein the compressed point cloud is generated based on the quantization result of the depth image of the target point cloud; the compressed point cloud is decoded to obtain the quantization result of the depth image of the target point cloud; the target point cloud is reconstructed based on the quantization result of the depth image of the target point cloud.
[0016] According to the point cloud data reconstruction method provided in this application, the quantization result of the depth image is obtained by: quantizing the pixels in the depth image using a first quantization step size to obtain a first quantization result of the depth image; predicting a second quantization step size based on the first quantization result using a quantization step size prediction model; wherein the quantization step size prediction model is a trained machine learning model; quantizing the difference between each element in the depth image and its corresponding element in the first quantization result based on the second quantization step size to obtain a second quantization result; and obtaining the quantization result of the depth image based on the first quantization result and the second quantization result.
[0017] This application also provides a point cloud data compression device, comprising the following modules: an acquisition module for acquiring a depth image of a target point cloud; a quantization module for quantizing the depth values in the depth image to obtain a quantization result of the depth image; and a generation module for generating a compressed target point cloud based on the quantization result of the depth image.
[0018] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the point cloud data compression method described above.
[0019] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the point cloud data compression method as described above.
[0020] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the point cloud data compression method described above.
[0021] The point cloud data compression method provided in this application acquires a depth image of the target point cloud; quantizes the depth values in the depth image to obtain the quantization result of the depth image; and generates a compressed target point cloud based on the quantization result of the depth image. By converting the 3D point cloud into a 2D depth image and then compressing the point cloud based on the quantization result of the depth image, the complexity of the point cloud compression algorithm is effectively reduced, and the encoding and decoding speed of the point cloud compression algorithm is accelerated. This allows for efficient lossless compression of point clouds, meeting the requirements of low-latency lossless transmission scenarios for point cloud data. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 is a flowchart illustrating the point cloud data compression method provided in this application.
[0024] Figure 2 is a flowchart illustrating the method for generating compressed target point clouds based on the quantization results of depth images provided in this application.
[0025] Figure 3 is a flowchart illustrating the point cloud data reconstruction method provided in this application.
[0026] Figure 4 is a flowchart illustrating the training method of the machine learning model used in the compression and reconstruction process of point cloud data provided in this application.
[0027] Figure 5 is a schematic diagram of the encoding and decoding process of the point cloud data provided in this application.
[0028] Figure 6 is a schematic diagram of the process of extracting and encoding latent vectors provided in this application.
[0029] Figure 7 is a schematic diagram of the point cloud data compression device provided in this application.
[0030] Figure 8 is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0032] The point cloud data compression method of this application is described below with reference to Figures 1-6.
[0033] Figure 1 is a flowchart illustrating the point cloud data compression method provided in this application. As shown in Figure 1, the method includes the following:
[0034] Step 101: Obtain the depth image of the target point cloud.
[0035] The target point cloud is the point cloud that needs to be compressed. The target point cloud can include, but is not limited to, radar point clouds, infrared point clouds, and laser point clouds.
[0036] In some embodiments, the depth image of the target point cloud can be a range image. A range image is a two-dimensional image generated based on the LiDAR scanning results. Each pixel in the image stores the distance (i.e., depth value) of a point in the point cloud from the LiDAR sensor. The horizontal coordinate of the pixel gives the azimuth angle of the point in the polar coordinate system, and the vertical coordinate corresponds to the pitch angle of the point in the polar coordinate system.
[0037] In practice, target point clouds can be acquired through various methods depending on the application scenario, and are not limited to the descriptions in this manual. For example, during the operation of an autonomous vehicle, onboard LiDAR can be used to continuously scan the surrounding environment to obtain target point cloud data.
[0038] Step 102: Quantize the depth values in the depth image to obtain the quantization result of the depth image.
[0039] In practice, depth values in depth images can be quantized in various ways to obtain the quantization results of the depth images.
[0040] In some embodiments, a first quantization step size is used to quantize the depth values in the depth image to obtain a first quantization result of the depth image; a second quantization step size is predicted based on the first quantization result using a quantization step size prediction model; wherein the quantization step size prediction model is a trained machine learning model; based on the second quantization step size, the difference between each element in the depth image and its corresponding element in the first quantization result is quantized to obtain a second quantization result; and the quantization result of the depth image is obtained based on the first quantization result and the second quantization result.
[0041] In practice, the value of the first quantization step can be determined based on experience or experimental results. For example, the first quantization step can be 1, 2, etc., and is not limited by the description in this specification.
[0042] In practical implementation, quantization step size prediction models can be constructed in various ways. For example, they can be built based on convolutional neural networks. For training methods of the quantization step size prediction model, please refer to the relevant content in Figure 4, which will not be elaborated here.
[0043] As an example only, the quantization process in the above embodiments can be implemented using the following formula: x1=floor(r) (1)
[0044] Where x1 is the first quantization result; r is the depth value with continuous values in the depth image of the target point cloud; the floor() function is used to round down.
[0045] The following steps were achieved using formula (1): The depth values in the depth image were quantized using a first quantization step size of 1 to obtain the first quantization result x1 of the depth image. s2=g θ (x1) (2) x2=round(clip[(r-x1) / s2,-Q N Q P ])×s2 (3)
[0046] Where x2 is the second quantization result, g θ This is the quantization step size prediction model; s2 is the predicted second quantization step size; round() is the rounding operation, and clip operation refers to restricting the result to [-Q]. N Q P Within the range.
[0047] The following steps were achieved using formula (2): using the quantization step size prediction model, the second quantization step size was predicted based on the first quantization result.
[0048] The following steps are achieved by formula (3): Based on the second quantization step size, the difference between each element in the depth image and its corresponding element in the first quantization result is quantized to obtain the second quantization result.
[0049] Finally, the first quantization result and the second quantization result are used as the quantization result of the depth image.
[0050] In the embodiments provided in this application, a first quantization result is obtained by performing preliminary quantization processing on the depth values in the depth image using a first quantization step size, thereby mapping continuous depth values to discrete numerical values. A second quantization step size is predicted based on the first quantization result using a quantization step size prediction model. Since the powerful reasoning and prediction capabilities of the machine learning model are fully utilized, the quantization step size can be dynamically adjusted based on the existing quantization result, thereby obtaining a second quantization step size that can effectively improve the accuracy of quantization. Based on the predicted second quantization step size, the difference between each element in the depth image and its corresponding element in the first quantization result is quantized to obtain a second quantization result. By quantizing the difference, detailed information in the depth image can be effectively captured while maintaining high compression. Furthermore, the lossless compression of the target point cloud can be efficiently completed by using the quantization result of the depth image that combines the first and second quantization results.
[0051] Step 103: Generate a compressed target point cloud based on the quantization results of the depth image.
[0052] In practice, compressed target point clouds can be generated based on the quantization results of depth images in various ways, without being limited by the description in this specification.
[0053] An example of generating a compressed target point cloud based on the quantization results of the depth image is shown in Figure 2, and will not be repeated here.
[0054] Figure 2 is a flowchart illustrating the method for generating compressed target point clouds based on the quantization results of depth images provided in this application. As shown in Figure 2, the method includes the following:
[0055] Step 201: Use a neural network to extract the latent vector of the quantization result of the depth image.
[0056] In practice, convolutional neural networks can be used to extract the latent vectors of the quantization results of depth images.
[0057] In some embodiments, to further improve compression efficiency, as shown in Figure 5, the hyper prior of the latent vector can be extracted based on the latent vector of the quantization result of the extracted depth image.
[0058] In the above embodiment, the latent vectors include a first latent vector (e.g., latent vector y in Figure 5) and a second latent vector (e.g., latent vector z in Figure 5), and the neural network includes a first feature extraction network and a second feature extraction network. As shown in Figure 6, the first latent vector can be extracted from the quantization result of the depth image using the first feature extraction network (e.g., using a convolutional neural network); and the second latent vector can be extracted from the first latent vector using the second feature extraction network (e.g., using a convolutional neural network).
[0059] In practice, the first and second feature extraction networks can be constructed in various ways. For example, they can be constructed based on convolutional neural networks.
[0060] The training process of the first feature extraction network and the second feature extraction network is shown in Figure 4 and will not be repeated here.
[0061] Step 202: Using the depth value probability prediction model, based on the quantization results of the latent vector, predict the first probability distribution of the depth value in the quantization results of the depth image.
[0062] The depth value probability prediction model is a trained machine learning model used to predict the probability distribution of depth values in the quantization results of the corresponding depth image based on the quantization results of the latent vectors.
[0063] In practical implementation, depth value probability prediction models can be constructed in various ways. For example, they can be based on convolutional neural networks. The quantization results of the latent vectors are used as input to the depth value probability prediction model, which outputs the first probability distribution of the depth values in the quantized results of the predicted depth image.
[0064] In the embodiment containing the second latent vector described in step 201, the first probability distribution can be predicted using a depth value probability prediction model based on the quantization result of the first latent vector.
[0065] In some embodiments, the depth value probability prediction model includes a first depth value probability prediction model, a second depth value probability prediction model, a third depth value probability prediction model, and a fourth depth value probability prediction model. Based on the chessboard context structure, the first quantization result of the depth image is divided into two parts to obtain a first context group and a second context group; based on the chessboard context structure, the second quantization result of the depth image is divided into two parts to obtain a third context group and a fourth context group.
[0066] Using a first depth value probability prediction model, based on the quantization results of the first latent vector, a third probability distribution of depth values in the first context group is predicted. Using a second depth value probability prediction model, based on the quantization results of the first latent vector and the first context group, a fourth probability distribution of depth values in the second context group is predicted. Using a third depth value probability prediction model, based on the quantization results of the first latent vector, the first context group, and the second context group, a fifth probability distribution of depth values in the third context group is predicted. Using a fourth depth value probability prediction model, based on the quantization results of the first latent vector, the first context group, the second context group, and the third context group, a sixth probability distribution of depth values in the fourth context group is predicted. Based on the third, fourth, fifth, and sixth probability distributions, a first probability distribution is obtained.
[0067] In the specific implementation process, a first depth value probability prediction model, a second depth value probability prediction model, a third depth value probability prediction model, and a fourth depth value probability prediction model can be constructed based on convolutional neural networks.
[0068] In the specific implementation process, the quantization result of the latent vector can be input into the first depth value probability prediction model, which outputs the third probability distribution of depth values in the first context group; the quantization result of the first latent vector and the first context group can be input into the second depth value probability prediction model, which outputs the fourth probability distribution of depth values in the second context group; the quantization result of the first latent vector, the first context group, and the second context group can be input into the third depth value probability prediction model, which outputs the fifth probability distribution of depth values in the third context group; the quantization result of the first latent vector, the first context group, the second context group, and the third context group can be input into the fourth depth value probability prediction model, which outputs the sixth probability distribution of depth values in the fourth context group; the third probability distribution, the fourth probability distribution, the fifth probability distribution, and the sixth probability distribution are combined to obtain the first probability distribution.
[0069] For the training process of the first depth value probability prediction model, the second depth value probability prediction model, the third depth value probability prediction model, and the fourth depth value probability prediction model, please refer to the relevant content in Figure 4, which will not be repeated here.
[0070] Step 203: Based on the first probability distribution, entropy encoding is performed on the quantization result of the depth image to obtain the compressed target point cloud.
[0071] Entropy coding is a lossless compression method that assigns codes of different lengths based on the probability distribution of symbols. Symbols with higher probabilities are assigned shorter codes, while symbols with lower probabilities are assigned longer codes.
[0072] Based on the high-precision first probability distribution predicted by a neural network, lossless compression of target point clouds can be achieved using a smaller bitstream.
[0073] Step 204: Entropy coding is performed on the quantization results of the latent vectors to obtain the entropy coding of the latent vectors. In the process of reconstructing the compressed target point cloud, the second probability distribution of the depth values in the quantization results of the depth image predicted by the entropy coding of the latent vectors is used to decode the compressed target point cloud.
[0074] In the embodiment containing the second latent vector described in step 201, the entropy encoding includes first entropy encoding and second entropy encoding. The entropy encoding of the latent vector is implemented in the following manner.
[0075] Using a probability prediction model of the latent vector element value distribution, the probability distribution of each element value in the first latent vector is predicted based on the quantization result of the second latent vector. Based on the probability distribution of each element value in the first latent vector, the quantization result of the first latent vector is entropy encoded to obtain the first entropy code. Using a probability model (e.g., the probability model of Huffman coding) or a coding table, the quantization result of the second latent vector is entropy encoded to obtain the second entropy code.
[0076] The latent vector element value distribution probability prediction model is a trained machine learning model, such as a trained convolutional neural network. For the training process of the latent vector element value distribution probability prediction model, please refer to the relevant content in Figure 4, which will not be elaborated here.
[0077] For a detailed description of how the second probability distribution of depth values in the quantization result of the depth image predicted by the entropy coding of the latent vector is used to decode the compressed target point cloud during the reconstruction process, please refer to the relevant content in Figure 3, which will not be repeated here.
[0078] In the embodiments provided in this application, a neural network is first used to extract latent vectors from the depth image. Then, a depth value probability prediction model is used to accurately and efficiently predict the first probability distribution of depth values in the depth image based on the quantization results of the latent vectors. Based on this probability distribution, the entropy coding algorithm can accurately compress the depth image with a smaller bitstream, ultimately obtaining the compressed target point cloud. This process leverages the powerful inference capabilities and data processing efficiency of the depth value probability prediction model to achieve low-latency, lossless compression of the point cloud.
[0079] Figure 3 is a flowchart illustrating the point cloud data reconstruction method provided in this application. As shown in Figure 3, the method includes the following:
[0080] Step 301: Obtain the compressed point cloud; wherein, the compressed point cloud is generated based on the quantization result of the depth image of the target point cloud.
[0081] In practice, compressed point clouds can be obtained in different ways depending on the application scenario.
[0082] As an example, in autonomous driving applications, during vehicle operation, the onboard LiDAR continuously scans roads, vehicles, pedestrians, and other objects, generating a large amount of point cloud data. This point cloud data is compressed using the methods shown in Figures 1 and 2, and then transmitted in real-time to the vehicle's central processing unit (CPU) or graphics processing unit (GPU) for decoding and reconstruction. Based on the reconstructed point cloud data, the vehicle can perform various processing and analysis tasks, such as recognizing road signs, detecting obstacles, and tracking other vehicles and pedestrians, thereby providing accurate navigation and obstacle avoidance capabilities for autonomous vehicles.
[0083] For details on the generation process of compressed point clouds, please refer to the relevant content in Figures 1 and 2, which will not be repeated here.
[0084] Step 302: Decode the compressed point cloud to obtain the quantization result of the depth image of the target point cloud.
[0085] In practice, the compressed point cloud can be decoded based on the generation method of the compressed point cloud to obtain the quantization result of the depth image of the target point cloud.
[0086] In some embodiments, the compressed point cloud is obtained by entropy encoding the quantization result of the depth image based on a first probability distribution of depth values in the quantization result of the depth image; the first probability distribution is predicted using a depth value probability prediction model based on the quantization result of the latent vector; the latent vector is extracted from the quantization result of the depth image using a neural network; during the generation of the compressed point cloud, the entropy encoding of the quantization result of the latent vector is simultaneously generated. In this embodiment, the compressed point cloud is decoded in the following manner to obtain the quantization result of the depth image of the target point cloud.
[0087] Based on the entropy encoding of the quantization result of the latent vector and the probability model or encoding table used when encoding the quantization result of the latent vector, the quantization result of the latent vector is decoded.
[0088] Using a depth value probability prediction model, the second probability distribution of depth values in the quantization results of the depth image is predicted based on the quantization results of the latent vectors.
[0089] Based on the second probability distribution, entropy decoding is performed on the compressed point cloud to obtain the quantization result of the depth image of the target point cloud.
[0090] In some embodiments, the latent vector includes a first latent vector and a second latent vector. Entropy coding includes first entropy coding and second entropy coding.
[0091] The first latent vector is extracted from the quantization result of the depth image using the first feature extraction network; the second latent vector is extracted from the first latent vector using the second feature extraction network.
[0092] The first entropy encoding is obtained by entropy encoding the quantization result of the first latent vector based on the probability distribution of each element value in the first latent vector. The probability distribution of each element value in the first latent vector is predicted based on the quantization result of the second latent vector using a latent vector element value distribution probability prediction model. The second entropy encoding is obtained by entropy encoding the quantization result of the second latent vector using a probability model or encoding table. The first probability distribution is predicted based on the quantization result of the first latent vector using a depth value probability prediction model.
[0093] In the above embodiments, the compressed point cloud is decoded to obtain the quantization result of the depth image of the target point cloud:
[0094] The second entropy code is decoded using the probability model or encoding table used to generate the second entropy code, and the quantization result of the second latent vector is obtained.
[0095] By using the probability prediction model of the latent vector element value distribution, the probability distribution of each element in the first latent vector is predicted based on the quantization result of the second latent vector.
[0096] By using the probability distribution of each element in the first hidden vector, the first entropy code is decoded to obtain the quantization result of the first hidden vector.
[0097] Using a depth value probability prediction model, the second probability distribution of depth values in the quantization result of the first latent vector is predicted based on the quantization result of the depth image.
[0098] Using the second probability distribution, entropy decoding is performed on the compressed point cloud to obtain the quantization result of the depth image of the target point cloud.
[0099] In some embodiments, the quantization result of the depth image is obtained in the following manner:
[0100] Using a first quantization step size, pixels in the depth image are quantized to obtain a first quantization result of the depth image; using a quantization step size prediction model, a second quantization step size is predicted based on the first quantization result; wherein, the quantization step size prediction model is a trained machine learning model; based on the second quantization step size, the difference between each element in the depth image and its corresponding element in the first quantization result is quantized to obtain a second quantization result; based on the first quantization result and the second quantization result, the quantization result of the depth image is obtained.
[0101] Based on the chessboard context structure, the first quantization result of the depth image is divided into two parts, resulting in the first context group and the second context group; based on the chessboard context structure, the second quantization result of the depth image is divided into two parts, resulting in the third context group and the fourth context group.
[0102] In the above embodiments, the depth value probability prediction model includes a first depth value probability prediction model, a second depth value probability prediction model, a third depth value probability prediction model, and a fourth depth value probability prediction model. The second probability distribution of depth values in the quantization result of the depth image can be predicted using the depth value probability prediction model based on the quantization result of the first latent vector.
[0103] The quantization result of the first latent vector is input into the first depth value probability prediction model, which outputs the seventh probability distribution of depth values in the first context group. The quantization result of the first latent vector and the predicted first context group are input into the second depth value probability prediction model, which outputs the eighth probability distribution of depth values in the second context group. The quantization result of the first latent vector, the first context group, and the second context group are input into the third depth value probability prediction model, which outputs the ninth probability distribution of depth values in the third context group. The quantization result of the first latent vector, the first context group, the second context group, and the third context group are input into the fourth depth value probability prediction model, which outputs the tenth probability distribution of depth values in the fourth context group. The seventh, eighth, ninth, and tenth probability distributions are combined to obtain the second probability distribution.
[0104] Step 303: Reconstruct the target point cloud based on the quantization results of the depth image of the target point cloud.
[0105] In practical implementation, the depth image can be reconstructed based on the quantization results of the depth image of the target point cloud.
[0106] In the embodiment described in step 302, based on the chessboard context structure, the first context group and the second context group can be combined to obtain the first quantization result; the third context group and the fourth context group can be combined to obtain the second quantization result. Then, the corresponding elements in the first quantization result and the second quantization result are added together to obtain the depth image of the target point cloud.
[0107] Finally, the target point cloud is reconstructed based on the depth image of the target point cloud.
[0108] Figure 4 is a flowchart illustrating the training method of the machine learning model used in the compression and reconstruction process of point cloud data provided in this application. As shown in Figure 4, the method includes the following steps.
[0109] Step 401: Obtain the training dataset.
[0110] In practice, different methods can be used to collect point cloud data as sample data for different application scenarios (such as autonomous driving scenarios). Then, lossless compressed point cloud data is obtained after compression and reconstruction of these sample data, and this data is used as the label for the corresponding sample data.
[0111] The training dataset consists of sample data and their corresponding labels.
[0112] Step 402: Using the training dataset, perform end-to-end training on the depth value probability prediction model, the latent vector element value distribution probability prediction model, the quantization step size prediction model, the first feature extraction network, and the second feature extraction network.
[0113] In the specific implementation process, loss functions (e.g., cross-entropy loss function) can be established for the depth value probability prediction model (including the first depth value probability prediction model, the second depth value probability prediction model, the third depth value probability prediction model and the fourth depth value probability prediction model), the latent vector element value distribution probability prediction model and the quantization step size prediction model, the first feature extraction network and the second feature extraction network.
[0114] In each training round, the sample data is used as the target point cloud, and the method provided in the embodiments of this application is applied to compress and reconstruct the sample data to obtain the predicted reconstructed point cloud. A loss function is used to evaluate the difference between the label corresponding to the sample data and the predicted reconstructed point cloud.
[0115] The model parameters are adjusted using a preset optimization algorithm (such as gradient descent) until the joint loss function converges or the preset number of training iterations are reached, resulting in the trained depth value probability prediction model, latent vector element value distribution probability prediction model, and quantization step size prediction model.
[0116] The point cloud data compression apparatus provided in this application is described below. The point cloud data compression apparatus described below can be referred to in correspondence with the point cloud data compression method described above.
[0117] Figure 7 is a schematic diagram of the point cloud data compression device provided in this application. As shown in Figure 7, the device 700 includes the following modules.
[0118] The acquisition module 710 is used to acquire the depth image of the target point cloud.
[0119] The quantization module 720 is used to quantize the depth values in the depth image to obtain the quantization result of the depth image.
[0120] The generation module 730 is used to generate a compressed target point cloud based on the quantization result of the depth image.
[0121] Figure 8 illustrates a schematic diagram of the physical structure of an electronic device. As shown in Figure 8, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. The processor 810, communication interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a point cloud data compression method. This method includes: acquiring a depth image of a target point cloud; quantizing the depth values in the depth image to obtain a quantization result of the depth image; and generating a compressed target point cloud based on the quantization result of the depth image.
[0122] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0123] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the point cloud data compression method provided by the above methods. The method includes: acquiring a depth image of a target point cloud; quantizing the depth values in the depth image to obtain a quantization result of the depth image; and generating a compressed target point cloud based on the quantization result of the depth image.
[0124] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for compressing point cloud data provided by the methods described above. The method includes: acquiring a depth image of a target point cloud; quantizing the depth values in the depth image to obtain a quantization result of the depth image; and generating a compressed target point cloud based on the quantization result of the depth image.
[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for compressing point cloud data, comprising: Obtain a depth image of the target point cloud; The depth values in the depth image are quantized to obtain the quantization result of the depth image; Based on the quantization results of the depth image, a compressed target point cloud is generated.
2. The point cloud data compression method according to claim 1, wherein, The step of quantizing the depth values in the depth image to obtain the quantization result of the depth image includes: Using the first quantization step size, the depth values in the depth image are quantized to obtain the first quantization result of the depth image; Using a quantization step size prediction model, a second quantization step size is predicted based on the first quantization result; wherein, the quantization step size prediction model is a trained machine learning model. Based on the second quantization step size, the difference between each element in the depth image and its corresponding element in the first quantization result is quantized to obtain the second quantization result; The quantization result of the depth image is obtained based on the first quantization result and the second quantization result.
3. The point cloud data compression method according to claim 1 or 2, wherein, The process of generating the compressed target point cloud based on the quantization result of the depth image includes: The latent vector of the quantization result of the depth image is extracted using a neural network; Using a depth value probability prediction model, a first probability distribution of depth values in the quantization result of the latent vector is predicted based on the quantization result of the latent vector; wherein, the depth value probability prediction model is a trained machine learning model used to predict the probability distribution of depth values in the quantization result of the corresponding depth image based on the quantization result of the latent vector. Based on the first probability distribution, the quantization result of the depth image is entropy encoded to obtain the compressed target point cloud; The quantization result of the latent vector is entropy encoded to obtain the entropy code of the latent vector. In the process of reconstructing the compressed target point cloud, the second probability distribution of the depth value in the quantization result of the depth image predicted by the entropy code of the latent vector is used to decode the compressed target point cloud.
4. The point cloud data compression method according to claim 3, wherein, The latent vector includes a first latent vector and a second latent vector; the neural network includes a first feature extraction network and a second feature extraction network; the entropy encoding includes a first entropy encoding and a second entropy encoding; the extraction of the latent vector of the quantization result of the depth image using the neural network includes: The first latent vector is extracted from the quantization result of the depth image using the first feature extraction network. The second feature extraction network is used to extract the second hidden vector from the first hidden vector; The method of using a depth value probability prediction model to predict the first probability distribution of depth values in the quantized result of the depth image based on the quantization result of the latent vector includes: Using the depth value probability prediction model, the first probability distribution is predicted based on the quantization result of the first latent vector; The step of entropy encoding the quantization result of the latent vector to obtain the entropy encoding of the latent vector includes: Using the probability prediction model of the latent vector element value distribution, the probability distribution of each element value in the first latent vector is predicted based on the quantization result of the second latent vector. Based on the probability distribution of each element value in the first hidden vector, the quantization result of the first hidden vector is entropy encoded to obtain the first entropy code; The quantization result of the second hidden vector is entropy encoded to obtain the second entropy code.
5. The point cloud data compression method according to claim 2, wherein, The depth value probability prediction model includes a first depth value probability prediction model, a second depth value probability prediction model, a third depth value probability prediction model, and a fourth depth value probability prediction model; the method further includes: Based on the chessboard context structure, the first quantization result of the depth image is divided into two parts to obtain the first context group and the second context group; Based on the chessboard context structure, the second quantization result of the depth image is divided into two parts to obtain the third context group and the fourth context group; The method of using a depth value probability prediction model to predict the first probability distribution of depth values in the quantized result of the depth image based on the quantization result of the latent vector includes: Using the first depth value probability prediction model, based on the quantization result of the first latent vector, the third probability distribution of depth values in the first context group is predicted; Using the second depth value probability prediction model, based on the quantization result of the first latent vector and the first context group, the fourth probability distribution of the depth value in the second context group is predicted; Using the third depth value probability prediction model, based on the quantization result of the first latent vector, the first context group, and the second context group, the fifth probability distribution of the depth value in the third context group is predicted; Using the fourth depth value probability prediction model, based on the quantization result of the first latent vector, the first context group, the second context group, and the third context group, the sixth probability distribution of the depth value in the fourth context group is predicted; The first probability distribution is obtained based on the third probability distribution, the fourth probability distribution, the fifth probability distribution, and the sixth probability distribution.
6. A method for reconstructing point cloud data, comprising: Obtain a compressed point cloud; wherein the compressed point cloud is generated based on the quantization result of the depth image of the target point cloud; The compressed point cloud is decoded to obtain the quantization result of the depth image of the target point cloud; The target point cloud is reconstructed based on the quantization results of the depth image of the target point cloud.
7. The method for reconstructing point cloud data according to claim 6, wherein, The quantization result of the depth image is obtained in the following way: Using a first quantization step size, the pixels in the depth image are quantized to obtain a first quantization result of the depth image; Using a quantization step size prediction model, a second quantization step size is predicted based on the first quantization result; wherein, the quantization step size prediction model is a trained machine learning model. Based on the second quantization step size, the difference between each element in the depth image and its corresponding element in the first quantization result is quantized to obtain the second quantization result; The quantization result of the depth image is obtained based on the first quantization result and the second quantization result.
8. A point cloud data compression device, comprising: The acquisition module is used to acquire depth images of the target point cloud; The quantization module is used to quantize the depth values in the depth image to obtain the quantization result of the depth image. The generation module is used to generate a compressed target point cloud based on the quantization results of the depth image.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the point cloud data compression method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the point cloud data compression method as described in any one of claims 1 to 7.