A Point Cloud Video Coding Method Based on Inter-Frame Implicit Correlation

By adopting the entropy minimization motion compensation method based on inter-frame implicit correlation in point cloud video encoding, a reference frame that minimizes conditional entropy is generated, which solves the problem of failing to effectively utilize inter-frame redundant information in the prior art, and achieves more efficient point cloud video compression.

CN116800979BActive Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310865197.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-14
Publication Date
2025-05-27
Estimated Expiration
2043-07-14

AI Technical Summary

Technical Problem

The prior art fails to effectively utilize inter-frame redundant information in point cloud video encoding, especially on dynamic frames, resulting in low compression efficiency.

Method used

Using an entropy minimization motion compensation method based on inter-frame implicit correlation, a reference frame with minimized conditional entropy is generated by voxelizing the point cloud and dividing it into small cubes, and a reference frame with minimized conditional entropy is generated for inter-frame entropy encoding.

Benefits of technology

Effectively utilize implicit correlation between frames, reduces the amount of data in point cloud video, improves compression rate, reduces bandwidth consumption, and improves compression performance on dynamic frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_20
    Figure SMS_20
  • Figure SMS_33
    Figure SMS_33
  • Figure SMS_62
    Figure SMS_62
Patent Text Reader

Abstract

A point cloud video encoding method based on inter-frame implicit correlation, which relates to the field of point cloud video compression, includes: First, motion compensation with minimized entropy: First, voxelize the point cloud; use a motion compensation method to generate a reference frame, which aligns the topological structure in the inter-frame implicit correlation while minimizing the conditional entropy; divide the frame into small cubes; use an index to evaluate the matching degree between two cubes; search for the cube with the best matching degree from the cubes of the previous frame for each cube of the current frame; splice each best matching cube to generate a reference frame; select the reference frame that can minimize the conditional entropy as the output of motion compensation with minimized entropy; Second, inter-frame entropy encoding. The present invention makes full use of the inter-frame redundant information of dynamic frames to perform lossless compression on point cloud videos, effectively reducing the consumption of the transmission bandwidth of the point cloud video stream; at the same time, it uses the inter-frame implicit correlation to compress the point cloud video, effectively reducing the amount of video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of point cloud video compression. Specifically, it relates to a point cloud video encoding method based on inter-frame implicit correlation. Background Art

[0002] Mainstream point cloud video codecs project point cloud videos onto 2D videos for encoding or directly encode point clouds. The V-PCC (Sebastian Schwarz et al. 2019. Emerging MPEG Standards for Point Cloud Compression. IEEE Journal on Emerging and Selected Topics in Circuits and Systems 9, 1 (2019), 133–148.) encoder projects the geometry and attributes of point cloud videos onto multiple 2D video tracks. Vues (Yu Liu et al. 2022. Vues: Practical Mobile Volumetric Video Streaming through Multiview Transcoding. In Proceedings of the 28th Annual International Conference on Mobile Computing And Networking (Sydney, NSW, Australia) (MobiCom’22). Association for Computing Machinery, New York, NY, USA, 514–527.) uses edge servers to transcode point cloud videos into 2D videos. Recently, there has also been work on directly encoding point cloud videos. Draco introduced by Google uses a kd-tree to encode point clouds. GROOT (Kyungjin Lee et al. 2020. GROOT: A Real-Time Streaming System of High-Fidelity Volumetric Videos. In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking (London, United Kingdom) (MobiCom’20). Association for Computing Machinery, New York, NY, USA, Article 57, 14 pages.) proposed a parallel octree to improve decoding efficiency.YuZu (Anlan Zhang et al. 2022. YuZu: Neural-Enhanced Volumetric Video Streaming. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22). USENIX Association, Renton, WA, 137–154.) uses 3D super-resolution technology to increase the point cloud density. AITransfer (Yakun Huang et al. 2021. AITransfer: Progressive AI-Powered Transmission for Real-Time Point Cloud Video Streaming. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM’21). Association for Computing Machinery, New York, NY, USA, 3989–3997.) encodes point clouds using an AI model. Generally speaking, these systems encode each frame of the point cloud independently without considering the inter-frame redundant information.

[0003] In recent years, there have also been some advances in the direction of inter-frame coding of point cloud videos. Kammerl (Julius Kammerl et al. 2012. Real-time compression of point cloud streams. In 2012 IEEE International Conference on Robotics and Automation. 778–785.) et al. proposed representing a point cloud frame as the difference from the previous frame. However, this method is only effective for point cloud frames full of static content. Many subsequent studies use entropy coding-based compression methods to compress the inter-frame redundant information. Most of these methods rely on explicit correlations between frames, that is, the repeated information at the same position or adjacent positions between adjacent frames. However, they are not designed specifically for point cloud video streaming and ignore the video dynamics. Summary of the Invention

[0004] The objective of the present invention is to provide a point cloud video coding method based on implicit inter-frame correlations.

[0005] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0006] A point cloud video coding method based on inter-frame implicit correlation of the present invention includes the following steps:

[0007] Step S1: Motion compensation with minimum entropy;

[0008] Step S1.1: First, voxelize the point cloud, that is, map the point cloud into a three-dimensional grid;

[0009] Step S1.2: Use a motion compensation method to generate a reference frame, which aligns the topological structure in the inter-frame implicit correlation and minimizes the conditional entropy;

[0010] Step S1.2.1: Divide the frame into small cubes;

[0011] Step S1.2.2: Use an index to evaluate the matching degree between two cubes;

[0012] Step S1.2.3: Search for the cube with the best matching degree from the cubes of the previous frame for each cube of the current frame;

[0013] Step S1.2.4: Stitch each best-matching cube to generate a reference frame;

[0014] Step S1.2.5: Select the reference frame that can minimize the conditional entropy as the output of motion compensation with minimum entropy:

[0015] Step S2: Inter-frame entropy coding;

[0016] Use the inter-frame entropy coding algorithm S4D and the reference frame that can minimize the conditional entropy generated in step S1 as the context to code the current frame.

[0017] Further, the specific operation process of step S1.1 is as follows:

[0018] The voxelized point cloud frame is represented by a three-dimensional array, and each member, i.e., the voxel, is empty or occupied; the definition of the conditional entropy H is as follows:

[0019]

[0020] where P 0 represents the probability that the voxel in the current frame is empty; P 1 represents the probability that the voxel in the current frame is occupied; and respectively represent the conditional probabilities that the voxel in the current frame is empty and occupied when the voxel in the previous frame is empty; and respectively represent the conditional probabilities that the voxel in the current frame is empty and occupied when the voxel in the previous frame is occupied.

[0021] Furthermore, the specific operation process of step S1.2.1 is as follows:

[0022] The current frame and the previous frame are represented by I t and I t-1 respectively. The current frame I t is divided into non - overlapping cubes with side length of M voxels. Each cube is located at the position (x j , y j , z j ).

[0023] Furthermore, the specific operation process of step S1.2.2 is as follows:

[0024] To find the best - matching cube for each cube in the previous frame I t-1 , an exhaustive search is performed in a search window centered at with side length of W voxels, where is the cube in the previous frame I with the same position as the cube t-1 . This search space is represented as a set of candidate cubes . The motion vector is defined as the vector from the best - matching cube to the cube . When a candidate cube has a high degree of matching with the cube , the voxels in the cube are predicted by the voxels in the candidate cube . The prediction results are represented using a confusion matrix as shown in the following table;

[0025]

[0026] Among them, indicates that the voxel in the cube is occupied, i.e., the value is 1, indicates that the voxel in the cube is empty, i.e., the value is 0, indicates that the voxel in the candidate cube is occupied, i.e., the value is 1, indicates that the voxel in the candidate cube ​In this case, the voxel is empty, i.e., the value is 0. True Positive (TP) represents a true positive example, False Positive (FP) represents a false positive example, False Negative (FN) represents a false negative example, and True Negative (TN) represents a true negative example;

[0027] The candidate cube and the cube The matching degree is represented by precision or recall. Precision refers to the proportion of voxels with positive true values among the voxels predicted as positive, that is Recall refers to the proportion of voxels with positive true values that are predicted as positive, that is n(TP) represents the number of occurrences of TP examples, n(FP) represents the number of occurrences of FP examples, and n(FN) represents the number of occurrences of FN examples; The F-score is used to balance these two metrics to evaluate the matching degree F of the two cubes β :

[0028]

[0029] where β represents the balance coefficient, and β balances the importance of precision and recall; precision represents the precision value, and recall represents the recall value.

[0030] Furthermore, the specific operation process of step S1.2.3 is as follows:

[0031] In the set of candidate cubes the best matching cube that generates the highest F-score with the cube is regarded as the optimal matching result; Use k candidate matching metrics F 1 , …, F k to search for k best matching cubes for each cube in the current frame respectively.

[0032] Furthermore, the specific operation process of step S1.2.4 is as follows:

[0033] For a certain candidate matching metric F k after obtaining all the best matching cubes use the best matching cube to replace the cube to generate a reference frame; Since k candidate matching metrics are used, k reference frames will be generated.

[0034] Furthermore, the specific operation process of step S1.2.5 is as follows:

[0035] Use a series of β values to generate corresponding multiple reference frames, 0 < β 1 < … < βk < +∞, and calculate the conditional entropy corresponding to each reference frame Select the β corresponding to the minimized conditional entropy k value, and this β k The reference frame corresponding to the value is the motion compensation output with minimized entropy.

[0036] The beneficial effects of the present invention are as follows:

[0037] A point cloud video coding method based on inter-frame implicit correlation of the present invention makes full use of the inter-frame redundant information of dynamic frames to perform lossless compression on the point cloud video, effectively reducing the consumption of the transmission bandwidth of the point cloud video stream; at the same time, the present invention uses the inter-frame implicit correlation to compress the point cloud video. The key lies in using a motion compensation method with minimized entropy to generate a reference frame, effectively reducing the conditional entropy between the reference frame and the current frame, and effectively reducing the amount of video data. The present invention has obvious advantages in the compression performance of the point cloud video compared with the prior art. Specific embodiments

[0038] The present invention finds that the explicit inter-frame correlation is significantly reduced in frames with strong dynamics. For this reason, the present invention finds an inter-frame implicit correlation, that is, the consistency of the topological structure between adjacent frames. The present invention finds that even in frames with strong dynamics, the inter-frame implicit correlation remains at a relatively high level. Therefore, the inter-frame implicit correlation has great potential to help compress the point cloud video. In order to make full use of the inter-frame implicit correlation to compress the point cloud video, the present invention adopts the widely used entropy coding as the basic encoder model, which can use a reference frame as auxiliary information, while in the prior art, the previous frame is directly used as the reference frame, and the coding effect is not good. At the same time, the smaller the inter-frame conditional entropy, the higher the theoretical upper limit of the compression ratio can be provided, and simply using the existing motion estimation method to align adjacent frames cannot effectively reduce the inter-frame conditional entropy.

[0039] For this reason, the present invention provides a point cloud video coding method based on inter-frame implicit correlation, which specifically includes the following steps:

[0040] Step S1: Motion compensation with minimized entropy;

[0041] The goal of this step is to generate a reference frame that can effectively improve the compression ratio, and the effectiveness of the compression can be measured by the conditional entropy defined on the reference frame and the current frame. The specific operation steps are as follows:

[0042] Step S1.1: Point cloud voxelization;

[0043] First, voxelize the point cloud, that is, map the point cloud into a three-dimensional grid. The voxelized point cloud frame is represented by a three-dimensional array, and each member (voxel) of it is empty (represented by 0) or occupied (represented by 1). The definition of the conditional entropy H is as follows:

[0044]

[0045] Among them, P 0 represents the probability that the voxel in the current frame is empty; P 1 represents the probability that the voxel in the current frame is occupied; and respectively represent the conditional probabilities that the voxel in the current frame is empty and occupied when the voxel in the previous frame is empty; and respectively represent the conditional probabilities that the voxel in the current frame is empty and occupied when the voxel in the previous frame is occupied.

[0046] Step S1.2: Use a motion compensation method to generate a reference frame that aligns the topological structure in the implicit inter-frame correlation and minimizes the conditional entropy H, which specifically includes the following 5 steps:

[0047] Step S1.2.1: Divide the frame into small cubes;

[0048] The current frame and the previous frame are respectively represented by I t and I t-1 . Divide the current frame I t into non-overlapping cubes with side length M voxels. Each cube is located at the position (x j , y j , z j ); divide the previous frame into non-overlapping cubes in the same way The cube is located at the position (x j , y j , z j ).

[0049] Step S1.2.2: Use an index to evaluate the matching degree between two cubes;

[0050] To find the best matching cube for each cube t-1 in the previous frame I in the search window centered at with side length W voxels for an exhaustive search, where is the cube in the previous frame I t-1 with the same position as the cube ; this search space is represented as a set of candidate cubes The motion vector is defined as pointing from the best matching cube to the cube

[0051] When a candidate cube has a high degree of matching with the cube , the voxels in the candidate cube can be used to predict the corresponding voxels in the cube . Regarding this process as a binary prediction, since each voxel can be 1 (Positive) or 0 (Negative), a confusion matrix is used to represent the prediction results, as shown in Table 1.

[0052] Table 1

[0053]

[0054] Among them, indicates that the voxel is occupied in the cube , that is, the value is 1, indicates that the voxel is empty in the cube , that is, the value is 0, indicates that the voxel is occupied in the candidate cube , that is, the value is 1, indicates that the voxel is empty in the candidate cube , that is, the value is 0. True Positive (TP) represents a true positive example, False Positive (FP) represents a false positive example, False Negative (FN) represents a false negative example, and True Negative (TN) represents a true negative example.

[0055] Use n(·) to represent the number of occurrences of a specific element in the confusion matrix. The matching degree of two cubes (the candidate cube and the cube ) can be represented by precision or recall. Specifically, precision refers to the proportion of positive voxels in the predicted positives, that is while recall refers to the proportion of positive voxels in the true positives that are predicted as positive, that is where represents the number of occurrences of TP examples, n(FP) represents the number of occurrences of FP examples, and n(FN) represents the number of occurrences of FN examples. Since these two metrics are not unified, the present invention uses the F-score to balance these two metrics and better evaluate the matching degree F of the two cubes β :

[0056]

[0057] Among them, β represents the balance coefficient, β balances the importance of precision and recall; precision represents the precision value, and recall represents the recall value.

[0058] Step S1.2.3: Search for the cube with the best matching degree from the cubes of the previous frame for each cube of the current frame;

[0059] In the set of candidate cubes the best matching cube that generates the highest F-score with the cube is regarded as the optimal matching result. However, when searching for the best matching cube, it is difficult to take into account minimizing the conditional entropy because the conditional entropy depends on the entire reference frame. For this reason, the present invention uses k candidate matching metrics F 1 , …, F k to search for k best matching cubes for each cube of the current frame respectively.

[0060] Step S1.2.4: Stitch together each best matching cube to generate a reference frame;

[0061] Since step S1.2.3 uses k candidate matching metrics, k reference frames will be generated. For a certain candidate matching metric F k after obtaining all the best matching cubes the best matching cube is used to replace the cube to generate a reference frame.

[0062] Step S1.2.5: Select the reference frame that can minimize the conditional entropy;

[0063] The conditional entropy depends on the reference frame and its corresponding matching degree F β metric. However, during the process of generating the reference frame, it is not certain which β value corresponds to the reference frame with the minimized conditional entropy. Therefore, the present invention simultaneously uses multiple β values to generate corresponding multiple reference frames and calculates the conditional entropy corresponding to each reference frame. Let H β represent the conditional entropy of the reference frame corresponding to the matching degree F β metric. Using a series of β values, 0 < β 1 < … < β k < +∞, and then calculating the corresponding conditional entropy The β value corresponding to the minimum conditional entropy k is selected, and the corresponding reference frame is the motion compensation output with the minimized entropy.

[0064] Step S2: Inter-frame entropy coding;

[0065] Use the inter-frame entropy coding algorithm (S4D) and the reference frame that can minimize the conditional entropy generated in step S1 as the context to encode the current frame. The specific operation steps are as follows:

[0066] S4D encodes each element of the three-dimensional array of the point cloud using context-adaptive binary arithmetic coding (CABAC). Specifically, for any element value (0 or 1) in the three-dimensional array of the current frame, S4D uses the element value at the same position in the three-dimensional array of the reference frame as the context for CABAC encoding.

[0067] The prior art directly uses the previous frame as the reference frame and only utilizes the explicit inter-frame correlation (the repeated information at the same position or adjacent positions between adjacent frames), resulting in poor compression efficiency and performance on frames with high dynamics. In contrast, the present invention specifically generates a reference frame corresponding to the minimized conditional entropy through step S1, which is very beneficial for inter-frame entropy coding. Meanwhile, the present invention effectively utilizes the implicit inter-frame correlation and can effectively improve the compression ratio on dynamic frames. The prior art uses multiple voxels of the current frame and the reference frame as the probability conditions in entropy coding, causing a high computational complexity and a decoding frame rate less than 1 FPS. The present invention only uses the left adjacent voxel of the voxel to be encoded and the voxel at the same position in the reference frame as the probability conditions in entropy coding, reducing the computational complexity. Therefore, the present invention is more suitable for mobile streaming systems.

[0068] To verify the effect of a point cloud video coding method based on implicit inter-frame correlation of the present invention, three publicly available datasets, Ricardo, Pizza, and Longdress, are used to conduct experimental comparisons between the prior art and the present invention. The results show that compared with the prior inter-frame encoder that utilizes explicit inter-frame correlation, the present invention reduces the bandwidth consumption by 23.15%, 1.06%, and 43.32% respectively on the three datasets. It can be seen that the present invention has obvious advantages in point cloud video compression performance compared with the prior art.

[0069] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A point cloud video encoding method based on inter-frame implicit correlation, characterized in that, it includes the following steps: Step S1: Motion compensation with minimum entropy; Step S1.1: First, voxelize the point cloud, that is, map the point cloud into a three-dimensional grid; Step S1.2: Use a motion compensation method to generate a reference frame, which aligns the topological structure in the inter-frame implicit correlation and minimizes the conditional entropy; Step S1.2.1: Divide the frame into small cubes; Step S1.2.2: Use an index to evaluate the matching degree between two cubes; Step S1.2.3: Search for the cube with the best matching degree from the cubes of the previous frame for each cube of the current frame; Step S1.2.4: Stitch each best-matching cube to generate a reference frame; Step S1.2.5: Select the reference frame that can minimize the conditional entropy as the output of motion compensation with minimum entropy; Step S2: Inter-frame entropy encoding; Use the inter-frame entropy encoding algorithm S4D and the reference frame that can minimize the conditional entropy generated in Step S1 as the context to encode the current frame.

2. The point cloud video encoding method based on inter-frame implicit correlation according to claim 1, characterized in that, the specific operation process of Step S1.1 is as follows: The voxelized point cloud frame is represented by a three-dimensional array, and each member, i.e., the voxel, is either empty or occupied; the definition of the conditional entropy H is as follows: Among them, P 0 represents the probability that the voxel in the current frame is empty; P 1 represents the probability that the voxel in the current frame is occupied; and respectively represent the conditional probabilities that the voxel in the current frame is empty and occupied when the voxel in the previous frame is empty; and respectively represent the conditional probabilities that the voxel in the current frame is empty and occupied when the voxel in the previous frame is occupied.

3. The point cloud video encoding method based on inter-frame implicit correlation according to claim 1, characterized in that, the specific operation process of Step S1.2.1 is as follows: The current frame and the previous frame are represented by I t and I t-1 respectively. The current frame I t is divided into non - overlapping cubes with side length M voxels. Each cube is located at the position (x j , y j , z j ).

4. The point cloud video encoding method based on inter-frame implicit correlation according to claim 3, characterized in that, the specific operation process of Step S1.2.2 is as follows: In order to find the best-matching cube for each cube in the previous frame I t-1 perform an exhaustive search within a search window centered at and with a side length of W voxels, where is the cube in the previous frame I that has the same position as the cube t-1 in the previous frame I ; The search space is represented as a set of candidate cubes The motion vector is defined as from the best matching cube Pointing to the cube When a candidate cube Matches the cube With a high degree of matching, the voxels in the candidate cube Are used to predict the corresponding voxels in the cube The prediction results are represented using a confusion matrix as shown in the following table; Among them, Positive in means that the voxel is occupied in the cube , that is, the value is 1. Negative in means that the voxel is empty in the cube , that is, the value is 0. Positive in means that the voxel is occupied in the candidate cube , that is, the value is 1. Negative in means that the voxel is empty in the candidate cube , that is, the value is 0. True Positive (TP) represents a true positive example, False Positive (FP) represents a false positive example, False Negative (FN) represents a false negative example, and True Negative (TN) represents a true negative example; The candidate cube The matching degree with the cube is represented by precision or recall. Precision refers to the proportion of voxels predicted as positive that are actually positive, i.e., Recall refers to the proportion of voxels that are actually positive and are predicted as positive, i.e., n(TP) represents the number of occurrences of TP examples, n(FP) represents the number of occurrences of FP examples, and n(FN) represents the number of occurrences of FN examples; the F-score is used to balance these two metrics to evaluate the matching degree F of the two cubes β : where β represents the balance coefficient, β balances the importance of precision and recall; precision represents the precision value, and recall represents the recall value.

5. The point cloud video encoding method based on inter-frame implicit correlation according to claim 4, characterized in that, the specific operation process of Step S1.2.3 is as follows: In the set of candidate cubes , the best matching cube that produces the highest F-score with the cube is regarded as the optimal matching result; k candidate matching metrics F 1 , …, F k are used to search for k best matching cubes for each cube of the current frame respectively.

6. The point cloud video encoding method based on inter-frame implicit correlation according to claim 5, characterized in that, the specific operation process of Step S1.2.4 is as follows: For a certain candidate matching metric F k , after obtaining all the best matching cubes , use the best matching cube to replace the cube to generate a reference frame; since k candidate matching metrics are used, k reference frames will be generated.

7. The point cloud video encoding method based on inter-frame implicit correlation according to claim 6, characterized in that, the specific operation process of Step S1.2.5 is as follows: Use a series of β values to generate a corresponding plurality of reference frames, where 0 < β 1 <… < β k < +∞, and calculate the conditional entropy corresponding to each reference frame Select the β corresponding to the minimized conditional entropy k value. The reference frame corresponding to this β k value is the motion compensation output with minimized entropy.

Citation Information

Patent Citations

  • Point cloud geometrical information inter-frame encoding and decoding method

    CN112565764A

  • Point cloud compression method, encoder, decoder and storage medium

    CN113766228A