Coding unit optimization division and dynamic point cloud compression method for dynamic point cloud compression
By training and division to identify the gradient and variance of the coding unit, filtering and predicting the optimal division mode, the problem of high computational complexity of coding unit division in the prior art is solved, and efficient dynamic point cloud compression is achieved.
Patent Information
- Application Number
- CN202510184217.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-19
AI Technical Summary
In the existing dynamic point cloud compression method, the calculation complexity of encoding unit division is high, which leads to a long time-consuming process, which is not conducive to real-time encoding applications.
By training the division recognition model, the attribute characteristics and geometric features of the coding unit are extracted, the gradient and variance are calculated, the division mode is initially screened using division rules, and the neural network model is used to predict the optimal division mode to reduce unnecessary division mode search.
It reduces the encoding calculation overhead, improves coding efficiency, shortens encoding time, and improves the real-time compression performance of dynamic point clouds.
Smart Images

Figure CN120034663A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point cloud coding, and in particular to a coding unit optimization division method and a dynamic point cloud compression method for dynamic point cloud compression. Background Art
[0002] With the rapid development of three-dimensional data acquisition technology, point cloud data has been widely used in autonomous driving, virtual reality, augmented reality and other fields. Compared with traditional two-dimensional image data, point cloud data can more comprehensively and accurately describe the shape, structure and dynamic changes of three-dimensional objects and scenes, thereby providing a more immersive interactive experience. However, due to the high precision and rich details of point cloud data, its data volume is extremely large, which brings huge challenges to storage and transmission. Therefore, how to efficiently compress point cloud data has become one of the research hotspots. To address this problem, MPEG proposed a video-based point cloud compression (V-PCC) scheme, such as Figure 1 As shown in the figure, the scheme projects 3D dynamic point clouds onto 2D video frames and compresses the projection data using existing video coding technologies (such as H.265 / HEVC, H.266 / VVC, etc.) to improve coding efficiency. V-PCC relies on traditional video encoders, which are mainly oriented to ordinary video data and fail to fully utilize the spatial structure characteristics of point cloud data, resulting in high computational complexity and affecting coding efficiency in real-time application scenarios.
[0003] In the V-PCC framework, the division method of coding units (CUs) directly affects the final compression efficiency and encoding quality. Currently, V-PCC uses H.265 / HEVC as the basic encoder by default, in which CU division adopts a quadtree (QT) structure, and the optimal block division method is determined by traversing different levels of division modes. With the evolution of video coding technology, the latest H.266 / VVC standard has significantly improved compression efficiency compared to H.265 / HEVC. H.266 / VVC adopts a more flexible multi-tree division mode, such as Figure 2As shown in the figure, the combined structures including quadtree (QT), binary tree (BT) and ternary tree (TT) make the CU segmentation more diversified, thereby improving the coding efficiency. In order to find the optimal CU partitioning of the current frame, H.266 needs to traverse all possible partitioning situations and calculate the rate-distortion cost (Rate-DistortionCost, RD-cost) of each partitioning scheme, and finally select the partitioning mode with the smallest RD-cost as the optimal scheme. Although this method improves the compression performance, it also significantly increases the amount of coding calculations, resulting in a long block partitioning process, which is not conducive to real-time coding applications. Therefore, in order to reduce the computational complexity of coding unit division and improve coding efficiency, it is necessary to study an efficient and fast partitioning method to reduce unnecessary partitioning mode searches while ensuring coding quality, thereby effectively reducing the coding calculation overhead and improving the real-time compression performance of dynamic point clouds. Summary of the invention
[0004] The present invention provides a coding unit optimization division method and a dynamic point cloud compression method for dynamic point cloud compression, which solves the defects of the existing division method that the coding calculation amount is significantly increased, the block division process is time-consuming, and it is not conducive to real-time coding applications.
[0005] The present invention can be achieved through the following technical solutions:
[0006] A coding unit optimization division method for dynamic point cloud compression comprises the following steps:
[0007] Step 1: Train the segmentation recognition model
[0008] Extract attribute features and geometric features of multiple coding units and construct data sets for training and testing multiple partition recognition models, which correspond to different partition modes respectively;
[0009] Step 2: Create a surveillance image dataset
[0010] The gradient and variance of the current coding unit are calculated to preliminarily screen the division mode of the current coding unit, and then the attribute characteristics and geometric features of the current coding unit are extracted. The trained division recognition model is selected to test the screened division modes one by one, and the optimal division mode is selected for the division of the current coding unit.
[0011] Furthermore, the gradient and variance of the current coding unit in the horizontal and vertical directions are calculated respectively, and the division mode is preliminarily screened using the following division rules:
[0012] If the vertical gradient is significantly greater than the horizontal gradient, the vertical division mode is disabled;
[0013] If the vertical gradient is not significantly greater than the horizontal gradient, the variance is used for auxiliary judgment;
[0014] If the horizontal gradient is significantly greater than the vertical gradient, the horizontal division mode is disabled;
[0015] If the horizontal gradient is not significantly greater than the vertical gradient, the variance is used for auxiliary judgment;
[0016] If the horizontal and vertical gradients are relatively balanced, all division modes are retained.
[0017] Further, let the horizontal gradient, vertical gradient, horizontal variance, and vertical variance of the current coding unit be G hor , G ver 、V hor 、V ver , then the gradient ratio Set the threshold value 1 ,th 2 , and th 1 <th 2 , the division rules are set as follows,
[0018] (VI) When and When , it means that the vertical gradient is significantly greater than the horizontal gradient, then the vertical division mode is directly disabled;
[0019] (VII) When When , it means that the vertical gradient is large but not significant enough, and the variance is used to assist judgment:
[0020] (3) When V hor >th 1 , indicating that the sub-block variance ratio in the horizontal direction is large, that is, the texture in the horizontal direction changes greatly, so the vertical division mode is disabled;
[0021] (4) If V hor ≤th 1 , indicating that the sub-block variance ratio in the horizontal direction is small, that is, the texture in the horizontal direction is relatively uniform, so all division modes are retained and the subsequent encoding process is entered;
[0022] (VIII) When G div >th 1 And G div >th 2 When , it means that the horizontal gradient is significantly greater than the vertical gradient, then the horizontal division mode is directly disabled;
[0023] (IX) When 1 <G div ≤th 2When , it means that the horizontal gradient is large but not significant enough, and the variance is used to assist judgment:
[0024] (3) If V ver >th 1 , indicating that the directional texture changes greatly, the horizontal division mode is disabled;
[0025] (4) If V ver ≤th 1 , indicating that the texture in the vertical direction is relatively uniform, so all division modes are retained;
[0026] (10) When , it means that the horizontal and vertical gradients are relatively balanced, so all division modes are retained.
[0027] Furthermore, a trained partition recognition model is selected to test the screened partition patterns one by one. If the output result of the current partition recognition model is less than the threshold, the next partition recognition model is selected to continue testing until the output result of the current partition recognition model is not less than the threshold. In this case, the partition pattern corresponding to the current partition recognition model is considered to be the optimal partition pattern.
[0028] Furthermore, the division modes include a binary tree horizontal division mode BTH, a binary tree vertical division mode BTV, a ternary tree horizontal division mode TTH, a ternary tree vertical division mode TTV and a non-division mode. The division recognition models corresponding to the first four division modes are BTH neural network model, BTV neural network model, TTH neural network model and TTV neural network model respectively.
[0029] Furthermore, the original V-PCC encoder is used to encode several frames of data on the dynamic point cloud under different quantization parameters. The final selected division mode is the optimal division mode. The attribute features and geometric features of the corresponding coding units are extracted to construct a data set for training and testing the division recognition model corresponding to the optimal division mode.
[0030] Furthermore, the attribute features include the aspect ratio, rate-distortion superiority, and directional consistency of the current coding unit, and the geometric features include the absolute average linearity and absolute average curvature of the point cloud area corresponding to the current coding unit.
[0031] A method for fast compression of dynamic point cloud based on the above-mentioned coding unit optimization division method for dynamic point cloud compression, when performing dynamic point cloud compression, the coding unit is divided using the above-mentioned coding unit optimization division method for dynamic point cloud compression.
[0032] The beneficial technical effects of the present invention are:
[0033] 1. By calculating the gradient and variance of the current coding unit, analyzing the directional characteristics of the point cloud coding unit, preliminarily updating the candidate list of partitioning modes, reducing redundant mode searches, and improving the rationality of partitioning decisions;
[0034] 2. Drawing on the experience of optimizing 2D video coding and combining the characteristics of 3D point cloud data, the attribute characteristics and geometric characteristics of the point cloud are calculated, and the neural network model is used for prediction. Based on the prediction results, it is determined whether the test of the current division mode can be skipped, that is, the test of the current division mode is terminated in advance. Such an early termination mechanism can significantly reduce the amount of calculation, while maintaining a high compression ratio and restoration quality, reducing coding complexity, effectively improving coding efficiency and significantly shortening coding time, and improving the running speed of V-PCC;
[0035] 3. It uses a lightweight neural network model with low computational overhead, which can be efficiently integrated into V-PCC's reference coding software, reducing dependence on hardware resources.
[0036] The embodiment of the present application adopts H.266 as the coding complexity optimization method of the two-dimensional encoder inside the V-PCC. As the latest video coding standard, H.266 (VVC) has higher coding efficiency, stronger adaptability to different types of two-dimensional video content, and greater flexibility in encoder architecture selection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of the encoding process of the existing VPCC;
[0038] Figure 2 Schematic diagram of six division modes of the existing H.266 / VVC encoding method;
[0039] Figure 3 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION
[0040] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] Aiming at the deficiencies in the existing technology, this paper deeply analyzes the process and characteristics of V-PCC dynamic point cloud coding, draws on the mature experience in the field of complexity optimization of traditional video coding, comprehensively considers the geometric characteristics of point cloud coding and the characteristics of traditional two-dimensional coding, and provides a coding unit optimization division method for dynamic point cloud compression based on texture features and machine learning. It reduces the time overhead while ensuring the prediction accuracy. Firstly, the directional characteristics of the current coding unit are evaluated by calculating the gradient and variance, and the candidate list of division modes is preliminarily optimized to reduce the computational complexity. Then, for the division mode in the optimized candidate list, the aspect ratio, rate-distortion advantage, segmentation direction consistency, absolute average linearity and absolute average curvature of the current coding unit are extracted as features, and the trained neural network model corresponding to the division mode is input. According to the network output results, it is predicted whether the test of the current division mode (including binary tree horizontal division BTH, binary tree vertical division BTV, ternary tree horizontal division TTH, ternary tree vertical division TTV) can be skipped in advance until the optimal division mode is found, thereby greatly improving the running speed of V-PCC while ensuring the quality of compressed point cloud.
[0042] For details, see the attached Figure 3 , a flow chart of a coding unit optimization division method for dynamic point cloud compression based on texture features and machine learning in an embodiment of the present application. The dynamic point cloud fast compression method based on texture features and machine learning in an embodiment of the present application comprises the following steps:
[0043] Step 1: Train the segmentation recognition model
[0044] In the process of encoding dynamic point clouds using V-PCC, a strategy of gradually refining partitions is adopted to encode point cloud videos. Specifically, the encoding starts from a larger basic coding unit, and according to the distribution of its texture complexity, it is dynamically determined whether it needs to be further divided and how to divide it, and it is divided into smaller coding units step by step. This partition coding method can adapt to the texture characteristics of different regions, and improve the coding efficiency while ensuring the reconstruction quality.
[0045] Therefore, we use the original V-PCC encoder to encode several frames of data, such as 10 frames, on the dynamic point cloud under different quantization parameters. The final selected division mode is the optimal division mode. The attribute features and geometric features of the corresponding coding units are extracted to construct a data set for training and testing the division recognition model corresponding to the optimal division mode.
[0046] Four partition recognition models are constructed corresponding to the binary tree horizontal partition mode BTH, the binary tree vertical partition mode BTV, the ternary tree horizontal partition mode TTH, and the ternary tree vertical partition mode TTV, respectively, namely BTH neural network model, BTV neural network model, TTH neural network model and TTV neural network model. Each neural network has three fully connected layers. Specifically, the number of neurons in the first fully connected layer is 15, the number of neurons in the second fully connected layer is 10, and the number of neurons in the third fully connected layer is 5.
[0047] To encode the dynamic point cloud video sequence, first divide a frame of image into several CTUs (Coding Tree Unit, CTU), and each CTU is treated as an independent coding unit for block partitioning. On the basis of following the original partitioning strategy of the coding standard, different partitioning modes are applied to each CTU for refinement operation. In this process, the attribute features and geometric features of each coding unit are extracted, and the final partitioning mode corresponding to each coding unit is recorded as the optimal partitioning mode. By extracting the attribute features, geometric features and final partitioning mode of each coding unit, a data set containing the features of the coding unit and the corresponding optimal partitioning mode is constructed. These features are used as input data to train the neural network, laying the foundation for the subsequent prediction of the partitioning mode.
[0048] The extracted data set is preprocessed and input into the corresponding neural network model for training to learn the final decision of the division of coding units under different modes. The four neural network models finally trained will be used to predict the early termination of the division mode in the encoding process to screen out the optimal division model.
[0049] Step 2: During the encoding process, obtain the grayscale image of the current coding unit and calculate the gradient features and variance features.
[0050] Extract the grayscale image of the current coding unit and calculate its directional gradient ratio G based on the grayscale image div And the sub-block variance ratio in the horizontal and vertical directions (V hor and V ver ), which is used to quantitatively evaluate the differences in texture complexity in different directions.
[0051] S21, obtain the grayscale image matrix G of the current coding unit, which is defined as:
[0052] G = {g i,j}for i∈[x,x+W],j∈[y,y+H]
[0053] Where W and H represent the width and height of the region, respectively, and g i,j is the grayscale value at position (i, j).
[0054] S22, gradient is used to measure the texture complexity of the current coding unit, extract the horizontal gradient G of the current coding unit hor , vertical gradient G ver , and the gradient ratio G is calculated div , and its calculation formula is defined as follows:
[0055]
[0056] S23, sub-block variance σ of the coding unit 2 Used to measure the texture complexity of the current coding unit, which is defined as:
[0057]
[0058] Among them, represents x i The pixel value in the sub-block, represents the sub-block mean, and N is the number of pixels.
[0059] The horizontal sub-block variance ratio V of the current coding unit is calculated using the following formula: hor And the vertical sub-block variance ratio V ver , which is defined as:
[0060]
[0061] in, represents the variance of the upper half of the sub-blocks of the coding unit, represents the variance of the lower half of the sub-blocks of the coding unit, represents the variance of the left half of the sub-block of the coding unit, Represents the variance of the right half of the sub-blocks in the coding unit.
[0062] Step 3: Based on the gradient and variance characteristics, the directionality of the partitioning pattern is evaluated, and the candidate partitioning patterns are preliminarily optimized. The candidate partitioning patterns: binary tree horizontal partitioning pattern BTH, binary tree vertical partitioning pattern BTV, ternary tree horizontal partitioning pattern TTH, ternary tree vertical partitioning pattern TTV and no partitioning pattern are made into a pattern candidate list, and the following partitioning rules are used for preliminary screening:
[0063] If the vertical gradient is significantly greater than the horizontal gradient, the vertical division mode is disabled;
[0064] If the vertical gradient is not significantly greater than the horizontal gradient, the variance is used for auxiliary judgment;
[0065] If the horizontal gradient is significantly greater than the vertical gradient, the horizontal division mode is disabled;
[0066] If the horizontal gradient is not significantly greater than the vertical gradient, the variance is used for auxiliary judgment;
[0067] If the horizontal and vertical gradients are relatively balanced, all division modes are retained.
[0068] Specifically, the horizontal gradient, vertical gradient, horizontal variance, and vertical variance of the current coding unit are G hor , G ver 、V hor 、V ver , then the gradient ratio Set the threshold value 1 ,th 2 , and th 1 <th 2 , the division rules are set as follows,
[0069] (a) When and When , it means that the vertical gradient is significantly greater than the horizontal gradient, then the vertical partitioning mode (BTV and TTV) is directly disabled to update the mode candidate list;
[0070] (ii) When When , it means that the vertical gradient is large but not significant enough, and the variance is used to assist judgment:
[0071] (1) When V hor >th 1 , indicating that the sub-block variance ratio in the horizontal direction is large, that is, the texture in the horizontal direction changes greatly, then the vertical partitioning mode (BTV and TTV) is disabled to update the mode candidate list;
[0072] (2) If V hor ≤th 1 , indicating that the sub-block variance ratio in the horizontal direction is small, that is, the texture in the horizontal direction is relatively uniform, so all division modes are retained and the subsequent encoding process is entered;
[0073] (III) When G div >th 1 And G div >th 2 When , it means that the horizontal gradient is significantly greater than the vertical gradient, then the horizontal division mode (BTH and TTH) is directly disabled to update the mode candidate list;
[0074] (IV) When 1 <G div ≤th 2 When , it means that the horizontal gradient is large but not significant enough, and the variance is used to assist judgment:
[0075] (5) If V ver>th 1 , indicating that the directional texture changes greatly, the horizontal division mode (BTH and TTH) is disabled and the mode candidate list is updated;
[0076] (6) If V ver ≤th 1 , indicating that the texture in the vertical direction is relatively uniform, so all division modes are retained;
[0077] (V) When , it means that the horizontal and vertical gradients are relatively balanced, so all division modes are retained.
[0078] Note: Keeping all the partition modes means not disabling any partition mode, that is, all the partition modes in the original mode candidate list, such as BTH, TTH, BTV, TTV, QT, non-split mode, etc., can participate in the subsequent encoding process.
[0079] Step 4: traverse the partition modes in the mode candidate list. When the partition mode is BTH, TTH, BTV and TTV, extract the attribute features and geometric features of the current coding unit, input the trained partition recognition model corresponding to the current partition mode for testing, and determine whether to terminate the test of the partition mode in advance or select the next partition mode to continue testing according to the predicted value output by the partition recognition model until the optimal partition mode is found;
[0080] If the current division mode is QT or non-division mode, the encoding continues according to the original V-PCC process.
[0081] S41, feature extraction of current coding unit:
[0082] The attribute features include the aspect ratio, rate-distortion superiority and directional consistency of the current coding unit, and the geometric features include the absolute average linearity and absolute average curvature of the point cloud area corresponding to the current coding unit. These features jointly characterize the multi-dimensional information of the current coding unit in texture, structure and geometry, and can fully reflect its coding complexity and partition characteristics.
[0083] Ⅰ. Absolute average linearity, as follows:
[0084] Step 1: For a point in the point cloud, select its neighborhood point set, which are selected by a fixed radius or a fixed number of neighbors.
[0085] Step 2: For these neighborhood points, use principal component analysis (PCA) to fit the local plane. PCA can obtain three principal component directions in three-dimensional space by eigendecomposing the covariance matrix of the point set, specifically:
[0086] 1. Let point P in the neighborhood j =(xj ,y j , z j ) corresponds to the coordinate P j =[x j ,y j , z j ] T , then the covariance matrix C of the neighborhood point set is:
[0087]
[0088] Among them is The center point (mean point) of the neighborhood point set.
[0089] 2. Perform eigendecomposition on the covariance matrix C to obtain three eigenvalues λ 1 'λ 2 'λ 3 and its corresponding eigenvector v 1 , v 2 , v 3 These three eigenvectors represent the three principal component directions in the point cloud, where λ 1 ≥λ 2 ≥λ 3 are the eigenvalues corresponding to the covariance matrix C in descending order.
[0090] 3. From PCA, we can get the main eigenvector v 1 The corresponding direction is the direction of maximum variance, v 2 is the direction of the second largest variance, v 3 is the direction of minimum variance. For a local point set, the principal component v 1 and the secondary principal component v 2 A fitting plane is determined. Using these eigenvectors and eigenvalues, a local plane can be fitted, defining the plane equation:
[0091] ax+by+cz+d=0
[0092] 4. For each point P i =(x i ,y i , z i ), calculate its distance D to the fitting plane i . Click P i The distance to the plane is given by:
[0093]
[0094] Where a, b, c are the components of the plane normal vector and d is the offset of the plane.
[0095] Step 3: Linearity measures the linearity of the local structure of the point cloud. We can get the linearity by calculating the standard deviation of the point set to the fitting plane. Let the distance from the point set to the plane be D 1 , D 2 , ..., D n , then the standard deviation σ plane The calculation formula is:
[0096]
[0097] in, is the mean distance of all points to the plane:
[0098] Step 4: Calculate the linearity L of the current point, specifically:
[0099]
[0100] where d max Indicates the maximum distance from a point in the point cloud to the fitted plane.
[0101] Step 5: Calculate the absolute average linearity of the current encoding region Specifically:
[0102]
[0103] Where L i represents the linearity of the i-th point, N is the total number of points, |L i | is the absolute value of the linearity at each point.
[0104] II. Absolute mean curvature Specifically:
[0105]
[0106] Among them, λ 1 , 2 , 3 It is the eigenvalue of the eigendecomposition of the covariance matrix C, corresponding to the main change direction of the local point cloud.
[0107] S42, determine the optimal division mode
[0108] After preprocessing, the above geometric features and attribute features are passed as input data to the trained partition recognition model. The neural network further predicts whether the current partition mode is the optimal choice for the coding unit by learning and analyzing the features. If not, it terminates early to provide data support for improving coding efficiency and performance.
[0109] Specifically, according to the current segmentation mode, the corresponding segmentation recognition model (BTH, BTV, TTH and TTV) is selected, and the extracted feature data is input to obtain a probability value, which is compared with the set threshold;
[0110] If the probability value is not less than the threshold, the current partitioning mode is determined to be the optimal partitioning mode, and encoding can continue in this mode according to the original V-PCC process;
[0111] If the probability value is less than the threshold, the test of the current partition mode is terminated and the test of the next partition mode in the mode candidate list is directly entered. In this way, when executing the partition mode test, it is possible to determine whether the current test is terminated without performing encoding or even a complete rate-distortion cost calculation, that is, to give the recognition result of the optimal partition mode, which is convenient and fast and can greatly save encoding time.
[0112] In addition, the present invention also provides a method for fast compression of dynamic point clouds. When performing dynamic point cloud compression, the coding unit is divided using the coding unit optimization division method for dynamic point cloud compression as described above.
[0113] The following simulation experiments are conducted to verify the encoding performance of the dynamic point cloud fast compression method based on texture features and machine learning proposed in this embodiment.
[0114] In order to evaluate the feasibility and effectiveness of the above method, the dynamic 3D point cloud coding reference software TMC2-v18.0 and the VVC reference software VTM-v13.0 were used as the test platform for independent execution. The test sequence includes 5 different test sequences provided by 8i: Soldier, Longdress, Loot, Queen, dancer and Basketball_player. The coding quantization parameter combination (QPs) is set to ([32, 42], [28, 37], [24, 32], [20, 27], [16, 22]), and the coding configuration is All Intra (AI) mode.
[0115] Table 1 Dynamic point cloud complexity optimization and encoding performance
[0116]
[0117] Experimental results show that the proposed method exhibits good comprehensive performance in dynamic point cloud coding. In terms of geometric error, the average changes of D1 and D2 are 0.2% and 0.4% respectively, and the accuracy of geometric data reconstruction remains basically stable. In terms of color error, the average changes of Luma, Cb and Cr are 1.1%, 1.2% and 0.6% respectively, indicating that the reconstruction of color components has high accuracy. At the same time, the encoding time (ΔT) is reduced by 40.10% on average, showing a significant time optimization effect. The overall results show that this method can effectively balance the quality of geometric and color reconstruction while improving the coding efficiency, and has good performance.
[0118] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A coding unit optimization division method for dynamic point cloud compression, characterized in that The following steps are involved: Step 1: Train the segmentation recognition model Extract attribute features and geometric features of multiple coding units and construct data sets for training and testing multiple partition recognition models, which correspond to different partition modes respectively; Step 2: Create a surveillance image dataset The gradient and variance of the current coding unit are calculated to preliminarily screen the division mode of the current coding unit, and then the attribute characteristics and geometric features of the current coding unit are extracted. The trained division recognition model is selected to test the screened division modes one by one, and the optimal division mode is selected for the division of the current coding unit.
2. The coding unit optimization division method for dynamic point cloud compression according to claim 1, characterized in that: The gradient and variance of the current coding unit in the horizontal and vertical directions are calculated respectively, and the division mode is preliminarily screened using the following division rules: If the vertical gradient is significantly greater than the horizontal gradient, the vertical division mode is disabled; If the vertical gradient is not significantly greater than the horizontal gradient, the variance is used for auxiliary judgment; If the horizontal gradient is significantly greater than the vertical gradient, the horizontal division mode is disabled; If the horizontal gradient is not significantly greater than the vertical gradient, the variance is used for auxiliary judgment; If the horizontal and vertical gradients are relatively balanced, all division modes are retained.
3. The coding unit optimization division method for dynamic point cloud compression according to claim 2, characterized in that: Let the horizontal gradient, vertical gradient, horizontal variance, and vertical variance of the current coding unit be denoted as G hor , G ver , V hor , V ver . Then the gradient ratio Set thresholds th1 and th2, where th1 < th2. The partitioning rules are set as follows: (a) When and When , it means that the vertical gradient is significantly greater than the horizontal gradient, then the vertical division mode is directly disabled; (ii) When When , it means that the vertical gradient is large but not significant enough, and the variance is used to assist judgment: (1) When V hor >th1, indicating that the sub-block variance ratio in the horizontal direction is large, that is, the texture in the horizontal direction changes greatly, and the vertical division mode is disabled; (2) If V hor ≤th1, indicating that the sub-block variance ratio in the horizontal direction is small, that is, the texture in the horizontal direction is relatively uniform, so all division modes are retained and the subsequent encoding process is entered; (III) When G div >th1 and G div When >th2, it means that the horizontal gradient is significantly greater than the vertical gradient, and the horizontal division mode is directly disabled; (IV) When th1 <G div When ≤th2, it means that the horizontal gradient is large but not significant enough, and the variance is used to assist in judgment: (1) If V ver >th1, indicating that the directional texture changes greatly, the horizontal division mode is disabled; (2) If V ver ≤th1, indicating that the texture in the vertical direction is relatively uniform, and all division modes are retained; (V) When , it means that the horizontal and vertical gradients are relatively balanced, so all division modes are retained.
4. The coding unit optimization division method for dynamic point cloud compression according to claim 1, characterized in that: Select the trained partition recognition model to test the screened partition patterns one by one. If the output result of the current partition recognition model is less than the threshold, select the next partition recognition model to continue testing until the output result of the current partition recognition model is not less than the threshold. In this case, the partition pattern corresponding to the current partition recognition model is considered to be the optimal partition pattern.
5. The coding unit optimization division method for dynamic point cloud compression according to claim 4, characterized in that: The division modes include binary tree horizontal division mode BTH, binary tree vertical division mode BTV, ternary tree horizontal division mode TTH, ternary tree vertical division mode TTV and no division mode. The division recognition models corresponding to the first four division modes are BTH neural network model, BTV neural network model, TTH neural network model and TTV neural network model respectively.
6. The coding unit optimization division method for dynamic point cloud compression according to claim 1, characterized in that: The original V-PCC encoder is used to encode several frames of data for the dynamic point cloud under different quantization parameters. The final selected partitioning mode is the optimal partitioning mode. The attribute features and geometric features of the corresponding coding units are extracted to construct a data set for training and testing the partitioning recognition model corresponding to the optimal partitioning mode.
7. The coding unit optimization division method for dynamic point cloud compression according to claim 6, characterized in that: The attribute features include the aspect ratio, rate-distortion superiority, and directional consistency of the current coding unit, and the geometric features include the absolute average linearity and absolute average curvature of the point cloud area corresponding to the current coding unit.
8. A method for fast compression of dynamic point cloud based on the coding unit optimization division method for dynamic point cloud compression according to claim 1, characterized in that: When performing dynamic point cloud compression, the coding unit is divided using the coding unit optimization division method for dynamic point cloud compression as described in claims 1-7.
Citation Information
Patent Citations
Video-based fast dynamic point cloud coding method and system
CN112601082A
Encoding methods, devices, electronic devices, and storage media based on cell characteristics
CN114938455A
Point cloud compression method, encoder, decoder and storage medium
CN116132671A
Dynamic point cloud low-complexity coding method and device based on machine learning prediction
CN118368419A
Methods for level partition of point cloud, and decoder
US20230237705A1
Cited By
VVC-based efficient point cloud projection video coding method
CN121644829A