Geometric video coding optimization method for dynamic point cloud

By constructing a geometric video encoding optimization method for dynamic point clouds, using geometric features and MPED score indicators for screening and optimization, the problems of low coding efficiency and poor reconstruction quality in the existing technology are solved, and more efficient coding and better reconstruction effects are achieved.

CN119996701APending Publication Date: 2025-05-13SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256568.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing dynamic point cloud encoding method has low encoding efficiency and poor reconstruction quality in complex scenarios and severe dynamic changes, making it difficult to meet the needs of high-precision point cloud processing.

Method used

A geometric video encoding optimization method is proposed. By extracting the geometric features of dynamic point clouds, screening them with MPED score indexes, building a feature fusion mathematical model, and optimizing rate distortion optimization function, thereby improving coding efficiency and reconstruction quality.

Benefits of technology

It effectively reduces the bit rate overhead, maintains high-quality geometric reconstruction, improves coding efficiency and reconstructs the visual effect of point clouds, especially in application scenarios where fine structure is required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996701A_ABST
    Figure CN119996701A_ABST
Patent Text Reader

Abstract

The invention discloses a geometric video coding optimization method for dynamic point clouds, which is characterized by comprising the following steps of: randomly extracting a plurality of frames from a dynamic point cloud video to be coded in advance, extracting geometric features, and screening the geometric features by combining an MPED score index so as to construct a feature fusion mathematical model; and when a V-PCC encoder is used for encoding, calculating the screened geometric features corresponding to each point in each frame of point cloud, and substituting the screened geometric features into the feature fusion mathematical model to obtain a feature fusion value so as to optimize a rate-distortion optimization function, thereby realizing encoding optimization. Compared with the existing V-PCC encoder, the encoding method provided by the invention can effectively improve the visual effect of the reconstructed point cloud, especially in application scenes requiring a fine structure, such as the fields of three-dimensional reconstruction, virtual reality, automatic driving and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of point cloud coding, and in particular to a geometric video coding optimization method for dynamic point clouds. Background Art

[0002] Dynamic point clouds, as a collection of three-dimensional points that change over time, are widely used in autonomous driving, virtual reality (VR), augmented reality (AR) and other fields. However, how to efficiently encode dynamic point clouds while ensuring high-quality geometric information to reduce storage and transmission costs has always been a research hotspot and technical difficulty that the industry is concerned about. Existing encoding methods often have problems with low encoding efficiency or poor reconstruction quality when faced with complex scenes and drastic dynamic changes.

[0003] The compression rate and reconstruction accuracy of point cloud coding directly affect its usability in various application scenarios. For example, in the point cloud compression process, if the compression rate is too high and the detail information is lost, the reconstructed model may have defects such as geometric structure deformation and texture blur, which will affect the accuracy of downstream tasks. The traditional dynamic point cloud compression framework V-PCC encoder was developed by MPEG. It projects the point cloud onto a 2D plane and encodes it into video frames, using existing video compression technologies (such as HEVC) to efficiently store and transmit point cloud data. Since it fails to fully utilize the geometric features of the point cloud in geometric video coding, it leads to low coding efficiency, which limits its application in high-precision point cloud processing. Summary of the invention

[0004] The present invention proposes a geometric video coding optimization method for dynamic point clouds, which makes full use of the spatial correlation of point clouds to optimize the coding process and improve coding efficiency. Compared with the existing V-PCC coding method, the present invention can effectively reduce the bit rate overhead while maintaining high-quality geometric reconstruction, thereby overcoming the defect of low coding efficiency of the existing method in dynamic scenes.

[0005] The present invention can be implemented by the following technical solutions:

[0006] A geometric video coding optimization method for dynamic point clouds is proposed. Several frames are randomly selected from the dynamic point cloud video to be encoded in advance, geometric features are extracted, and then the geometric features are screened in combination with the MPED score index to construct a feature fusion mathematical model.

[0007] When the V-PCC encoder is used for encoding, the screened geometric features corresponding to each point in each frame of the point cloud are calculated and substituted into the feature fusion mathematical model to obtain the feature fusion value to optimize the rate-distortion optimization function, thereby achieving coding optimization.

[0008] Furthermore, when constructing the feature fusion mathematical model, for each randomly selected frame of point cloud, the eigenvalues ​​of the five geometric features corresponding to each point are first calculated, and then the five MPED scores of the five geometric features corresponding to each frame of point cloud are calculated. Then, the MPED scores are sorted from small to large, and the three geometric features with the highest MPED scores are selected. Finally, the average distribution of the MPED scores corresponding to the three selected geometric features is used as the weight coefficient to construct the feature fusion mathematical model. Its expression is as follows:

[0009] Q=ω1f1+ω2f2+ω3f3

[0010] Among them, f1, f2, and f3 represent the three selected geometric features, and ω1, ω2, and ω3 represent the weight coefficients corresponding to the three selected geometric features f1, f2, and f3.

[0011] Further, the method for screening geometric features comprises the following steps:

[0012] Step 1: The kth geometric feature T of all points in the i-th frame point cloud ik The corresponding eigenvalues ​​are arranged in descending order, and the top 50% of the points are selected to form the target point cloud. The geometric features T of the target point cloud and the original frame point cloud are calculated. ik The corresponding MPED score is MPED ik , k=1,2,3,4,5, i=1,2,…,N, N represents the number of randomly selected point cloud frames;

[0013] Step 2: Repeat step 1 to calculate the five geometric features T of each frame point cloud ik The corresponding MPED scores are ik , a total of five MPED scores;

[0014] Step 3: Repeat steps 1 and 2 to calculate the MPED scores corresponding to all randomly selected frame point clouds;

[0015] Step 4: Calculate the geometric feature T using the following formula ik The corresponding MPED score sum S k , a total of five, the sum of these five MPED scores S k Sort in ascending order and select the top three MPED score sums S k The corresponding geometric features are used to construct the feature fusion mathematical model.

[0016]

[0017] Where M represents the number of geometric features.

[0018] Furthermore, the weight coefficient ω corresponding to the geometric features f1, f2, and f3 in the feature fusion mathematical model is calculated using the following formula: k ,

[0019]

[0020] At this time, M=3 is the number of geometric features screened out.

[0021] Furthermore, the rate-distortion optimization function corresponding to each frame of point cloud is calculated using the following formula.

[0022] J=D+λR

[0023] λ=λ0×coeff

[0024]

[0025] Among them, Q s It represents the fusion feature value of the sth point in a CTU, q represents the number of CTUs in the point cloud of the current frame, p represents the number of points in a CTU, and λ0 represents the original Language multiplier.

[0026] Furthermore, the five geometric characteristics include anisotropy, linearity, sphericity, curvature, and flatness.

[0027] The beneficial technical effects of the present invention are:

[0028] 1. Through the proposed geometric feature fusion model of dynamic point cloud, the spatial geometric feature correlation of dynamic point cloud is fully considered, and a more efficient encoding strategy is implemented on the existing dynamic point cloud encoder V-PCC.

[0029] 2. By calculating the fusion weight coefficients of the point cloud geometric features, the adaptive bit rate allocation is optimized in the HEVC coding framework. More bit rates can be allocated to areas with larger eigenvalues, and less bit rates can be allocated to areas with smaller eigenvalues, thereby optimizing the coding quality and efficiency overall.

[0030] 3. Compared with the existing V-PCC encoder, the encoding method of the present invention can effectively improve the visual effect of reconstructed point cloud, especially in application scenarios that require fine structures, such as three-dimensional reconstruction, virtual reality, and autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is the overall flow chart of the present invention;

[0032] Figure 2 It is a schematic diagram of the overall process of the present invention;

[0033] Figure 3It is a schematic diagram of the columnar distribution of the MPED scores of five geometric features corresponding to the randomly selected partial frame point cloud of the present invention. DETAILED DESCRIPTION

[0034] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] In view of the deficiencies in the prior art, the present invention provides a geometric video coding optimization method based on dynamic point cloud by deeply analyzing the process and characteristics of V-PCC dynamic point cloud coding. Through the feature fusion model, the geometric video of the dynamic point cloud is compressed and optimized. At the same time, HEVC is used for coding adjustment to achieve adaptive bit rate control. Finally, D1, D2, PCQM, and MPED point cloud evaluation indicators are used for objective score evaluation, which fully demonstrates that the method of the present invention can improve the coding quality of dynamic point cloud, reduce the bit rate used for coding, and improve the corresponding bd-rate.

[0036] See also Figure 1-2 As shown, the present invention provides a geometric video coding optimization method based on dynamic point cloud, which randomly extracts several frames from the dynamic point cloud video to be encoded in advance, extracts geometric features, and then combines the MPED score index to perform geometric feature screening to construct a feature fusion mathematical model;

[0037] Then, when the V-PCC encoder is used for encoding, the filtered geometric features corresponding to each point in each frame of the point cloud are calculated and substituted into the feature fusion mathematical model to obtain the feature fusion value to optimize the rate-distortion optimization function, thereby achieving coding optimization.

[0038] The details are as follows:

[0039] Step 1: Construct a feature fusion mathematical model

[0040] S11. Randomly extract a number of frames from the dynamic point cloud video to be encoded, such as 25 frames, i.e., N=25, and for each frame, extract the geometric features of each point on it, including but not limited to the statistical properties of the point cloud, such as anisotropy, linearity, sphericity, curvature, flatness, etc.

[0041] Curvature represents the degree of local surface change of the point cloud and is usually defined as:

[0042]

[0043] In the formula, λ min is the minimum eigenvalue of the covariance matrix, λ1, λ2, λ3 are the three eigenvalues ​​of the covariance matrix, and satisfy λ1≥λ2≥λ3≥0. The higher the curvature value, the more likely the point is to be located at the edge or sharp area.

[0044] Anisotropy refers to the non-uniformity of local point distribution and is usually defined as:

[0045]

[0046] Where λ1 and λ3 are the maximum and minimum eigenvalues ​​of the covariance matrix, respectively. The larger the anisotropy value, the more the distribution of the point cloud tends to be linear or planar.

[0047] Linearity represents the linear characteristics of the point cloud distribution in a local neighborhood and is defined as:

[0048]

[0049] Where λ1 and λ2 are the largest and second largest eigenvalues ​​of the covariance matrix, respectively. When the linearity value is large, the point cloud distribution is closer to a straight line.

[0050] Planarity represents the planar characteristics of the point cloud distribution in a local neighborhood and is defined as:

[0051]

[0052] Where λ2 and λ3 are the second largest and smallest eigenvalues ​​of the covariance matrix, respectively. When the flatness value is larger, the point cloud distribution is closer to a plane.

[0053] Sphericity represents the spherical characteristics of the point cloud distributed in a local neighborhood and is defined as:

[0054]

[0055] Where λ1 and λ3 are the maximum and minimum eigenvalues ​​of the covariance matrix, respectively. When the spherical value is larger, the point cloud distribution is closer to a sphere.

[0056] S12. Use MPED scores to perform feature screening, screen individual features and eliminate weakly correlated features.

[0057] By calculating the MPED score for each of the 25 randomly selected point clouds, we obtain five MPED scores corresponding to the above five features, and then perform feature screening based on the MPED scores to retain features with high correlation.

[0058] For a certain geometric feature such as curvature, firstly, the feature values ​​corresponding to all points in a frame point cloud are sorted in descending order, and the top 50% of the points are selected to form the target point cloud. Based on the target point cloud and the original frame point cloud (including 100% of the points), the corresponding MPED score is calculated. Since each geometric feature uses an independent calculation formula, a frame point cloud can generate five independent MPED scores corresponding to five geometric features, which are grouped as a group, for a total of 25 groups of MPED scores;

[0059] Then, the 25 MPED scores corresponding to the same geometric feature are summed up to obtain the cumulative MPED score of each geometric feature on randomly selected frames. Finally, the three features with the highest cumulative scores are selected for subsequent encoding optimization, while the two groups with higher cumulative MPED scores are discarded, because the lower MPED score indicates that the target point cloud is more correlated with the source point cloud, which reflects better encoding quality.

[0060] The method for screening geometric features is as follows:

[0061] S121, the kth geometric feature T of all points in the i-th frame point cloud ik The corresponding eigenvalues ​​are arranged in descending order, and the top 50% of the points are selected to form the target point cloud. The geometric features T of the target point cloud and the original frame point cloud are calculated. ik The corresponding MPED score is MPED ik , k=1,2,3,4,5,i=1,2,...,N, where N=25;

[0062] The calculation method of each MPED score can be obtained by referring to the following paper:

[0063] MPED_Quantifying_Point_Cloud_Distortion_Based_on_Multiscale_Potential_Energy_Discrepancy.

[0064] Specifically, the source point cloud and the target point cloud are first divided into multiple local neighborhoods. For each neighborhood, a "neighborhood center" (which can be selected by high-frequency points or other strategies) is selected and set as the zero potential energy surface, where the source point cloud is a frame of 25 randomly selected frames of point cloud, and the target point cloud is the feature values ​​corresponding to all points in the frame point cloud, sorted in descending order, and the top 50% of the points are selected.

[0065] For every point x in the neighborhood i , calculate its point potential PPE, defined as:

[0066] E(x i )=m(xi )·g(x i )·h(x i )

[0067] Among them, m(x i ) represents the difference from the point attribute (such as color) (usually calculated by color difference and nonlinearly transformed according to Stevens' power law), g(x i ) is a weighting factor related to the geometric distance, which is used to enhance the sensitivity to location information (it meets the requirement of "the farther the distance, the smaller the weight"). i ) is the point x i The distance to the neighborhood center.

[0068] For a neighborhood center c, the total point potential of all points in its neighborhood is

[0069]

[0070] Among them, N(c;k) represents the set of K nearest neighbor points centered at c.

[0071] At the same scale (i.e., fixed neighborhood size K), the total point potential of the source point cloud and the target point cloud in the corresponding neighborhood are calculated respectively, and then the absolute difference between the two is calculated for each neighborhood, that is, the single-scale potential difference (PED) is formed:

[0072]

[0073] Among them, C is the selected neighborhood center set.

[0074] In order to take into account both local and global information, MPED calculates PED at multiple scales (i.e., using different neighborhood sizes K), and then averages the PED values ​​at each scale to obtain the final MPED score:

[0075]

[0076] In this way, the MPED score integrates the overall "potential energy" differences in geometry and attributes (such as color) of point clouds at different scales, thereby more comprehensively reflecting the distortion of the point cloud.

[0077] S122, repeat S121, calculate the five geometric features T of each frame point cloud ik The corresponding MPED scores are ik , a total of five MPED scores;

[0078] S123, repeat S121 and S122, calculate the MPED scores corresponding to all randomly selected frame point clouds, and the calculation results are as follows: Figure 3 As shown;

[0079] S124. Calculate each geometric feature T using the following formula: ik The corresponding MPED score sum S k , a total of five, the sum of these five MPED scores S k Sort in ascending order and select the top three MPED score sums S k The corresponding geometric features are used to construct the feature fusion mathematical model.

[0080]

[0081] For the MPED score, the smaller the value, the more similar the target point cloud is to the source point cloud, and the higher the quality. Taking into account the visual or perceptual errors in point cloud reconstruction, not just the geometric errors, the two geometric features corresponding to the larger sum of scores, such as anisotropy and planarity, are eliminated according to the above MPED score sum calculation, and only the three geometric features with smaller sum of scores are retained.

[0082] S13. Construction of feature fusion mathematical model: The weight coefficients corresponding to the three retained geometric features are calculated according to the sum of the MPED scores, and they are fused with the three retained features to more comprehensively describe the characteristics of the dynamic point cloud. The feature fusion mathematical model including the three retained features is obtained.

[0083] First, as described above, the sum of the MPED scores for each screened geometric feature is calculated:

[0084]

[0085] In the formula, S k is the sum of the MPED scores of the kth geometric features, N = 25 is the number of randomly selected frames, and M = 3 is the number of retained geometric features. ik is the MPED score of the kth geometric feature of the i-th frame point cloud.

[0086] Then calculate the average MPED score for each geometric feature:

[0087]

[0088] In the formula, is the average MPED score of the k-th geometric feature.

[0089] To assign weights to the feature fusion model, the following rules are followed:

[0090] 1. The smaller the MPED score, the greater the weight;

[0091] 2. The sum of weights is 1;

[0092] The calculation formula of its weight coefficient is as follows:

[0093]

[0094] In the formula, is the importance of the jth geometric feature (inversely proportional to MPED), is a normalization factor that ensures the weights sum to 1.

[0095] The final feature fusion mathematical model expression is:

[0096] Q=ω1f1+ω2f2+ω3f3

[0097] Among them, f1, f2, and f3 represent the three selected geometric features, ω1, ω2, and ω3 represent the weight coefficients corresponding to the three selected geometric features f1, f2, and f3, and Q is the fusion feature value of each point in the point cloud.

[0098] Step 2: Rate-distortion optimization function optimization

[0099] According to the fused eigenvalues, the rate-distortion optimization function coefficients in the encoding process are dynamically adjusted to balance the encoding rate and reconstruction quality and achieve the optimal encoding strategy.

[0100] At present, dynamic point cloud coding requires the use of V-PCC framework to map the three-dimensional point cloud to the two-dimensional plane through the related rule algorithm. For the mapped two-dimensional geometric video, the HEVC coding protocol standard is used to adaptively modify the independent variable coefficient of the rate-distortion optimization function to complete the encoding of the dynamic point cloud and finally reduce the bit rate. The core formula of the rate-distortion optimization function is:

[0101] J=D+λR

[0102] Where J is the rate-distortion cost, D is the distortion, usually expressed as the difference between the original video and the reconstructed video (such as mean square error, MSE), R is the encoding bit rate (Rate), which indicates the number of bits required to encode the block of data, and λ is the Language multiplier, which is used to control the trade-off between distortion and bit rate.

[0103] This step is equivalent to allocating a coefficient after point cloud feature fusion on the basis of the existing Language multiplier, so as to facilitate the subsequent adaptive bit rate allocation. The specific implementation method is as follows:

[0104] For a CTU frame in HEVC, based on the 64*64 division, the arithmetic mean of the fusion eigenvalue Q in each CTU in the point cloud geometry video is calculated, and then the geometric mean of all CTUs in the entire frame is calculated. For areas with larger eigenvalues, the weight factor is less than 1 and more bitrate is allocated. For areas with smaller eigenvalues, the weight factor is greater than 1 and less bitrate is allocated. Therefore, the calculation process of the weight coefficient coeff is as follows:

[0105]

[0106]

[0107] Among them, θ s is the arithmetic mean of the fused eigenvalues ​​in a CTU, Q s It represents the fusion feature value of the sth point in a CTU, q represents the number of CTUs in the point cloud of the current frame, p represents the number of points in a CTU, and λ0 represents the original Language multiplier.

[0108] therefore:

[0109] λ=λ0×coeff

[0110] In the formula, λ0 is the original Language multiplier, and λ is the adjusted Language multiplier.

[0111] The adaptive bit rate allocation of the rate-distortion optimization function can be completed by using the coeff weight coefficient calculated based on the feature fusion model.

[0112] In order to verify the feasibility of the optimization method of the present invention, we independently executed it on the dynamic point cloud coding reference software TMC2-v18.0 and HEVC reference software HM16.20-SCM8.8 as the test platform. The test sequence includes 6 different dynamic point cloud test sequences provided by 8i: Soldier, Longdress, Loot, Queen, dancer and Basketball_player. The bit rate of the geometric video is saved by an average of 4.47%. At the five bit rate levels of r1, r2, r3, r4 and r5, the bd-rate indicators of D1, D2, MPED and PCQM point cloud quality evaluation are improved by an average of 0.53%, 0.41%, 0.17% and 1.03% respectively, as shown in Table 1.

[0113] Table 1 Dynamic point cloud geometry video optimization and encoding performance

[0114]

[0115] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A geometric video coding optimization method for dynamic point clouds, characterized in that: In advance, several frames are randomly selected from the dynamic point cloud video to be encoded, and geometric features are extracted. Then, the geometric features are screened in combination with the MPED score index to construct a feature fusion mathematical model. When the V-PCC encoder is used for encoding, the screened geometric features corresponding to each point in each frame of the point cloud are calculated and substituted into the feature fusion mathematical model to obtain the feature fusion value to optimize the rate-distortion optimization function, thereby achieving coding optimization.

2. The geometric video coding optimization method for dynamic point cloud according to claim 1, characterized in that: When constructing the feature fusion mathematical model, for each randomly selected frame of point cloud, first calculate the eigenvalues ​​of the five geometric features corresponding to each point, then calculate the five MPED scores of the five geometric features corresponding to each frame of point cloud, and then sort the MPED scores from small to large, and select the three geometric features with the highest MPED scores. Finally, the average distribution of the MPED scores corresponding to the three selected geometric features is used as the weight coefficient to construct the feature fusion mathematical model. Its expression is as follows Q=ω1f1+ω2f2+ω3f3 Among them, f1, f2, and f3 represent the three selected geometric features, and ω1, ω2, and ω3 represent the weight coefficients corresponding to the three selected geometric features f1, f2, and f3.

3. The geometric video coding optimization method for dynamic point cloud according to claim 2, characterized in that The method for screening geometric features includes the following steps: Step 1: The kth geometric feature T of all points in the i-th frame point cloud ik The corresponding eigenvalues ​​are arranged in descending order, and the top 50% of the points are selected to form the target point cloud. The geometric features T of the target point cloud and the original frame point cloud are calculated. ik The corresponding MPED score is MPED ik , k = 1, 2, 3, 4, 5, i = 1, 2, ..., N, N represents the number of randomly selected point cloud frames; Step 2: Repeat step 1 to calculate the five geometric features T of each frame point cloud ik The corresponding MPED scores are ik , a total of five MPED scores; Step 3: Repeat steps 1 and 2 to calculate the MPED scores corresponding to all randomly selected frame point clouds; Step 4: Calculate the geometric feature T using the following formula ik The corresponding MPED score sum S k , a total of five, the sum of these five MPED scores S k Sort in ascending order and select the top three MPED score sums S k The corresponding geometric features are used to construct the feature fusion mathematical model. Where M represents the number of geometric features.

4. The geometric video coding optimization method for dynamic point cloud according to claim 3, characterized in that: Use the following formula to calculate the weight coefficient ω corresponding to the geometric features f1, f2, and f3 in the feature fusion mathematical model: k , At this time, M=3 is the number of geometric features screened out.

5. The geometric video coding optimization method for dynamic point cloud according to claim 4, characterized in that: Use the following formula to calculate the rate-distortion optimization function corresponding to each frame of point cloud. J=D+λR λ=λ0×coeff Among them, Q s It represents the fusion feature value of the sth point in a CTU, q represents the number of CTUs in the point cloud of the current frame, p represents the number of points in a CTU, and λ0 represents the original Language multiplier.

6. The geometric video coding optimization method for dynamic point cloud according to claim 1, characterized in that: The five geometric characteristics include anisotropy, linearity, sphericity, curvature, and flatness.