Fast Determination Method and System for Inter-Frame Coding Mode of Dynamic 3D Point Cloud Compression

By judging the preset conditions of geometric and texture map frames in dynamic 3D point cloud compression, and skipping unnecessary inter prediction mode calculations, the problem of high computing complexity is solved, and encoding efficiency and real-time transmission is improved.

CN115499660BActive Publication Date: 2025-07-18NANHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211185785.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-07-18
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

During the existing dynamic 3D point cloud compression process, the computational complexity of inter-frame mode selection leads to limited real-time transmission and user experience, and the existing video compression methods cannot be effectively applied to dynamic 3D point clouds.

Method used

By judging whether the geometric map and texture map frames of the dynamic 3D point cloud encoding unit meet preset conditions, unnecessary inter prediction mode calculation is skipped, and first, second, third and fourth prediction distortion transformations are used to determine whether the conditions are met to skip the remaining inter-frame mode.

Benefits of technology

It significantly reduces the computational complexity during dynamic 3D point cloud compression, while ensuring encoding quality, improving encoding efficiency and real-time transmission capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115499660B_ABST
    Figure CN115499660B_ABST
Patent Text Reader

Abstract

Fast determination method and system for inter-frame coding mode of dynamic 3D point cloud compression, which relates to the field of point cloud compression technology. In the present invention, it is determined whether the geometric map frame / texture map frame satisfies the first and second preset conditions to predict whether the best mode is the SKIP / Merge mode, and it is determined whether the geometric map frame / texture map frame satisfies the third and fourth preset conditions to predict whether the best mode is the Inter_2N×2N mode. As long as the preset conditions can be satisfied, the calculation of the remaining inter-frame prediction modes can be avoided, so as to achieve the purpose of saving time. Compared with the existing methods, the present invention can not only ensure the coding quality of the target coding, but also significantly reduce the computational time complexity in the process of making coding mode decisions for the target coding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of point cloud compression, and particularly to a method and system for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression. Background Art

[0002] With the latest development of visual capture technology, the capture and digitization technology of three-dimensional world scenes has gradually become clear. Among them, point cloud is one of the main 3D data representation forms. In addition to providing spatial coordinates, the point cloud can also provide attributes related to the points in the 3D world (such as color or reflectivity). The point cloud in the original format requires a large amount of storage memory or transmission bandwidth. For example, a classic dynamic point cloud for entertainment usually contains about 1 million points per frame. Without compression, the total bandwidth for 30 frames per second is 3.6 Gbps. In addition, the emergence of high-resolution point cloud capture technology, on the contrary, puts forward higher requirements for the size of the point cloud. In order to promote the point cloud technology and apply it to actual production, compression technology needs to be developed urgently.

[0003] In order to effectively compress dynamic 3D point clouds, MPEG (Moving Picture Experts Group) is formulating a method for video compression (High Efficiency Video Coding, HEVC) for dynamic 3D point cloud compression, that is, video-based point cloud compression V-PCC (Video-based Point Cloud Compression). Since in the HEVC compression coding process, when selecting an inter-frame prediction mode for each coding unit (CU), the rate distortion cost (RDcost) calculations of modes such as SKIP / Merge mode, Inter_2N×2N mode, Inter_N×N mode, Inter_2N×N mode, Inter_N×2N mode, Inter2N×nU mode, Inter_2N×nD mode, Inter_nL×2N mode, Inter_nR×2N mode, etc. will be carried out in sequence, and finally the optimal prediction mode is selected by comparing the RDcost. However, selecting the optimal mode by calculating the RDcost also brings huge computational complexity to the compression coding, seriously hindering the real-time transmission of dynamic 3D point clouds. In addition, due to the very large differences between the geometric and color videos generated by mapping dynamic 3D point clouds and natural videos, the methods for quickly determining the HEVC inter-frame coding mode proposed for natural videos cannot be well used in the optimization of video-based dynamic 3D point cloud compression coding. Therefore, there is an urgent need to develop a method for quickly determining the inter-frame coding mode in dynamic 3D point cloud compression to solve the problem of network real-time transmission and improve the user experience.

[0004] It can be seen that how to reduce the computational complexity of inter-frame mode selection in the process of video-based dynamic 3D point cloud compression is a technical problem that those skilled in the art need to solve. Summary of the Invention

[0005] One of the purposes of the present invention is to provide a method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression to reduce the computational complexity in mode selection.

[0006] To solve the above technical problems, the present invention adopts the following technical solutions: A method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression, which includes making the following judgments on the target coding CU after calculating the rate-distortion cost in the Merge mode:

[0007] If the target coding CU is a geometric graph CU, perform a first prediction distortion transformation on the target coding CU, and determine whether the first preset condition is satisfied. If so, skip the calculation of subsequent inter-frame prediction modes; when the target coding CU does not satisfy the first preset condition, after the calculation of the inter-frame Inter 2N×2N mode is completed, perform a third prediction distortion transformation on the target coding CU, and determine whether the third preset condition is satisfied. If so, skip the calculation of subsequent inter-frame modes.

[0008] If the target coding CU is a texture graph CU, perform a second prediction distortion transformation on the target coding CU, and determine whether the second preset condition is satisfied. If so, skip the calculation of subsequent inter-frame modes; when the target coding CU does not satisfy the second preset condition, after the calculation of the inter-frame Inter 2N×2N mode is completed, perform a fourth prediction distortion transformation on the target coding CU, and determine whether the fourth preset condition is satisfied. If so, skip the calculation of subsequent inter-frame modes.

[0009] In the above method, specifically, in the intra-frame compression coding configuration of dynamic 3D point cloud compression, after calculating the rate-distortion cost in the Merge mode, determine whether the current dynamic 3D point cloud compression coding frame does not belong to the placeholder frame and is not an I frame;

[0010] If so, sequentially extract the CUs to be compressed and coded w×h to obtain the target coding CU w×h and judge the geometric graph and the texture graph by the chrominance component;

[0011] Obtain the target coding CU w×h and the prediction distortion Err w×h (i, j) of the luminance component;

[0012] If the current coding CU w×h is a geometric graph frame, perform a first prediction distortion transformation on the prediction distortion Err w×h (i, j) to obtain Tra1 w×h(i, j);

[0013] Among them,

[0014] Among them,

[0015] In the formula, w and h are the width and height of the target coded CU w×h respectively, and Ori w×h (i, j) is the original luminance value of the target coded CU w×h respectively, and Pre w×h (i, j) is the predicted luminance value of the target coded CU w×h obtained after frame - by - frame prediction mode calculation. i and j are the abscissa and ordinate of the pixel point of the target coded CU w×h respectively, and QP is the quantization parameter of the current compression - coded frame;

[0016] Bisect the current CU w×h region horizontally and vertically. Horizontally, two sub - matrices are obtained, and vertically, two sub - matrices are obtained. Calculate the maximum Tra1 w×h (i, j) variance value among the four sub - blocks to obtain maxVar B ;

[0017] Bisect the current CU w×h region horizontally and vertically into four equal parts. Horizontally, four sub - matrices are obtained, and vertically, four sub - matrices are obtained. Calculate the maximum Tra1 w×h (i, j) variance value among the eight sub - blocks to obtain maxVar Q ; Then determine whether the target coded CU w×h meets the first preset condition;

[0018] If it does not meet the condition, after the Inter 2N×2N prediction mode, obtain the prediction distortion Err w×h (i, j) of the target coded CU w×h and perform the third prediction distortion transformation on it to obtain Tra3 w×h (i, j);

[0019] Among them,

[0020] Among them,

[0021] In the formula, w and h are the width and height of the target coded CU w×h respectively, and Ori w×h (i, j) is the original luminance value of the target coded CU, and Pre w×h (i, j) is the target coded CU w×hThe predicted luminance value obtained after frame - based prediction mode calculation, where i and j are the horizontal and vertical coordinates of the pixel points of the target encoded CU w×h , QP is the quantization parameter of the current compressed encoded frame;

[0022] Bisect the current CU w×h region horizontally and vertically respectively. Horizontally, two sub - matrices are obtained, and vertically, two sub - matrices are obtained. Calculate the maximum Tra3 w×h (i, j) variance value for the corresponding positions in the four sub - blocks to obtain maxVar B ;

[0023] Bisect the current CU w×h region horizontally and vertically respectively. Horizontally, four sub - matrices are obtained, and vertically, four sub - matrices are obtained. Calculate the maximum Tra3 w×h (i, j) variance value for the corresponding positions in the eight sub - blocks to obtain maxVar Q ; Determine whether the target encoded CU w×h satisfies the third preset condition;

[0024] If so, skip the calculation of the remaining frame - based prediction modes;

[0025] If the current encoded CU w×h is a texture map frame, perform the second prediction distortion transformation on it to obtain Tra2 w×h (i, j);

[0026] Among them,

[0027] Among them,

[0028] In the formula, w and h are the width and height of the target encoded CU w×h , Ori w×h (i, j) is the original luminance value of the target encoded CU w×h , Pre w×h (i, j) is the predicted luminance value of the target encoded CU w×h obtained after frame - based prediction mode calculation, where i and j are the horizontal and vertical coordinates of the pixel points of the target encoded CU w×h , QP is the quantization parameter of the current compressed encoded frame;

[0029] Bisect the current CU w×h region horizontally and vertically respectively. Horizontally, two sub - matrices are obtained, and vertically, two sub - matrices are obtained. Calculate the maximum Tra2 w×h (i, j) variance value for the corresponding positions in the four sub - blocks to obtain maxVar B ;

[0030] For the current CU w×h The area is divided into four equal parts horizontally and vertically. Four sub-matrices are obtained horizontally and four sub-matrices are obtained vertically. The corresponding maximum Tra2 w×h (i, j) variance value is obtained to get maxVar Q ; Determine whether the target encoded CU w×h satisfies the second preset condition;

[0031] If not, after the Inter 2N×2N prediction mode, perform the third prediction distortion transformation on the prediction distortion Err w×h (i, j) to obtain Tra4 w×h (i, j);

[0032] Among them,

[0033] Among them,

[0034] In the formula, w and h are the width and height of the target encoded CU w×h respectively, Ori w×h (i, j) is the original luminance value of the target encoded CU, Pre w×h (i, j) is the predicted luminance value of the target encoded CU w×h obtained after frame-inter prediction mode calculation, and i and j are the horizontal and vertical coordinates of the pixel points of the target encoded CU w×h respectively, and QP is the quantization parameter of the current compressed encoded frame;

[0035] For the current CU w×h The area is divided into two equal parts horizontally and vertically. Two sub-matrices are obtained horizontally and two sub-matrices are obtained vertically. The corresponding maximum Tra4 w×h (i, j) variance value is obtained to get maxVar B ;

[0036] For the current CU w×h The area is divided into four equal parts horizontally and vertically. Four sub-matrices are obtained horizontally and four sub-matrices are obtained vertically. The corresponding maximum Tra4 w×h (i, j) variance value is obtained to get maxVar Q ; Determine whether the target encoded CU w×h satisfies the fourth preset condition;

[0037] If so, skip the RDcost calculation of the remaining frame-inter modes;

[0038] Among them, the expression of the first preset condition is:

[0039] POC > 4 && POC != 8 && POC != 16 and

[0040] and

[0041] The expression of the second preset condition is:

[0042] and

[0043] The expression of the third preset condition is:

[0044] POC > 4 && POC != 8 && POC != 16 and

[0045] and

[0046] The expression of the fourth preset condition is:

[0047] and

[0048] In the formula is the adaptive adjustment threshold, and

[0049]

[0050]

[0051] In the formula are the variances calculated respectively after dividing Tra w×h (i, j) horizontally into two sub - matrices; are the variances calculated respectively after dividing Tra w×h (i, j) vertically into two sub - matrices; are the variances calculated respectively after dividing Tra w×h (i, j) horizontally into four sub - matrices; are the variances calculated respectively after dividing Tra w×h (i, j) vertically into four sub - matrices, where Tra w×h (i, j) is Tra1 w×h (i, j) or Tra2 w×h (i, j) or Tra3 w×h (i, j) or Tra4 w×h (i, j).

[0052] Preferably, after the process of determining whether the current dynamically 3D point cloud compression - encoded frame is not an I - frame, it further includes: if not, then no fast determination of the inter - frame prediction mode is performed on the encoded frame.

[0053] More preferably, after the process of determining whether the target encoded CU meets the first preset condition, it further includes: if so, the target encoding skips the calculation of the remaining inter-frame prediction modes.

[0054] More preferably, after determining whether the target encoded CU meets the second preset condition, it further includes: if so, the target encoding skips the calculation of the remaining inter-frame prediction modes.

[0055] More preferably, after the process of determining whether the target encoded CU meets the third preset condition, it further includes: if not, the target encoding continues to perform the calculation of the remaining inter-frame prediction modes.

[0056] More preferably, after the process of determining whether the target encoded CU meets the fourth preset condition, it further includes: if not, the target encoding continues to perform the calculation of the remaining inter-frame prediction modes.

[0057] In addition, the present invention also provides a system for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression, which includes:

[0058] An encoded CU judgment module: used for making a judgment on whether the target encoded CU is a geometric graph CU or a texture graph CU after calculating the rate-distortion cost of the Merge mode of the target encoded CU.

[0059] A first coding mode determination module: used for processing the target encoded CU in the form of a geometric graph CU, including performing a first prediction distortion transformation on the target encoded CU and determining whether it meets the first preset condition. If so, the calculation of the subsequent inter-frame prediction modes is skipped; when the target encoded CU does not meet the first preset condition, after the calculation of the inter-frame Inter 2N×2N mode is completed, a third prediction distortion transformation is performed on the target encoded CU and it is determined whether it meets the third preset condition. If so, the calculation of the subsequent inter-frame modes is skipped.

[0060] A second coding mode determination module: used for processing the target encoded CU in the form of a texture graph CU, including performing a second prediction distortion transformation on the target encoded CU and determining whether it meets the second preset condition. If so, the calculation of the subsequent inter-frame modes is skipped; when the target encoded CU does not meet the second preset condition, after the calculation of the inter-frame Inter2N×2N mode is completed, a fourth prediction distortion transformation is performed on the target encoded CU and it is determined whether it meets the fourth preset condition. If so, the calculation of the subsequent inter-frame modes is skipped.

[0061] This system can operate through the above-mentioned method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression.

[0062] The present invention predicts the best mode by determining whether the geometric frame / texture frame meets the first and second preset conditions; similarly, the best mode can be predicted by determining whether the geometric frame / texture frame meets the third and fourth preset conditions. As long as the preset conditions can be met, the calculation of the remaining inter-frame prediction modes can be avoided, thereby achieving the purpose of saving time. Therefore, through the method provided by the present invention, not only can the coding quality of the target coding be ensured, but also the computational time complexity in the process of compression coding of the target coding can be significantly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic flow chart of a method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression according to an embodiment of the present invention;

[0064] Figure 2 It is an example of binary partitioning and describes the variance value in the formula of the present invention

[0065] Figure 3 It is an example of quaternary partitioning and describes the variance value in the formula of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] For the convenience of those skilled in the art, the present invention will be further described below in conjunction with the embodiments and the drawings. The content mentioned in the embodiments does not limit the present invention.

[0067] As Figure 1 shown, a method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression includes the following steps:

[0068] Step S1: Determine whether the current dynamic 3D point cloud compression coding frame is an I frame. If so, do not skip any inter-frame mode, otherwise execute S2;

[0069] Step S2: Sequentially extract the CUs to be compression-coded, and obtain the chrominance components of the target coding CU to judge the geometric map and the texture map;

[0070] Step S3: Determine whether the current frame is a geometric frame. If so, execute step S4, otherwise execute step S7;

[0071] Step S4: Calculate the residual Err w×h (i, j) of the original image and the predicted image and perform the first prediction distortion transformation to obtain Tra1 w×h (i, j);

[0072] Step S5: Perform on Tra1 w×h(i, j) enters the corresponding CU region for horizontal and vertical bisection and quartering, and calculates the maximum variances of the sub-blocks corresponding to the bisection and quartering respectively to obtain maxVar B and maxVar Q , determines whether maxVar B and maxVar Q meet the first preset condition. If so, execute step S6; otherwise, execute S9;

[0073] Step S6: Skip the remaining inter-frame prediction modes;

[0074] Step S7: Calculate the residual Err w×h (i, j) of the original image and the predicted image and perform a second prediction distortion transformation to obtain Tra2 w×h (i, j);

[0075] Step S8: Bisect and quarter the CU region corresponding to Tra2 w×h (i, j) horizontally and vertically, and calculate the maximum variances of the sub-blocks corresponding to the bisection and quartering respectively to obtain maxVar B and maxVar Q ; determine whether maxVar B and maxVar Q meet the second preset condition. If so, execute step S6; otherwise, execute step S12;

[0076] Step S9: Perform a third prediction distortion transformation on the geometric graph to obtain Tra3 w×h (i, j);

[0077] Step S10: Bisect and quarter the CU region corresponding to Tra3 w×h (i, j) horizontally and vertically, and calculate the maximum variances of the sub-blocks corresponding to the bisection and quartering respectively to obtain maxVar B and maxVar Q ; determine whether maxVar B and maxVar Q meet the third preset condition. If so, execute step S11; otherwise, continue with the remaining inter-frame prediction modes;

[0078] Step S11: Skip the remaining inter-frame prediction modes;

[0079] Step S12: Perform a fourth prediction distortion transformation on the geometric graph to obtain Tra4 w×h (i, j);

[0080] Step S13: For Tra4 w×hBisect and quarter the CU region corresponding to (i, j) horizontally and vertically, and calculate the maximum variance of the corresponding sub-blocks of the bisection and quartering to obtain maxVar B and maxVar Q ; Determine whether maxVar B and maxVar Q meet the fourth preset condition. If so, execute step S11; otherwise, continue with the remaining inter-frame prediction modes;

[0081] Among them,

[0082] Among them,

[0083] Among them,

[0084] Among them,

[0085] Among them,

[0086] In the formula, w and h are the width and height of the target encoded CU w×h , Ori w×h (i, j) is the original luminance value of the target encoded CU, Pre w×h (i, j) is the predicted luminance value of the target encoded CU w×h obtained after frame-inter prediction mode calculation, i and j are the abscissa and ordinate of the pixel points of the target encoded CU w×h respectively, QP is the quantization parameter of the current compression-encoded frame. Hereinafter, Tra3 w×h (i, j) and Tra4 w×h (i, j) are collectively referred to as Tra w×h (i, j);

[0087] The expression of the first preset condition is:

[0088] POC > 4 && POC!= 8 && POC!= 16 and

[0089] and

[0090] The expression of the second preset condition is:

[0091] and

[0092] The expression of the third preset condition is:

[0093] POC > 4 && POC!= 8 && POC!= 16 and

[0094] And

[0095] The fourth preset condition expression is:

[0096] And

[0097] In the formula is the adaptive adjustment threshold, and

[0098]

[0099]

[0100] In the formula are the variances respectively calculated after horizontally dividing Tra w×h (i, j) into two sub-matrices; are the variances respectively calculated after vertically dividing Tra w×h (i, j) into two sub-matrices; are the variances respectively calculated after horizontally dividing Tra w×h (i, j) into four sub-matrices; are the variances respectively calculated after vertically dividing Tra w×h (i, j) into four sub-matrices.

[0101] Using steps S4 - S8, it is possible to predict whether the best mode is the SKIP / Merge mode by determining whether the geometric picture frame / texture picture frame meets the first and second preset conditions. Similarly, steps S9 - S13 can predict whether the best mode is the Inter_2N×2N mode by determining whether the geometric picture frame / texture picture frame meets the third and fourth preset conditions. As long as the preset conditions can be met, the calculation of the remaining inter-frame prediction modes can be avoided, thus achieving the purpose of saving time. Therefore, through the method provided in this embodiment, not only can the encoding quality of the target encoded CU w×h be ensured, but also the computational time complexity in the encoding mode decision process for the target CU w×h can be significantly reduced.

[0102] It can be seen that in this embodiment, because the transform value Tra1 w×h of the prediction distortion of the current encoded CU w×h (i, j) can reflect the best inter-frame prediction mode that the current encoded CU w×h should select. During the inter-frame prediction process of the target encoded CU w×h , four judgment criteria are set to determine the target encoded CU w×hWhether it is possible to skip the remaining inter-frame prediction modes, that is, by judging the target coding CU w×h whether the prediction distortion transformation value satisfies the first preset condition, the second preset condition, the third preset condition, and the fourth preset condition to judge whether the target coding CU w×h can skip the calculation of the remaining inter-frame prediction modes, and, during the partitioning process of the target coding CU w×h an equilibrium between coding time savings and coding rate increase can be achieved by adaptively adjusting the threshold Obviously, since this method can terminate the steps of performing multiple inter-frame prediction modes on the target coding CU w×h in advance, therefore, the method for quickly determining the inter-frame coding mode provided by this embodiment can significantly reduce the time complexity during the inter-frame mode selection process of the target coding CU w×h

[0103] On the above basis, this embodiment further explains and optimizes the technical solution. As a preferred implementation manner, in the above step S1: whether the current dynamic 3D point cloud compression coding frame is an I frame. If the current dynamic 3D point cloud compression coding frame is an I frame, the method for quickly determining the mode is not performed on the current coding frame.

[0104] Those skilled in the art should know that in the dynamic 3D point cloud compression inter-frame compression coding configuration, it includes both placeholder frames and geometric frames and texture frames, and includes both I frames, P frames, and B frames. The proposed method only performs a method for quickly determining the mode for B frames of geometric and texture maps.

[0105] Through the method provided by this embodiment, the integrity during the mode determination process of the target coding CU w×h can be further ensured.

[0106] ​Based on the technical content disclosed in the above embodiments, using the dynamic 3D point cloud compression reference software TCM2-V7.0 as the test platform, and its corresponding HEVC reference software is HM-16.20+SCM-8.8. Execute the fast determination method disclosed in this embodiment on a PC with an Inter(R) Core(TM) i7_9700 CPU and 16GB RAM, and use this to evaluate the feasibility and effectiveness of this method. The dynamic 3D point cloud sequences for general testing include "loot", "redandblack", "soldier", "queen", "longdress", and each dynamic 3D point cloud includes 32 frames. The encoding configuration is Random Access (RA), and five combinations of encoding quantization parameters (quality parameters, QP) are adopted, which are (32, 42), (28, 37), (24, 32), (20, 27), (16, 22), and are used for the compression of geometric video and texture video in the dynamic 3D point cloud. For the evaluation of the geometric video compression performance of the dynamic 3D point cloud, the rate change situation (BD-rate) of point-to-point PSNR (D1) and point-to-plane PSNR (D2) is adopted. For the evaluation of the color video compression performance of the dynamic 3D point cloud, the BD-rate of Luma, Cb, and Cr is adopted. The encoding time saving is defined as:

[0107]

[0108] In the formula, i represents 5 different QP values, and T O is the compression encoding time of the original test model, and T P is the compression encoding time after applying the present invention to the original test model, and TS GP and TS TP respectively represent the encoding time saving of the P-frame of the geometric video of the dynamic 3D point cloud and the encoding time saving of the P-frame of the texture video.

[0109] Please refer to Table 1. Table 1 shows the performance comparison results of the method provided in this embodiment on the TCM2-V7.0 test platform, where

[0110] Table 1. Performance comparison results of the method provided in this embodiment on the TCM2-V7.0 test platform (unit: %)

[0111]

[0112] Please refer to Table 2. Table 2 shows the performance comparison results of the method provided in this embodiment on the TCM2-V7.0 test platform, where

[0113] Table 2. Performance comparison results of the method provided in this embodiment on the TCM2-V7.0 test platform (unit: %)

[0114]

[0115] Please refer to Table 3. Table 3 shows the performance comparison results of the method provided in this embodiment on the TCM2-V7.0 test platform, where

[0116] Table 3. Performance comparison results of the method provided in this embodiment on the TCM2-V7.0 test platform (unit: %)

[0117]

[0118] Please refer to Table 4. Table 4 shows the performance comparison results of the method provided in this embodiment on the TCM2-V7.0 test platform, where

[0119] Table 4. Performance comparison results of the method provided in this embodiment on the TCM2-V7.0 test platform (unit: %)

[0120]

[0121] From the comparison results of the BD-rate and encoding time savings of the geometric video and texture video of the dynamic 3D point cloud in Table 1, Table 2, Table 3, and Table 4, it can be seen that when the adaptive adjustment threshold is set to different values, the method provided in this embodiment can respectively save the encoding time of the geometric video and color video on average by (39.91%, 47.48%), (45.47%, 50.95%), (51.41%, 58.54%), and (55.51%, 64.13%). The BD-rate of the geometric video (D1 and D2) and the BD-rate of the texture video (Luma) increase on average by (-0.18%, -0.25%, 0.09%), (-0.30%, -0.39%, 0.20%), (-0.13%, -0.24%, 0.58%), and (0.21%, 0.05%, 1.45%). As the adaptive adjustment threshold increases, the compression encoding time savings of the geometric video and color video of the dynamic 3D point cloud are more, and the BD-rate of D1, D2, and Luma increases.

[0122] In summary, without substantially reducing the geometric and texture video compression quality of dynamic 3D point cloud compression, the fast determination method provided by the present invention can greatly reduce the computational complexity of the inter-frame prediction mode. At the same time, the method of the present invention is simply designed and does not refer to the spatio-temporal domain information of the CU, and can be integrated into the parallel compression coding framework to further reduce the dynamic 3D point cloud compression coding time.

[0123] The above has introduced in detail the fast determination method for the inter-frame coding mode of dynamic 3D point cloud compression provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression, characterized in that: After calculating the rate - distortion cost of the target - coded CU in Merge mode, the following judgments are made on the target - coded CU: If the target - coded CU is a geometric - map CU, perform the first predictive distortion transformation on the target - coded CU, and determine whether the first preset condition is met. If so, skip the calculation of subsequent inter - frame prediction modes; when the target - coded CU does not meet the first preset condition, after the calculation of the inter - frame Inter 2N×2N mode is completed, perform the third predictive distortion transformation on the target - coded CU, and determine whether the third preset condition is met. If so, skip the calculation of subsequent inter - frame modes; If the target - coded CU is a texture - map CU, perform the second predictive distortion transformation on the target - coded CU, and determine whether the second preset condition is met. If so, skip the calculation of subsequent inter - frame modes; when the target - coded CU does not meet the second preset condition, after the calculation of the inter - frame Inter 2N×2N mode is completed, perform the fourth predictive distortion transformation on the target - coded CU, and determine whether the fourth preset condition is met. If so, skip the calculation of subsequent inter - frame modes; In the intra - compression coding configuration of dynamic 3D point - cloud compression, after calculating the rate - distortion cost in Merge mode, determine whether the current dynamic 3D point - cloud compression - coded frame does not belong to the placeholder - map frame and does not belong to the I - frame; If both are, then sequentially extract the CUs that need to be compression-coded w×h , and obtain the target coded CU w×h to judge the geometry map and the texture map of the chrominance component Obtain the target coded CU w×h Prediction distortion Err of the luminance component w×h (i,j); If the current coded CU w×h is a geometric picture frame, then perform a first predicted distortion transformation on the predicted distortion Err w×h (i, j) to obtain Tra1 w×h (i, j); Among them, Among them, Wherein, w and h are the width and height of the target coding CU w×h respectively, Ori w×h (i, j) is the original luminance value of the target coding CU w×h respectively, Pre w×h (i, j) is the predicted luminance value of the target coding CU w×h obtained after frame - inter prediction mode calculation, and i and j are the abscissa and ordinate of the pixel point of the target coding CU w×h respectively, and QP is the quantization parameter of the current compression - coded frame; For the current CU w×h Bisect the area horizontally and vertically respectively. Horizontally, two sub-matrices are obtained, and vertically, two sub-matrices are obtained. Calculate the maximum Tra1 corresponding to the four sub-blocks w×h (i,j) variance value to obtain maxVar B ; For the current CU w×h The area is equally divided into four parts horizontally and vertically. Four sub-matrices are obtained horizontally and four sub-matrices are obtained vertically. The maximum Tra1 w×h variance value of (i, j) is obtained to get maxVar Q ; Then determine whether the target coded CU w×h meets the first preset condition; If not satisfied, after the Inter 2N×2N prediction mode, obtain the target coding CU w×h of the prediction distortion Err w×h (i,j), perform the third prediction distortion transformation on it to obtain Tra3 w×h (i,j); Among them, Among them, Wherein, w and h are the width and height of the target coded CU, respectively w×h Ori w×h (i, j) is the original luminance value of the target coded CU, and Pre w×h (i, j) is the target coded CU w×h The predicted luminance value obtained after frame - by - frame prediction mode calculation, and i and j are the abscissa and ordinate of the pixel point of the target coded CU, respectively w×h The quantization parameter of the current compressed coded frame; For the current CU w×h Bisect the region horizontally and vertically. Horizontally, two sub - matrices are obtained, and vertically, two sub - matrices are obtained. Calculate the maximum Tra3 w×h (i,j) variance value among the four sub - blocks to obtain maxVar B ; For the current CU w×h Divide the area horizontally and vertically into four equal parts respectively. Four sub-matrices are obtained horizontally and four sub-matrices are obtained vertically. Calculate the maximum Tra3 w×h (i, j) variance value among the eight sub-blocks to obtain maxVar Q ; Determine whether the target coded CU w×h satisfies the third preset condition; If so, skip the calculation of the remaining inter - frame prediction modes; If the current coded CU w×h is a texture picture frame, perform a second predictive distortion transform on it to obtain Tra2 w×h (i,j); Among them, Among them, Wherein, w and h are the width and height of the target coded CU respectively w×h , Ori w×h (i, j) is the original luminance value of the target coded CU w×h , Pre w×h (i, j) is the predicted luminance value of the target coded CU w×h obtained after inter-frame prediction mode calculation, and i and j are the abscissa and ordinate of the pixel point of the target coded CU respectively w×h , and QP is the quantization parameter of the current compression-coded frame; For the current CU w×h Bisect the region horizontally and vertically respectively. Horizontally, two sub-matrices are obtained, and vertically, two sub-matrices are obtained. Calculate the maximum Tra2 w×h in the corresponding positions of the four sub-blocks, and obtain maxVar B ; For the current CU w×h Divide the area horizontally and vertically into four equal parts respectively. Four sub-matrices are obtained horizontally and four sub-matrices are obtained vertically. Calculate the maximum Tra2 corresponding to the eight sub-blocks w×h (i, j) variance value to obtain maxVar Q ; Determine whether the target coded CU w×h satisfies the second preset condition; If not satisfied, after the Inter 2N×2N prediction mode, perform a third prediction distortion transformation on the prediction distortion Err w×h (i,j) to obtain Tra4 w×h (j,j); Among them, Among them, Wherein, w and h are the width and height of the target encoded CU, respectively w×h Ori w×h (i, j) is the original luminance value of the target encoded CU, and Pre w×h (i, j) is the target encoded CU w×h The predicted luminance value obtained after inter-frame prediction mode calculation, and i and j are the horizontal and vertical coordinates of the pixel points of the target encoded CU, respectively w×h The quantization parameter of the current compression encoded frame; For the current CU w×h Bisect the region horizontally and vertically respectively. Horizontally, two sub - matrices are obtained, and vertically, two sub - matrices are obtained. Calculate the maximum Tra4 corresponding to the four sub - blocks w×h (i,j) variance value to obtain maxVar B ; For the current CU w×h Divide the area horizontally and vertically into four equal parts respectively. Four sub - matrices are obtained horizontally and four sub - matrices are obtained vertically. Calculate the corresponding maximum Tra4 w×h (i, j) variance value to obtain maxVar Q ; Determine whether the target coded CU w×h satisfies the fourth preset condition; If so, skip the RDcost calculation of the remaining inter - frame modes; Among them, the expression of the first preset condition is: POC>4&&POC!=8&&POC!=16 and and The expression of the second preset condition is: and The expression of the third preset condition is: POC>4&&POC!=8&&POC!=16 and and The expression of the fourth preset condition is: and where is the adaptive adjustment threshold, and where are the variances obtained by separately calculating after horizontally dividing Tra w×h (i,j) into two sub - matrices; are the variances obtained by separately calculating after vertically dividing Tra w×h (i,j) into two sub - matrices; are the variances obtained by separately calculating after horizontally dividing Tra w×h (i,j) into four sub - matrices; are the variances obtained by separately calculating after vertically dividing Tra w×h (i,j) into four sub - matrices, where Tra w×h (i,j) is Tra1 w×h (i,j) or Tra2 w×h (i,j) or Tra3 w×h (i,j) or Tra4 w×h (i,j).

2. The method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression according to claim 1, wherein: After the process of determining whether the current dynamic 3D point - cloud compression - coded frame is not an I - frame, it also includes: if not, do not perform fast determination of the inter - frame prediction mode for the coded frame.

3. The method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression according to claim 1, wherein: After the process of determining whether the target - coded CU meets the first preset condition, it also includes: if so, the target - coded CU skips the calculation of the remaining inter - frame prediction modes.

4. The method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression according to claim 1, wherein: After determining whether the target - coded CU meets the second preset condition, it also includes: if so, the target - coded CU skips the calculation of the remaining inter - frame prediction modes.

5. The method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression according to claim 1, characterized in that: After the process of determining whether the target - coded CU meets the third preset condition, it also includes: if not, the target - coded CU continues to perform the calculation of the remaining inter - frame prediction modes.

6. The method for quickly determining an inter-frame coding mode for dynamic 3D point cloud compression according to claim 1, wherein: After the process of determining whether the target - coded CU meets the fourth preset condition, it also includes: if not, the target - coded CU continues to perform the calculation of the remaining inter - frame prediction modes.

Citation Information

Patent Citations

  • Fast inter-frame predictive coding method in video coding

    CN110351557A

  • HEVC prediction mode determining method and device thereof

    KR1020140056599A