Point cloud projection video low-complexity coding optimization method based on depth structure prediction
Through the method based on deep partition prediction, the encoding efficiency of point cloud projection video is improved, the problem of high encoding complexity in the existing technology is solved, and more efficient point cloud compression and transmission is achieved.
Patent Information
- Application Number
- CN202510231123.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-03
AI Technical Summary
When encoding dynamic 3D point clouds, the existing technology has high encoding complexity, resulting in unsatisfactory compression efficiency and making it difficult to achieve real-time transmission.
The rapid encoding method of point cloud projection video based on depth partition prediction is adopted to improve the accuracy of deep structure prediction through feature enhancement and deep learning models and reduce the coding time complexity.
It effectively reduces the time complexity of point cloud video encoding, improves the balance of encoding efficiency and quality, and achieves more efficient point cloud compression and transmission.
Smart Images

Figure CN120091144A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point cloud compression encoding, and specifically to a fast encoding method for point cloud projection video based on depth partition prediction, providing a technical option for the optimal compression scheme of dynamic point clouds. Background Art
[0002] With the rapid development of multimedia technology and three-dimensional acquisition technology, three-dimensional point clouds have gradually become an important medium for storing and presenting information. The presentation of 3D point clouds requires a large amount of data, including geometric, color, and normal information, which is the main obstacle to the real-time transmission of 3D point clouds. To effectively compress dynamic point clouds, MPEG has developed two 3D point cloud compression standards: geometry-based point cloud compression (G-PCC) and video-based point cloud compression (V-PCC). Among them, in V-PCC, 3D point cloud frames are projected onto a 2D plane to generate occupancy sequences, geometry sequences, and attribute sequences, and mature 2D video encoders are used to meet the compression requirements. The occupancy sequence is used to reflect whether the 2D projection area corresponds to valid point cloud data, the geometry sequence is used to reflect information such as position and normal, and the attribute sequence is used to reflect information such as color and texture. Among the point cloud compression standards released by standardization organizations in recent years, the compression efficiency of V-PCC is already leading. However, a frame of dynamic 3D point cloud will be projected to generate four 2D images with the same resolution (corresponding to the near layer and far layer of the geometry frame and the attribute frame respectively). When using H.265 / H.266 encoders at the same resolution, the encoding complexity of a frame of dynamic 3D point cloud will increase exponentially compared to natural video frames. Therefore, it is urgent to reduce the encoding complexity of V-PCC to better achieve the compression and transmission of dynamic point clouds.
[0003] In fact, for the projection video sequence of 3D point clouds, a large amount of work has focused on fast algorithms for encoding unit partitioning based on the rate-distortion optimization (RDO) strategy, and certain breakthroughs have been made. At the same time, the introduction of learning-based methods will provide more possibilities for the optimal point cloud compression scheme.
[0004] The existing fast encoding methods for V-PCC at the encoding end can be roughly divided into traditional optimization methods and learning-based optimization methods. Traditional methods usually refer to the ideas and methods of 2D video encoding fast algorithms and combine the correlation characteristics between the occupancy map, attribute map, and geometry map in V-PCC. For example, using occupancy information to prematurely terminate the encoding unit (CU) partitioning, accelerating the decision-making through the isotropy of occupancy information and pixels in 2D and 3D spaces, and determining the partitioning strategy according to the detailed classification of block types, etc. Learning-based methods attempt to deeply explore the available information, including extracting cross-projection information, directly encoding features, distortion prediction errors, and other features that have not been fully concerned in traditional optimization methods. There are also some schemes that directly predict the partitioning method of encoding units.
[0005] At present, the method based on occupancy information has made some progress in early skipping / terminating encoding, but it fails to deeply explore the internal correlations among geometry, attributes, and occupancy information, as well as between near and far-layer frames. At the same time, most of the methods based on depth prediction refer to the deep learning-accelerated 2D video compression scheme, that is, the coding unit division is attributed to image classification or regression problems. In addition, in the intra-frame configuration of V-PCC, the near-layer frame and the far-layer frame are respectively encoded as an I-frame and a P-frame, and every two frames form a group of pictures (GOP). Due to the disorder of the projected frames and the difference in the coding structure from natural video frames, directly applying the coding acceleration scheme of natural videos to V-PCC is not always effective.
[0006] In summary, although the fast algorithms for video coding based on depth prediction are relatively mature, the effects of most methods integrated into V-PCC are not ideal. Accordingly, the present invention proposes an optimization scheme, aiming to make more full use of coding information, reduce the coding complexity of V-PCC through innovative technical means and deep learning models, and is expected to achieve a more efficient coding process. Summary of the Invention
[0007] Aiming at the above technical problems existing in the prior art, the object of the present invention is: 1) fully analyze the correlations and respective statistical characteristics among different types of point cloud projected frames, and apply them to the depth prediction task for feature enhancement; 2) effectively improve the accuracy of depth structure prediction for different types of point cloud projected frames through feature aggregation; 3) design a coding acceleration algorithm for point cloud projected videos to reduce the coding time complexity by avoiding unnecessary mode decision and partition rate distortion cost calculation.
[0008] Considering the time complexity bottleneck faced by point cloud projected video coding, and aiming at the problems of insufficient utilization of point cloud projection information and insufficient accuracy of coding unit depth prediction, the present invention proposes a low-complexity point cloud video coding optimization method based on a depth prediction network. After predicting a more accurate optimal partition structure of coding tree units (CTUs) for near and far-layer geometry and attributes through feature enhancement, the result is applied to the coding acceleration optimization algorithm designed by us to achieve a better balance between coding efficiency and quality.
[0009] The technical solution adopted by the present invention is as follows:
[0010] Step 1. Extract the occupancy video, geometry video, and attribute video of the dynamic point cloud sequence at different bitrates through the standard encoder TMC2 of V-PCC for information extraction and model training.
[0011] Step 2. Extract the feature information of the near-layer frame for feature enhancement to obtain the near-layer enhanced features corresponding to the occupancy map, attribute map, and geometry map
[0012] Step 21. Considering that attribute frames and geometric frames often contain a large number of unoccupied areas, and geometric frames and attribute frames correspond to the same occupied frame. The statistical characteristics of the partitioning complexity of occupied blocks (all valid point cloud data), boundary blocks (partially containing valid point cloud data) and unoccupied blocks (containing no valid point cloud data) are different, so the occupancy information should be used as a key feature in the prediction tasks of geometric frames and attribute frames. The present invention extracts the binary occupancy matrix corresponding to the projected CTU as the near-layer occupancy enhancement feature. Used to reflect the partition complexity of unoccupied blocks and boundary blocks.
[0013] Step 22. The smooth texture in the geometric frame corresponds to the flow area on the surface of the 3D point cloud, and the 3D surface fluctuations of the point cloud are usually accompanied by changes in attributes, which are also reflected in the texture changes of the attribute frame. This means that there is a significant mapping relationship between the attribute frame and the geometric frame. Therefore, the present invention uses the texture change trend inside the geometric CTU to further reflect the impact of the texture complexity of the attribute CTU on its depth strategy, and expresses it in the form of regional pixel variance as a near-layer attribute enhancement feature
[0014] Step 23. During the generation of the geometric frame, for areas with high-density multi-directional edges, a smaller CU size is required for filling, so that more points and edges are generated in some geometric CTUs. Edge detection can better highlight the overall structure of the geometric CTU and remove the influence of non-critical pixel areas. The present invention uses edge detection operators to generate edge image blocks of geometric CTUs as near-layer geometric enhancement features.
[0015] Step 3. Extract the feature information of the far-layer frame for feature enhancement to obtain the far-layer enhanced features corresponding to the occupancy map and geometry / attribute map
[0016] Step 31. Unlike natural video frames, the far-layer frames in the point cloud projection sequence have a strong dependence on their near-layer frames, and the spatial co-location CTU division method of the far-layer frames is usually simpler. Compared with the near-layer frames, the far-layer frames should focus on the similarities and differences between them and the near-layer frames, so as to reduce the encoding complexity. In order to reflect the temporal correlation between the far and near layers, the present invention uses the residual brightness CTU obtained by motion compensation prediction of the current CTU as the far-layer enhancement feature.
[0017] Step 32. Although the depth division structure of the distant frame is simpler, it is still necessary to distinguish whether the simple depth structure is caused by insufficient effective point cloud data or too high similarity with the near frame. Therefore, the binary occupancy matrix is also used as an enhanced feature of the distant frame.
[0018] Step 4. The prediction of the depth partition structure at the CTU level is essentially an image regression prediction task. Based on the prediction results, the reasonable partition depth of each region within the CTU is determined. The present invention applies the feature information representing different types of frames to enhance the features of the projected image blocks, and establishes a model prediction for different branches to predict the 4×4 partition depth structure corresponding to a 64×64-sized CTU, so as to effectively improve the prediction accuracy and reduce the temporal redundancy of point cloud video coding. The model proposed by the present invention includes a feature enhancement module, an aggregation learning module, and a prediction output module.
[0019] Step 41. Establish a feature enhancement module, which is the main difference among the prediction models of different types of projected frames. According to the content mentioned in the previous step, the present invention respectively enhances the features of the luminance projected CTUs belonging to different types of frames, and uses the depth structure ground truth obtained and processed during the encoding process as the learning label. This module performs pre-learning and channel fusion operations on multiple input features to enhance the feature expression ability of the prediction model.
[0020] Step 42. Establish an aggregation learning module, which increases the number of channels and gradually reduces the size of the feature map through multi-branch convolution and channel attention mechanism, so as to more effectively capture important features. Since the prediction task complexities of the far-layer frames and the near-layer frames are different, the structure of this module for the far-layer frames is relatively simpler.
[0021] Step 43. Establish a prediction output module, which simulates the quadtree partition form in HEVC through hierarchical convolution. The weights of these convolutional layers are not shared, and the number of channels is gradually reduced to generate the final prediction result. Compared with the existing scheme of predicting the maximum depth of the projected frame CU, the present invention directly predicts the optimal partition structure of the CTU, thereby further avoiding the calculation of redundant rate-distortion costs.
[0022] Step 44. Correct the predicted depth map. The depth structure of the CTU should follow the depth consistency within a 16×16-sized region. Since the floating-point result predicted by the model has been rounded, regional errors may occur. Therefore, the prediction result is further corrected to conform to the depth consistency to avoid the coding loss generated in this link.
[0023] Step 5. Utilize the prediction results of the depth partition structure and combine with an optimized coding algorithm to implement a low-complexity point cloud projection video coding acceleration process.
[0024] Step 51. Deploy the trained model in the 2D video encoder called by the TMC2 to infer and generate the predicted CTU depth structure diagram.
[0025] Step 52. Accelerate the encoding of the geometric video. Make special mode decisions for unoccupied blocks and far-layer frames, utilize the predicted depth map to skip unnecessary encoding mode decisions and sub-CU recursive partitioning processes, and terminate the CU partitioning in a timely manner.
[0026] Step 53. Accelerate the encoding of the attribute video. Since the partitioning complexity of attribute frames is generally higher than that of geometric frames, compared with geometric frames, only the partitioning of unoccupied blocks is terminated in advance and the far-layer frames are only restricted to perform inter-frame encoding mode decisions, and the remaining blocks perform mode decision skipping and partitioning termination according to the predicted depth structure.
[0027] The beneficial effects of the present invention are as follows:
[0028] 1. The present invention strengthens the features of the depth prediction task of the point cloud projection frame. Compared with the natural video solution and some existing point cloud video solutions, it can further improve the accuracy of the depth prediction of the point cloud video frame, thereby reducing the encoding loss generated in this part.
[0029] 2. The present invention processes different types of point cloud projection frames separately. Compared with the existing solution that uses the same processing method for different types of frames, it can specifically improve the accuracy of the prediction task.
[0030] 3. The present invention designs an acceleration algorithm for the point cloud geometric video and the attribute video by applying the predicted depth map, and can effectively balance the encoding acceleration and the encoding loss. Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. The following drawings only show some embodiments of the present invention, so they should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings. The proportional relationships of the various components in the drawings of this specification do not represent the proportional relationships in actual material selection and design, where:
[0032] Figure 1 is the V-PCC optimization framework diagram for the full-frame intra-coding scenario proposed by the present invention;
[0033] Figures 2a - 2c is the depth partitioning structure prediction model diagram for different types of projection frames of the present invention;
[0034] Figure 3 is the encoder algorithm optimization flowchart proposed by the present invention;
[0035] Figure 4Schematic diagram of the accuracy of the model of the present invention and two existing deep structure prediction models applied to the depth prediction tasks of the near-layer attribute frame (A), near-layer geometry frame (G), and far-layer frame (F) of point cloud projection. Detailed implementation manners
[0036] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Generally, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0037] To further optimize the time complexity of point cloud video coding in V-PCC, the present invention adopts a depth prediction model that conforms to the current coding framework, and optimizes modules such as mode decision and depth partitioning in point cloud projection sequence coding through means such as model training, model integration, and rate-distortion decision optimization. The optimization process is as Figure 1 shown. The specific embodiments are implemented as follows:
[0038] (1) Enhanced feature extraction of point cloud projection frames
[0039] Main feature: Projected luminance CTU. Predict the partition depth map of the CTU, which can also be regarded as an evaluation of the local texture complexity. The luminance component can reflect the texture features to a greater extent than the chrominance component. The present invention uses the luminance CTU of each projection frame as the main input for the prediction task.
[0040] Occupancy enhanced feature: Each projection CTU corresponds to an occupancy block with a minimum precision of 4*4, which represents whether there are pixel occupancies in each minimum-precision 2D area of the point cloud projection. Generally speaking, the division complexity trend from high to low is boundary block, occupancy block, and unoccupied block. Accordingly, the present invention uses the binary-processed occupancy block as the enhanced feature for the far-layer frame and the near-layer frame.
[0041]
[0042] where V() represents the numerical value in the matrix, P i represents the pixel value in the projection CTU, and O represents the set of pixels containing valid point cloud data.
[0043] Near-layer attribute enhanced feature: The present invention uses the pixel variance of the geometric region to reflect the mapping relationship between the geometry frame and the attribute frame,
[0044]
[0045] where w and h respectively represent the width and height of the current region, represents the geometric pixel value, while
[0046] represents the regional average value of the geometric pixels.
[0047]
[0048] Here, λ represents the occupancy type of the current block, 0 indicates that the current block is an unoccupied block, and 1 indicates that the current block is an occupied block or a boundary block. represents the geometric luminance CTU.
[0049] Far - layer enhancement feature: The present invention uses the residual luminance CTU obtained by performing inter - frame prediction on the current far - layer CTU with reference to the near - layer CTU as the input feature of the far - layer frame.
[0050]
[0051] Here and represent the original pixel value and the predicted pixel value of the far - layer CTU respectively.
[0052] (2) Depth partition structure prediction model of the point - cloud projection frame
[0053] The present invention provides a depth prediction network for accurate prediction of different types of frames. This network designs slightly different network architectures for the near - layer attribute frame, the near - layer geometric frame, and the far - layer frame, as Figures 2a - 2c shown.
[0054] Feature enhancement module: The main feature is pre - learned and down - sampled through the residual layer and the convolutional layer. Other enhancement features are pre - learned through up - sampling, the residual layer, etc., and then channel fusion is performed when the feature size matches that of the main input. For features with relatively simple data types such as binary occupancy blocks, channel fusion is performed directly without pre - learning. The aggregated input features will jointly participate in subsequent multi - scale feature learning.
[0055] Aggregation learning module: Inspired by the Inception - SE model, a dual - branch multi - size convolutional module is established. While increasing the number of channels, the feature map size is reduced to 4 * 4 through the max - pooling layer. This module extracts depth features through multiple paths: one path uses a small convolutional kernel for dimension enhancement, and the other paths use larger - size convolutional kernels to further learn features, and a channel attention mechanism module is added at the end.
[0056] Prediction Output Module: Referring to the quadtree (QT) partitioning method in H.265, the output of the aggregation learning module in the present invention is divided into 4 tensors of the same size and processed in four independent convolutional layers, where the weights of these convolutional layers are not shared. The outputs of each hierarchical convolution are merged and then passed through three convolutional layers to generate the final predicted depth map.
[0057] Model Training Strategy: Since the training complexities of different scenarios are not the same, and the training complexities decrease in order as attribute near-layer frames, geometric near-layer frames, and far-layer frames, the model details and training conditions also vary slightly in different scenarios. The loss function uses a multi-scale variant of the L1 function, fully considering the characteristic of local consistency of the depth map, integrating max-min pooling operations to enhance local consistency and make it more suitable for the network model of the present invention. The specific expression of the loss function is as follows:
[0058]
[0059] where k_size represents the pooling kernel size, true represents the label ground truth, and pred represents the prediction output.
[0060] Correction of Predicted Depth Map: The predicted 4×4 depth map can be divided into 4 2×2 regions, and the following corrections are made according to the regional consistency within each region. When the depth within a 2×2 region is 3, the minimum depth within the region should be at least 2; when the maximum depth within a 2×2 region is 2, the depth values within the region should all be 2; when the maximum depth within a 2×2 region is 1, the depth values within the region should all be 1; when the 4×4 region contains both depth 0 and other values, the depth values within the 4×4 region should be corrected to all 1.
[0061] (3) Design of Coding Acceleration Algorithm
[0062] The trained model will be integrated and deployed in the HM encoder called when encoding the projection geometry / attribute sequence by TMC2, and a method of parallel prediction and coding is adopted. The process is as Figure 3 shown, where Dp represents the predicted depth and Dt represents the current depth.
[0063] Point Cloud Video Coding Acceleration Algorithm:
[0064] ① Refer to Figure 3 , and obtain the predicted depth map before making the depth partitioning decision to determine the optimal partitioning structure of the current CTU.
[0065] ② After determining that the current CU belongs to the near layer or the far layer, if the current block is an unoccupied block, the coding mode of the near-layer block is confirmed as INTRA_2N×2N, and the far-layer block is confirmed as SKIP, and the current block partitioning ends.
[0066] ③Far - layer geometry fast mode decision: If the current block is a boundary block or an occupied block, skip the intra - coding mode decision and perform the inter - frame mode decision only when the predicted depth matches the current depth. If the maximum depth has been reached, only make decisions for skip and INTER_2N×2N. If the maximum depth has not been reached, perform all inter - frame mode decisions.
[0067] ④Far - layer attribute fast mode decision: Different from the geometry frame, at the maximum depth, only avoid the mode decision for INTER_N×N.
[0068] ⑤Near - layer fast mode decision: If the current block is a boundary block or an occupied block, perform the intra - frame mode decision only when the predicted depth matches the current depth. If the maximum depth has been reached, avoid the mode decision related to INTRA_N×N. If the maximum depth has not been reached, perform all intra - frame mode decisions.
[0069] ⑥Fast block partitioning: Compare the predicted depth with the depth of the current CU to be encoded. When the predicted depth is greater, directly partition downward and skip the remaining calculation process. When the predicted depth is smaller, immediately terminate the CU partitioning process. Calculate the rate - distortion cost for further partitioning only when the current depth matches the predicted depth to reduce a large amount of time redundancy.
[0070] ⑦End the CU partitioning and encoding at the current depth.
[0071] The effects of the present invention can be further illustrated by the following simulation experiments:
[0072] I. Depth structure prediction accuracy
[0073] The present invention uses a total of twelve different dynamic point - cloud sequences from the 8iVFBv2 and UVG - VPC datasets, and extracts GOPs at intervals for model training. The test sequences use the GOPs not involved in training in the selected sequences and other sequences in the dataset. The model proposed by the present invention has significantly better effects on the point - cloud projection frames at different coding rates compared with the existing partitioning depth prediction models, as Figure 4 shown. Figure 4 compares the accuracy of the model of the present invention and two existing depth - structure prediction models in the depth - prediction tasks for the near - layer attribute frames (A), near - layer geometry frames (G), and far - layer frames (F) of point - cloud projections, respectively.
[0074] II. Analysis of the encoding performance of the scheme
[0075] Table 1 preliminarily shows the encoding - time savings and encoding losses after applying the point - cloud video fast - encoding scheme of the present invention. The test point - cloud sequences are casualsquat and flowerwave, and the calculation method is as follows:
[0076]
[0077] T S represents the degree of time saving. Anchor means testing under the original coding conditions of TMC2-18.0, and test means using the solution of the present invention as the test conditions. Then, the commonly used video coding metric BDBR( rate) is used to measure the coding performance loss of the fast coding optimization solution.
[0078] Flowerwave_vox10 Time saving Coding loss Attribute video 69.02% 1.31% Geometric video 70.75% 1.57% Casualsquat_vox10 Time saving Coding loss Attribute video 62.51% 2.41% Geometric video 66.03% 1.68%
[0079] Table 1 Preliminary test results of coding performance.
[0080] As can be seen from Table 1, the present invention effectively reduces the time complexity of the point cloud video coding process and increases the coding loss as little as possible.
[0081] As described above, the above are only the preferred embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be thought of by those skilled in the art within the technical scope disclosed by the present invention without creative labor should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope defined by the claims.
Claims
1. A low-complexity coding optimization method for point cloud projection video based on deep structure prediction, characterized in that The steps include: Step 1. Extract occupancy video, geometry video and attribute video of dynamic point cloud sequence at different bit rates through V-PCC's standard encoder TMC2; Step 2: Extract feature information of near-layer frames for feature enhancement, and obtain near-layer enhanced features corresponding to occupancy maps, attribute maps, and geometric maps. Step 3. Extract the feature information of the far-layer frame for feature enhancement to obtain the far-layer enhanced features corresponding to the occupancy map and geometry / attribute map Step 4. Feature information of different types of frames is used to enhance the features of the projected image blocks, and a 4*4 partitioned deep structure corresponding to a 64*64 size CTU is established for model prediction of different branches. The deep structure includes a feature enhancement module, an aggregation learning module, and a prediction output module; Step 5. Use the deep partitioning structure prediction results combined with the optimized encoding algorithm to achieve low-complexity point cloud projection video encoding acceleration.
2. The method for optimizing low-complexity coding of point cloud projection video based on deep structure prediction according to claim 1, characterized in that: The step 2 is specifically as follows: Step 21. Extract the binary occupancy matrix corresponding to the projected CTU as the near-layer occupancy enhancement feature Used to reflect the partition complexity of unoccupied blocks and boundary blocks; Step 22. Use the texture change trend inside the geometric CTU to further reflect the impact of the texture complexity of the attribute CTU on its depth strategy, and express it in the form of regional pixel variance as a near-layer attribute enhancement feature Step 23. Use the edge detection operator to generate the edge image block of the geometric CTU as the near-layer geometric enhancement feature 3. The method for optimizing low-complexity coding of point cloud projection video based on deep structure prediction according to claim 1, characterized in that: The step 3 is as follows: Step 31. Use the residual brightness CTU obtained by motion compensation prediction of the current CTU as the far-layer enhancement feature Step 32. Use the binary occupancy matrix as the enhanced feature of the distant frame 4. The method for optimizing low-complexity coding of point cloud projection video based on deep structure prediction according to claim 1, characterized in that: The step 4 comprises the following steps: Step 41. Establish a feature enhancement module to perform feature enhancement on the brightness projection CTUs belonging to different types of frames respectively, and use the deep structure truth value obtained and processed during the encoding process as the learning label; Perform pre-learning and channel fusion operations on multiple input features to enhance the feature expression ability of the prediction model; Step 42. Establish an aggregate learning module to increase the number of channels and gradually reduce the size of the feature map through multi-branch convolution and channel attention mechanism; Step 43. Establish a prediction output module, simulate the quadtree partitioning form in HEVC through layered convolution, and gradually reduce the number of channels to generate the final prediction result; Step 44: Correct the predicted depth map and further correct the predicted result to make it meet the depth consistency.
5. The method for optimizing low-complexity coding of point cloud projection video based on deep structure prediction according to claim 1, characterized in that: The step 5 is specifically as follows: Step 51. Deploy the trained model in the 2D video encoder called by TMC2 for inference to generate the predicted CTU depth structure map; Step 52: Accelerate the encoding of the geometric video, make special mode decisions for unoccupied blocks and far-layer frames, use the predicted depth map to skip unnecessary coding mode decisions and sub-CU recursive division processes, and terminate the CU division in a timely manner; Step 53. Encoding the attribute video is accelerated. Compared with the geometric frame, only the division of unoccupied blocks is terminated in advance and the distant layer frame is limited to only inter-frame coding mode decision. The remaining blocks are skipped and the division is terminated according to the predicted depth structure.
Citation Information
Cited By
Lightweight multi-source information combined V-PCC inter-frame mode rapid selection method
CN121486590A
VVC intra-frame prediction fast mode decision-making method suitable for V-PCC
CN121967685A