Video-based Point Cloud Compression (V-PCC) Component Synchronization

By using output delay synchronization technology in PCC decoder, the problem of low video compression efficiency in the prior art is solved, and the effect of improving the video compression ratio without affecting the image quality is achieved.

CN114503164BActive Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080070550.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-07
Filing Date
2020-10-06
Publication Date
2025-06-27
Estimated Expiration
2040-10-06

AI Technical Summary

Technical Problem

With limited network resources and the need for higher video quality, it is difficult for the prior art to improve the video compression ratio without affecting the image quality.

Method used

By designing a PCC decoder that receives point cloud code streams, performs caches based on time, and uses output delay synchronization to reduce buffer memory size, thereby improving compression efficiency.

Benefits of technology

This method improves data synchronization through output delay synchronization, reduces buffer memory size, improves video compression efficiency, and maintains image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114503164B_ABST
    Figure CN114503164B_ABST
Patent Text Reader

Abstract

A method implemented by a PCC decoder, the method comprising: the PCC decoder receiving a point cloud bitstream; the PCC decoder performing buffering on the point cloud bitstream based on time, the performing including determining the time based on a delay and a delay offset; the PCC decoder decoding the point cloud bitstream based on the buffering. A method is implemented by a PCC decoder, the method comprising: the PCC decoder receiving a point cloud bitstream; the PCC decoder performing buffering on the point cloud bitstream based on a delay, the delay being based on a first delay and a second delay; the PCC decoder decoding the point cloud bitstream based on the buffering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments generally relate to a PCC, and more particularly to a V-PCC component synchronization. Background Art

[0002] Even if a video is relatively short, a large amount of video data is required to describe it, which may cause difficulties when the data is to be streamed or otherwise transmitted over a communication network with limited bandwidth capacity. Therefore, video data is usually compressed first and then transmitted through modern telecommunication networks. Since memory resources may be limited, the size of the video may also be a problem when storing the video in a storage device. Video compression devices typically use software and / or hardware on the source side to encode the video data and then transmit or store it, thereby reducing the amount of data required to represent the digital video image. Then, the compressed data is received on the destination side by a video decompression device that decodes the video data. In the context of limited network resources and the growing demand for higher video quality, improved compression and decompression techniques are needed that can increase the compression ratio with little impact on image quality. Summary of the Invention

[0003] A first aspect relates to a method implemented by a PCC decoder, the method comprising: the PCC decoder receiving a point cloud bitstream; the PCC decoder performing buffering on the point cloud bitstream based on time, the performing including determining the time based on a delay and a delay offset; the PCC decoder decoding the point cloud bitstream based on the buffering.

[0004] This embodiment provides a solution, wherein the decoded V-PCC components are output at a corresponding component decoder (referred to as CP A) and transmitted to a buffer, wherein the synchronization process is to prepare the data for reconstruction at CP B. Output delay synchronization is used in the synchronization process. Output delay synchronization improves synchronization, thereby reducing the buffer memory size.

[0005] Optionally, in any of the above aspects, the time is further based on a removal time.

[0006] Optionally, in any of the above aspects, the time is further based on a ClockTick.

[0007] Optionally, in any of the above aspects, the time is further based on a first expression of the delay and the delay offset.

[0008] Optionally, in any of the above aspects, the time is further based on a second expression, and the second expression is a product of the ClockTick and the first expression.

[0009] Optionally, in any of the above aspects, the time is further based on the sum of the removal time and the second expression.

[0010] Optionally, in any of the above aspects, the point cloud bitstream includes a plurality of components.

[0011] Optionally, in any of the above aspects, the component includes an occupancy map.

[0012] Optionally, in any of the above aspects, the component includes geometric data.

[0013] Optionally, in any of the above aspects, the component includes attribute data.

[0014] Optionally, in any of the above aspects, the component includes an atlas frame.

[0015] Optionally, in any of the above aspects, the time is further based on the number of the components.

[0016] Optionally, in any of the above aspects, the time is DpbDabOutputTime.

[0017] Optionally, in any of the above aspects, the delay is PicAtlasDpbOutputDelay.

[0018] Optionally, in any of the above aspects, the delay offset is DpbDabDelayOffset.

[0019] Optionally, in any of the above aspects, DpbDabDelayOffset is equal to the difference between MaxInitialDelay and PicAtlasDpbOutputDelay.

[0020] Optionally, in any of the above aspects, the method further includes: storing the point cloud bitstream; displaying an image or video from the point cloud bitstream.

[0021] A second aspect relates to a method implemented by a PCC decoder, the method including: the PCC decoder receiving a point cloud bitstream; the PCC decoder performing buffering on the point cloud bitstream based on a delay, the delay being based on a first delay and a second delay; the PCC decoder decoding the point cloud bitstream based on the buffering.

[0022] Optionally, in any of the above aspects, the delay is further based on the maximum value of the first delay and the second delay.

[0023] Optionally, in any of the above aspects, the delay is MaxInitialDelay.

[0024] Optionally, in any of the above aspects, the first delay is MaxInitDelay.

[0025] Optionally, in any of the above aspects, the second delay is PicAtlasDpbOutputDelay.

[0026] Optionally, in any of the above aspects, the buffer is further based on DpbDabDelayOffset, and DpbDabDelayOffset = MaxInitialDelay - PicAtlasDpbOutputDelay.

[0027] Optionally, in any of the above aspects, the method further includes: storing the point cloud bitstream; displaying an image or video from the point cloud bitstream.

[0028] Any of the above embodiments can be combined with any of the other above embodiments to create new embodiments. These and other features will be more clearly understood from the following detailed description in conjunction with the accompanying drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] To more fully understand the present invention, reference is now made to the following brief description taken in conjunction with the accompanying drawings and specific embodiments, in which like reference numerals represent like components.

[0030] Figure 1 is a flowchart of an exemplary method for decoding a video signal.

[0031] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system for video decoding.

[0032] Figure 3 is a schematic diagram of an exemplary video encoder.

[0033] Figure 4 is a schematic diagram of an exemplary video decoder.

[0034] Figure 5 is an example of a point cloud media that can be decoded according to the PCC mechanism.

[0035] Figure 6 is an example of a patch created from a point cloud.

[0036] Figure 7A shows an exemplary occupancy frame associated with a set of patches.

[0037] Figure 7B shows an exemplary geometry frame associated with a set of patches.

[0038] Figure 7CShows an exemplary set of atlas frames associated with a set of slices.

[0039] Figure 8 Is a schematic diagram of an exemplary compliance test mechanism.

[0040] Figure 9 Is a schematic diagram of an exemplary HRD for performing a compliance test on a PCC bitstream.

[0041] Figure 10 Is a schematic diagram of an exemplary PCC bitstream.

[0042] Figure 11 Is a schematic diagram of an exemplary video decoding device.

[0043] Figure 12 Is a schematic diagram of component synchronization in point cloud reconstruction.

[0044] Figure 13 Is a schematic diagram showing the maximum delay calculation for a single picture.

[0045] Figure 14 Is a schematic diagram showing the maximum delay calculation for multiple pictures.

[0046] Figure 15 Is a flowchart of a method for decoding a bitstream provided by a first embodiment.

[0047] Figure 16 Is a flowchart of a method for decoding a bitstream provided by a second embodiment. Detailed implementation

[0048] First, it should be understood that although the following provides illustrative implementations of one or more embodiments, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The present invention should in no way be limited to the illustrative embodiments, diagrams, and techniques described below, including the exemplary designs and embodiments illustrated and described herein, but may be modified within the full scope of the appended claims and their equivalents.

[0049] The following abbreviations apply:

[0050] ASIC: Application-Specific Integrated Circuit

[0051] AU: Access Unit

[0052] BT: Binary Tree

[0053] CAB: Coded Atlas Buffer

[0054] CABAC: Context-Adaptive Binary Arithmetic Coding

[0055] CAVLC: Context-Adaptive Variable-Length Coding

[0056] Cb: Blue Difference Chroma

[0057] CPA: Conformance Point A

[0058] CPB: Conformance Point B

[0059] CPU: Central Processing Unit

[0060] Cr: Red Difference Chroma

[0061] CTB: Coding Tree Block

[0062] CTU: Coding Tree Unit

[0063] CU: Coding Unit

[0064] DAB: Decoded Atlas Buffer

[0065] DC: Direct Current

[0066] DCT: Discrete Cosine Transform

[0067] DMM: Depth Modeling Mode

[0068] DPB: Decoded Picture Buffer

[0069] DSP: Digital Signal Processor

[0070] DST: Discrete Sine Transform

[0071] EO: Electrical-to-Optical

[0072] FIFO: First-In, First-Out

[0073] FPGA: Field-Programmable Gate Array

[0074] HEVC: High Efficiency Video Coding

[0075] HRD: Hypothetical Reference Decoder

[0076] HSS: Hypothetical Stream Scheduler

[0077] ID: Identifier

[0078] I / O: Input / Output

[0079] NAL: Network Abstraction Layer

[0080] OE: Optical-to-Electrical

[0081] PCC: Point Cloud Compression

[0082] PIPE: Probability Interval Partitioning Entropy

[0083] PU: Prediction Unit

[0084] QT: Quad Tree

[0085] RAM: Random-Access Memory

[0086] RDO: Rate-Distortion Optimization

[0087] RGB: Red, Green, and Blue

[0088] ROM: Read-Only Memory

[0089] SAD: Sum of Absolute Differences

[0090] SAO: Sample Adaptive Offset

[0091] SBAC: Syntax-Based Arithmetic Coding

[0092] SEI: Supplemental Enhancement Information

[0093] SPS: Sequence Parameter Set

[0094] SRAM: Static Random-Access Memory

[0095] SSD: Sum of Squared Differences

[0096] TCAM: Ternary Content-Addressable Memory

[0097] TT: Triple Tree

[0098] TU: Transform Unit

[0099] TX / RX: Transceiver Unit

[0100] V-PCC: Video-Based PCC

[0101] 2D: Two-Dimensional

[0102] 3D: Three-Dimensional

[0103] Unless otherwise specified, the following terms are defined as follows. Terms may be described differently in different contexts. Therefore, the following definitions should be regarded as supplementary information and should not be regarded as restricting any other definitions provided.

[0104] An encoder is a device for compressing point cloud data into a bitstream through an encoding process. A decoder is a device for reconstructing point cloud data from the bitstream for display through a decoding process. Point cloud / point cloud representation is a set of points (e.g., samples) in 3D space, where each point may include a position and (optional) attributes such as color. A bitstream is a series of bits that includes point cloud data, which is compressed for transmission between the encoder and the decoder. In the context of PCC, the bitstream includes a series of bits of encoded V-PCC components.

[0105] A V-PCC component or more generally a PCC component can be atlas data, occupancy map data, geometric data, or a specific type of attribute data associated with a V-PCC point cloud. An atlas can be a collection of 2D bounding boxes or tiles projected into a rectangular frame corresponding to a 3D bounding box in 3D space, where each 2D bounding box represents a subset of the point cloud. An occupancy map can be a 2D array corresponding to the atlas, and the values of the occupancy map indicate whether each sample position in the atlas corresponds to a valid 3D point in the point cloud representation. A geometry map can be a 2D array created by aggregating geometric information associated with each tile, where the geometric information / data can be a set of Cartesian coordinates associated with a point cloud frame. An attribute can be a scalar or vector attribute optionally associated with each point in the point cloud, and can refer to color, reflectance, surface normal, timestamp, or material ID. The complete set of atlas data, occupancy map, geometry map, or attributes associated with a specific time instance can be referred to as an atlas frame, occupancy map frame, geometry frame, and attribute frame, respectively. Atlas data, occupancy map data, geometric data, or attribute data can be components of the point cloud and can thus be referred to as atlas components, occupancy map components, geometry components, and attribute frame components, respectively.

[0106] An AU can be a set of NAL units that are associated with each other according to specified classification rules and related to a specific output time. An encoded component can be data that has been compressed for inclusion in a bitstream. A decompressed component can be data from a bitstream or sub-bitstream that has been reconstructed as part of the decoding process or as part of an HRD compliance test. An HRD can be a decoder model that runs on an encoder, which examines the variability of the bitstream produced by the encoding process to verify compliance with specified constraints. An HRD compliance test can determine whether the encoded bitstream complies with the standard. A compliance point can be a point in the decoding / reconstruction process where the HRD performs an HRD compliance check to verify whether the decompressed or reconstructed data complies with the standard. An HRD parameter can be a syntax element that initializes or defines the operating conditions of the HRD. An SEI message can be a syntax structure that has a specified semantics for conveying information not required for the decoding process in order to determine the values of samples in a decoded image. A buffering period SEI message can be an SEI message that includes data representing an initial removal delay related to the CAB in the HRD. An atlas frame timing SEI message can include data representing a removal delay related to the CAB and an output delay related to the DAB in the HRD. A reconstructed point cloud can be a point cloud generated from data from a PCC bitstream. The reconstructed point cloud should approximate the point cloud encoded into the PCC bitstream.

[0107] A decoding unit can be any encoded component in a bitstream or sub-bitstream stored in a buffer for decoding. A CAB removal delay can be the amount of time a component can remain in the CAB before being removed. An initial CAB removal delay can be the amount of time a component in the first AU in a bitstream or sub-bitstream can remain in the CAB before being removed. A DAB can be a FIFO buffer in the HRD that includes decoded atlas frames in decoding order for use during a PCC bitstream compliance test. A DAB output delay can be the amount of time a decoded component can remain in the DAB before being output (e.g., as part of a reconstructed point cloud).

[0108] V-PCC is a mechanism for efficiently decoding 3D objects represented by point clouds with different attributes. Specifically, V-PCC is used to encode or decode these point clouds for display as part of a video sequence. Point clouds are captured over time and included in PCC frames. The PCC frames are divided into PCC components, and then the PCC components are encoded. The position of each valid point in the cloud at a given moment is stored as a geometry map in a geometry frame. Colors are stored as attribute frames. Specifically, slices at a given moment are packed into an atlas frame. Slices typically do not cover the entire atlas frame. Therefore, an occupancy frame is also generated, which represents which parts of the atlas frame include valid slice data. Optionally, attributes of the points (such as transparency, opacity, and / or other data) can be included in the attribute frame. Thus, each PCC frame can be encoded to include multiple frames describing different components of the point cloud at the corresponding moment. In addition, different components can be decoded by using different encoding and decoding (codec) systems.

[0109] Figure 1 It is a flowchart of an exemplary operation method 100 for decoding a video signal. Specifically, the video signal is encoded at an encoder. The encoding process compresses the video signal by using various mechanisms to reduce the video file size. The smaller file size enables the compressed video file to be sent to a user while reducing the associated bandwidth overhead. Then, a decoder decodes the compressed video file to reconstruct the original video signal for display to an end user. The decoding process typically corresponds to the encoding process to facilitate the decoder to consistently reconstruct the video signal.

[0110] In step 101, the video signal is input into the encoder. For example, the video signal can be an uncompressed video file stored in a memory. As another example, the video file can be captured by a video capture device (e.g., a camera) and encoded to support live streaming of the video. The video file can include an audio component and a video component. The video component includes a series of image frames. When these image frames are viewed in sequence, they give a visual effect of motion. The frames include a luminance component or luminance samples, which are pixels represented by light, and include a chrominance component or chrominance samples, which are pixels represented by color. In some examples, these frames can also include depth values to support 3D viewing.

[0111] In step 103, the video is segmented into blocks. The segmentation includes subdividing the pixels in each frame into square or rectangular blocks for compression. For example, in HEVC, a frame can first be divided into CTUs, which are blocks with a predefined size (e.g., 64×64 pixels). These CTUs include luminance samples and chrominance samples. Coding trees can be used to divide the CTUs into blocks, and then these blocks are repeatedly subdivided until a configuration that supports further encoding is obtained. For example, the luminance component of a frame can be subdivided until each block includes relatively uniform luminance values. Additionally, the chrominance component of a frame can be subdivided until each block includes relatively uniform color values. Therefore, the segmentation mechanism varies depending on the content of the video frame.

[0112] In step 105, the image blocks obtained by segmentation in step 103 are compressed using various compression mechanisms. For example, inter-frame prediction or intra-frame prediction can be used. Inter-frame prediction takes advantage of the fact that objects in a general scene tend to appear in consecutive frames. Therefore, blocks that describe objects in a reference frame do not need to be repeatedly described in adjacent frames. Specifically, an object (e.g., a table) can remain in a fixed position in multiple frames. Thus, the table is described once, and adjacent frames can refer back to this reference frame. Pattern matching mechanisms can be used to match objects across multiple frames. Additionally, due to reasons such as object movement or camera movement, moving objects can be represented across multiple frames. In a specific example, a video can show a car moving across the screen over multiple frames. Motion vectors can be used to describe such movement. A motion vector is a 2D vector that provides the offset of the coordinates of an object in one frame to the coordinates of that object in a reference frame. Therefore, inter-frame prediction can encode an image block in the current frame as a set of motion vectors representing the offset of the image block in the current frame from the corresponding block in the reference frame.

[0113] Intra prediction encodes blocks in a common frame. Intra prediction exploits the fact that luminance and chrominance components tend to cluster in a frame. For example, a patch of green in a part of a tree is often adjacent to several similar patches of green. Intra prediction uses multiple directional prediction modes (e.g., 33 modes in HEVC), a planar mode, and a DC mode. These directional modes indicate that the samples of the current block are similar or identical to the samples of adjacent blocks in the corresponding direction. The planar mode indicates that a series of blocks on a row / column (e.g., a plane) can be interpolated based on adjacent blocks on the edge of that row. The planar mode actually represents a smooth transition of light / color across the row / column by using a relatively constant slope of varying values. The DC mode is used for boundary smoothing and indicates that the block is similar or identical to the average of the samples of all adjacent blocks that are related to the angular directions of the directional prediction modes. Thus, an intra prediction block can represent an image block as various relational prediction mode values rather than as actual values. Additionally, an inter prediction block can represent an image block as motion vector values rather than as actual values. In either case, the prediction block may not accurately represent the image block in some cases. All differences are stored in a residual block. A transform can be applied to the residual block to further compress the file.

[0114] In step 107, various filtering techniques can be applied. In HEVC, the filters are applied according to an in-loop filtering scheme. The block-based prediction described above may produce a blocky image at the decoder. Additionally, the block-based prediction scheme can encode a block and then reconstruct the encoded block for subsequent use as a reference block. The in-loop filtering scheme iteratively applies a noise reduction filter, a deblocking filter, an adaptive loop filter, and an SAO filter to the block / frame. These filters reduce block artifacts, enabling the accurate reconstruction of the encoded file. Additionally, these filters reduce artifacts in the reconstructed reference blocks, making it less likely for artifacts to produce other artifacts in subsequent blocks encoded based on the reconstructed reference blocks.

[0115] Once the video signal is segmented, compressed, and filtered, in step 109, the resulting data is encoded into a bitstream. The bitstream includes the data described above and any indication data needed to support proper video signal reconstruction at the decoder. For example, this data can include segmentation data, prediction data, residual blocks, and various flags that provide decoding instructions to the decoder. The bitstream can be stored in memory for transmission to the decoder upon request. The bitstream can also be broadcast or multicast to multiple decoders. Creating the bitstream is an iterative process. Thus, step 101, step 103, step 105, step 107, and step 109 can be performed continuously or simultaneously over multiple frames and blocks. Figure 1 The steps shown can be in another suitable order.

[0116] In step 111, the decoder receives the bitstream and starts the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax data and video data. In step 111, the decoder uses the syntax data in the bitstream to determine the segmented parts of the frame. The segmentation should match the result of the block segmentation in step 103. The entropy encoding / decoding used in step 111 is described below. The encoder makes many choices during the compression process. For example, it selects a block segmentation scheme from several possible choices based on the spatial localization of the values in the input image. Indicating the exact choice may use a large number of bits. A "bit" is a binary value that serves as a variable (e.g., a bit value that may vary depending on the content). Entropy encoding enables the encoder to discard any options that are clearly not suitable for a particular situation, leaving behind a set of available options. Then, a codeword is assigned to each available option. The length of the codeword depends on the number of available options (i.e., one binary symbol corresponds to two options, two binary symbols correspond to three to four options, and so on). Then, the encoder encodes the codewords of the selected options. This scheme reduces the codewords because the codewords are as large as expected, uniquely indicating a selection from a small subset of the allowable options rather than uniquely indicating a selection from a potentially large set of all possible options. Then, the decoder decodes this selection by determining the set of allowable options in a similar manner to the encoder. By determining the set of allowable options, the decoder can read the codeword and determine the choice made by the encoder.

[0117] In step 113, the decoder performs block decoding. Specifically, the decoder uses an inverse transform to generate residual blocks. Then, the decoder uses the residual blocks and the corresponding prediction blocks to reconstruct the image blocks according to the segmentation. The prediction blocks may include intra-prediction blocks and inter-prediction blocks generated by the encoder in step 105. Then, the reconstructed image blocks are placed in the frame of the reconstructed video signal according to the segmentation data determined in step 111. The syntax for step 113 may also be indicated in the bitstream through the entropy encoding described above.

[0118] In step 115, the frame of the reconstructed video signal is filtered in a manner similar to step 107 at the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter can be applied to the frame to remove block artifacts. Once the frame is filtered, in step 117, the video signal can be output to a display for the end user to view.

[0119] Figure 2FIG. 0 is a schematic diagram of an exemplary encoding and decoding (codec) system 200 for video decoding. Specifically, the codec system 200 provides functions to support the implementation of the operation method 100. The codec system 200 is generally used to describe components used at both the encoder and the decoder. The codec system 200 receives a video signal and segments the video signal to obtain a segmented video signal 201, as described in conjunction with steps 101 and 103 in the operation method 100. Then, when acting as an encoder, the codec system 200 compresses the segmented video signal 201 into an encoded bitstream, as described in conjunction with steps 105, 107, and 109 in method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described in conjunction with steps 111, 113, 115, and 117 in the operation method 100. The codec system 200 includes a general-purpose decoder control component 211, a transform, scaling, and quantization component 213, an intra-frame estimation component 215, an intra-frame prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control and analysis component 227, an in-loop filter component 225, a decoded image buffer component 223, and a header format and CABAC component 231. The black lines represent the movement of data to be decoded, while the dashed lines represent the movement of control data that controls the operation of other components. All components in the codec system 200 can exist in the encoder. The decoder can include a subset of the components in the codec system 200. For example, the decoder can include the intra-frame prediction component 217, the motion compensation component 219, the scaling and inverse transform component 229, the in-loop filter component 225, and the decoded image buffer component 223.

[0120] The segmented video signal 201 is a captured video sequence that has been segmented into pixel blocks through a coding tree. The coding tree uses various partitioning modes to subdivide the pixel blocks into smaller pixel blocks. Then, these blocks can be further subdivided into smaller blocks. These blocks can be referred to as nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is called the depth of the node / coding tree. The partitioned blocks can be included in a CU. For example, a CU can be a subpart of a CTU, including a luminance block, a Cr block, and a Cb block, as well as the corresponding syntax instructions of the CU. The partitioning modes can include BT, TT, and QT, which are used to divide a node into two, three, or four child nodes with different shapes, respectively, depending on the partitioning mode used. The segmented video signal 201 is forwarded to the general-purpose decoder control component 211, the transform, scaling, and quantization component 213, the intra-frame estimation component 215, the filter control and analysis component 227, and the motion estimation component 221 for compression.

[0121] The general decoder control component 211 is used to make decisions related to encoding images in a video sequence into a bitstream according to application constraints. For example, the general decoder control component 211 manages the optimization of the bitrate / bitstream size relative to the reconstruction quality. These decisions can be made according to the storage space / bandwidth availability and the image resolution request. The general decoder control component 211 also manages the buffer utilization according to the transmission speed to alleviate buffer underflow and overflow problems. To solve these problems, the general decoder control component 211 manages the segmentation, prediction, and filtering performed by other components. For example, the general decoder control component 211 can dynamically increase the compression complexity to increase the resolution and bandwidth utilization, or decrease the compression complexity to decrease the resolution and bandwidth utilization. Therefore, the general decoder control component 211 controls other components in the codec system 200 to balance the video signal reconstruction quality and the bitrate issue. The general decoder control component 211 generates control data, which is used to control the operations of other components. The control data is also forwarded to the header format and CABAC component 231 for encoding into the bitstream, thereby indicating the parameters used by the decoder for decoding.

[0122] The segmented video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter-frame prediction. The frames or slices of the segmented video signal 201 can be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-frame predictive coding on the received video blocks relative to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform multiple coding rounds to select a suitable coding mode for each video data block, and so on.

[0123] The motion estimation component 221 and the motion compensation component 219 can be highly integrated, but are described separately for conceptual purposes. Motion estimation performed by the motion estimation component 221 is a process of generating motion vectors, where these motion vectors are used to estimate the motion of video blocks. For example, a motion vector can represent the displacement of a decoded object relative to a prediction block. A prediction block is a block that is found to highly match the block to be decoded in terms of pixel difference. A prediction block can also be referred to as a reference block. Such pixel difference can be determined by SAD, SSD, or other difference metrics. HEVC uses several coding objects, including CTU, CTB, and CU. For example, a CTU can be divided into CTBs, and a CTB can then be divided into CBs, and CBs are included in CUs. A CU can be encoded as a PU including prediction data or a TU including transform residual data of the CU. The motion estimation component 221 uses rate-distortion analysis as part of the rate-distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. of the current block / frame, and can select the reference blocks, motion vectors, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance the quality of video reconstruction (e.g., the amount of data loss caused by compression) and decoding efficiency (e.g., the size of the final encoding).

[0124] In some examples, the codec system 200 can calculate values at sub-integer pixel positions of a reference image stored in the decoded picture buffer component 223. For example, the video codec system 200 can interpolate values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, the motion estimation component 221 can perform motion search relative to integer pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy. The motion estimation component 221 calculates the motion vectors of PUs of video blocks in an inter-coded slice by comparing the position of the PU with the position of the prediction block of the reference image. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header format and CABAC component 231 for encoding and outputs them as motion data to the motion compensation component 219.

[0125] The motion compensation performed by the motion compensation component 219 may involve obtaining or generating a prediction block according to the motion vector determined by the motion estimation component 221. Similarly, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. When receiving the motion vector of the PU of the current video block, the motion compensation component 219 may locate the prediction block pointed to by the motion vector. Then, the pixel values of the prediction block are subtracted from the pixel values of the current video block being decoded to obtain a pixel difference, thereby forming a residual video block. Generally, the motion estimation component 221 performs motion estimation with respect to the luminance component, and the motion compensation component 219 uses the motion vector calculated according to the luminance component for both the chrominance component and the luminance component. The prediction block and the residual block are forwarded to the transform scaling and quantization component 213.

[0126] The segmented video signal 201 is also sent to the intra estimation component 215 and the intra prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra estimation component 215 and the intra prediction component 217 may be highly integrated, but are described separately for conceptual purposes. The intra estimation component 215 and the intra prediction component 217 perform intra prediction on the current block with respect to each block in the current frame to replace the inter prediction performed between frames by the motion estimation component 221 and the motion compensation component 219 as described above. Specifically, the intra estimation component 215 determines an intra prediction mode to encode the current block. In some examples, the intra estimation component 215 selects a suitable intra prediction mode from a plurality of tested intra prediction modes to encode the current block. Then, the selected intra prediction mode is forwarded to the header format and CABAC component 231 for encoding.

[0127] For example, the intra estimation component 215 performs rate-distortion analysis on various tested intra prediction modes to calculate rate-distortion values, and selects the intra prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to produce the encoded block, and determines the bit rate (e.g., the number of bits) used to produce the encoded block. The intra estimation component 215 calculates a ratio based on the distortion and rate of various encoded blocks to determine the intra prediction mode that exhibits the best rate-distortion value of the block. Additionally, the intra estimation component 215 can be used to decode depth blocks of a depth map using DMM according to RDO.

[0128] When implemented on the encoder, the intra prediction component 217 can generate a residual block from a prediction block according to a selected intra prediction mode determined by the intra estimation component 215, or when implemented on the decoder, it can read the residual block from the bitstream. The residual block includes the difference between the prediction block and the original block, represented as a matrix. Then, the residual block is forwarded to the transform scaling and quantization component 213. The intra estimation component 215 and the intra prediction component 217 can operate on the luminance component and the chrominance component.

[0129] The transform scaling and quantization component 213 is used to further compress the residual block. The transform scaling and quantization component 213 applies transforms such as DCT, DST, or conceptually similar transforms to the residual block, generating a video block including residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms can also be used. The transform can convert the residual information from the pixel value domain to the transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also used to scale the transform residual information according to frequency, etc. This scaling involves applying a scaling factor to the residual information so as to quantify different frequency information at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also used to quantize the transform coefficients to further reduce the bitrate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 can then perform a scan on the matrix including the quantized transform coefficients. The quantized transform coefficients are forwarded to the header format and CABAC component 231 for encoding into the bitstream.

[0130] The scaling and inverse transform component 229 performs operations opposite to those of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, inverse transform, and / or dequantization to reconstruct the residual block in the pixel domain, e.g., for subsequent use as a reference block. This reference block can become the prediction block for another current block. The motion estimation component 221 and / or the motion compensation component 219 can calculate the reference block by adding the residual block back to the corresponding prediction block for motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce the artifacts generated during scaling, quantization, and transformation. These artifacts may make the prediction inaccurate (and generate additional artifacts) when predicting subsequent blocks.

[0131] The filter control analysis component 227 and the in-loop filter component 225 apply a filter to the residual block and / or the reconstructed image block. For example, the transform residual block from the scaling and inverse transform component 229 can be combined with the corresponding prediction block from the intra prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Then, a filter can be applied to the reconstructed image block. In some examples, the filter can instead be applied to the residual block. As Figure 2Among other components, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and can be implemented together, but are described separately for conceptual purposes. Filters applied to the reconstructed reference block are applied to specific spatial regions, and these filters include multiple parameters to adjust the way these filters are used. The filter control analysis component 227 analyzes the reconstructed reference block to determine where these filters need to be used and sets the corresponding parameters. This data is forwarded as filter control data to the header format and CABAC component 231 for encoding. The in-loop filter component 225 applies these filters based on the filter control data. These filters can include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. These filters can be applied, for example, in the spatial domain / pixel domain (e.g., for a reconstructed pixel block) or in the frequency domain.

[0132] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded image buffer component 223 for subsequent use in motion estimation, as described above. When operating as a decoder, the decoded image buffer component 223 stores the reconstructed and filtered blocks and forwards them as part of the output video signal to a display. The decoded image buffer component 223 can be any storage device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0133] The header format and CABAC component 231 receives data from various components of the codec system 200 and encodes this data into an encoded bitstream for transmission to a decoder. Specifically, the header format and CABAC component 231 generates various headers to encode control data (e.g., general control data and filter control data). In addition, prediction data (including intra prediction data and motion data) as well as residual data in the form of quantized transform coefficient data are both encoded into the bitstream. The final bitstream includes all the information required for the decoder to reconstruct the original segmented video signal 201. This information can also include an intra prediction mode index table (also referred to as a codeword mapping table), definitions of the coding context for various blocks, an indication of the most likely intra prediction mode, an indication of segmentation information, etc. This data can be encoded using entropy coding. For example, the information can be encoded by using CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. After entropy coding, the encoded bitstream can be sent to another device (e.g., a video decoder) or archived for subsequent transmission or retrieval.

[0134] Figure 3It is a block diagram of an exemplary video encoder 300. The video encoder 300 can be used to implement the encoding function of the codec system 200 and / or implement steps 101, 103, 105, 107, and / or 109 in the operation method 100. The encoder 300 divides the input video signal to obtain the divided video signal 301, which is substantially similar to the divided video signal 201. Then, the divided video signal 301 is compressed by components in the encoder 300 and encoded into a bitstream.

[0135] Specifically, the divided video signal 301 is forwarded to the intra prediction component 317 for intra prediction. The intra prediction component 317 can be substantially similar to the intra estimation component 215 and the intra prediction component 217. The divided video signal 301 is also forwarded to the motion compensation component 321 for inter prediction based on the reference blocks in the decoded picture buffer component 323. The motion compensation component 321 can be substantially similar to the motion estimation component 221 and the motion compensation component 219. The predicted blocks and residual blocks from the intra prediction component 317 and the motion compensation component 321 are forwarded to the transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 can be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding predicted blocks (along with the relevant control data) are forwarded to the entropy coding component 331 for encoding in the bitstream. The entropy coding component 331 can be substantially similar to the header format and CABAC component 231.

[0136] The transformed and quantized residual blocks and / or the corresponding predicted blocks are also forwarded from the transform and quantization component 313 to the inverse transform and dequantization component 329 to be reconstructed into reference blocks for use by the motion compensation component 321. The inverse transform and dequantization component 329 can be substantially similar to the scaling and inverse transform component 229. According to an example, the in-loop filter in the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks. The in-loop filter component 325 can be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 can include multiple filters, as described in connection with the in-loop filter component 225. Then, the filtered blocks are stored in the decoded picture buffer component 323 to be used as reference blocks for the motion compensation component 321. The decoded picture buffer component 323 can be substantially similar to the decoded picture buffer component 223.

[0137] Figure 4It is a block diagram of an exemplary video decoder 400. The video decoder 400 can be used to implement the decoding function of the codec system 200 and / or implement steps 111, 113, 115, and / or 117 in the operation method 100. For example, the decoder 400 receives a bitstream from the encoder 300 and generates a reconstructed output video signal according to the bitstream for display to the end user.

[0138] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is used to perform an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE decoding, or other entropy decoding techniques. For example, the entropy decoding component 433 can use the header information to provide context for parsing additional data encoded as codewords in the bitstream. The decoded information includes any information required to decode the video signal, such as general control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients in the residual blocks. The quantized transform coefficients are forwarded to the inverse transform and inverse quantization component 429 to reconstruct the residual blocks. The inverse transform and inverse quantization component 429 can be similar to the inverse transform and inverse quantization component 329.

[0139] The reconstructed residual blocks and / or prediction blocks are forwarded to the intra prediction component 417 to be reconstructed into image blocks according to the intra prediction operation. The intra prediction component 417 can be similar to the intra estimation component 215 and the intra prediction component 217. Specifically, the intra prediction component 417 uses the prediction mode to locate the reference block in the frame and adds the residual block to the above result to reconstruct the intra prediction image block. The reconstructed intra prediction image blocks and / or residual blocks and the corresponding inter prediction data are forwarded to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 and the in-loop filter component 425 can be substantially similar to the decoded image buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks, and / or prediction blocks. This information is stored in the decoded image buffer component 423. The reconstructed image blocks from the decoded image buffer component 423 are forwarded to the motion compensation component 421 for inter prediction. The motion compensation component 421 can be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses the motion vectors of the reference blocks to generate prediction blocks and applies the residual blocks to the above result to reconstruct the image blocks. The obtained reconstructed blocks can also be forwarded to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 continues to store other reconstructed image blocks. These reconstructed image blocks can be reconstructed into frames through the segmentation information. These frames can also be placed in a sequence. This sequence is output as a reconstructed output video signal to the display.

[0140] Figure 5This is an example of the point cloud media 500 that can be decoded according to the PCC mechanism. Thus, when the method 100 is executed, the point cloud media 500 can be encoded by an encoder (such as the codec system 200 and / or the encoder 300) and reconstructed by a decoder (such as the codec system 200 and / or the decoder 400).

[0141] Figures 1 to 4 The mechanisms described in to Figures 1 to 4 generally assume that 2D frames are being decoded. However, the point cloud media 500 is a point cloud that changes over time. Specifically, the point cloud media 500 (which can also be referred to as a point cloud and / or a point cloud representation) is a set of points in 3D space. These points can also be called samples. Each point can be associated with multiple types of data. For example, each point can be described by a position. The position is a location in 3D space and can be described as a set of Cartesian coordinates. Additionally, each point can include a color. The color can be described by luminance (e.g., light) and chrominance (e.g., color). The color can be described by red (R), green (G), and blue (B) values (represented as (R, G, B)) or luminance (Y), blue projection (U), and red projection (V) (represented as (Y, U, V)). These points can also include other attributes. An attribute is an optional scalar or vector characteristic that can be associated with each point in the point cloud. Attributes can include reflectivity, transparency, surface normal, timestamp, material ID.

[0142] Since each point in the point cloud media 500 can be associated with multiple types of data, therefore, according to Figures 1 to 4 the mechanisms described in to Figures 1 to 4 , several support mechanisms are used to prepare the point cloud media 500 for compression. For example, the point cloud media 500 can be classified into frames, where each frame includes all the data related to the point cloud for a specific state or a certain moment. Thus, Figure 5 a single frame of the point cloud media 500 is described. Then, the point cloud media 500 is decoded frame by frame. The point cloud media 500 can be surrounded by a 3D bounding box 501. The 3D bounding box 501 is a 3D rectangular prism whose dimensions are designed to enclose all the points of the point cloud media 500 for the corresponding frame. It should be noted that in the case where the point cloud media 500 includes a disjoint set, multiple 3D bounding boxes 501 can be employed. For example, the point cloud media 500 can describe two unconnected shapes, in which case the 3D bounding boxes 501 will be placed around each shape. The points in the 3D bounding box 501 are processed as described below.

[0143] Figure 6An example of a slice 603 created from the point cloud 600. The point cloud 600 is a single frame of the point cloud media 500. In addition, the point cloud 600 is surrounded by a 3D bounding box 601 that is substantially similar to the 3D bounding box 501. Thus, when the method 100 is executed, the point cloud 600 can be encoded by an encoder (such as the codec system 200 and / or the encoder 300) and reconstructed by a decoder (such as the codec system 200 and / or the decoder 400).

[0144] The 3D bounding box 601 includes six faces and thus includes six 2D rectangular frames 602, each 2D rectangular frame 602 being positioned on one face of the 3D bounding box 601 (e.g., the top face, the bottom face, the left face, the right face, the front face, and the back face). By projecting the point cloud 600 onto the corresponding 2D rectangular frame 602, the point cloud 600 can be converted from 3D data to 2D data. In this way, the slice 603 is created. The slice 603 is a 2D representation of the 3D point cloud, where the slice 603 includes a representation of the point cloud 600 visible from the corresponding 2D rectangular frame 602. It should be noted that the representation of the point cloud 600 from the 2D rectangular frame 602 may include multiple non-overlapping components. Thus, the 2D rectangular frame 602 may include multiple slices 603. Therefore, the point cloud 600 can be represented by more than six slices 603. The slice 603 may also be referred to as an atlas, atlas data, atlas information, and / or atlas component. By converting the 3D data into a 2D format, the point cloud 600 can be encoded according to video coding mechanisms such as inter-frame prediction and / or intra-frame prediction.

[0145] Figures 7A to 7C Shows a mechanism for encoding a 3D point cloud that has been converted to 2D information as described in Figure 6 Specifically, Figure 7A Shows an exemplary occupancy frame 710 associated with a set of slices (such as slice 603). The occupancy frame 710 is decoded in binary form. For example, zero indicates that a portion of the bounding box 601 is not occupied by one of the slices 603. Those portions of the bounding box 601 represented by zero do not participate in the reconstruction of the volume representation (e.g., the point cloud 600). In contrast, one indicates that a portion of the bounding box 601 is occupied by one of the slices 603. Those portions of the bounding box 601 represented by one participate in the reconstruction of the volume representation (e.g., the point cloud 600). In addition, Figure 7B Shows an exemplary geometry frame 720 associated with a set of slices (such as slice 603). The geometry frame 720 provides or describes the contour or topographic map of each slice 603. Specifically, the geometry frame 720 represents the distance of each point in the slice 603 from the planar surface of the bounding box 601 (e.g., the 2D rectangular frame 602). In addition, Figure 7CAn exemplary atlas frame 730 associated with a set of tiles (e.g., tile 603) is shown. The atlas frame 730 provides or describes samples of the tile 603 within the bounding box 601. The atlas frame 730 may include, for example, color components of points in the tile 603. The color components may be based on the RGB color model or other color models. The occupancy frame 710, the geometry frame 720, and the atlas frame 730 may be used to decode the point cloud 600 and / or the point cloud media 500. Thus, when the method 100 is executed, the occupancy frame 710, the geometry frame 720, and the atlas frame 730 may be encoded by an encoder (e.g., the codec system 200 and / or the encoder 300) and reconstructed by a decoder (e.g., the codec system 200 and / or the decoder 400).

[0146] Various tiles created by projecting 3D information onto a 2D plane can be packed into a rectangular (or square) video frame. This approach can be advantageous because various video codecs are preconfigured to decode such video frames. Thus, the PCC codec can use other video codecs to decode in tile blocks. As Figure 7A shown, tiles can be packed into a frame. The tiles can be packed by any algorithm. For example, tiles can be packed into the frame according to size. In a particular example, the tiles are included from largest to smallest. The largest tile can be placed first in any open space, and once a size threshold is exceeded, smaller tiles fill in the gaps. As Figure 7A shown, this packing scheme results in blank spaces that do not include tile data. To avoid encoding the blank spaces, the occupancy frame 710 is used. The occupancy frame 710 includes all occupancy data of the point cloud at a particular moment. Specifically, the occupancy frame 710 includes one or more occupancy maps (also referred to as occupancy data, occupancy information, and / or occupancy components). The occupancy map is defined as a 2D array corresponding to an atlas (tile set), and the values of the atlas indicate whether each sample position in the atlas corresponds to a valid 3D point in the point cloud representation. As Figure 7A shown, the occupancy map includes a valid data region 713. The valid data region 713 indicates the presence of the atlas / tile data in the corresponding position within the occupancy frame 710. The occupancy map also includes an invalid data region 715. The invalid data region 715 indicates the absence of the atlas / tile data in the corresponding position within the occupancy frame 710.

[0147] Figure 7BDescribes the geometric frame 720 of point cloud data. The geometric frame 720 includes one or more geometric maps 723 (also referred to as geometric data, geometric information, and / or geometric components) of the point cloud at a specific moment. The geometric map 723 is a 2D array created by aggregating the geometric information associated with each slice, where the geometric information / data is a set of Cartesian coordinates associated with the point cloud frame. Specifically, the slices are all projected from points in 3D space. This projection has the effect of removing 3D information from the slices. The geometric map 723 retains the 3D information removed from the slices. For example, each sample in the slice is obtained from a point in 3D space. Therefore, the geometric map 723 can include the 3D coordinates associated with each sample in each slice. Thus, the geometric map 723 can be used by the decoder to map / transform the 2D slices back to 3D space to reconstruct the 3D point cloud. Specifically, the decoder can map each slice sample to the appropriate 3D coordinates to reconstruct the point cloud.

[0148] Figure 7C Describes the atlas frame 730 of point cloud data. The atlas frame 730 includes one or more atlases 733 (also referred to as atlas data, atlas information, atlas components, and / or slices) of the point cloud at a specific moment. The atlases 733 are a set of 2D bounding boxes projected into a rectangular frame corresponding to a 3D bounding box in 3D space, where each 2D bounding box / slice represents a subset of the point cloud. Specifically, the atlases 733 include the slices created when the 3D point cloud as described Figure 6 is projected into 2D space. Thus, the atlases 733 / slices include the image data (e.g., color and light values) associated with the point cloud and the corresponding moment. The atlases 733 correspond to Figure 7A the occupancy map and Figure 7B the geometric map 723. Specifically, the atlases 733 include the data in the valid data region 713 and do not include the data in the invalid data region 715. In addition, the geometric map 723 includes the 3D information of the samples in the atlases 733.

[0149] It should also be noted that the point cloud can include attributes (also referred to as attribute data, attribute information, and / or attribute components). These attributes can be included in the atlas frame. The atlas frame can include all the data regarding the corresponding attributes of the point cloud at a specific moment. Examples of the attribute frame are not shown because the attributes can include a wide variety of different data. Specifically, the attributes can be any scalar or vector specific associated with each point in the point cloud, such as reflectivity, surface normal, timestamp, material ID, etc. In addition, the attributes are optional (e.g., user-defined) and can vary according to the application. However, when used, the point cloud attributes can be included in the attribute frame in a manner similar to the atlases 733, geometric map 723, and occupancy map.

[0150] Thus, the encoder can compress the point cloud frame into an atlas frame 730 of the atlas 733, a geometry frame 720 of the geometry map 723, an occupancy frame 710 of the occupancy map, and (optionally) an attribute frame of the attributes. The atlas frame 730, the geometry frame 720, the occupancy frame 710, and / or the attribute frame can be further compressed, for example, by different encoders, for transmission to the decoder. The decoder can decompress the atlas frame 730, the geometry frame 720, the occupancy frame 710, and / or the attribute frame. Then, the decoder can use the atlas frame 730, the geometry frame 720, the occupancy frame 710, and / or the attribute frame to reconstruct the point cloud frame to determine the reconstructed point cloud at the corresponding moment. Then, the reconstructed point cloud frames can be sequentially included to reconstruct the original point cloud sequence (e.g., for display and / or for data analysis). In a specific example, it can be achieved by using the techniques described in conjunction with Figures 1 to 4 to encode and decode the atlas frame 730 and / or the atlas 733 (e.g., by using VVC, HEVC, and / or AVC codecs).

[0151] Figure 8 is a schematic diagram of an exemplary compliance test mechanism 800. The compliance test mechanism 800 can be used by an encoder (e.g., the codec system 200 and / or the encoder 300) to verify whether the PCC bitstream complies with the standard, and thus, the PCC bitstream can be decoded by a decoder (e.g., the codec system 200 and / or the decoder 400). For example, the compliance test mechanism 800 can be used to check whether the point cloud media 500 and / or the slice 603 have been encoded as the occupancy frame 710, the geometry frame 720, the atlas frame 730, and / or the attribute frame in a manner that can be correctly decoded when performing the method 100.

[0152] The compliance test mechanism 800 can test whether the PCC bitstream complies with the standard. A PCC bitstream that complies with the standard should always be decodable by any decoder that also complies with the standard. A PCC bitstream that does not comply with the standard may not be decodable. Therefore, a PCC bitstream that fails the compliance test mechanism 800 should be re-encoded, for example, by using different settings. The compliance test mechanism 800 includes a type-I compliance test 881 and a type-II compliance test 883, which can also be referred to as compliance point A and compliance point B, respectively. The type-I compliance test 881 checks the compliance of the components of the PCC bitstream. The type-II compliance test 883 checks the compliance of the reconstructed point cloud. The encoder is typically used to perform the type-I compliance test 881 and can optionally perform the type-II compliance test 883.

[0153] Before performing the compliance test mechanism 800, the encoder encodes the compressed V-PCC bitstream 801 as described above. Then, the encoder can perform the compliance test mechanism 800 on the compressed V-PCC bitstream 801 using the HRD. The compliance test mechanism 800 divides the compressed V-PCC bitstream 801 into components. Specifically, the compressed V-PCC bitstream 801 is divided into a compressed atlas sub-bitstream 830, a compressed occupancy map sub-bitstream 810, a compressed geometry sub-bitstream 820, and an (optional) compressed attribute sub-bitstream 840. Each of the above compressed sub-bitstreams includes a sequence of encoded atlas frames 730, encoded geometry frames 720, occupancy frames 710, and (optional) attribute frames respectively.

[0154] Entropy decompression or video decompression 860 is performed on the sub-streams. Entropy decompression or video decompression 860 is a mechanism opposite to the component-specific compression. The compressed atlas sub-bitstream 830, the compressed occupancy map sub-bitstream 810, the compressed geometry sub-bitstream 820, and the compressed attribute sub-bitstream 840 may be encoded by one or more codecs. Therefore, entropy decompression or video decompression 860 includes applying a decoder to each sub-bitstream based on the encoder used to create the corresponding sub-bitstream. Entropy decompression or video decompression 860 reconstructs the decompressed atlas sub-bitstream 831, the decompressed occupancy map sub-bitstream 811, the decompressed geometry sub-bitstream 821, and the decompressed attribute sub-bitstream 841. Each of the above decompressed self-bitstreams is from the compressed atlas sub-bitstream 830, the compressed occupancy map sub-bitstream 810, the compressed geometry sub-bitstream 820, and the compressed attribute sub-bitstream 840 respectively. The decompressed sub-bitstream / component is the data from the sub-bitstream that has been reconstructed as part of the decoding process or in this case as part of the HRD compliance test.

[0155] Apply a Type-I compliance test 881 to the decompressed atlas sub-bitstream 831, decompressed occupancy map sub-bitstream 811, decompressed geometry sub-bitstream 821, and decompressed attribute sub-bitstream 841. The Type-I compliance test 881 examines each component (the decompressed atlas sub-bitstream 831, decompressed occupancy map sub-bitstream 811, decompressed geometry sub-bitstream 821, and decompressed attribute sub-bitstream 841) to ensure that the corresponding component complies with the standards used by the codec for encoding and decoding that component. For example, the Type-I compliance test 881 can verify whether the hardware resources of the normalization quantity can decompress the corresponding component without buffer overflow or underflow. In addition, the Type-I compliance test 881 can check that the component's blocking HRD correctly reconstructs the encoding errors of the corresponding component. In addition, the Type-I compliance test 881 can examine each corresponding component to ensure that all standard requirements are met and all standard prohibitions are omitted. When all components pass the corresponding tests, the Type-I compliance test 881 is satisfied, and when any one component fails the corresponding test, the Type-I compliance test 881 is not satisfied. Any component that passes the Type-I compliance test 881 should be decodable on any decoder that also complies with the corresponding standards. Therefore, the Type-I compliance test 881 can be used when encoding the compressed V-PCC bitstream 801.

[0156] Although the Type-I compliance test 881 ensures that the components are decodable, the Type-I compliance test 881 does not guarantee that the decoder can reconstruct the original point cloud from the corresponding components. Therefore, the compliance test mechanism 800 can also be used to perform a Type-II compliance test 883. The decompressed occupancy map sub-bitstream 811, decompressed geometry sub-bitstream 821, and decompressed attribute sub-bitstream 841 are forwarded for transformation 861. Specifically, the transformation 861 can transform the chroma format, resolution, and / or frame rate of the decompressed occupancy map sub-bitstream 811, decompressed geometry sub-bitstream 821, and decompressed attribute sub-bitstream 841 as needed to match the chroma format, resolution, and / or frame rate of the decompressed atlas sub-bitstream 831.

[0157] The result of transformation 861 and the decompressed atlas sub-bitstream 831 are forwarded to geometric reconstruction 862. At geometric reconstruction 862, the occupancy map from the decompressed occupancy map sub-bitstream 811 is used to determine the location of valid atlas data. Then, geometric reconstruction 862 can obtain geometric data from the decompressed geometric sub-bitstream 821 from any location that includes the valid atlas data. Then, a rough point cloud can be reconstructed using the geometric data, and the reconstructed point cloud is forwarded to duplicate point removal 863. For example, during the creation of 2D slices from a 3D cloud, some cloud points can be viewed from multiple directions. When this occurs, the same point will be projected as a sample into multiple slices. Then geometric data is generated based on the samples, so the geometric data includes duplicate data for these points. When the geometric data indicates that multiple points are located at the same position, duplicate point removal 863 merges such duplicate data to create a single point. The resulting reconstructed geometric map 871 is a mirror image of the geometric map of the original encoded point cloud. Specifically, the reconstructed geometric map 871 includes the 3D position of each point from the encoded point cloud.

[0158] The reconstructed geometric map 871 is forwarded for smoothing 864. Specifically, the reconstructed geometric map 871 may include certain features that appear sharp due to noise generated during the encoding process. Smoothing 864 can use one or more filters to remove such noise in order to create a smoothed geometric map 873, which is an exact representation of the original encoded point cloud. Then, the smoothed geometric map 873, together with the atlas data from the decompressed atlas sub-bitstream 831 and the attribute data from transformation 861, is forwarded to attribute reconstruction 865. Attribute reconstruction 865 colors the points located in the smoothed geometric map 873 using the colors from the atlas / slice data. Attribute reconstruction 865 also applies any attributes to the points. The resulting reconstructed cloud 875 is a mirror image of the original encoded point cloud. The reconstructed cloud 875 may include color or other attribute noise caused by the encoding process. Therefore, the reconstructed cloud 875 is forwarded for color smoothing 866, which applies one or more filters to the luminance, chrominance, or other attribute values to smooth such noise. Color smoothing 866 can then output the reconstructed point cloud 877. If lossless encoding is used, the reconstructed point cloud 877 should be an exact representation of the original encoded point cloud. Otherwise, the reconstructed point cloud 877 is extremely close to the original encoded point cloud, with a difference not exceeding a predefined tolerance.

[0159] Apply a type-II compliance test 883 to the reconstructed point cloud 877. The type-II compliance test 883 examines the reconstructed point cloud 877 to ensure that the reconstructed point cloud 877 complies with the V-PCC standard and thus can be decoded by a decoder compliant with the V-PCC standard. For example, the type-II compliance test 883 can verify whether the hardware resources of a standardized amount can reconstruct the reconstructed point cloud 877 without buffer overflow or underflow. In addition, the type-II compliance test 883 can check for coding errors in the reconstructed point cloud 877 that prevent the HRD from correctly reconstructing the reconstructed point cloud 877. In addition, the type-II compliance test 883 can check each decompressed component and / or any intermediate data to ensure that all standard requirements are met and all standard prohibitions are omitted. When the reconstructed point cloud 877 and any intermediate components pass the corresponding tests, the type-II compliance test 883 is satisfied; when the reconstructed point cloud 877 or any intermediate component fails the corresponding test, the type-II compliance test 883 is not satisfied. When the reconstructed point cloud 877 passes the type-II compliance test 883, the reconstructed point cloud 877 should be decodable on any decoder that also complies with the V-PCC standard. Therefore, compared with the type-I compliance test 881, the type-II compliance test 883 can provide a more robust verification of the compressed v-PCC bitstream 801.

[0160] Figure 9 FIG. is a schematic diagram of an exemplary HRD 900 for performing a compliance test on a PCC bitstream, for example, by using a compliance test mechanism 800. The PCC bitstream may include point cloud media 500 and / or slices 603 encoded as occupancy frames 710, geometry frames 720, atlas frames 730, and / or attribute frames. Thus, the HRD 900 can be used by the codec system 200 and / or the encoder 300 as part of a method 100, where the codec system 200 and / or the encoder 300 encode the bitstream for decoding by the codec system 200 and / or the decoder 400. Specifically, the HRD 900 can examine the PCC bitstream and / or its components before the PCC bitstream is forwarded to the decoder. In some examples, when encoding the PCC bitstream, the PCC bitstream can be continuously forwarded through the HRD 900. If a part of the PCC bitstream does not meet the relevant constraints, the HRD 900 will indicate such non-compliance to the encoder, which causes the encoder to re-encode the corresponding part of the bitstream using a different mechanism. In some examples, the HRD 900 can be used to perform checks on atlas sub-bitstreams and / or reconstructed point clouds. In some examples, the occupancy map component, geometry component, and attribute component can be encoded by other codecs. Thus, sub-bitstreams including the occupancy map component, geometry component, and attribute component can be examined by other HRDs. Therefore, in some examples, multiple HRDs including the HRD 900 can be used to fully check the compliance of the PCC bitstream.

[0161] The HRD 900 includes the HSS 941. The HSS 941 is a component for performing a hypothetical transfer mechanism. The hypothetical transfer mechanism is used to check the compliance of a bitstream, a sub-bitstream, and / or a decoder with reference to the time and data stream of the PCC bitstream 951 input to the HRD 900. For example, the HSS 941 may receive the PCC bitstream 951 or its sub-bitstream output from an encoder. Then, the HSS 941 may manage the compliance test process for the PCC bitstream 951 by using, for example, the compliance test mechanism 800. In a specific example, the HSS 941 may control the rate at which the coded atlas data moves through the HRD 900 and verify that the PCC bitstream 951 does not include non-compliant data. The HSS 941 may forward the PCC bitstream 951 to the CAB 943 at a predefined rate. To implement the HRD 900, any unit (e.g., AU and / or NAL unit) including the coded video in the PCC bitstream 951 may be referred to as a decoded atlas unit 953. In some examples, the decoded atlas unit 953 may include only atlas data. In other examples, the decoded atlas unit 953 may include other PCC components and / or a set of data for reconstructing the point cloud. Thus, in the same example, the decoded atlas unit 953 may generally be referred to as a decoded unit. The CAB 943 is a FIFO buffer in the HRD 900. The CAB 943 includes the decoded atlas unit 953, which includes atlas data, geometry data, occupancy data, and / or attribute data arranged in decoding order. The CAB 943 stores such data for use during the PCC bitstream compliance test / check.

[0162] The CAB 943 forwards the decoded atlas unit 953 to the decoding process component 945. The decoding process component 945 is a component that complies with the PCC standard or other standards for decoding the PCC bitstream and / or its sub-bitstream. For example, the decoding process component 945 may simulate the decoder used by an end user. For example, the decoding process component 945 may perform a type-I compliance test through the decoded atlas component and / or a type-II compliance test by reconstructing the point cloud data. The decoding process component 945 decodes the decoded atlas unit 953 at a rate achievable by an exemplary standardized decoder. If the decoding process component 945 cannot decode the decoded atlas unit 953 fast enough to prevent the CAB 943 from overflowing, the PCC bitstream 951 does not meet the standard and should be re-encoded. Similarly, if the decoding process component 945 decodes the decoded atlas unit 953 too fast and the CAB 943 runs out of data (e.g., buffer underflow), the PCC bitstream 951 does not meet the standard and should be re-encoded.

[0163] The decoding process component 945 decodes the decoded atlas unit 953 to create the decoded atlas frame 955. The decoded atlas frame 955 may include a complete atlas data set of a PCC frame in the case of a type I compliance test or a frame of a reconstructed point cloud in the context of a type II compliance test. The decoded atlas frame 955 is forwarded to the DAB 947. The DAB 947 is a FIFO buffer in the HRD 900 that includes decoded atlas frames / decompressed atlas frames and / or reconstructed point cloud frames (based on context) arranged in decoding order for use during PCC bitstream compliance testing. The DAB 947 may be substantially similar to the decoded picture buffer components 223, 323, and / or 423. To support inter-frame prediction, the frames identified to be used as reference atlas frames 956 obtained from the decoded atlas frame 955 are returned to the decoding process component 945 to support further decoding. The DAB 947 outputs the atlas data 957 (or reconstructed point cloud, based on context) frame by frame. Thus, the HRD 900 can determine whether the decoding is satisfactory and whether the PCC bitstream 951 and / or its components meet the constraints.

[0164] Figure 10 is a schematic diagram showing an exemplary PCC bitstream 1000 for use in initializing an HRD (e.g., HRD 900) to support HRD compliance testing (e.g., compliance test mechanism 800). For example, the bitstream 1000 may be generated by the codec system 200 and / or the encoder 300 for decoding by the codec system 200 and / or the decoder 400 according to the method 100. Additionally, the bitstream 1000 may include point cloud media 500 and / or slices 603 encoded as occupancy frames 710, geometry frames 720, atlas frames 730, and / or attribute frames. Additionally, the bitstream 1000 may be checked for compliance by the HRD (e.g., HRD 900) using a compliance test mechanism (e.g., compliance test mechanism 800).

[0165] The PCC bitstream 1000 includes a sequence of PCC AUs 1010. The PCC AU 1010 includes sufficient components to reconstruct a single PCC frame captured at a particular moment. For example, the PCC AU 1010 may include an atlas frame 1011, an occupancy map frame 1013, and a geometry map frame 1015 that may be substantially similar to the atlas frame 730, the occupancy frame 710, and the geometry frame 720 in the atlas, respectively. The PCC AU 1010 may also include an attribute frame 1017 that includes all the attributes related to the point cloud at a certain moment encoded in the PCC AU 1010. Such attributes may include scalar or vector characteristics optionally associated with each point in the point cloud, such as color, reflectance, surface normal, timestamp, material ID, etc. The PCC AU 1010 may be defined as a set of NAL units that are associated with each other according to specified classification rules and related to a specific output time. Thus, the data is located in the PCC AU 1010 in the NAL units. The NAL unit is a data container of a packet size. For example, the size of a single NAL unit is typically designed to enable network transmission. The NAL unit may include a header representing the NAL unit type and a payload including the related data.

[0166] The PCC bitstream 1000 also includes various data structures to support the decoding of the PCC AUs 1010, such as as part of the decoding process and / or as part of the HRD process. For example, the PCC bitstream 1000 may include various parameter sets having parameters for decoding one or more PCC AUs 1010. In a specific example, the PCC bitstream 1000 may include an atlas SPS 1020. The atlas SPS 1020 is a syntax structure that includes syntax elements applied to zero or more complete encoded atlas sequences, and the content of these syntax elements is determined by the syntax elements found in the atlas SPS 1020 and referenced by the syntax elements found in each tile group header. For example, the atlas SPS 1020 may include parameters related to the entire sequence of the atlas frame 1011.

[0167] The PCC bitstream 1000 also includes various SEI messages. The SEI message is a syntax structure with a specified semantics that conveys information not required for the decoding process to determine the values of the samples in the decoded image. Thus, the SEI message can be used to transmit data not directly related to the decoding of the PCC AU 1010. In the example shown, the PCC bitstream 1000 includes a buffering period SEI message 1030 and an atlas frame timing SEI message 1040.

[0168] In the example shown, when performing a compliance test on the PCC bitstream 1000, the picture set SPS 1020, the buffering period SEI message 1030, and the picture set frame timing SEI message 1040 are used to initialize and manage the functions of the HRD. For example, the HRD parameter 1021 may be included in the picture set SPS 1020. The HRD parameter 1021 is a syntax element that initializes and / or defines the operating conditions of the HRD. For example, the HRD parameter 1021 can be used to specify the compliance points for the HRD compliance check performed at the HRD, such as the type I compliance test 881 or the type II compliance test 883. Therefore, the HRD parameter 1021 can be used to indicate whether the HRD compliance check should be performed on the decompressed PCC components or the reconstructed point cloud. For example, the HRD parameter 1021 can be set to a first value to indicate that the HRD compliance check should be performed on the decompressed attribute component, the decompressed picture set component, the decompressed occupancy map component, and the decompressed geometry component (e.g., the attribute frame 1017, the picture set frame 1011, the occupancy map frame 1013, and the geometry map frame 1015, respectively). In addition, the HRD parameter 1021 can be set to a first value to indicate that the HRD compliance check should be performed on the reconstructed point cloud generated from the PCC components (e.g., reconstructed from the entire PCC AU 1010).

[0169] The buffering period SEI message 1030 is an SEI message that includes data representing the initial deletion delay associated with the CAB in the HRD (e.g., CAB 943). The initial CAB deletion delay is the amount of time that a component in the first AU in the bitstream (e.g., PCC AU 1010) or the first AU in the sub-bitstream (e.g., picture set frame 1011) can remain in the CAB before being deleted. For example, the HRD can start deleting any decoding units associated with the first PCC AU 1010 from the CAB in the HRD during the HRD compliance check according to the initial delay specified in the buffering period SEI message 1030. Therefore, the buffering period SEI message 1030 includes data sufficient to initialize the HRD compliance test process to start the HRD compliance test at the encoded PCC AU 1010 associated with the buffering period SEI message 1030. Specifically, the buffering period SEI message 1030 can indicate that the HRD should start the compliance test at the first PCC AU 1010 in the PCC bitstream 1000.

[0170] The picture set frame timing SEI message 1040 is an SEI message that includes data representing the deletion delay associated with a CAB (e.g., CAB 943) and the output delay associated with a DAB (e.g., DAB 947) in the HRD. The CAB deletion delay is the amount of time that a component (e.g., any corresponding component) can be retained in the CAB before deletion. The CAB deletion delay can be encoded with reference to the initial CAB deletion delay specified by the buffering period SEI message 1030. The DAB output delay is the amount of time that a decompressed component / decoded component (e.g., any corresponding component) can be retained in the DAB before output (e.g., as part of the reconstructed point cloud). Thus, the HRD can delete decoding units from the CAB in the HRD during compliance checking as specified by the picture set frame timing SEI message 1040. In addition, the HRD can set the output delay of the DAB in the HRD as specified by the picture set frame timing SEI message 1040.

[0171] Thus, the encoder can encode the HRD parameters 1021, the buffering period SEI message 1030, and the picture set frame timing SEI message 1040 into the PCC bitstream 1000 during the encoding process. Then, the HRD can read the HRD parameters 1021, the buffering period SEI message 1030, and the picture set frame timing SEI message 1040 to obtain sufficient information to perform a compliance check on the PCC bitstream 1000, such as the compliance test mechanism 800. In addition, the decoder obtains the HRD parameters 1021, the buffering period SEI message 1030, and / or the picture set frame timing SEI message 1040 from the PCC bitstream 1000 and infers that an HRD check has been performed on the PCC bitstream 1000 by the presence of this data. Thus, the decoder can infer that the PCC bitstream 1000 is decodable and can thus decode the PCC bitstream 1000 according to the HRD parameters 1021, the buffering period SEI message 1030, and / or the picture set frame timing SEI message 1040.

[0172] The PCC bitstream 1000 can have different sizes and can be transmitted from an encoder to a decoder at different rates over a transmission network. For example, when using an HEVC-based encoder, a volumetric sequence with a length of about one hour can be encoded into a PCC bitstream 1000 with a file size between 15 and 70 gigabytes. Compared with an HEVC encoder, a VVC-based encoder can further reduce the file size by about 30% to 35%. Thus, a volumetric sequence with a length of one hour encoded with a VVC encoder may result in a file with a size of about 10 to 49 gigabytes. The PCC bitstream 1000 can be transmitted at different rates according to the state of the transmission network. For example, the PCC bitstream 1000 can be transmitted over the network at a bitrate of 5 to 20 megabytes per second. Similarly, the encoding and decoding processes described herein can be performed, for example, at a rate faster than one megabyte per second.

[0173] Figure 11 is a schematic diagram of an exemplary video decoding device 1100. The video decoding device 1100 is suitable for implementing the disclosed embodiments. The video decoding device 1100 includes a downlink port 1120, an uplink port 1150, and a TX / RX 1110, which includes a transmitter or a transmitting module and includes a receiver or a receiving module for transmitting data over a network. The video decoding device 1100 also includes a processor 1130 or a processing module, which includes a logic unit or a CPU for processing data and a memory 1132 or a storage module for storing data. The video decoding device 1100 may also include electronic components, OE components, EO components, or wireless communication components coupled to the uplink port 1150 or the downlink port 1120 for transmitting data over an electrical, optical, or wireless communication network. The video decoding device 1100 may also include an I / O device 1160 for data communication with a user. The I / O device 1160 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 1160 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0174] The processor 1130 is implemented by hardware and software. The processor 1130 can be implemented as one or more CPU chips, one or more cores (e.g., multi-core processors), one or more FPGAs, one or more ASICs, and one or more DSPs. The processor 1130 communicates with the downlink port 1120, Tx / Rx 1110, the uplink port 1150, and the memory 1132. The processor 1130 includes a decoding module 1114. The decoding module 1114 implements the disclosed embodiments described herein (e.g., methods 100, 1200, and 1300), which may use the point cloud media 500 divided into a group of slices 603 and encoded into the occupancy frames 710, geometric frames 720, and atlas frames 730 in the PCC bitstream 1000. In addition, the decoding module 1114 can implement the HRD 900, which performs the compliance testing mechanism 800 on the PCC bitstream 1000. The decoding module 1114 can also implement any other methods / mechanisms described herein. In addition, the decoding module 1114 can implement the codec system 200, the encoder 300, and / or the decoder 400. Alternatively, the decoding module 1114 can be implemented as instructions stored in the memory 1132 and executed by the processor 1130 (e.g., implemented as a computer program product stored on a non-transitory medium).

[0175] The memory 1132 includes one or more types of memories, such as magnetic disks, tape drives, solid-state drives, ROM, RAM, flash memory, TCAM, SRAM, etc. The memory 1132 can be used as an overflow data storage device to store these programs when a program is selected for execution, as well as the instructions and data read during the execution of the program.

[0176] A point cloud is a volumetric representation of space on a regular 3D grid. The voxels in the point cloud have x, y, and z coordinates and can have RGB color components, reflectivity, or other attributes. The data representation in V-PCC depends on the 3D to 2D transformation and is described as a set of planar 2D images with four types of data, where these four types of data are called components: occupancy map, geometric data, attribute data, and atlas frame. The occupancy map is a binary image representation of occupied or unoccupied blocks in the 2D projection. The geometric data is a height map of the slice data, describing the per-point difference in the distance from the slice projection plane. The attribute data is a 2D texture map representing the corresponding component of the attribute value at the corresponding 3D point of the point cloud. The atlas frame is the metadata information required to perform the 2D to 3D transformation.

[0177] The atlas frame lacks the information required for V-PCC bitstream sub-component synchronization and buffering. Specifically, when implementing the buffering process without an appropriate buffering model, the decoded information must be stored in undefined memory locations. The size of the required memory may impose a limitation on the decoding device. Therefore, it is desirable to limit the size of the buffer memory required.

[0178] Embodiments of V-PCC component synchronization are disclosed. This embodiment provides a solution in which the decoded V-PCC components are output at a corresponding component decoder (referred to as CP A) and transmitted to a buffer, where the synchronization process is to prepare the data for reconstruction at CP B. Output delay synchronization is used in the synchronization process. Output delay synchronization improves synchronization, thereby reducing the buffer memory size.

[0179] Figure 12 It is a schematic diagram 1200 of component synchronization in point cloud reconstruction. The schematic diagram 1200 shows a V-PCC component 1210, which may include an occupancy map, geometric data, attribute data, and an atlas frame from a bitstream. After CP A, each V-PCC component 1210 is delayed for a period corresponding to the output delay synchronization 1220 of the V-PCC component 1210. Then, the V-PCC component 1210 is buffered in the V-PCC unit buffer 1230 and undergoes PCC reconstruction 1240. Finally, at CP B, the reconstructed point cloud 1250 is output.

[0180] When the number of pictures is 1 and when the number of pictures is N, the output delay synchronization 1220 can be calculated differently. When the number of pictures is 1, the output time delay of DAB / DPB in the picture / atlas in the timing SEI of access unit n is modified as follows:

[0181] Let vpccComponentNum be the number of V-PCC components, then:

[0182]

[0183] PicAtlasDpbOutputDelay[n] is the value of the i-th component of pic_dpb_output_delay / atlas_dpb_output_delay in the picture / atlas timing SEI message associated with access unit n.

[0184] ClockSubTick = ClockTick / (tick_divisor_minus2 + 2)

[0185] DpbDabOutputTime[n] can be similar to AuNominalRemovalTime, AuCpbCabRemovalTime[n] can be similar to AuNominalRemovalTime[firstAtlasInCurrBuffPeriod], PicAtlasDpbOutputDelay can be similar to AuCabRemovalDelayVal, and DpbDabDelayOffset can be similar to CabDelayOffset. Therefore, the above equation can be:

[0186] AuNominalRemovalTime[n] = AuNominalRemovalTime[firstAtlasInCurrBuffPeriod] + ClockTick * (AuCabRemovalDelayVal - CabDelayOffset),

[0187] where AuNominalRemovalTime[n] is the nominal removal time when access unit n is not the first access unit of the buffering period for access unit n to be removed from the CAB, AuNominalRemovalTime[firstAtlasInCurrBuffPeriod] is the nominal removal time of the first access unit of the current buffering period, AuCabRemovalDelayVal is the value of AuCabRemovalDelayVal derived from aft_cab_removal_delay_minus1 in the atlas timing SEI message associated with access unit n, and CabDelayOffset is set to be equal to the value of the buffering period SEI message syntax element bp_cab_delay_offset or set to 0.

[0188] Figure 13Schematic diagram 1300 shows the maximum delay calculation for a single picture. In schematic diagram 1300, component 0 has sequential coding, such that the frames of component 0 are cached in the same order as they are displayed. Specifically, the DPB caches frame 0, then caches frame 1, then caches frame 2, then caches frame 3, then caches frame 4, and finally caches frame 5. Similarly, the display device displays frame 0, then displays frame 1, then displays frame 2, then displays frame 3, then displays frame 4, and finally displays frame 5. Component 1 has out-of-order coding, so the frames of component 1 are cached in a different order than they are displayed. Specifically, the DPB caches frame 0, then caches frame 2, then caches frame 1, then caches frame 3, then caches frame 5, and finally caches frame 4. However, the display device can only start displaying when frame 2 is decoded, so there is a delay of 1 frame in the decoding order. After this delay, the display device displays frame 0, then displays frame 1, then displays frame 2, then displays frame 3, then displays frame 4, and finally displays frame 5. Component 2 also has out-of-order coding, so the frames of component 2 are cached in a different order than they are displayed. Specifically, the DPB caches frame 0, then caches frame 4, then caches frame 2, then caches frame 1, then caches frame 3, and finally caches frame 5. However, the display device can only start displaying when frame 4 is decoded, so there is a delay of 2 frames in the decoding order. After this delay, the display device displays frame 0, then displays frame 1, then displays frame 2, then displays frame 3, then displays frame 4, and finally displays frame 5.

[0189] The delay of component 2 is the largest, so it determines the maximum offset of 2 frames. The delay of component 0 is 0 frames, while the maximum offset is 2 frames, so an additional offset of 2 frames is added to component 0. The delay of component 1 is 1 frame, while the maximum offset is 2 frames, so an additional offset of 1 frame is added to component 1.

[0190] When the number of pictures is 1, the DAB / DPB output time delay in the picture / atlas at the timing SEI of access unit n is modified as follows:

[0191]

[0192] Figure 14 Schematic diagram 1400 shows the maximum delay calculation for multiple pictures. Schematic diagram 1400 is similar to schematic diagram 1300. Specifically, component 0 has sequential coding, while components 1 and 2 have out-of-order coding. In addition, the delay of component 2 is the largest. However, different from schematic diagram 1300, components 1 and 2 in schematic diagram 1400 have 2 pictures, namely picture 0 and Figure 1 . In addition, component 2 determines the maximum offset of 3 frames. In addition, the display device cannot display a frame until the decoder has decoded all the pictures of that frame.

[0193] When the number of pictures is N, the DAB / DPB output time delay in the picture / atlas at the timing SEI of access unit n is modified as follows:

[0194]

[0195] DpbDabDelayOffset is derived as follows:

[0196]

[0197]

[0198] DpbDabDelayOffset[n][i]=MaxInitialDelay-PicAtlasDpbOu in utDelay[n][i])

[0199] The first instance of MaxInitialDelay may be similar to removalDelay, and the second instance of MaxInitDelay may be similar to auCabRemovalDelayDelta. PicAtlasDpbOutputDelay may be similar to InitCabRemovalDelay.

[0200] Therefore, the above equation can be:

[0201] removalDelay=Max(auCabRemovalDelayDelta, Ceil((InitCabRemovalDelay[SchedSelIdx]÷90000+offsetTime)÷ClockTick),

[0202] Wherein, removalDelay is defined as shown, auCabRemovalDelayDelta is the value of the syntax element (bp_atlas_cab_removal_delay_delta_minusl+1) in the buffering period SEI message associated with access unit n, and InitCabRemovalDelayOffset[SchedSelIdx] is set equal to the value of the syntax element bp_nal_initial_alt_cab_removal_offset[SchedSelIdx] of the buffering period SEI message, or InitCabRemovalDelay[SchedSelIdx] is set equal to the value of the syntax element bp_nal_initial_cab_removal_delay[SchedSelIdx] of the buffering period SEI message.

[0203] Figure 15 It is a flowchart of a method 1500 for decoding a bitstream provided by the first embodiment. The decoder 400 can implement the method 1500. The decoder 400 can be a PCC decoder. In step 1510, a point cloud bitstream is received. In step 1520, caching is performed on the point cloud bitstream or point cloud bitstream components based on time. This execution includes determining time based on latency and latency offset. Finally, in step 1530, the point cloud bitstream is decoded based on the cache.

[0204] The method 1500 can implement additional embodiments. For example, time is also based on removal time. Time is also based on ClockTick. Time is also based on a first expression of latency and latency offset. An expression is a mathematical concept that describes a meaningful combination of components. For example, the first expression is (PicAtlasDpbOutputDelay[n][i] + DpbDabDelayOffset[n][i]) above. In other embodiments, the first expression is the difference between PicAtlasDpbOutputDelay[n][i] and DpbDabDelayOffset[n][i] or between similar components. Time is also based on a second expression, which is the product of ClockTick and the first expression. Time is also based on the sum of removal time and the second expression. The point cloud bitstream includes multiple components. The components include occupancy maps. The components include geometric data. The components include attribute data. The components include atlas frames. Time is also based on the number of components. Time is DpbDabOutputTime. Latency is PicAtlasDpbOutputDelay. Latency offset is DpbDabDelayOffset. DpbDabDelayOffset is equal to the difference between MaxInitialDelay and PicAtlasDpbOutputDelay.

[0205] Figure 16 It is a flowchart of a method 1600 for decoding a bitstream provided by the second embodiment. The decoder 400 can implement the method 1600. The decoder 400 can be a PCC decoder. In step 1610, a point cloud bitstream is received. In step 1620, caching is performed on the point cloud bitstream or point cloud bitstream components based on latency. The latency is based on a first latency and a second latency. Finally, in step 1630, the point cloud bitstream is decoded based on the cache.

[0206] Method 1600 may implement additional embodiments. For example, the delay is also based on the maximum of the first delay and the second delay. The delay is MaxInitialDelay. The first delay is MaxInitDelay. The second delay is PicAtlasDpbOutputDelay. The cache is also based on DpbDabDelayOffset, where DpbDabDelayOffset = MaxInitialDelay - PicAtlasDpbOutputDelay.

[0207] In one embodiment, a receiving module receives a point cloud bitstream. A processing module performs caching on the point cloud bitstream based on time. This performance includes determining time based on a delay and a delay offset. The processing module decodes the point cloud bitstream based on the cache.

[0208] Unless otherwise specified, the term "about" means a range including ±10% of the subsequent number. Although the present invention provides several embodiments, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present invention. The examples of the present invention should be considered illustrative rather than restrictive, and the present invention is not limited to the details given in this text. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0209] In addition, technologies, systems, subsystems, and methods described and shown as discrete or separate in various embodiments may be combined or integrated with other systems, components, technologies, or methods without departing from the scope of the present invention. What is shown or described as being coupled may be directly coupled or may be indirectly coupled or communicate through some interface, device, or intermediate component in an electrical, mechanical, or other manner. Other variations, substitutions, and change examples can be determined by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

1. A method implemented by a point cloud compression (PCC) decoder, characterized in that: It includes: The PCC decoder receives a point cloud bitstream including components; The PCC decoder performs buffering on the point cloud bitstream based on time, and the performing includes determining the time based on a delay and a delay offset; The PCC decoder decodes the point cloud bitstream based on the buffering.

2. The method according to claim 1, characterized in that The time is also based on a removal time.

3. The method according to claim 1, characterized in that, The time is also based on ClockTick.

4. The method according to claim 1, wherein The time is also based on a first expression of the delay and the delay offset.

5. The method according to claim 1, characterized in that The time is also based on a second expression, and the second expression is the product of ClockTick and the first expression.

6. The method according to claim 1, characterized in that The time is also based on the sum of the removal time and the second expression.

7. The method according to claim 1, wherein The point cloud bitstream includes multiple components.

8. The method according to claim 1, wherein The components include occupancy maps.

9. The method according to claim 1, wherein The components include geometric data.

10. The method according to claim 1, characterized in that The components include attribute data.

11. The method according to claim 1, characterized in that, The components include atlas frames.

12. The method according to claim 1, wherein The time is also based on the number of the components.

13. The method according to claim 1, wherein The time is DpbDabOutputTime.

14. The method according to claim 1, wherein The delay is PicAtlasDpbOutputDelay.

15. The method according to claim 1, wherein The delay offset is DpbDabDelayOffset.

16. The method according to claim 1, wherein DpbDabDelayOffset is equal to the difference between MaxInitialDelay and PicAtlasDpbOutputDelay.

17. The method according to any one of claims 1 to 16, characterized in that, It further includes: Storing the point cloud bitstream; Displaying an image or video from the point cloud bitstream.

18. A point cloud compression (PCC) decoder, characterized in that, It includes: A memory for storing instructions; A processor coupled to the memory and for executing the instructions to perform the method according to any one of claims 1 to 17.

19. A computer program product, characterized in that, The computer program product includes computer-executable instructions stored in a non-transitory medium; when executed by a processor, the computer-executable instructions cause a point cloud compression (PCC) decoder to perform the method according to any one of claims 1 to 17.

20. A method implemented by a point cloud compression (PCC) decoder, characterized in that: It includes: The PCC decoder receives a point cloud bitstream; The PCC decoder performs buffering on the point cloud bitstream based on a delay, and the delay is based on a first delay and a second delay; The PCC decoder decodes the point cloud bitstream based on the buffering.

21. The method according to claim 20, wherein The delay is also based on the maximum value of the first delay and the second delay.

22. The method according to claim 20, wherein The delay is MaxInitialDelay.

23. The method according to claim 20, wherein The first delay is MaxInitDelay.

24. The method according to claim 20, wherein The second delay is PicAtlasDpbOutputDelay.

25. The method according to claim 20, characterized in that The buffering is also based on DpbDabDelayOffset, and DpbDabDelayOffset = MaxInitialDelay – PicAtlasDpbOutputDelay.

26. The method according to any one of claims 20 to 25, characterized in that, It further includes: Store the point cloud bitstream; Display an image or video from the point cloud bitstream.

27. A point cloud compression (PCC) decoder, characterized in that, Comprising: A memory for storing instructions; A processor coupled to the memory and operative to execute the instructions to perform the method according to any one of claims 20 to 26.

28. A computer program product, characterized in that, The computer program product comprises computer-executable instructions stored in a non-transitory medium; when executed by a processor, the computer-executable instructions cause a point cloud compression (PCC) decoder to perform the method according to any one of claims 20 to 26.

29. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when run on an electronic device, cause the electronic device to perform the method according to any one of claims 1-17 or 20-26.