Improved Attribute Layers and Signaling in Point Cloud Coding
By enabling different codecs to be used for various attributes within PCC video frames, the method addresses the inefficiencies in existing video coding technologies, achieving improved compression, resource usage, and video quality.
Patent Information
- Application Number
- JP2021514382
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-14
- Filing Date
- 2019-09-10
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2039-09-10
AI Technical Summary
Existing video coding technologies face challenges in efficiently compressing and decompressing point cloud coding (PCC) video frames, particularly in managing multiple attributes like texture, reflectance, transparency, and normal, which require different codecs for optimal compression.
The proposed solution involves a method where a video decoder receives a bitstream containing encoded PCC frames, parses it to identify the codecs used for each attribute, and decodes the frames accordingly. This allows different codecs to be employed for different PCC attributes within the same sequence of frames, enhancing coding flexibility and efficiency.
This approach optimizes processor resource usage, improves compression and coding efficiency, and reduces memory and network resource usage by allowing the most efficient codecs to be used for each attribute, thereby supporting higher video quality and more complex PCC frames.
Smart Images

Figure 0007693995000009 
Figure 0007693995000010 
Figure 0007693995000011
Abstract
Description
Technical Field
[0001] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 731,693, titled "High-Level Syntax Designs for Point Cloud Coding," filed on September 14, 2018, by Ye-Kui Wang et al., which is incorporated herein by reference.
[0002] The present disclosure generally relates to video coding, and more specifically to coding of video attributes for point cloud coding (PCC) video frames.
Background Art
[0003] The amount of video data required to represent relatively short videos is substantial and can pose difficulties when the data is streamed or otherwise communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. Also, the size of the video can be a problem when the video is stored on a storage device because memory resources may be limited. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio without sacrificing much of the image quality are desirable because network resources are limited and the demand for higher video quality is constantly increasing.
Summary of the Invention
[0004] According to one embodiment, the present disclosure includes a method implemented by a video decoder. The method includes receiving, by a receiver, a bitstream including an encoded sequence of a plurality of point cloud coding (PCC) frames, wherein the encoded sequence of the plurality of PCC frames is GeometryRepresents a plurality of PCC attributes including texture, and one or more of reflectance, transparency, and normal, and receiving each coded PCC frame represented by one or more PCC network abstraction layer (NAL) units. The method further includes parsing a bitstream by a processor to obtain an indication of one of a plurality of video coders (codecs) used to code the corresponding PCC attribute for each PCC attribute. The method further includes decoding the bitstream by the processor based on the indicated video codec for the PCC attribute. In some video coding systems, an entire sequence of PCC frames is coded using a single codec. A PCC frame may include a plurality of PCC attributes. Some video codecs may be more efficient at coding some PCC attributes than others. This embodiment allows different video codecs to code different PCC attributes for the same sequence of PCC frames. This embodiment also provides various syntax elements to support coding flexibility when PCC frames in a sequence use a plurality of PCC attributes (e.g., three or more). By providing more attributes, the encoder can code more complex PCC frames. Further, the decoder can decode and display more complex PCC frames. Further, by enabling different codecs to be employed for different attributes, the coding process can be optimized based on codec selection. This may reduce the amount of processor resources used at both the encoder and the decoder. Further, this may support increased compression and coding efficiency, which reduces memory usage and network resource usage while transmitting the bitstream between the encoder and the decoder.
[0005] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that each sequence of the PCC frame is associated with a sequence-level data unit that includes sequence-level parameters, and the sequence-level data unit includes a first syntax element indicating that a first attribute is encoded by a first video codec and that a second attribute is encoded by a second video codec.
[0006] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the first syntax element is an identified_codec_for_attribute element included in a group of frame headers in the bitstream.
[0007] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the first attribute is organized into a plurality of streams, and the second syntax element indicates stream membership for a data unit of the bitstream associated with the first attribute.
[0008] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the first attribute is organized into a plurality of layers, and the third syntax element indicates layer membership for a data unit of the bitstream associated with the first attribute.
[0009] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the second syntax element is a num_streams_for_attribute element included in a group of frame headers in the bitstream, and the third syntax element is a num_layers_for_attribute element included in a group of frame headers in the bitstream.
[0010] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the fourth syntax element indicates that the first of the plurality of layers includes data associated with an irregular point cloud.
[0011] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the fourth syntax element is the regular_points_flag element included in a group of frame headers in the bitstream.
[0012] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the bitstream is decoded into a decoded sequence of PCC frames, and further includes transferring, by a processor, the decoded sequence of PCC frames towards a display for presentation.
[0013] According to one embodiment, the present disclosure includes a method implemented in a video encoder. The method comprises encoding, by a processor, a plurality of PCC attributes of a sequence of PCC frames into a bitstream with a plurality of codec (codec), the plurality of PCC attributes being Geometry, including texture, and one or more of reflectance, transmittance, and normal, and each coded PCC frame is represented by one or more PCC network abstraction layer (NAL) units, including encoding. The method further includes encoding, by a processor, an indication of one of the video codecs used to encode the corresponding PCC attribute for each PCC attribute. The method further includes transmitting, by a transmitter, a bitstream towards a decoder. In some video coding systems, an entire sequence of PCC frames is encoded using a single codec. A PCC frame may include multiple PCC attributes. Some video codecs may be more efficient at encoding some PCC attributes than others. This embodiment enables different video codecs to encode different PCC attributes for the same sequence of PCC frames. This embodiment also provides various syntax elements to support coding flexibility when PCC frames within a sequence use multiple PCC attributes (e.g., three or more). By providing more attributes, the encoder can encode more complex PCC frames. Further, the decoder can decode and display more complex PCC frames. Further, by enabling different codecs to be employed for different attributes, the coding process can be optimized based on codec selection. This may reduce the amount of processor resources used at both the encoder and the decoder. Further, this may support increased compression and coding efficiency, which reduces memory usage and network resource usage while transmitting the bitstream between the encoder and the decoder.
[0014] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that each sequence of the PCC frame is associated with a sequence-level data unit that includes sequence-level parameters, and the sequence-level data unit includes a first syntax element indicating that a first PCC attribute has been encoded by a first video codec and that a second PCC attribute has been encoded by a second video codec.
[0015] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the first syntax element is an identified_codec_for_attribute element included in a group of frame headers in the bitstream.
[0016] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the first attribute is organized into a plurality of streams, and the second syntax element indicates stream membership for a data unit of the bitstream associated with the first attribute.
[0017] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the first attribute is organized into a plurality of layers, and the third syntax element indicates layer membership for a data unit of the bitstream associated with the first attribute.
[0018] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the second syntax element is a num_streams_for_attribute element included in a group of frame headers in the bitstream, and the third syntax element is a num_layers_for_attribute element included in a group of frame headers in the bitstream.
[0019] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the fourth syntactic element indicates that the first layer of the plurality of layers includes data associated with an irregular point cloud.
[0020] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the fourth syntactic element is the regular_points_flag element included in a group of frame headers in a bitstream.
[0021] In one embodiment, the present disclosure includes a processor, a receiver coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, and the transmitter are configured to execute the method of any of the foregoing aspects.
[0022] In one embodiment, the present disclosure includes a non-transitory computer-readable medium including a computer program product used by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that when executed by a processor, the video coding device executes the method of any of the foregoing aspects.
[0023] In one embodiment, the present disclosure is encoding a plurality of PCC attributes of a sequence of PCC frames into a bitstream with a plurality of coders / decoders (codecs), the plurality of PCC attributes Geometry, texture, and one or more of reflectance, transmittance, and normal, and each encoded PCC frame is represented by one or more PCC network abstraction layer (NAL) units, and includes an encoder including first attribute encoding means and second attribute encoding means for performing encoding. The encoder further includes syntax encoding means for encoding an indication of one of the video codecs used to encode the corresponding PCC attribute for each PCC attribute. The encoder further includes transmission means for transmitting the bitstream towards the decoder.
[0024] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that the encoder is further configured to execute the method of any of the foregoing aspects.
[0025] In one embodiment, the present disclosure is to receive a bitstream including an encoded sequence of a plurality of PCC frames, wherein the encoded sequence of the plurality of PCC frames Geometry , texture, and represents a plurality of PCC attributes including one or more of reflectance, transparency, and normal, and each encoded PCC frame is represented by one or more PCC network abstraction layer (NAL) units, and includes a decoder including receiving means for performing reception. The decoder further includes parsing means for parsing the bitstream to obtain an indication of one of the plurality of video codec decoders (codecs) used to code the corresponding PCC attribute for each PCC attribute. The decoder further includes decoding means for decoding the bitstream based on the indicated video codec for the PCC attribute.
[0026] Optionally, in any of the foregoing aspects, another embodiment of the aspect provides that it is further configured to execute the method of any of the foregoing aspects.
[0027] For clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0028] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0029] To understand the present disclosure more fully, reference is made to the following brief description in connection with the accompanying drawings and detailed description. Here, like reference numerals represent like parts.
[0030]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
[0031] First, exemplary implementations of one or more embodiments are provided below, but it should be understood that the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The present disclosure should in no way be limited to the exemplary embodiments, drawings, and techniques below, but includes the exemplary designs and embodiments illustrated and described herein, and may be modified within the scope of the appended claims, along with the full scope of their equivalents.
[0032] Many video compression techniques can be used to reduce the size of video files with minimal data loss. For example, video compression techniques can include performing temporal (e.g., intra-picture) prediction and / or spatial (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be divided into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in adjacent blocks within the same picture. Video blocks in an inter-coded (P or B) slice of a picture may be coded by employing spatial prediction with respect to reference samples in adjacent blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture is called a frame, and a reference picture is called a reference frame. Spatial or temporal prediction results in a prediction block representing the image block. Residual data represents the pixel difference between the original image block and the prediction block. Thus, an inter-coded block is coded according to a motion vector pointing to a block of reference samples forming the prediction block and residual data indicating the difference between the coded block and the prediction block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain. This results in residual transform coefficients, which may be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of transform coefficients. Entropy coding may be applied to achieve more compression. Such video compression techniques are discussed in more detail below.
[0033] To ensure that the encoded video is accurately decoded, the video is encoded and decoded according to the corresponding video coding standard. Video coding standards include International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262, or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or Advanced Video Coding (AVC) also known as ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC) also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multi-View Video Coding (MVC) and Multi-View Video Coding Plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC), 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVT) of ITU-T and ISO / IEC has started the development of a video coding standard called Versatile Video Coding (VVC), which is included in the Working Draft (WD), which includes JVET-K1001-v4 and JVET-K1002-v1.
[0034] PCC is a mechanism for encoding videos of 3D objects. A point cloud is a set of data points in 3D space. Such data points include parameters that determine, for example, a spatial position and color. Point clouds may be used in various applications such as real-time 3D immersive telepresence, content virtual reality (VR) viewing with interactive parallax, 3D free-viewpoint sports relay broadcasting, geographic information systems, cultural heritage, autonomous navigation based on large-scale 3D dynamic maps, automotive applications, etc. The ISO / IEC MPEG codec for PCC may operate on reversible and / or irreversible compression point clouds having substantial coding efficiency and robustness to network environments. By using this codec, it becomes possible to manipulate point clouds as a form of computer data, store them in various storage media, transmit and receive them via a network, and distribute them over a broadcast channel. The PCC coding environment is classified into PCC category 1, PCC category 2, and PCC category 3. The present disclosure is directed to PCC category 2 related to MPEG output documents N17534 and N17533. The design of the PCC category 2 codec aims to utilize other video codecs to compress the point cloud data as a set of different video sequences, Geometry and texture information. For example, two video sequences, one representing the point cloud data Geometry information and the other representing the texture information, can be generated and compressed using one or more video codecs. Additional metadata (e.g., occupancy maps and auxiliary patch information) that supports the interpretation of the video sequences can also be generated and compressed separately.
[0035] The PCC system includes position data GeometryThe texture PCC attributes including PCC attributes and color data may be supported. However, some video applications may also include other types of data such as reflectance, transparency, normal vectors, etc. Some of these types of data may be coded more efficiently using one codec than another. However, the PCC system may require that the entire PCC stream, and thus all PCC attributes, be coded by the same codec. Furthermore, the PCC attributes may be divided into multiple layers. Such layers may then be combined and / or coded into one or more PCC attribute streams. For example, the layers of attributes can be coded according to a temporarily interleaved coding scheme, where the first layer is coded in a PCC access unit (AU) having an even value in the picture output order, and the second layer is coded in a PCC AU having an odd value in the picture output order. Since there may be streams from 0 to 4 for each attribute and various combinations of such layers of streams, proper identification of the streams and layers may be an issue. However, the PCC system may not be able to determine how many layers are coded or combined with a PCC attribute stream in a given PCC bitstream. Furthermore, the PCC system may not have a mechanism for indicating the manner of combination of the layers and / or the correspondence between such layers and the PCC attribute streams. Finally, patches are employed to code PCC video data. For example, a three-dimensional (3D) PCC object can be represented as a set of two-dimensional (2D) patches. This enables PCC to operate with a video codec designed to code 2D video frames. However, some points in the point cloud may not be captured by the patches in some cases. For example, isolated points in 3D space may be difficult to code as part of a patch.The only patch that makes sense in such a case is a 1 pixel x 1 pixel patch that contains a single point, which, in the case of many such points, significantly increases the signaling overhead. Instead, an irregular point cloud can be used, which is a special patch that contains multiple isolated points. To signal the attributes for an irregular point cloud patch, a different approach than for other patch types is used. However, the PCC system may not be able to indicate that the PCC attribute layer carries an irregular point cloud / patch.
[0036] Disclosed herein is a mechanism for improving PCC by addressing the above problems. In one embodiment, the PCC system may employ different codecs to encode different PCC attributes. Specifically, separate syntax elements can be employed to identify the video codec for each attribute. In another embodiment, the PCC system explicitly signals the number of layers that are encoded and / or combined to represent each PCC attribute stream. Additionally, the PCC system may employ syntax elements to signal the mode used to encode and / or combine the layers of PCC attributes within a PCC attribute stream. Further, the PCC system may employ syntax elements to specify the layer index of the layer associated with each data unit of the corresponding PCC attribute stream. In yet another embodiment, a flag can be employed for each PCC attribute layer to indicate whether the PCC attribute layer carries any irregular point cloud points. Such embodiments can be used alone or in combination. Further, such embodiments enable the PCC system to employ more complex coding mechanisms in a manner that is recognizable by a decoder and thus decodable by the decoder. These and other examples are described in detail below.
[0037] FIG. 1 is a flowchart of an exemplary operating method 100 for coding a video signal. Specifically, the video signal is coded by an encoder. The coding process compresses the video signal by employing various mechanisms to reduce the video file size. A smaller file size enables the transmission of the compressed video file to the user, while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file and reconstructs the original video signal for display to the end user. The decoding process generally mirrors the coding process to enable the decoder to consistently reconstruct the video signal.
[0038] In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and coded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that give the impression of visual movement when viewed in sequence. The frames include pixels represented in terms of light, referred to herein as the luminance component (or luminance samples), and color, referred to as the chrominance component (or color samples). In some examples, the frames may also include depth values to support three-dimensional viewing.
[0039] In step 103, the video is divided into blocks. The division includes further dividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H Part 2, a frame can first be divided into Coding Tree Units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels × 64 pixels). A CTU contains both luminance samples and chrominance samples. By adopting a coding tree, the CTU can be divided into blocks, and the blocks can be recursively further divided until a configuration that supports further encoding is achieved. For example, the luminance component of a frame may be further divided until the individual blocks contain relatively uniform illumination values. Additionally, the chrominance component of a frame may be further divided until the individual blocks contain relatively uniform color values. Thus, the division mechanism varies depending on the content of the video frame.
[0040] In step 105, various compression mechanisms are employed to compress the image blocks divided in step 103. For example, inter prediction and / or intra prediction may be employed. Inter prediction is designed to utilize the fact that objects in a common scene tend to appear in consecutive frames. Therefore, blocks representing objects in the reference frame do not need to be repeatedly described in adjacent frames. Specifically, an object such as a table can remain in a fixed position over a plurality of frames. Thus, the table is described once, and adjacent frames can refer to the reference frame. A pattern matching mechanism may be employed to match objects over multiple frames. Furthermore, a moving object may be represented over a plurality of frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over a plurality of frames. To explain such movement, motion vectors may be employed. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in the reference frame. In this way, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame.
[0041] Intra prediction encodes blocks within a common frame. Intra prediction utilizes the fact that luminance and chrominance components tend to concentrate in a frame. For example, a green patch in a part of a tree tends to be positioned adjacent to a similar green patch. Intra prediction employs multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. The directional mode indicates that the current block is similar / same as the samples of adjacent blocks in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the end of the row. The planar mode effectively shows a smooth transition of light / color across the row / column by adopting a relatively constant slope when changing values. The DC mode is used for boundary smoothing and indicates that the block is similar / same as the average value associated with the samples of all adjacent blocks associated with the angular direction of the directional prediction mode. Thus, an intra prediction block can represent an image block as various relational prediction mode values instead of the actual values. Additionally, an inter prediction block can represent an image block as motion vector values instead of the actual values. In either case, the prediction block may not accurately represent the image block in some cases. Any differences are stored in the residual block. Transformation may be applied to the residual block to further compress the file.
[0042] In step 107, various filtering techniques may be applied. In HEVC, the filters are applied according to an in-loop filtering scheme. The block-based prediction discussed above may result in the generation of a blocky image at the decoder. Further, the block-based prediction scheme may encode a block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering scheme repeatedly applies a noise reduction filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to the blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Further, these filters reduce artifacts in the reconstructed reference blocks, and the artifacts are less likely to generate additional artifacts in subsequent blocks encoded based on the reconstructed reference blocks.
[0043] When the video signal is segmented, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes the data discussed above and any signaling data desirable to support proper video signal reconstruction at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission towards the decoder if required. The bitstream may also be broadcast and / or multicast towards multiple decoders. The generation of the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may occur continuously and / or simultaneously over many frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of discussion and is not intended to limit the video coding process to a particular order.
[0044] The decoder receives the bitstream and starts the decoding process at step 111. Specifically, the decoder adopts an entropy decoding method that converts the bitstream into corresponding syntax and video data. The decoder uses the syntax data from the bitstream to determine the partition for the frame at step 111. The partition splitting should match the result of the block partition splitting at step 103. The entropy encoding / decoding employed at step 111 is described here. The encoder makes many choices during the compression process, such as selecting a block splitting method from several possible options based on the spatial positioning of the values within the input image. Precise signaling of the options may use a number of bins. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can vary depending on the context). Entropy coding enables the encoder to discard any option that is clearly not executable in a particular case, leaving a set of acceptable options. A codeword is assigned to each acceptable option. The length of the codeword is based on the number of acceptable options (e.g., one bin for two options, two bins for three to four options, etc.), and the encoder then encodes the codeword for the selected option. This scheme reduces the size of the codeword because the codeword desirably uniquely indicates a selection from a small subset of acceptable options, as opposed to uniquely indicating a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of acceptable options in a manner similar to the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.
[0045] In step 113, the decoder performs block decoding. Specifically, the decoder employs inverse transformation to generate the residual block. Then, the decoder employs the residual block and the corresponding prediction block to reconstruct the image block according to the segmentation. The prediction block may include both an intra prediction block and an inter prediction block as generated by the encoder in step 105. Then, the reconstructed image block is positioned within the frame of the reconstructed video signal according to the segmentation data determined in step 111. The syntax for step 113 may also be signaled within the bitstream via entropy coding as discussed above.
[0046] In step 115, filtering is performed on the frame of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. Once the frame is filtered, the video signal is output to the display in step 117 and can be viewed by the end user.
[0047] FIG. 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functionality to support the implementation of the operation method 100. The codec system 200 is generalized to show components employed in both the encoder and the decoder. The codec system 200 receives and splits a video signal as discussed with respect to steps 101 and 103 of the operation method 100, thereby resulting in a split video signal 201. Next, when acting as an encoder, the codec system 200 compresses the split video signal 201 into a coded bitstream as discussed with respect to steps 105, 107, and 109 of method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream as discussed with respect to steps 111, 113, 115, and 117 of the operation method 100. The codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, a header formatting and context adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In FIG. 2, the black lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of other components. All components of the codec system 200 may be present within the encoder. The decoder may include a subset of the components of the codec system 200. For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. Here, these components will be described.
[0048] The segmented video signal 201 is a captured video sequence that is segmented into blocks of pixels by a coding tree. The coding tree employs various split modes to split the blocks of pixels into smaller blocks of pixels. These blocks can then be further split into even smaller blocks. The blocks are sometimes referred to as nodes on the coding tree. A large parent node is split into smaller child nodes. The number of times a node is further split is called the depth of the node / coding tree. In some cases, the split blocks can be included in a coding unit (CU). For example, a CU can be a sub-part of a CTU that includes a luminance block, a red chrominance (Cr) block, and a blue chrominance (Cb) block, along with the syntax instructions of the corresponding CU. The split modes may include a binary tree (BT), a triple tree (TT), and a quad tree (QT) that are employed to split a node into two, three, or four child nodes of different shapes, respectively, depending on the split mode employed. The segmented video signal 201 is transferred for compression to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221.
[0049] The general coder control component 211 is configured to make decisions related to coding the images of the video sequence into a bitstream according to the constraints of the application. For example, the general coder control component 211 manages the optimization of bitrate / bitstream size versus reconstructed quality. Such decisions may be made based on the availability of memory area / bandwidth and the image resolution requirements. The general coder control component 211 also manages the utilization of the buffer in light of the transmission speed in order to mitigate buffer underrun and overrun problems. To manage these problems, the general coder control component 211 manages the partitioning, prediction, and filtering by other components. For example, the general coder control component 211 can dynamically increase the complexity of compression to increase the resolution and increase the use of bandwidth, or can decrease the complexity of compression to decrease the resolution and the use of bandwidth. Therefore, the general coder control component 211 controls other components of the codec system 200 to balance concerns about bitrate and the quality of the reconstructed video signal. The general coder control component 211 generates control data for controlling the operation of other components. The control data is also transferred to the header formatting that is encoded in the bitstream with signal parameters for decoding by the decoder and to the CABAC component 231.
[0050] The divided video signal 201 is also transmitted to the motion estimation component 221 and the motion compensation component 219 for inter prediction. The frames or slices of the divided video signal 201 may be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter prediction coding of the received video blocks with respect to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may execute a plurality of coded paths, for example, to select an appropriate coding mode for each block of video data.
[0051] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are illustrated separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating motion vectors, and the motion vectors estimate the motion of video blocks. The motion vectors may indicate, for example, the displacement of the coded object relative to the prediction block. The prediction block is a block that is found to closely match the block to be coded with respect to the pixel difference. The prediction block may also be referred to as a reference block. Such pixel differences can be determined by the sum of the absolute difference (SAD), the sum of squared differences (SSD), or other difference metrics. HEVC employs several coded objects including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into CTBs, and a CTB can be divided into CUs for inclusion in a CB. A CU can be coded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing the transform residual data of the CU. The motion estimation component 221 uses rate distortion analysis as part of the rate distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, or may select the reference block, motion vector, etc. having the best rate distortion characteristics. The best rate distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final encoding).
[0052] In some examples, the codec system 200 may calculate the values of the sub-integer pixel positions of the reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 can interpolate the values of the 1 / 4 pixel position, 1 / 8 pixel position, or other fractional pixel positions of the reference picture. Thus, the motion estimation component 221 may perform a motion search for all pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy. The motion estimation component 221 calculates the motion vector for the PUs of the video blocks in the inter-coded slice by comparing the position of the PU with the position of the predicted block of the reference picture. The motion estimation component 221 outputs the calculated motion vector as motion data for header formatting for encoding and to the CABAC component 231, and outputs the motion to the motion compensation component 219.
[0053] The motion compensation performed by the motion compensation component 219 may involve fetching or generating a predicted block based on the motion vector determined by the motion estimation component 221. Also, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving the motion vector for the PU of the current video block, the motion compensation component 219 may position the predicted block pointed to by the motion vector. Then, by subtracting the pixel values of the predicted block from the pixel values of the current video block to be encoded, a residual video block is formed and pixel difference values are formed. Generally, the motion estimation component 221 performs motion estimation for the luminance component, and the motion compensation component 219 uses the motion vector calculated based on the luminance component for both the chrominance components and the luminance component. The predicted block and the residual block are transferred for conversion by the scaling and quantization component 213.
[0054] The segmented video signal 201 is also sent to the intra picture estimation component 215 and the intra picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra picture estimation component 215 and the intra picture prediction component 217 may be highly integrated, but are illustrated separately for conceptual purposes. The intra picture estimation component 215 and the intra picture prediction component 217 perform intra prediction of the current block for a block within the current frame, instead of inter prediction performed by the inter-frame motion estimation component 221 and the motion compensation component 219 as described above. In particular, the intra picture estimation component 215 determines the intra prediction mode to be used to encode the current block. In some examples, the intra picture estimation component 215 selects an appropriate intra prediction mode to encode the current block from a plurality of tested intra picture prediction modes. The selected intra prediction mode is then transferred to the header formatting and CABAC component 231 for encoding.
[0055] For example, the intra picture estimation component 215 calculates rate distortion values using rate distortion analysis for various tested intra prediction modes and selects the intra prediction mode having the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original non-coding block encoded to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate distortion value for the block. Additionally, the intra picture estimation component 215 may be configured to encode depth blocks of the depth map using a depth modeling mode (DMM) based on rate distortion optimization (RDO).
[0056] The intra-picture prediction component 217 may generate a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215 when implemented in an encoder, or read a residual block from a bitstream when implemented in a decoder. The residual block includes the difference in values between the prediction block represented as a matrix and the original block. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luminance component and the chrominance component.
[0057] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), a conceptually similar transform, etc. to the residual block to generate a video block containing residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms can also be used. The transform may transform the residual information from the pixel value domain to a transform domain, such as a frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix containing the quantized transform coefficients. The quantized transform coefficients are transferred to the header formatting and CABAC component 231 and encoded in the bitstream.
[0058] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct the residual block in the pixel domain and use it, for example, as a reference block that can later become the prediction block of another current block. The motion estimation component 221 and / or the motion compensation component 219 may calculate the reference block by adding the residual block to the corresponding prediction block for use in motion estimation of subsequent blocks / frames. The filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transformation. Otherwise, such artifacts can cause inaccurate predictions (generate additional artifacts) when subsequent blocks are predicted.
[0059] The filter control analysis component 227 and the in-loop filter component 225 apply a filter to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters to adjust how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data for encoding to the header formatting and CABAC component 231. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., the reconstructed pixel block) or the frequency domain depending on the example.
[0060] When operating as a symbolizer, the filtered and reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as discussed above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks as part of the output video signal and transfers them towards the display. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0061] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into an encoded bitstream for transmission to a decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information may also include an intra prediction mode index table (also called a codeword mapping table), the definition of encoding contexts for various blocks, an indication of the most likely intra prediction mode, an indication of partition information, etc. Such data may be encoded by employing entropy coding. For example, the information may be encoded by using context-adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy coding (PIPE), or another entropy coding technique. In accordance with the entropy coding, the encoded bitstream may be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.
[0062] Figure 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 may be employed to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 splits the input video signal, resulting in a split video signal 301, which is substantially similar to the split video signal 201. The split video signal 301 is then compressed by components of the encoder 300 and encoded into a bitstream.
[0063] Specifically, the divided video signal 301 is transferred to the intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. Also, the divided video signal 301 is transferred to the motion compensation component 321 for inter prediction based on the reference blocks in the decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to the transform and quantization component 313 for the transform and quantization of the residual blocks. The transform and quantization component 313 may be substantially the same as the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (along with the associated control data) are transferred to the entropy coding component 331 for coding into the bitstream. The entropy coding component 331 may be substantially the same as the header formatting and CABAC component 231.
[0064] The transformed and quantized residual blocks and / or corresponding prediction blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into the reference blocks used by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially the same as the scaling and inverse transform component 229. The in-loop filter within the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks, for example, depending on the example. The in-loop filter component 325 may be substantially the same as the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as discussed for the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 may be substantially the same as the decoded picture buffer component 223.
[0065] FIG. 4 is a block diagram showing an exemplary video decoder 400. The video decoder 400 may be employed to implement the decoding functions of the codec system 200 and / or to perform the steps 111, 113, 115, and / or 117 of the operation method 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to an end user.
[0066] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may employ header information to provide a context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0067] The reconstructed residual block and / or prediction block is transferred to the intra-picture prediction component 417 to reconstruct the image block based on the intra prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 adopts a prediction mode to position the reference block within the frame, applies the residual block to the result, and reconstructs the intra prediction image block. The reconstructed intra prediction image block and / or residual block and the corresponding inter prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225 respectively. The in-loop filter component 425 filters the reconstructed image block, residual block and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reconstructed image block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 adopts the motion vector from the reference block to generate a prediction block, applies the residual block to the result, and reconstructs the image block. The resulting reconstructed block may be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks that can be reconstructed into the frame via the partition information. Such a frame may be placed in the sequence. This sequence is output towards the display as the reconstructed output video signal.
[0068] FIG. 5 is an example of a point cloud medium 500 that can be encoded according to the PCC mechanism. A point cloud is a set of data points in space. The point cloud may be generated by a 3D scanner that measures a number of points on the outer surface of surrounding objects. The point cloud may be Geometry described from viewpoints such as attributes, texture attributes, reflectance attributes, transparency attributes, and normal attributes. Each attribute can be encoded by a codec such as the video codec system 200, the encoder 300, and / or the decoder 400 as part of method 100. Specifically, each attribute of the PCC frame can be separately encoded by the encoder, decoded by the decoder, and recombined again to recreate the PCC frame.
[0069] The point cloud medium 500 includes three bounding boxes 502, 504, and 506. Each of the bounding boxes 502, 504, and 506 represents a portion or segment of a 3D image from the current frame. The bounding boxes 502, 504, and 506 include a 3D image of a person, but other objects may be included in the bounding box in an actual application. Each bounding box 502, 504, and 506 includes an x-axis, a y-axis, and a z-axis that respectively indicate the number of pixels occupied by the 3D image in the x, y, and z directions. For example, the x-axis and the y-axis indicate about 400 pixels (e.g., about 0 to 400 pixels), and the z-axis indicates about 1000 pixels (e.g., about 0 to 1000 pixels).
[0070] Each of the boundary boxes 502, 504, and 506 includes one or more patches 508 represented by the cubes or boxes of FIG. 5. Each patch 508 includes a portion of the entire object within one of the boundary boxes 502, 504, or 506 and may be described or represented by patch size information. The patch information may include, for example, two-dimensional (2D) coordinates and / or three-dimensional (3D) coordinates that describe the position of the patch 508 within the boundary box 502, 504, or 506. The patch information may also include other parameters. For example, the patch information may include a parameter such as a normalAxis that is inherited from the reference patch information to the current patch information. That is, one or more parameters from the patch information of the reference frame may be inherited to the patch information of the current frame. Additionally, one or more metadata portions from the reference frame (e.g., patch rotation, scale parameter, material identifier, etc.) may be inherited to the current frame. The patch 508 may sometimes be interchangeably referred to herein as a 3D patch or a patch data unit. A list of the patches 508 within each boundary box 502, 504, or 506 may be generated and stored in a patch buffer in descending order from the largest patch to the smallest patch. This patch may then be encoded by an encoder and / or decoded by a decoder.
[0071] The patch 508 can describe various attributes of the point cloud medium 500. Specifically, the position of each pixel on the x-axis, y-axis, and z-axis is the geometry of that pixel. The patch 508 including the positions of all the pixels within the current frame can be encoded to capture the Geometry attributes for the current frame of the point cloud medium 500. Further, each pixel may include color values in the red, blue, and green (RGB) and / or luminance and chrominance (YUV) spectra. The patch 508 including the colors of all the pixels within the current frame can be encoded to capture the texture attributes for the current frame of the point cloud medium 500.
[0072] Additionally, each pixel may (or may not) include some reflectivity. Reflectivity is the amount of light (e.g., colored light) projected from a pixel to adjacent pixels. Shiny objects have high reflectivity and thus spread the light / color of their corresponding pixels to other nearby pixels. On the other hand, matte objects may have little or no reflectivity and may not affect the color / light levels of adjacent pixels. A patch 508 that includes the reflectivity of all pixels within the current frame can be encoded to capture the reflectivity attribute for the current frame of the point cloud medium 500. Some pixels may be partially or completely transparent (e.g., glass, transparent plastic, etc.). Transparency is the amount of light / color of adjacent pixels that can pass through the current pixel. A patch 508 that includes the transparency levels of all pixels within the current frame can be encoded to capture the transparency attribute for the current frame of the point cloud medium 500. Further, the points of the point cloud medium may generate a surface. The surface can be associated with a normal vector, which is a vector perpendicular to the surface. The normal vector may be useful when explaining the movement and / or interaction of objects. Thus, in some cases, the user may wish to encode the normal vector for the surface to support additional functionality. A patch 508 that includes the normal vector for the surface within the current frame can be encoded to capture the normal attribute for the current frame of the point cloud medium 500.
[0073] Geometry, texture, reflectivity, transparency, and normal attributes can, for example, include data that describes some or all of the data points within the point cloud medium 500. For example, reflectivity, transparency, and normal attributes are optional and thus may occur individually or in combination, even within the same bitstream, in some examples of a point cloud medium 500 and not in others. Thus, the number of patches 508, and even the number of attributes, may vary from frame to frame and from video to video based on the filmed subject, video settings, and the like.
[0074] Figure 6 is an example of data segmentation and packing for a point cloud medium frame 600. Specifically, the example of Figure 6 shows a 2D representation of the patches 508 of the point cloud medium 500. The point cloud medium frame 600 includes a bounding box 602 corresponding to the current frame from a video sequence. The bounding box 602 is 2D, in contrast to the 3D bounding boxes 502, 504, and 506 of Figure 5. As shown, the bounding box 602 includes a number of patches 604. Patches 604 may be interchangeably referred to herein as 2D patches or patch data units. In summary, the patches 604 of Figure 6 are a representation of the image within the bounding box 504 of Figure 5. Thus, the 3D image within the bounding box 504 of Figure 5 is projected onto the bounding box 602 via the patches 604. The portion of the bounding box 602 that does not include one of the patches 604 is referred to as a void 606. The void 606 may also be referred to as a gap, empty sample, and the like.
[0075] Taking the above into consideration, it should be noted that the video-based point cloud compression (PCC) codec solution is based on the segmentation of 3D point cloud data (e.g., patch 508 in FIG. 5) into 2D projection patches (e.g., patch 604 in FIG. 6). In fact, the coding methods or processes described above may be beneficially implemented for various types of technologies such as, for example, immersive six degrees of freedom (6 DoF), dynamic augmented reality / virtual reality (AR / VR) objects, cultural heritage, geographic information system (GIS), computer-aided design (CAD), autonomous navigation, and the like.
[0076] The position of each patch (e.g., one of the patches 604 in FIG. 6) within the boundary box (e.g., boundary box 602) can be determined by only the size of the patch. For example, the largest of the patches 604 in FIG. 6 is first projected onto the boundary box 602 starting from the upper left corner (0, 0). After the largest of the patches 604 is projected onto the boundary box 602, the next largest patch 604 is projected (or alternatively filled) onto the boundary box 602, and this continues until the smallest of the patches 604 is projected onto the boundary box 602. Again, in this process, only the size of each patch 604 is considered. In some cases, patches 604 with smaller sizes may occupy the space between larger patches and will be located closer to the upper left corner of the boundary box 602 than the larger patches 604. During encoding, this process may be repeated for each relevant attribute until the patches for each attribute within the frame are encoded into one or more corresponding attribute streams. Then, the group of data units within the attribute streams used to recreate a single frame can be stored in the bitstream within the PCC AU. In the decoder, these attribute streams are retrieved from the PCC AU and decoded to generate the patches 604. Then, such patches 604 can be combined to recreate the PCC medium. Thus, the point cloud media frame 600 can be encoded by a codec such as the video codec system 200, the encoder 300, and / or the decoder 400 as part of method 100 for compressing the point cloud media 500 for transmission.
[0077] FIG. 7 is a schematic diagram illustrating an exemplary PCC video stream 700 having an extended attribute set. For example, the PCC video stream 700 may be generated when the point cloud media frame 600 from the point cloud media 500 is encoded according to method 100 by employing, for example, the video codec system 200, the encoder 300, and / or the decoder 400.
[0078] The PCC video stream 700 includes the sequence of PCC AUs 710. The PCC AU 710 contains sufficient data to reconstruct a single PCC frame. The data is arranged in the PCC AU 710 within the NAL unit 720. The NAL unit 720 is a data container of packet size. For example, a single NAL unit 720 is generally sized to enable simple network transmission. The NAL unit 720 may include a header indicating the type of the NAL unit 720 and a payload containing the associated video data. The PCC video stream 700 is designed for an extended attribute set and thus includes some attribute-specific NAL units 720.
[0079] The PCC video stream 700 includes a Group of Frames (GOF) header 721, an auxiliary information frame 722, an occupancy map frame 723, Geometry a NAL unit 724, a texture NAL unit 725, a reflection NAL unit 726, a transparency NAL unit 727, and a normal NAL unit 728, each of which is a type of NAL unit 720. The GOF header 721 includes various syntax elements that describe the corresponding PCC AU 710, the frames associated with the corresponding PCC AU 710, and / or other NAL units 720 within the PCC AU 710. The PCC AU 710 may optionally include a single GOF header 721 or may not include the GOF header 721. The auxiliary information frame 722 may include metadata related to the frame, such as information related to patches used to encode attributes. The occupancy map frame 723 may include additional metadata related to the frame, such as an occupancy map indicating the regions of the frame occupied by data versus regions of empty frames. The remaining NAL units 720 contain attribute data for the PCC AU 710. Specifically, Geometry the NAL unit 724, the texture NAL unit 725, the reflection NAL unit 726, the transparency NAL unit 727, and the normal NAL unit 728 each GeometryIt includes an attribute, a texture attribute, a reflection attribute, a transparency attribute, and a normal attribute.
[0080] As described above, the attribute can be encoded into a stream. For example, there may be a stream from 0 to 4 for each attribute. The stream may include logically separated parts of the PCC video data. For example, attributes for different objects may be encoded into multiple attribute streams of the same type (e.g., the first stream for the first 3D bounding box, the second attribute stream for the second 3D bounding box, etc.). In another example, attributes associated with different frames may be encoded into multiple attribute streams (e.g., a transparency attribute stream for even frames, a transparency attribute stream for odd frames). In yet another example, patches may be placed in layers to represent a 3D object. Thus, separate layers may be included in separate streams (e.g., the first texture attribute stream for the top layer, the second texture attribute stream for the second layer, etc.). Regardless of the example, the PCC AU710 may include zero, one, or multiple NAL units for the corresponding attribute. Geometry As described above, the attribute can be encoded into a stream. For example, there may be a stream from 0 to 4 for each attribute. The stream may include logically separated parts of the PCC video data. For example, attributes for different objects may be encoded into multiple attribute streams of the same type (e.g., the first stream for the first 3D bounding box, the second attribute stream for the second 3D bounding box, etc.). In another example, attributes associated with different frames may be encoded into multiple attribute streams (e.g., a transparency attribute stream for even frames, a transparency attribute stream for odd frames). In yet another example, patches may be placed in layers to represent a 3D object. Thus, separate layers may be included in separate streams (e.g., the first texture attribute stream for the top layer, the second texture attribute stream for the second layer, etc.). Regardless of the example, the PCC AU710 may include zero, one, or multiple NAL units for the corresponding attribute.
[0081] This disclosure supports an increase in flexibility for coding various attributes (e.g., Geometry as included in the NAL unit 724, the texture NAL unit 725, the reflection NAL unit 726, the transparency NAL unit 727, and / or the normal NAL unit 728). In a first example, different codecs can be employed to code different PCC attributes. As a specific example, a first codec is employed to Geometry for GeometryIt can be encoded into the NAL unit 724, and the second codec is adopted to encode the reflection of the PCC video into the reflected NAL unit 726. As another example, when coding the PCC video, up to five codecs (for example, one codec for each attribute) can be adopted. Then, the codec used for the attribute can be signaled as a syntax element in the PCC video stream 700, for example, in the GOF header 721.
[0082] Furthermore, as described above, the PCC attributes may adopt many combinations of layers and / or streams. Therefore, in order to enable the decoder to determine the combination of layers and / or streams for each attribute when decoding, syntax elements (for example, within the GOF header 721) can be used to signal the combination of layers and / or streams used by the encoder when encoding each attribute. Additionally, syntax elements (for example, within the GOF header 721) can be used to signal the mode used to code and / or combine the layers of the PCC attributes in the PCC attribute stream. Additionally, syntax elements (for example, within the GOF header 721) can be used to specify the layer index of the layer associated with each NAL unit 720 corresponding to the PCC attribute stream. For example, using the GOF header 721, Geometry the number of layers and streams related to the attribute, the way such layers and streams are arranged, and each Geometry layer index for each NAL unit 724 can be signaled so that the decoder can assign each Geometry NAL unit 724 to the appropriate layer when decoding the PCC frame.
[0083] Finally, a flag (e.g., within the GOF header 721) can indicate whether any PCC attribute layer contains any irregular point cloud points. An irregular point cloud is a set of one or more data points that are discontinuous with adjacent data points and thus cannot be represented by a 2D patch such as patch 604. Instead, such points are represented as part of an irregular point cloud patch that includes the coordinates and / or transformation parameters associated with such irregular point cloud points. Since the irregular point cloud is represented using a different data structure than the 2D patch, the flag enables the decoder to properly recognize the presence of the irregular point cloud and select an appropriate mechanism for decoding such data.
[0084] The following is an exemplary mechanism for implementing the above-described aspects. Definitions: A video NAL unit is a PCC NAL unit for which the PccNalUnitType is equal to GMTRY_NALU, TEXTURE_NALU, REFLECT_NALU, TRANSP_NALU, or NORMAL_NALU.
[0085] Bitstream Format: This section specifies the relationship between the NAL unit stream and the byte stream, either of which is called a bitstream. The bitstream can be in one of two formats: the NAL unit stream format or the byte stream format. The NAL unit stream format is conceptually a more basic type and includes a sequence of syntax structures called PCC NAL units. This sequence is ordered in decoding order. There are constraints imposed on the decoding order (and content) of PCC NAL units in the NAL unit stream. The byte stream format can be constructed from the NAL unit stream format by arranging the NAL units in decoding order and prefixing each NAL unit with a start code prefix and zero or more zero-value bytes to form a byte stream. The NAL unit stream format can be extracted from the byte stream format by searching for the position of the unique start code prefix pattern within this byte stream. The byte stream format is similar to the formats adopted in HEVC and AVC.
[0086] The PCC NAL unit header syntax may be implemented as described in Table 1 below.
Table 1
[0087]
Table 2
[0088] The PCC profile and level syntax may be implemented as described in Table 3 below.
Table 3
[0089] The semantics of the PCC NAL unit header may be implemented as follows. The forbidden_zero_bit may be set equal to 0. The pcc_nal_unit_type_plus1 - 1 specifies the value of the variable PccNalUnitType, which specifies the type of the RBSP data structure included in the PCC NAL unit as specified in Table 4 below. The variable NalUnitType is specified as follows: PccNalUnitType = pcc_nal_unit_type_plus1 - 1 (7 - 1) PCC NAL units with nal_unit_type in the range of UNSPEC25 to UNSPEC30 for which the semantics are not specified do not affect the decoding process specified in this specification. It should be noted that PCC NAL units in the range of UNSPEC25 to UNSPEC30 may be used as determined by the application. The decoding process for these values of PccNalUnitType is not specified in this disclosure. Since different applications may use these PCC NAL unit types for different purposes, special attention should be paid to the design of the encoder that generates PCC NAL units with these PccNalUnitType values and the design of the decoder that interprets the content of PCC NAL units with these PccNalUnitType values. This disclosure does not define any management for these values. These PccNalUnitType values may only be suitable for use in contexts where usage collisions (e.g., different definitions of the meaning of the content of PCC NAL units for the same PccNalUnitType value) are not important and are impossible, for example, defined or managed in a control application or transport specification, or managed by controlling the environment in which the bitstream is distributed.
[0090] For purposes other than determining the amount of data within a PCC AU of a bitstream, the decoder may ignore (delete from the bitstream and discard) the content of all PCC NAL units that use a reserved value of PccNalUnitType. This requirement may allow for future definition of compatible extensions to this disclosure. [Table 4]
[0091] The identified video codec (e.g., HEVC or AVC) is indicated by the group of frame header NAL units present in the first PCC AU of each cloud point stream (CPS). The pcc_stream_id specifies the PCC stream identifier (ID) of the PCC NAL unit. When PccNalUnitType is equal to GOF_HEADER, AUX_INFO, or OCP_MAP, the value of pcc_stream_id is set to zero. In the definition of one or more sets of PCC profiles and levels, the value of pcc_stream_id may be restricted to be less than 4.
[0092] The order of PCC NAL units and their association to their PCC AUs is described below. A PCC AU consists of zero or one group of frame header NAL units, one auxiliary information frame NAL unit, one occupancy map frame NAL unit, and GeometryIt includes one or more video AUs that carry data units of PCC attributes such as texture, reflection, transparency, or normal. video_au(i,j) indicates a video AU with a pcc_stream_id equal to j for a PCC attribute whose PCC attribute ID is equal to attribute_type[i]. The video AUs present in the PCC AU may be ordered as follows. When attributes_first_ordering_flag is equal to 1, for any two video AUs video_au(i1,j1) and video_au(i2,j2) present in the PCC AU, the following applies. If i1 is less than i2, regardless of the values of j1 and j2, video_au(i1,j1) shall precede video_au(i2,j2). Otherwise, if i1 is equal to i2 and j1 is greater than j2, video_au(i1,j1) shall follow video_au(i2,j2).
[0093] Otherwise (e.g., if attributes_first_ordering_flag is equal to 0), the following applies to two video AUs video_au(i1,j1) and video_au(i2,j2) present in the PCC AU. If j1 is less than j2, then regardless of the values of i1 and i2, video_au(i1,j1) shall precede video_au(i2,j2). Otherwise, if j1 is equal to j2 and i1 is greater than i2, then video_au(i1,j1) shall follow video_au(i2,j2). The above order of video AUs results in the following. If attributes_first_ordering_flag is equal to 1, the order of video AUs, if present, within the PCC AU is (in the listed order) as follows. Here, within the PCC AU, all PCC NAL units of each PCC attribute, if present, are consecutive in decoding order without being interleaved with PCC NAL units of other PCC attributes. That is, video_au(0,0), video_au(0,1),..., video_au(0,num_streams_for_attribute[0]), video_au(1,0), video_au(1,1),..., video_au(1,num_streams_for_attribute[1]), video_au(num_attributes-1,0), video_au(num_attributes-1,1),..., video_au(num_attributes-1,num_streams_for_attribute[1]). Otherwise (attributes_first_ordering_flag is equal to 0), the order of video AUs, if present, within the PCC AU is (in the listed order) as follows, and within the PCC AU, all PCC NAL units of each specific pcc_stream_id value, if present, are consecutive in decoding order without being interleaved with PCC NAL units of other pcc_stream_id values.That is, video_au(0,0), video_au(1,0),..., video_au(num_attributes - 1,0), video_au(0,1), video_au(1,1),..., video_au(num_attributes - 1,1), video_au(0,num_streams_for_attribute[1]), video_au(1,num_streams_for_attribute[1]),..., video_au(num_attributes - 1,num_streams_for_attribute[1]).
[0094] The association of NAL units to video AUs and the order of NAL units within a video AU are specified by the identified video codec, e.g., the specification of HEVC or AVC. The identified video codec is indicated in the frame header NAL unit present in the first PCC AU of each CPS.
[0095] The first PCC AU of each CPS starts with a group of frame header NAL units, and each group of frame header NAL units specifies the start of a new PCC AU.
[0096] Other PCC AUs start with auxiliary information frame NAL units. In other words, an auxiliary information frame NAL unit starts a new PCC AU if it is not preceded by a group of frame header NAL units.
[0097] The semantics of a group of frame header RBSPs are as follows. num_attributes is the maximum number of PCC attributes that can be carried in a CPS ( Geometry, etc. for the texture). Note that in the definition of one or more sets of PCC profiles and levels, the value of num_attributes can be restricted to 5 or less. When attributes_first_ordering_flag is equal to 0, within a PCC AU, all PCC NAL units of each PCC attribute, if present, are consecutive in decoding order without being interleaved with the PCC NAL units of other PCC attributes. When attributes_first_ordering_flag is set equal to 0, within a PCC AU, all PCC NAL units of each specific pcc_stream_id value, if present, are consecutive in decoding order without being interleaved with the PCC NAL units of other pcc_stream_id values. attribute_type[i] specifies the PCC attribute type of the i-th PCC attribute. The interpretation of different PCC attribute types is specified in Table 5 below. In the definition of one or more sets of PCC profiles and levels, the values of attribute_type[0] and attribute_type[1] may be restricted to be equal to 0 and 1 respectively.
Table 5
[0098] identified_codec_for_attribute[i] specifies the identified video codec used for the coding of the i-th PCC attribute, as shown in Table 6 below.
Table 6
[0099] num_streams_for_attribute[i] specifies the maximum number of PCC streams for the i-th PCC attribute. Note that in the definition of one or more sets of PCC profiles and levels, the value of num_streams_for_attribute[i] may be restricted to 4 or less. num_layers_for_attribute[i] specifies the number of attribute layers for the i-th PCC attribute. Note that in the definition of one or more sets of PCC profiles and levels, the value of num_layer_for_attribute[i] may be restricted to 4 or less. max_attribute_layer_idx[i][j] specifies the maximum value of the attribute layer index of the PCC stream when pcc_stream_id is equal to j for the i-th PCC attribute. The value of max_attribute_layer_idx[i][j] should be smaller than num_layer_for_attribute[i]. attribution_layers_combination_mode[i][j] specifies the attribute layer combination mode for the attribute layers carried in the PCC stream when pcc_stream_id is equal to j for the i-th PCC attribute. The interpretation of different values of attribution_layers_combination_mode[i][j] is specified in Table 7 below.
Table 7
[0100] When attribution_layers_combination_mode[i][j] exists and is equal to 0, the variable attrLayerIdx[i][j] indicating the attribute layer index for the attribute layers of the PCC stream when pcc_stream_id is equal to j for the i-th PCC attribute, and the PCC NAL unit of the attribute layer carried in the video AU in the state equal to PicOrderCntVal as specified in the specification of the video codec where the picture order count value is identified, are derived as follows.
Number
[0101] When regular_points_flag[i][j] is equal to 1, it specifies that the attribute layer carries the normal points of the point cloud signal in the state where the layer index is equal to j for the i-th PCC attribute. When regular_points_flag[i][j] is equal to 0, it specifies that the attribute layer carries the irregular points of the point cloud signal in the state where the layer index is equal to j for the i-th PCC attribute. Note that in the definition of one or more sets of PCC profiles and levels, the value of regular_points_flag[i][j] may be restricted to zero. frame_width is Geometry and indicates the frame width of the texture video in pixels. The frame width should be a multiple of occupancyResolution. frame_height is Geometry and indicates the frame height of the texture video in pixels. The frame height should be a multiple of occupancyResoluton. occupancy_resolution is the Geometry horizontal and vertical resolutions at which the patches are packed into the texture video, in pixels. occupancy_resolution should be an even multiple value of occupancyPrecision. radius_to_smoothing indicates the radius for detecting neighbors for smoothing. The value of radius_to_smooting should be in the range from 0 to 255.
[0102] neighbor_count_smooting indicates the maximum number of neighbors used for smoothing. The value of neighbor_count_smooting should be in the range of 0 to 255. radius2_boundary_detection indicates the radius for boundary point detection. The value of radius2_boundary_detection should be in the range of 0 to 255. threshold_smooting indicates the smoothing threshold. The value of threshold_smooting should be in the range of 0 to 255. lossless_geometry indicates Geometry lossless coding. When the value of lossless_geometry is equal to 1, it indicates that the point cloud Geometry information is losslessly encoded. When the value of lossless_geometry is equal to 0, it indicates that the point cloud Geometry information is lossily encoded. lossless_texture indicates lossless texture coding. When the value of lossless_texture is equal to 1, it indicates that the point cloud texture information is losslessly encoded. When the value of lossless_texture is equal to 0, it indicates that the point cloud texture information is lossily encoded. lossless_geometry_444 indicates Geometry whether to use 4:2:0 for the frame or use the 4:4:4 video format. When the value of lossless_geometry_444 is equal to 1, Geometry it indicates that the video is encoded in the 4:4:4 format. When the value of lossless_geometry_444 is equal to 0, Geometry it indicates that the video is encoded in the 4:2:0 format.
[0103] absolute_d1_coding indicates how the layers other than the layer closest to the projection plane are Geometry encoded. When absolute_d1_coding is equal to 1, it indicates the actual Geometry value for the layers other than the layer closest to the projection planeGeometry Indicates being coded for the layer. When absolute_d1_coding is equal to 0, for layers other than the layer closest to the projection plane Geometry Indicates differential coding of the layer. bin_agithmetic_coding indicates whether binary arithmetic coding is used. When the value of bin_agithmetic_coding is equal to 1, it indicates that binary arithmetic coding is used for all syntax elements. When the value of bin_agithmetic_coding is equal to 0, it indicates that non-binary arithmetic coding is used for some syntax elements. gof_header_extension_flag, when equal to 0, specifies that the gof_header_extension_data_flag syntax element does not exist in the group of frame header RBSP syntax structures. gof_header_extension_flag, when equal to 1, specifies that the gof_header_extension_data_flag syntax element exists in the group of frame header RBSP syntax structures. The decoder may ignore all data following the value 1 of gof_header_extension_flag in the group of frame header NAL units. gof_header_extension_data_flag may have any value, and the presence and value of the flag do not affect decoder compliance. The decoder may ignore all gof_header_extension_data_flag syntax elements.
[0104] The PCC profile and level semantics are as follows. pcc_profile_idc indicates the profile to which the CPS conforms. pcc_pl_reserved_zero_19bits is equal to 0 in the bitstream conforming to this version of the present disclosure. Other values of pcc_pl_reserved_zero_19bits are reserved for future use by ISO / IEC. The decoder may ignore the value of pcc_pl_reserved_zero_19bits. pcc_level_idc indicates the level to which the CPS conforms. When the HEVC bitstream for the PCC attribute type equal to attribute_type[i] extracted as specified by the sub-bitstream extraction process is decoded by the conforming HEVC decoder, in the active SPS, hevc_ptl_12bytes_attribute[i] may be equal to the 12-byte value from general_profile_idc to general_level_idc. When the AVC bitstream for the PCC attribute type equal to attribute_type[i] extracted as specified by the sub-bitstream extraction process is decoded by the conforming AVC decoder, in the active SPS, avc_pl_3ytes_attribute[i] may be equal to the 3-byte value from profile_idc to level_idc.
[0105] The sub-bitstream extraction process is as follows. The inputs to this process are the PCC bitstream inBitstream, the target PCC attribute type targetAttrType, and the target PCC stream ID value targetStreamId. The output of this process is a sub-bitstream. It may be a requirement for bitstream compliance for the input bitstream that any output sub-bitstream that is the output of the process specified in this section, having a conforming PCC bitstream inBitstream, a targetAttrType indicating any type of PCC attribute present in inBitstream, and a targetStreamId less than or equal to the maximum PCC stream ID value of the PCC streams present in inBitstream for the attribute type targetAttrType, is a video bitstream that is conformant for each identified video codec specification for the attribute type targetAttrType.
[0106] The output sub-bitstream is derived by the following ordered steps. Depending on the value of targetAttrType, the following applies. If targetAttrType is equal to ATTR_GEOMETRY, all PCC NAL units where PccNalUnitType is not equal to GMTRY_NALU or pcc_stream_id is not equal to targetStreamId are deleted. Otherwise, if targetAttrType is equal to ATTR_TEXTURE, all PCC NAL units where PccNalUnitType is not equal to TEXTURE_NALU or pcc_stream_id is not equal to targetStreamId are deleted. Otherwise, if targetAttrType is equal to ATTR_REFLECT, all PCC NAL units where PccNalUnitType is not equal to REFLECT_NALU or pcc_stream_id is not equal to targetStreamId are deleted. Otherwise, if targetAttrType is equal to ATTR_TRANSP, all PCC NAL units where PccNalUnitType is not equal to TRANSP_NALU or pcc_stream_id is not equal to targetStreamId are deleted. Otherwise, if targetAttrType is equal to ATTR_NORMAL, all PCC NAL units where PccNalUnitType is not equal to NORMAL_NALU or pcc_stream_id is not equal to targetStreamId are deleted. For each PCC NAL unit, the first byte may be deleted.
[0107] In a first set of alternative embodiments of the method summarized above, the PCC NAL unit header is designed to use more bits for pcc_stream_id and allow for more than four streams for each attribute. In that case, an additional type is added to the header of the PCC NAL unit.
[0108] FIG. 8 is a schematic diagram showing an exemplary mechanism 800 for encoding PCC attributes 841 and 842 having a plurality of codecs 843 and 844. For example, mechanism 800 can be employed to encode and / or decode the attributes of PCC video stream 700. Thus, mechanism 800 can be employed to encode and / or decode point cloud media frame 600 based on point cloud media 500. In this way, mechanism 800 may be used for the encoder 300 to generate a bitstream from a PCC sequence and also for the decoder 400 when reconstructing a PCC sequence from the bitstream. Thus, mechanism 800 can be employed by codec system 200 and may further be employed to support method 100.
[0109] Mechanism 800 can be applied to a plurality of PCC attributes 841 and 842. For example, PCC attributes 841 and 842 may be any two attributes selected from the group including Geometry attributes, texture attributes, reflectance attributes, transparency attributes, and normal attributes. As shown in FIG. 8, mechanism 800 shows an encoding process when proceeding from left to right and a decoding process when proceeding from right to left. Codecs 843 and 844 may be any two codecs such as HEVC, AVC, VVC, or any version thereof. A particular codec 843 and 844, or a version thereof, may be more efficient than others when encoding particular PCC attributes 841 and 842. In this example, codec 843 is used to encode attribute 841 and codec 844 is used to encode attribute 842, respectively. The results of such encoding are combined to generate a PCC video stream 845 including both PCC attributes 841 and 842. In the decoder, codec 843 is used to decode attribute 841 and codec 844 is used to decode attribute 842, respectively. Then, the decoded attributes 841 and 842 can be combined again to generate a decoded PCC video stream 845.
[0110] The advantage of adopting mechanism 800 is that the most efficient codecs 843 and 844 can be selected for the corresponding attributes 841 and 842. Mechanism 800 is not limited to two attributes 841 and 842, and two codecs 843 and 844. For example, each attribute ( Geometry , texture, reflectivity, transparency, and normal) can be encoded by a separate codec. To ensure that appropriate codecs 843 and 844 can be selected to decode the corresponding attributes 841 and 842, the encoder may signal the codecs 843 and 844, and their respective correspondences to the attributes 841 and 842. For example, the encoder can include syntax elements in the GOF header to indicate the codec-attribute correspondence. Then, the decoder can read the relevant syntax, select the correct codecs 843 and 844 for the attributes 841 and 842, and decode the PCC video stream 845. As a specific example, the identified_codec_for_attribute syntax element can be adopted to indicate the codecs 843 and 844 for the attributes 841 and 842, respectively.
[0111] FIG. 9 is a schematic diagram 900 showing examples of attribute layers 931, 932, 933, and 934. For example, the attribute layers 931, 932, 933, and 934 can be employed to carry the attributes of the PCC video stream 700. Thus, the layers 931, 932, 933, and 934 can be used when encoding and / or decoding the point cloud media frame 600 based on the point cloud media 500. In this way, the layers 931, 932, 933, and 934 may be used by the encoder 300 to generate a bitstream from the PCC sequence, or may be used by the decoder 400 when reconstructing the PCC sequence from the bitstream. Thus, the layers 931, 932, 933, and 934 can be employed by the codec system 200 and may further be employed to support the method 100. Additionally, the attribute layers 931, 932, 933, and 934 may be used to carry one or more of the attributes 841 and 842.
[0112] Attribute layers 931, 932, 933, and 934 are groupings of data related to attributes that can be stored and / or modified independently of other groups of data related to the same attributes. In this way, each of the attribute layers 931, 932, 933, and 934 can be changed and / or represented without affecting the remaining attribute layers 931, 932, 933, and / or 934. In some examples, the attribute layers 931, 932, 933, and / or 934 may be visually represented on top of each other, as shown in FIG. 9. For example, a texture that covers the entire object (e.g., a more general one) can be stored in the attribute layer 931, and more detailed textures (e.g., more specific ones) can be included in the attribute layers 932, 933, and / or 934. In another example, the attribute layers 931 and / or 932 may be applied to odd-numbered frames, and the attribute layers 933 and / or 934 may be applied to even-numbered frames. This may allow some layers to be omitted in response to a change in the frame rate. Each attribute may have 0 to 4 attribute layers 931, 932, 933, and / or 934. To signal the configuration adopted, the encoder may adopt a syntax element such as num_layers_for_attribute[i] in sequence-level data such as the GOF header. The decoder can read the syntax element and determine the number of attribute layers 931, 932, 933, and / or 934 adopted for each attribute. Additional syntax elements, such as attribution_layers_combination_mode[i][j], attrLayerIdx[i][j], etc., can also be adopted to indicate the combination of attribute layers adopted in the PCC video stream and the index of each layer used by the corresponding attributes, respectively.
[0113] As yet another example, some of the attribute layers (e.g., attribute layers 931, 932, and 933) can carry data regarding regular patches, while other attribute layers (e.g., attribute layer 934) carry data associated with an irregular point cloud patch. This is useful because an irregular point cloud may be described using different data than a regular cloud patch. To signal that a particular layer carries data associated with an irregular point cloud, the encoder can encode another syntax element in the sequence level data. As a particular example, the regular_points_flag of the GOF header can be used to indicate that an attribute layer carries at least one irregular point cloud point. The decoder can then read the syntax element and decode the corresponding attribute layer accordingly.
[0114] FIG. 10 is a schematic diagram 1000 showing examples of attribute streams 1031, 1032, 1033, and 1034. For example, the attribute streams 1031, 1032, 1033, and 1034 can be employed to carry the attributes of the PCC video stream 700. Thus, the attribute streams 1031, 1032, 1033, and 1034 can be employed when encoding and / or decoding the point cloud media frame 600 based on the point cloud media 500. In this way, the attribute streams 1031, 1032, 1033, and 1034 may be used for the encoder 300 to create a bitstream from the PCC sequence and may also be used when the decoder 400 reconstructs the PCC sequence from the bitstream. Thus, the attribute streams 1031, 1032, 1033, and 1034 can be employed by the codec system 200 and may further be employed to support the method 100. Additionally, the attribute streams 1031, 1032, 1033, and 1034 may be used to carry one or more of the attributes 841 and 842. Further, the attribute streams 1031, 1032, 1033, and 1034 can be employed to carry the attribute layers 931, 932, 933, and 934.
[0115] Attribute streams 1031, 1032, 1033, and 1034 are sequences of attribute data over time. Specifically, attribute streams 1031, 1032, 1033, and 1034 are sub-streams of the PCC video stream. Each attribute stream 1031, 1032, 1033, and 1034 carries a sequence of attribute-specific NAL units and thus acts as a storage and / or transmission data structure. Each attribute stream 1031, 1032, 1033, and 1034 may carry one or more attribute layers 931, 932, 933, and 934 of data. For example, attribute stream 1031 can carry attribute layers 931 and 932, while attribute stream 1032 carries attribute layers 931 and 932 (attribute streams 1033 and 1034 are omitted). In another example, each attribute stream 1031, 1032, 1033, and 1034 carries a single corresponding attribute layer 931, 932, 933, and 934. In other examples, some of the attribute streams 1031, 1032, 1033, and 1034 carry multiple attribute layers 931, 932, 933, and 934, while other attribute streams 1031, 1032, 1033, and 1034 carry a single attribute layer 931, 932, 933, and 934 or are omitted. As can be seen, many combinations and permutations of attribute streams 1031, 1032, 1033, and 1034, as well as attribute layers 931, 932, 933, and 934, can occur. Thus, the encoder can adopt a syntax element such as num_streams_for_attribute in sequence-level data such as the GOF header to indicate the number of attribute streams 1031, 1032, 1033, and 1034 used to encode each attribute. The decoder can then use such information, in combination with, for example, attribute layer information, to decode the attribute streams 1031, 1032, 1033, and 1034 to reconstruct the PCC sequence.
[0116] FIG. 11 is a flowchart of an exemplary method 1100 for encoding a PCC video sequence having a plurality of codecs. For example, method 1100 can compile data into a bitstream according to mechanism 800 while using attribute layers 931, 932, 933, and 934, and / or streams 1031, 1032, 1033, and / or 1034. Also, method 1100 may specify the mechanism used to encode the attributes within the GOF header. Further, method 1100 may generate a PCC video stream 700 by encoding a point cloud media frame 600 based on the point cloud media 500. Additionally, method 1100 may be employed by codec system 200 and / or encoder 300 while performing the encoding steps of method 100.
[0117] Method 1100 may start when an encoder receives a sequence of PCC frames including point cloud media. The encoder may decide to encode such frames, for example, in response to receiving a user command. In method 1100, the encoder may decide that a first attribute should be encoded by a first codec while a second attribute should be encoded by a second codec. This decision may be made based on predetermined conditions and / or based on user input, for example, when the first codec is more efficient for the first attribute and the second codec is more efficient for the second attribute. Thus, in step 1101, the encoder encodes the first attribute of the sequence of PCC frames into a bitstream having the first codec. Further, in step 1103, the encoder encodes the second attribute of the sequence of PCC frames into the bitstream with a second codec different from the first codec.
[0118] In step 1105, the encoder encodes various syntax elements into a bitstream along with the encoded video data. For example, the syntax elements can be encoded into a sequence-level data unit including sequence-level parameters to indicate to the decoder the decisions made during encoding so that the PCC frame can be properly reconstructed. Specifically, the encoder encodes a sequence-level data unit to include a first syntax element indicating that the first attribute is encoded by the first codec and that the second attribute is encoded by the second codec. As a specific example, the PCC frame may include a plurality of attributes including the first attribute and the second attribute. Also, the plurality of attributes of the PCC frame may include Geometry , texture, and one or more of reflectance, transparency, and normal. Further, the first syntax element may be the identified_codec_for_attribute element included in the GOF header in the bitstream.
[0119] In some examples, the first attribute may be organized into a plurality of streams. In such a case, a second syntax element can be used to indicate the stream membership for the data unit of the bitstream associated with the first attribute. In some examples, the first attribute may also be organized into a plurality of layers. In such a case, a third syntax element may indicate the layer membership for the data unit of the bitstream associated with the first attribute. As a specific example, the second syntax element may be the num_streams_for_attribute element, and the third syntax element may be the num_layers_for_attribute element, each of which may be included in a group of frame headers in the bitstream. In yet another example, a fourth syntax element may be used to indicate that the first layer of the plurality of layers includes data associated with an irregular point cloud. As a specific example, the fourth syntax element may be the regular_points_flag element included in a group of frame headers in the bitstream.
[0120] By including such information in the sequence level data, the decoder may have sufficient information to decode the PCC video sequence. Thus, the encoder may, in step 1107, transmit a bitstream based on the first attribute encoded by the first codec, the second attribute encoded by the second codec, and other attributes and / or syntax elements described herein to support the generation of the decoded sequence of PCC frames.
[0121] FIG. 12 is a flowchart of an exemplary method 1200 for decoding a PCC video sequence with multiple codecs. For example, method 1200 may read data from a bitstream according to mechanism 800 while using attribute layers 931, 932, 933, and 934, and / or streams 1031, 1032, 1033, and / or 1034. Also, method 1200 may determine the mechanism used to encode attributes by reading the GOF header. Further, method 1200 may read the PCC video stream 700 to reconstruct the point cloud media frame 600 and the point cloud media 500. Additionally, method 1200 may be employed by codec system 200 and / or decoder 400 while performing the decoding steps of method 100.
[0122] Method 1200 may start when, in step 1201, a decoder receives a bitstream including a series of PCC frames. The decoder can then analyze the bitstream or a part thereof in step 1205. For example, the decoder can analyze the bitstream to obtain a sequence-level data unit including sequence-level parameters. The sequence-level data unit may include various syntax elements that describe the encoding process. Thus, the decoder can analyze the video data from the bitstream and use the syntax elements to determine an appropriate process for decoding the video data.
[0123] For example, the sequence-level data unit can include a first syntax element indicating that a first attribute is encoded by a first codec and that a second attribute is encoded by a second codec. As a specific example, a PCC frame may include a plurality of attributes including a first attribute and a second attribute. Also, the plurality of attributes of the PCC frame may Geometry , texture, and one or more of reflectance, transparency, and normal. Additionally, the first syntax element may be the identified_codec_for_attribute element included in the GOF header in the bitstream.
[0124] In some examples, the first attribute may be organized into multiple streams. In such cases, a second syntax element can be employed to indicate the stream membership for data units of the bitstream associated with the first attribute. In some examples, the first attribute may also be organized into multiple layers. In such cases, a third syntax element may indicate the layer membership for data units of the bitstream associated with the first attribute. As a specific example, the second syntax element may be the num_streams_for_attribute element, and the third syntax element may be the num_layers_for_attribute element, and each may be included in a group of frame headers in the bitstream. In yet another example, a fourth syntax element may be used to indicate that the first layer of the multiple layers includes data associated with an irregular point cloud. As a specific example, the fourth syntax element may be the regular_points_flag element included in a group of frame headers in the bitstream.
[0125] Thus, the decoder can decode the first attribute by the first codec and the second attribute by the second codec in step 1207 to generate the decoded sequence of the PCC frame. The decoder may also use other attributes and / or syntax elements as described herein when determining the appropriate mechanism to employ when decoding various attributes of the PCC video sequence based on the codec.
[0126] FIG. 13 is a schematic diagram of an exemplary video coding device 1300. The video coding device 1300 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 1300 includes a transceiver unit 1310 that includes a downstream port 1320, an upstream port 1350, and / or a transmitter and / or receiver for communicating data upstream and / or downstream via a network. The video coding device 1300 also includes a processor 1330 that includes a logic unit and / or a central processing unit (CPU), and a memory 1332 for storing data. The video coding device 1300 may also include wireless communication components coupled to the upstream port 1350 and / or the downstream port 1320 and / or an electrical, optical-electrical (OE) component, an electro-optical (EO) component, for communicating data via an electrical, optical, or wireless communication network. The video coding device 1300 may also include an input and / or output (I / O) device 1360 for communicating data with a user. The I / O device 1360 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The input-output device 1360 may also include input devices such as a keyboard, a mouse, a trackball, etc. and / or corresponding interfaces for interacting with such output devices.
[0127] Processor 1330 is implemented by hardware and software. Processor 1330 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and a digital signal processor (DSP). Processor 1330 communicates with downstream port 1320, Tx / Rx 1310, upstream port 1350, and memory 1332. Processor 1330 includes a coding module 1314. Coding module 1314 implements the disclosed embodiments described above, such as methods 100, 1100, 1200, 1500, and 1600, and mechanism 800, which may employ point cloud media 500, point cloud media frame 600, and / or PCC video stream 700 encoded in layers 931-934 and / or streams 1031-1034. Also, coding module 1314 may implement any other method / mechanism described herein. Further, coding module 1314 may implement codec system 200, encoder 300, and / or decoder 400. For example, coding module 1314 can employ an extended set of attributes for PCC having multiple streams and layers and can signal the use of such an attribute set within the sequence level data to support decoding. Thus, coding module 1314 enables video coding device 1300 to provide additional functionality and / or flexibility when coding PCC video data. In this way, coding module 1314 improves the functionality of video coding device 1300 and addresses problems specific to video coding technology. Further, coding module 1314 results in converting video coding device 1300 to a different state. Alternatively, coding module 1314 can be implemented as instructions stored in memory 1332 and executed by processor 1330 (e.g., as a computer program product stored on a non-transitory medium).
[0128] The memory 1332 includes one or more memory types such as a disk, a tape drive, a solid state drive, a read only memory (ROM), a random access memory (RAM), a flash memory, a ternary content addressable memory (TCAM), and a static random access memory (SRAM). The memory 1332 may be used as an overflow data storage device to store a program when the program is selected for execution and to store instructions and data read during program execution.
[0129] FIG. 14 is a schematic diagram of an exemplary system 1400 for coding a PCC video sequence having a plurality of codecs. The system 1400 includes a video coder 1402 including a first attribute coder module 1401 for coding a first attribute of a sequence of PCC frames into a bitstream with a first codec. The video coder 1402 further includes a second attribute coder module 1403 for coding a second attribute of a sequence of PCC frames into a bitstream with a second codec different from the first codec. The video coder 1402 further includes a syntax coder module 1405 for coding a sequence level data unit including sequence level parameters into a bitstream, the sequence level data unit including a first syntax element indicating that the first attribute has been coded by the first codec and that the second attribute has been coded by the second codec. The video coder 1402 further includes a transmission module 1407 for transmitting a bitstream to support generation of a decoded sequence of PCC frames based on the first attribute coded by the first codec and the second attribute coded by the second codec. The modules of the video coder 1402 can also be employed to perform any of the steps / items described above with respect to method 1100 and / or 1500.
[0130] System 1400 also includes a video decoder 1410 that includes a receiving module 1411 for receiving a bitstream including a sequence of PCC frames. The video decoder 1410 further includes an analysis module 1413 for analyzing the bitstream to obtain a sequence-level data unit including sequence-level parameters, where the sequence-level data unit includes a first syntax element indicating that a first attribute of the PCC frame is encoded by a first codec and that a second attribute of the PCC frame is encoded by a second codec. The video decoder 1410 further includes a decoding module 1415 for decoding the first attribute by the first codec and the second attribute by the second codec to generate a decoded sequence of PCC frames. The modules of the video decoder 1410 can also be employed to perform any of the steps / items described above with respect to method 1200 and / or 1600.
[0131] FIG. 15 is a flowchart of another exemplary method 1500 for encoding a PCC video sequence with multiple codecs. For example, method 1500 can organize data into a bitstream according to mechanism 800 while using attribute layers 931, 932, 933, and 934, and / or streams 1031, 1032, 1033, and / or 1034. Method 1500 may also specify the mechanism used to encode the attributes within the GOF header. Additionally, method 1500 may generate a PCC video stream 700 by encoding a point cloud media frame 600 based on a point cloud media 500. Optionally, method 1500 may be employed by codec system 200 and / or encoder 300 while performing the encoding steps of method 100.
[0132] In step 1501, a plurality of PCC attributes are encoded into a bitstream as part of a sequence of PCC frames. The PCC attributes are encoded with multiple codecs. The PCC attributes are Geometryand includes texture. The PCC attributes also include one or more of reflectance, transparency, and normal. Each coded PCC frame is represented by one or more PCC NAL units. In step 1503, an indication is coded for each PCC attribute. This indication indicates the video codec used to code the corresponding PCC attribute. In step 1505, the bitstream is sent towards the decoder.
[0133] FIG. 16 is a flowchart of another exemplary method 1600 for decoding a PCC video sequence with multiple codecs. For example, method 1600 can read data from the bitstream according to mechanism 800 while using attribute layers 931, 932, 933, and 934, and / or streams 1031, 1032, 1033, and / or 1034. Also, method 1600 may determine the mechanism used to code the attributes by reading the GOF header. Further, method 1600 may read the PCC video stream 700 to reconstruct the point cloud media frame 600 and the point cloud media 500. Additionally, method 1600 may be employed by the codec system 200 and / or the decoder 400 during the decoding steps of method 100.
[0134] In step 1601, a bitstream is received. The bitstream includes a coded sequence of multiple PCC frames. The coded sequence of PCC frames represents multiple PCC attributes. The PCC attributes Geometry and texture. The PCC attributes also include one or more of reflectance, transparency, and normal. Each coded PCC frame is represented by one or more PCC NAL units. In step 1603, the bitstream is parsed to obtain an indication of the codec used to code the corresponding PCC attribute for each PCC attribute. In step 1605, the bitstream is decoded based on the video codec indicated for the PCC attribute.
[0135] The first component is directly coupled to the second component when there are no intervening components other than a line, trace, or another medium between the first component and the second component. The first component is indirectly coupled to the second component when there are intervening components other than a line, trace, or other medium between the first component and the second component. The terms "coupled" and its variations include both direct and indirect coupling. The use of the term "about" means a range that includes ±10% of the number that follows, unless otherwise specified.
[0136] Although several embodiments have been provided in this disclosure, it will be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. This example is illustrative and not restrictive, and its intention is not limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0137] Additionally, the techniques, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of this disclosure. Other examples of changes, substitutions, and modifications may be ascertainable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.
Claims
1. A method implemented by a decoder, comprising: receiving, by the decoder, a bitstream including at least one coded sequence of a point cloud coding (PCC) frame, a plurality of attributes of the at least one coded sequence of the PCC frame, and a first syntax element, wherein the plurality of attributes includes a first attribute and a second attribute, each attribute being one of texture, reflectance, transparency, or normal, and the first syntax element indicates that the first attribute is coded by a first video codec and the second attribute is coded by a second video codec, and each coded PCC frame is represented by one or more PCC network abstraction layer (NAL) units; analyzing, by the decoder, the bitstream to determine, for the plurality of attributes, a plurality of video codecs respectively used to code the plurality of attributes; decoding, by the decoder, the bitstream based on the plurality of video codecs for the plurality of attributes.
2. The method according to claim 1, wherein the bitstream further includes a second syntax element, and the second syntax element indicates that the first attribute is of a first type and the second attribute is of a second type, and the first type and the second type are different among texture, reflectance, transparency, and normal.
3. The method according to claim 1, wherein the first syntax element is a first array, and each element of the first array specifies a video codec for a corresponding attribute in the plurality of attributes.
4. The first attribute is organized into a plurality of streams, and the bitstream further includes a third syntax element indicating stream membership for data units of the bitstream associated with the first attribute, and the third syntax element is a num_streams_for_attribute element included in a group of frame headers in the bitstream. The method according to any one of claims 1 to 3.
5. The first attribute is organized into a plurality of layers, and the bitstream further includes a fourth syntax element indicating layer membership for data units of the bitstream associated with the first attribute. The method according to any one of claims 1 to 3.
6. The fourth syntax element is a num_layers_for_attribute element included in a group of frame headers in the bitstream. The method according to claim 5.
7. The bitstream further includes a fifth syntax element indicating that the first layer of the plurality of layers includes data associated with an irregular point cloud. The method according to claim 5 or 6.
8. The fifth syntax element is a regular_points_flag element included in a group of frame headers in the bitstream. The method according to claim 7.
9. The bitstream is decoded into a decoded sequence of PCC frames, and the method further includes transferring, by the decoder, the decoded sequence of PCC frames to a display for presentation. The method according to any one of claims 1 to 8.
10. A method implemented in an encoder, Encoding, by the encoder, a plurality of attributes of a sequence of point cloud coding (PCC) frames into a bitstream by a plurality of video coders (codecs), wherein the plurality of attributes include a first attribute and a second attribute, each attribute being one of texture, reflectance, transmittance, or normal, and the bitstream includes a first syntax element indicating that the first attribute is encoded by a first video codec and the second attribute is encoded by a second video codec, and each encoded PCC frame is represented by one or more PCC network abstraction layer (NAL) units; Encoding, by the encoder, instructions for a plurality of video coders respectively used to encode the plurality of attributes with respect to the plurality of attributes; Transmitting, by the encoder, the bitstream towards a decoder. A method comprising the above steps.
11. The method according to claim 10, wherein the bitstream further includes a second syntax element, the second syntax element indicating that the first attribute is of a first type and the second attribute is of a second type, and the first type and the second type are different among texture, reflectance, transparency, and normal.
12. The method according to claim 10, wherein the first syntax element is a first array, and each element of the first array specifies a video codec for a corresponding attribute in the plurality of attributes.
13. The method according to any one of claims 10 to 12, wherein the first attribute is organized into a plurality of streams, and the bitstream further includes a third syntax element indicating stream membership for data units of the bitstream associated with the first attribute, and the third syntax element is a num_streams_for_attribute element included in a group of frame headers in the bitstream.
14. The method according to any one of claims 10 to 12, wherein the first attribute is organized into a plurality of layers, and the bitstream further includes a fourth syntax element indicating layer membership for data units of the bitstream associated with the first attribute.
15. The method according to claim 14, wherein the fourth syntax element is a num_layers_for_attribute element included in a group of frame headers in the bitstream.
16. The method according to claim 14 or 15, wherein the bitstream further includes a fifth syntax element indicating that the first layer of the plurality of layers includes data associated with an irregular point cloud.
17. The method according to claim 16, wherein the fifth syntax element is a regular_points_flag element included in a group of frame headers in the bitstream.
18. A decoding device, A non-transitory memory storage device configured to store video data in the form of a bitstream, and A decoder configured to execute the method according to any one of claims 1 to 9.
19. A non-transitory memory storage device configured to store video data in the form of a bitstream, and An encoding device configured to execute the method according to any one of claims 10 to 17.
20. A non-transitory computer-readable medium including computer-executable instructions stored on the non-transitory computer-readable medium, wherein when the computer-executable instructions are executed by a processor, a video coding device executes the method according to any one of claims 1 to 17. Claims 21 A decoder including a processing circuit for executing the method according to any one of claims 1 to 9. Claims 22 An encoder including a processing circuit for executing the method according to any one of claims 10 to 17. Claims 23 An encoder, encoding a plurality of attributes of a sequence of point cloud coding (PCC) frames into a bitstream by a plurality of video coders (codecs), wherein the plurality of attributes include a first attribute and a second attribute, and each attribute is one of texture, reflectivity, transmittance, or normal, and the bitstream includes a first syntax element indicating that the first attribute is encoded by a first video codec and the second attribute is encoded by a second video codec, and each encoded PCC frame is represented by one or more PCC network abstraction layer (NAL) units, attribute encoding means for encoding; syntax encoding means for encoding instructions of a plurality of video coders respectively used for encoding the plurality of attributes with respect to the plurality of attributes; transmission means for transmitting the bitstream towards a decoder. Claims 24 The encoder according to claim 23, wherein the encoder is further configured to execute the method according to any one of claims 10 to 17. Claims 25 Receiving a bitstream including at least one coded sequence of a point cloud coding (PCC) frame, a plurality of attributes of the at least one coded sequence of the PCC frame, and a first syntax element, wherein the plurality of attributes includes a first attribute and a second attribute, each attribute is one of texture, reflectance, transparency, or normal, and the first syntax element indicates that the first attribute is coded by a first video codec and the second attribute is coded by a second video codec, and each coded PCC frame is represented by one or more PCC network abstraction layer (NAL) units, and receiving means for performing the receiving; Analyzing means for analyzing the bitstream to determine a plurality of video codecs respectively used to code the plurality of attributes for the plurality of attributes; A decoder comprising: decoding means for decoding the bitstream based on the plurality of video codecs for the plurality of attributes. Claim 26 The decoder according to claim 25, wherein the decoder is further configured to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
WO2020027317A1