High-level grammar design for point cloud decoding

By introducing high-level syntax design in point cloud decoding, using PCC network abstract layer unit and frame group header NAL unit to carry video decoder information, the problem of delay and start code competition in point cloud decoding is solved, and the decoding efficiency and adaptability are improved.

CN115665104BActive Publication Date: 2025-08-22HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210978056.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-06-26
Filing Date
2019-04-11
Publication Date
2025-08-22
Estimated Expiration
2039-04-11

AI Technical Summary

Technical Problem

The existing point cloud decoding technology has problems with excessive end-to-end delay caused by data unit interleaving and the bitstream format is not suitable for the start code competition prevention of MPEG-2 transmission streams, and lacks the carrying of key decoding information.

Method used

Adopting advanced syntax design, the above problems are solved by using the payload of the PCC network abstract layer unit to carry the information of the video decoder, and the configuration file and level information are carried through the frame group header NAL unit.

Benefits of technology

Improves the efficiency and adaptability of video decoding, reduces end-to-end delay, and supports competition prevention of start codes, improving the performance of video decoder.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115665104B_ABST
    Figure CN115665104B_ABST
Patent Text Reader

Abstract

A point cloud coding (PCC) method implemented by a video decoder includes: receiving a coded bitstream, the coded bitstream including a frame group header disposed outside a coded representation of a texture component or a geometry component, wherein the frame group header identifies a video decoder for decoding the texture component or the geometry component; and decoding the coded bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 201980043451.4, and the original application date is April 11, 2019. The entire content of the original application is incorporated into this application by reference. Technical Field

[0002] The present disclosure relates to point cloud coding, and more particularly to a high-level syntax for point cloud coding. Background Art

[0003] Point clouds are widely used in applications such as the entertainment industry, smart car navigation, geospatial detection, and three-dimensional (3D) modeling and visualization of real-world objects. Given the non-uniformly sampled geometry of point clouds, a compact representation for storing and transmitting this data is useful. Compared to other 3D representations, irregular point clouds are more general and applicable to a wider range of sensors and data acquisition strategies. For example, when performing 3D rendering in a virtual reality world or remote rendering in a telepresence environment, the rendering of virtual graphics and real-time instructions is processed as a dense point cloud dataset. Summary of the Invention

[0004] A first aspect of the present disclosure relates to a point cloud coding (PCC) method implemented by a video decoder. The method comprises: receiving a coded bitstream, the coded bitstream including a group of frames header disposed outside a coded representation of a texture component or a geometry component, the group of frames header identifying a video decoder for decoding the texture component or the geometry component; and decoding the coded bitstream.

[0005] A second aspect relates to a point cloud coding (PCC) method implemented by a video encoder, comprising: generating a coded bitstream, the coded bitstream including a frame group header disposed outside a coded representation of a texture component or a geometry component, the frame group header identifying a video decoder for decoding the texture component or the geometry component; and transmitting the coded bitstream to a decoder.

[0006] The method provides a high-level syntax design that addresses one or more issues associated with point cloud decoding as described below. As a result, the video decoding process and video decoder are improved and more efficient, among other things.

[0007] In a first implementation of the method according to the first aspect or the second aspect, the coded representation is included in a payload of a PCC network abstraction layer (NAL) unit.

[0008] In a second embodiment of the method according to the first or second aspect or any one of the above embodiments according to the first or second aspect, the payload of the PCC NAL unit includes a geometry component corresponding to the video decoder.

[0009] In a third embodiment of the method according to the first or second aspect or any one of the above embodiments according to the first or second aspect, the payload of the PCC NAL unit includes a texture component corresponding to the video decoder.

[0010] In a fourth implementation of the method according to the first or second aspect or any one of the above implementations according to the first or second aspect, the video decoder is high efficiency video coding (HEVC).

[0011] In a fifth implementation of the method according to the first or second aspect or any one of the above implementations according to the first or second aspect, the video decoder is advanced video coding (AVC).

[0012] In a sixth implementation of the method according to the first or second aspect or any one of the above implementations of the first or second aspect, the video decoder is universal video coding (VVC).

[0013] In a seventh implementation of the method according to the first or second aspect or any one of the above implementations of the first or second aspect, the video decoder is essential video coding (EVC).

[0014] In an eighth implementation of the method according to the first or second aspect or any one of the above implementations according to the first or second aspect, the geometry component comprises a set of coordinates associated with the point cloud frame.

[0015] In a ninth embodiment of the method according to the first or second aspect or any one of the above embodiments according to the first or second aspect, the set of coordinates are Cartesian coordinates.

[0016] In a tenth implementation of the method according to the first aspect or the second aspect or any one of the above implementations of the first aspect or the second aspect, the texture component includes a set of brightness sample values ​​of the point cloud frame.

[0017] A third aspect relates to a coding device, comprising: a receiver for receiving a picture for encoding or receiving a bit stream for decoding; a transmitter connected to the receiver, the transmitter being used to send the above-mentioned bit stream to a decoder or send the decoded image to a display; a memory connected to at least one of the receiver or the transmitter, the memory being used to store instructions; and a processor connected to the memory, the processor being used to execute the instructions stored in the above-mentioned memory to perform the method of any one of the above-mentioned aspects or embodiments.

[0018] The decoding apparatus utilizes a high-level syntax design to solve one or more problems associated with point cloud decoding as described below. As a result, the video decoding process and the video decoder are improved and more efficient, among other things.

[0019] In a first embodiment of the device according to the third aspect, the device further comprises a display for displaying an image.

[0020] A fourth aspect relates to a system comprising an encoder and a decoder communicating with the encoder, wherein the encoder or decoder comprises the decoding device of any one of the above aspects or embodiments.

[0021] The system utilizes a high-level syntax design to address one or more issues associated with point cloud decoding as described below. As a result, the video decoding process and the video decoder are improved and more efficient, among other things.

[0022] The fifth aspect relates to a decoding device, comprising: a receiving device for receiving a picture for encoding or receiving a bit stream for decoding; a sending device connected to the receiving device, the sending device for sending the bit stream to a decoder or sending the decoded image to a display device; a storage device connected to at least one of the receiving device or the sending device, the storage device for storing instructions; and a processing device connected to the storage device, the processing device for executing the instructions stored in the storage device to execute the method of any one of the aforementioned aspects or embodiments.

[0023] The decoding apparatus utilizes a high-level syntax design to solve one or more problems associated with point cloud decoding as described below. As a result, the video decoding process and the video decoder are improved and more efficient, among other things.

[0024] For greater clarity, any one of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments within the scope of the present disclosure.

[0025] The above and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0027] Figure 1 is a block diagram illustrating an example coding system that may utilize context modeling techniques.

[0028] Figure 2 is a block diagram illustrating an example video encoder in which context modeling techniques may be implemented.

[0029] Figure 3 is a block diagram illustrating an example video decoder in which context modeling techniques may be implemented.

[0030] Figure 4 is a schematic diagram of an embodiment of a data structure compatible with PCC.

[0031] Figure 5 This is an embodiment of a point cloud decoding method implemented by a video decoder.

[0032] Figure 6 is an embodiment of a point cloud decoding method implemented by a video encoder.

[0033] Figure 7 is a schematic diagram of an example video decoding device.

[0034] Figure 8 is a schematic diagram of an embodiment of a decoding device. DETAILED DESCRIPTION

[0035] First, it should be understood that although exemplary implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of currently known or existing technologies. The present disclosure should not be limited to the exemplary implementations, drawings, and techniques shown below (including the exemplary designs and implementations shown and described herein), but may be modified within the scope of the appended claims and their full scope of equivalents.

[0036] Video coding standards include International Telecommunications Union Telecommunication Standardization Sector (ITU-T) H.261, International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) Moving Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC), also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes its extensions, such as scalable video coding (SVC), multiview video coding (MVC), multiview video coding plus depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes its extensions, such as scalable HEVC (SHVC), multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).

[0037] A point cloud is a set of data points in 3D space. Each data point consists of a set of parameters that determine its location (e.g., X, Y, Z), color (e.g., R, G, B or Y, U, V), and possibly other properties such as transparency, reflectivity, and acquisition time. Typically, each point in the cloud has the same number of additional properties. Point clouds can be used in a variety of applications, such as real-time 3D immersive telepresence, virtual reality (VR) viewing of content with interactive parallax, 3D free-viewpoint sports replay broadcasts, geographic information systems, cultural heritage, autonomous navigation based on large-scale 3D dynamic maps, and automotive applications.

[0038] In 2016, the ISO / IEC Moving Picture Experts Group began developing a new point cloud decoder standard for point cloud decoding of lossless and lossy compressed point cloud data. This point cloud decoder standard boasts high decoding efficiency and robustness to network environments. This decoder standard enables point clouds to be processed as a form of computer data, stored on various storage media, transmitted and received over existing and future networks, and distributed over existing and future broadcast channels.

[0039] Recently, point cloud coding (PCC) work has been divided into three categories: PCC Category 1, PCC Category 2, and PCC Category 3. Two separate working drafts are under development: one for PCC Category 2 (PCC Cat2) and the other for PCC Categories 1 and 3 (PCC Cat13). The latest working draft (WD) for PCC Cat2 is included in MPEG output document N17534, and the latest WD for PCC Cat13 is included in MPEG output document N17533.

[0040] The key design concept behind the PCC Cat2 decoder in PCC Cat2 WD is to compress the point cloud data into a set of different video sequences, leveraging existing video decoders to compress the geometry and texture information of dynamic point clouds. Specifically, the video decoder can generate and compress two video sequences: one representing the geometry of the point cloud data and the other representing the texture information. Additional metadata used to interpret the two video sequences, namely, occupancy maps and auxiliary patch information, can also be generated and compressed separately.

[0041] However, the existing design of PCC has flaws. For example, the data units belonging to a time instance (i.e., an access unit (AU)) are discontinuous in decoding order. In PCC Cat 2WD, the data units of texture information, geometry information, auxiliary information, and occupancy map of each AU are interleaved in frame groups. That is, the geometry data of all frames in the group are together. Similarly, the same applies to texture data, etc. In PCC Cat 13WD, the geometry data units and general attributes of each AU are interleaved at the level of the entire PCC bitstream (i.e., when only one group of frames has the same length as the entire PCC bitstream, it is the same as PCC Cat2 WD). The interleaving of data units belonging to an AU will result in a large end-to-end delay, which is at least equal to the length of the frame group within the presentation duration in the application system.

[0042] Another shortcoming relates to the bitstream format. The bitstream format allows for start code pattern emulation, such as 0x0003, making it unsuitable for transport over MPEG-2 transport streams (TS), which require start code emulation prevention. For PCC Category 2, when decoding geometry and texture components using HEVC or AVC, start code emulation prevention is currently only implemented in group_of_frames_geometry_video_payload() and group_of_frames_texture_video_payload(). For PCC Category 13, there is no start code emulation prevention anywhere in the bitstream.

[0043] In PCC Cat2 WD, some decoder information for the geometry bitstream and texture bitstream (e.g. which decoder, decoder profile, level, etc.) is hidden in multiple instances of the structures group_of_frames_geometry_video_payload() and group_of_frames_texture_video_payload(). In addition, information such as profile and level, which indicate the ability to decode auxiliary information and occupancy map components and point cloud reconstruction, is missing.

[0044] This document discloses a high-level syntax design that addresses one or more of the aforementioned issues associated with point cloud decoding. As explained more fully below, the present disclosure utilizes a type indicator in the data unit header (also known as the PCC Network Access Layer (NAL) header) to specify the content type in the payload of the PCC NAL unit. Furthermore, the present disclosure utilizes a frame group header NAL unit to carry the frame group header parameters. The frame group header NAL unit can also be used to signal the profile and level of each geometry or texture bitstream.

[0045] Figure 1 FIG. 1 is a block diagram illustrating an embodiment of a decoding system 10 that can utilize PCC video decoding technology. Figure 1As shown, decoding system 10 includes a source device 12 that provides encoded video data that is subsequently decoded by a destination device 14. In particular, source device 12 can provide the video data to destination device 14 via a computer-readable medium 16. Source device 12 and destination device 14 can include any of a variety of devices, including desktop computers, notebook computers (e.g., laptop computers), tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, so-called "smart" tablets, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 12 and destination device 14 can be configured for wireless communication.

[0046] The destination device 14 may receive the encoded video data to be decoded via a computer-readable medium 16. The computer-readable medium 16 may include any type of medium or device capable of transferring the encoded video data from the source device 12 to the destination device 14. In one example, the computer-readable medium 16 may include a communication medium to enable the source device 12 to send the encoded video data directly to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and sent to the destination device 14. The communication medium may include any wireless communication medium or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet network, such as a local area network, a wide area network, or a global network, such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that facilitates communication from the source device 12 to the destination device 14.

[0047] In some examples, the encoded data can be output from output interface 22 to a storage device. Similarly, the encoded data can be accessed from the storage device via the input interface. The storage device can include any of a variety of distributed or locally accessible data storage media, such as a hard drive, a Blu-ray disc, a digital video disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device can correspond to a file server or another intermediate storage device that can store the encoded video generated by source device 12. Destination device 14 can access the stored video data from the storage device by streaming or downloading. The file server can be any type of server capable of storing encoded video data and sending the encoded video data to destination device 14. Example file servers include a network server (e.g., for a website), a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Destination device 14 can access the encoded video data via any standard data connection, including an internet connection. These data connections may include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber lines (DSL), cable modems, etc.) suitable for accessing encoded video data stored on a file server, or a combination of wireless and wired connections. The transmission of the encoded video data from the storage device may be streaming, downloading, or a combination of streaming and downloading.

[0048] The technology of the present disclosure is not limited to wireless applications or settings. The technology can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission, such as dynamic adaptive streaming over HTTP (DASH) based on HTTP (hypertext transfer protocol), digital video encoded on a data storage medium, decoding digital video stored on a data storage medium, or other applications. In some embodiments, the decoding system 10 can be used to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0049] exist Figure 1 In the embodiment of the present invention, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to the present disclosure, the video encoder 20 of source device 12 and / or the video decoder 30 of destination device 14 can be used to apply video decoding techniques. In other examples, the source device and destination device can include other components or arrangements. For example, source device 12 can receive video data from an external video source such as an external camera. Similarly, destination device 14 can be connected to an external display device rather than including an integrated display device.

[0050] Figure 1 The decoding system 10 shown is merely an example. The techniques for video decoding can be performed by any digital video encoding and / or decoding device. Although the techniques of the present disclosure are typically performed by a video decoding device, these techniques can also be performed by a video encoder / decoder, commonly referred to as a "codec." Furthermore, the techniques of the present disclosure can also be performed by a video preprocessor. The video encoder and / or decoder can be a graphics processing unit (GPU) or similar device.

[0051] Source device 12 and destination device 14 are merely examples of encoding devices that generate encoded video data for transmission to destination device 14. In some examples, source device 12 and destination device 14 can operate in a substantially symmetrical manner, such that both source device 12 and destination device 14 include video encoding and decoding components. Thus, encoding system 10 can support one-way or two-way video transmission between video devices 12, 14, such as for video streaming, video playback, video broadcasting, or video telephony.

[0052] Video source 18 of source device 12 may include a video capture device (e.g., a video camera), a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. Alternatively, video source 18 may generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video.

[0053] In some cases, where video source 18 is a video camera, source device 12 and destination device 14 may form a so-called camera phone or videophone. However, as described above, the techniques described in this disclosure may be generally applicable to video decoding and may be applied to wireless and / or wired applications. In each case, captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video information may then be output by output interface 22 onto computer-readable medium 16.

[0054] Computer-readable medium 16 may include transient media such as wireless broadcast or wired network transmission, or storage media such as hard disks, flash drives, optical disks, digital video disks, Blu-ray disks, or other computer-readable media (i.e., non-transient storage media). In some embodiments, a network server (not shown) may receive encoded video data from source device 12 and provide the encoded video data to destination device 14, for example, via network transmission. Similarly, a computing device at a media production facility such as a disc stamping facility may receive encoded video data from source device 12 and produce a disc containing the encoded video data. Therefore, in different examples, computer-readable medium 16 may be understood to include one or more computer-readable media in different forms.

[0055] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information of the computer-readable medium 16 may include syntax information defined by the video encoder 20 and used by the video decoder 30, including syntax elements that describe characteristics and / or processing of blocks and other coding units such as groups of pictures (GOPs). The display device 32 displays the decoded video data to a user, and the display device may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0056] The video encoder 20 and the video decoder 30 may operate in accordance with a video coding standard, such as the High Efficiency Video Coding (HEVC) standard currently under development, and may conform to the HEVC Test Model (HM). Alternatively, the video encoder 20 and the video decoder 30 may operate in accordance with other proprietary or industry standards, such as the H.264 standard of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) (or H.265 / HEVC for Moving Picture Experts Group (MPEG)-4, Part 10, Advanced Video Coding (AVC)) or extensions of these standards. However, the technology of the present disclosure is not limited to any particular coding standard. Other examples of video coding standards include MPEG-2 and ITU-T H.263. Although in Figure 1 Although not shown, in some aspects, video encoder 20 and video decoder 30 can be integrated with an audio encoder and decoder, respectively, and can include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software to handle the encoding of audio and video in a common data stream or in separate data streams. Where applicable, the MUX-DEMUX units can conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).

[0057] The video encoder 20 and the video decoder 30 may each be implemented as any of a variety of applicable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When these techniques are implemented in part in software, the device may store the software's instructions in an applicable non-transitory computer-readable medium and implement these instructions in hardware using one or more processors to perform the techniques of the present disclosure. The video encoder 20 and the video decoder 30 may each be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device. A device including the video encoder 20 and / or the video decoder 30 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular telephone.

[0058] Figure 2is a block diagram illustrating an example of a video encoder 20 that may implement video decoding techniques. The video encoder 20 may perform intra-frame decoding and inter-frame decoding on video blocks within a video slice. Intra-frame decoding relies on spatial prediction to reduce or eliminate spatial redundancy in video within a given video frame or picture. Inter-frame decoding relies on temporal prediction to reduce or eliminate temporal redundancy in video within adjacent frames or pictures of a video sequence. Intra-frame mode (I-mode) may refer to any of several spatial-based decoding modes. Inter-frame modes, such as uni-directional prediction (also known as uniprediction) (P-mode) or bi-prediction (also known as bi-prediction) (B-mode), may refer to any of several temporal-based decoding modes.

[0059] like Figure 2 As shown, the video encoder 20 receives a current video block within a video frame to be encoded. Figure 2 In the example of FIG. 5 , the video encoder 20 includes a mode selection unit 40, a reference frame memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy decoding unit 56. The mode selection unit 40 further includes a motion compensation unit 44, a motion estimation unit 42, an intra-prediction (also known as intra prediction) unit 46, and a segmentation unit 48. For video block reconstruction, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform unit 60, and a summer 62. A deblocking filter (DFI) may also be included. Figure 2 The output of summer 50 is filtered by a filter (not shown) to remove blocking artifacts from the reconstructed video. If necessary, a deblocking filter is typically used to filter the output of summer 62. In addition to the deblocking filter, additional filters (in-loop or post-loop) may also be used. For the sake of clarity, such filters are not shown in the figure, but if necessary, the filters described above can be used to filter the output of summer 50 (as in-loop filters).

[0060] During the encoding process, video encoder 20 receives a video frame or slice to be encoded. The frame or slice may be divided into multiple video blocks. Motion estimation unit 42 and motion compensation unit 44 perform inter-frame prediction decoding on the received video block relative to one or more blocks in one or more reference frames to provide temporal prediction. Optionally, intra-frame prediction unit 46 may perform intra-frame prediction decoding on the received video block relative to one or more neighboring blocks in the same frame or slice as the block to be encoded to provide spatial prediction. Video encoder 20 may perform multiple decoding passes, for example, to select an appropriate decoding mode for each block of video data.

[0061] In addition, the segmentation unit 48 can segment the video data block into sub-blocks based on an evaluation of a previous segmentation scheme in a previous decoding pass. For example, the segmentation unit 48 can first segment a frame or slice into a largest coding unit (LCU) and segment each LCU into sub-coding units (sub-CUs) based on rate-distortion analysis (e.g., rate-distortion optimization). The mode selection unit 40 can further generate a quadtree data structure that indicates the segmentation of the LCU into sub-CUs. The leaf node CU of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs).

[0062] This disclosure uses the term "block" to refer to any of CU, PU, ​​or TU in the case of HEVC, or similar data structures in the case of other standards (for example, macroblocks and their sub-blocks in H.264 / AVC). A CU includes a decoding node, a PU associated with the decoding node, and a TU. The size of the CU corresponds to the size of the decoding node and is square. The size of the CU can range from 8×8 pixels to a maximum tree block size of 64×64 pixels or larger. Each CU can contain one or more PUs and one or more TUs. The syntax data associated with the CU can describe, for example, the partitioning of the CU into one or more PUs. Whether the CU is skipped or the CU is encoded in direct mode, intra-frame prediction mode, or inter-prediction (also called inter prediction) mode, the partitioning mode used may be different. The PU can be partitioned into non-square shapes. The syntax data associated with the CU can also describe, for example, the partitioning of the CU into one or more TUs according to a quadtree. The shape of the TU can be square or non-square (for example, rectangular).

[0063] Mode select unit 40 may select one of the coding modes (intra or inter) based on the error result, for example, and provide the resulting intra or inter coded block to summer 50 to generate residual block data and to summer 62 to reconstruct the coded block for use as a reference frame. Mode select unit 40 also provides syntax elements, such as motion vectors, intra mode indicators, partition information, and other such syntax information, to entropy coding unit 56.

[0064] Motion estimation unit 42 and motion compensation unit 44 can be highly integrated but are shown separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors, which are used to estimate the motion of video blocks. For example, a motion vector may indicate the displacement of a PU of a video block within a current video frame or picture relative to a prediction block within a reference frame (or other coding unit), where the prediction block within the reference frame is relative to the current block being encoded within the current frame (or other coding unit). A prediction block is a block that is found to closely match the block to be encoded in terms of pixel difference, which can be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics. In some embodiments, video encoder 20 may calculate values ​​for sub-integer pixel positions of a reference picture stored in reference frame memory 64. For example, video encoder 20 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Thus, motion estimation unit 42 may perform motion searches relative to full-pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.

[0065] Motion estimation unit 42 calculates a motion vector for a PU of a video block in an inter-coded slice by comparing the position of the PU with the position of a prediction block of a reference picture. The reference picture may be selected from a first reference picture list (list 0) or a second reference picture list (list 1), each of which identifies one or more reference pictures stored in reference frame memory 64. Motion estimation unit 42 sends the calculated motion vector to entropy encoding unit 56 and motion compensation unit 44.

[0066] Motion compensation performed by motion compensation unit 44 may include obtaining or generating a prediction block based on the motion vector determined by motion estimation unit 42. Also, in some examples, motion estimation unit 42 and motion compensation unit 44 may be functionally integrated. Upon receiving the motion vector for the PU of the current video block, motion compensation unit 44 may locate the prediction block to which the motion vector points in one of the reference picture lists. Summer 50 forms a residual video block by subtracting the pixel values ​​of the prediction block from the pixel values ​​of the current video block being encoded, thereby forming pixel difference values ​​as described below. Typically, motion estimation unit 42 performs motion estimation with respect to the luma component, and motion compensation unit 44 applies the motion vector calculated based on the luma component to the chroma and luma components. Mode select unit 40 may also generate syntax elements associated with video blocks and video slices for use by video decoder 30 when decoding video blocks of a video slice.

[0067] Intra-prediction unit 46 may perform intra-prediction on the current block as an alternative to the inter-prediction performed by motion estimation unit 42 and motion compensation unit 44, as described above. In particular, intra-prediction unit 46 may determine an intra-prediction mode to use for encoding the current block. In some examples, intra-prediction unit 46 may encode the current block using different intra-prediction modes, e.g., during separate encoding passes, and intra-prediction unit 46 (or, in some examples, mode selection unit 40) may select an applicable intra-prediction mode to use from the tested modes.

[0068] For example, the intra-frame prediction unit 46 can use rate-distortion analysis for various tested intra-frame prediction modes to calculate rate-distortion values ​​and select the intra-frame prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis is generally used to determine the amount of distortion (or error) between a coded block and the original uncoded block encoded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. The intra-frame prediction unit 46 can calculate a ratio based on the distortion and rate of the various coded blocks to determine which intra-frame prediction mode exhibits the best rate-distortion value for the block.

[0069] In addition, the intra prediction unit 46 can be used to encode the depth block of the depth map using a depth modeling mode (DMM). The mode selection unit 40 can use rate-distortion optimization (RDO) to determine whether the available DMM mode produces better decoding results than the intra prediction mode and other DMM modes. The data of the texture image corresponding to the depth map can be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 can also be used to perform inter-frame prediction on the depth block of the depth map.

[0070] After selecting an intra-prediction mode for a block (e.g., one of the conventional intra-prediction modes or the DMM mode), intra-prediction unit 46 may provide information indicating the selected intra-prediction mode for the block to entropy decoding unit 56. Entropy decoding unit 56 may encode the information indicating the selected intra-prediction mode. Video encoder 20 may include configuration data in the transmitted bitstream that may include a plurality of intra-prediction mode index tables and a plurality of modified intra-prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, and an indication of the most probable intra-prediction mode, the intra-prediction mode index tables, and the modified intra-prediction mode index tables for each context.

[0071] Video encoder 20 forms a residual video block by subtracting the prediction data from mode select unit 40 from the original video block being encoded. Summer 50 represents one or more components that perform this subtraction operation.

[0072] Transform processing unit 52 applies a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform, to the residual block, thereby generating a video block comprising residual transform coefficient values. Transform processing unit 52 may perform other transforms conceptually similar to the DCT. That is, a wavelet transform, an integer transform, a sub-band transform, or other types of transforms may also be used.

[0073] Transform processing unit 52 applies a transform to the residual block, thereby generating a block having residual transform coefficients. This transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then scan the matrix comprising the quantized transform coefficients. Optionally, entropy coding unit 56 may perform the scan.

[0074] After quantization, entropy decoding unit 56 performs entropy decoding on the quantized transform coefficients. For example, entropy decoding unit 56 can perform context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) decoding, or another entropy decoding technique. In the case of context-based entropy decoding, the context can be based on neighboring blocks. After entropy decoding unit 56 performs entropy decoding, the encoded bitstream can be sent to another device (e.g., video decoder 30) or archived for later transmission or retrieval.

[0075] Inverse quantization unit 58 and inverse transform unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual block in the pixel domain, for example, for subsequent use as a reference block. Motion compensation unit 44 may calculate a reference block by adding the residual block to a prediction block of one of the frames in reference frame memory 64. Motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values ​​for motion estimation. Summer 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reconstructed video block for storage in reference frame memory 64. Motion estimation unit 42 and motion compensation unit 44 may use the reconstructed video block as a reference block for inter-coding a block in a subsequent video frame.

[0076] Figure 3 is a block diagram illustrating an example of a video decoder 30 that may implement video decoding techniques. Figure 3 In the example of FIG. 1 , video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra-frame prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and a summer 80. In some examples, video decoder 30 may perform operations generally similar to those described with respect to video encoder 20 ( Figure 2 ) is a decoding pass that is the opposite of the encoding pass described in . The motion compensation unit 72 may generate prediction data based on the motion vector received from the entropy decoding unit 70, and the intra-prediction unit 74 may generate prediction data based on the intra-prediction mode indicator received from the entropy decoding unit 70.

[0077] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements from video encoder 20. Entropy decoding unit 70 of video decoder 30 entropy decodes the bitstream to generate quantization coefficients, motion vectors, or intra-prediction mode indicators, as well as other syntax elements. Entropy decoding unit 70 forwards the motion vectors and other syntax elements to motion compensation unit 72. Video decoder 30 may receive syntax elements at the video slice level and / or at the video block level.

[0078] In the case where the video slice is encoded as an intra-coded (I) slice, intra-prediction unit 74 may generate prediction data for the video block of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks of the current frame or picture. In the case where the video frame is encoded as an inter-coded (e.g., B, P, or GPB) slice, motion compensation unit 72 generates a prediction block for the video block of the current video slice based on the motion vectors and other syntax elements received from entropy decoding unit 70. The prediction block may be generated from one of the reference pictures in one of the reference picture lists. Video decoder 30 may construct the reference frame lists, i.e., List 0 and List 1, using a default construction technique based on the reference pictures stored in reference frame memory 82.

[0079] Motion compensation unit 72 parses the motion vectors and other syntax elements to determine prediction information for the video blocks of the current video slice and uses the prediction information to generate a prediction block for the current video block being decoded. For example, motion compensation unit 72 uses some of the received syntax elements to determine the prediction mode used to encode the video blocks of the video slice (e.g., intra or inter prediction), the inter-prediction slice type (e.g., B slice, P slice, or GPB slice), construction information for one or more reference picture lists for the slice, motion vectors for each inter-coded video block of the slice, inter-prediction status for each inter-coded video block of the slice, and other information used to decode the video blocks in the current video slice.

[0080] Motion compensation unit 72 may also perform interpolation based on interpolation filters. Motion compensation unit 72 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters used by video encoder 20 during encoding of the video block. In this case, motion compensation unit 72 may determine the interpolation filters used by video encoder 20 from received syntax elements and use the interpolation filters to produce a prediction block.

[0081] Data of the texture image corresponding to the depth map may be stored in the reference frame memory 82. The motion compensation unit 72 may also be configured to perform inter-frame prediction on a depth block of the depth map.

[0082] Based on the above content, some basic concepts of the present disclosure are discussed.

[0083] To address the first issue mentioned above, for PCC Cat2, data units associated with a time instance (e.g., an access unit) should be placed consecutively in the bitstream in decoding order. Once these data units are placed consecutively in the bitstream in decoding order, the identification of the type of each data unit can allow the identification to route each data unit to the correct decoder component. This design should also avoid violating the main design concept behind the PCC Cat2 decoder, which is to use existing video decoders to compress the geometric and texture information of dynamic point clouds.

[0084] In order to be able to utilize existing video decoders (e.g., HEVC as an example) to compress geometry information and texture information separately while having a single, self-contained PCC Cat2 bitstream, the following aspects should be explicitly specified: (1) extraction / construction of a compliant HEVC bitstream of the geometry component from the PCC Cat2 bitstream; (2) extraction / construction of a compliant HEVC bitstream of the texture component from the PCC Cat2 bitstream; and (3) signaling / indication of the compliance point (i.e., profile, layer, and level) of the compliant HEVC bitstream for each extracted geometry component and texture component.

[0085] In order to solve the above problems and meet all the above limitations, the present disclosure provides two sets of alternative methods related to PCC high-level syntax.

[0086] In the first group of methods, all video decoders have a common high-level syntax that can be used to decode the geometry and texture components of PCC Category 2. This group of methods is summarized as follows.

[0087] Figure 4 A data structure 400 compatible with PCC is shown. The data structure 400 may represent a portion of a bitstream generated by an encoder and received by a decoder. As shown, each data unit 402 (which may be referred to as a PCC NAL unit) is added with a data unit header 404 (which may be referred to as a PCC NAL unit header). Figure 4 In the data structure 400, one data unit 402 and one data unit header 404 are shown. However, in actual applications, the data structure 400 may include any number of data units 402 and data unit headers 404. In practice, a bitstream including the data structure 400 may include a sequence of data units 402, each data unit 402 including a data unit header 404.

[0088] The data unit header 404 may include, for example, one or two bytes. In one embodiment, each data unit 402 forms a PCC NAL unit. The data unit 402 includes a payload 406. In one embodiment, the data unit 402 may also include supplemental enhancement information (SEI) messages, sequence parameter sets, picture parameter sets, slice information, etc.

[0089] In one embodiment, the payload 406 of the data unit 402 can be an HEVC unit or an AVC NAL unit. In one embodiment, the payload 406 can contain data for a geometry component or a texture component. In one embodiment, the geometry component is a set of Cartesian coordinates associated with a point cloud frame. In one embodiment, the texture component is a set of luma sample values ​​for the point cloud frame. When HEVC is used, the data unit 402 can be considered a PCC NAL unit containing an HEVC NAL unit as the payload 406. When AVC is used, the data unit 402 can be considered a PCC NAL unit containing an AVC NAL unit as the payload 406.

[0090] In one embodiment, the design of the data unit header 404 (eg, PCC NAL unit header) is summarized as follows.

[0091] First, the data unit header 404 includes a type indicator. The type indicator may be, for example, 5 bits. The type indicator specifies the type of content carried in the payload 406. For example, the type indicator may specify that the payload 406 contains geometric information or texture information.

[0092] In one embodiment, some reserved data units (similar to data unit 402, but reserved for future use) can be used for PCC Category 13 data units. Therefore, the design of this disclosure is also applicable to PCC Category 13. Thus, PCC Category 2 and PCC Category 13 can be unified into a single decoder standard specification.

[0093] As described above, the current bitstream format allows for contention of a start code pattern that indicates, for example, the start of a new NAL unit or PCC NAL unit. The start code pattern may be, for example, 0x0003. Because the current bitstream format allows for contention of start code patterns, a start code may be inadvertently represented. The present disclosure provides PCC NAL unit syntax and semantics (see below) to address this issue. The PCC NAL unit syntax and semantics described herein can ensure contention for the start code of each PCC NAL unit and are not affected by its content. Therefore, the last byte of a one-byte or two-byte data unit header 404 (e.g., the data unit header itself if the data unit header is one byte) cannot be 0x00.

[0094] In addition, the frame group header 408 (also known as a frame group header NAL unit) is designed to carry the frame group header parameters. Furthermore, the frame group header NAL unit includes signaling of other global information, such as the profile and level for each geometry or texture bitstream. In one embodiment, a profile is a specified subset of syntax or a subset of decoding tools. In one embodiment, a level is a defined set of constraints on the values ​​that syntax elements and variables can take. In one embodiment, the combination of a bitstream's profile and level represents the specific decoding capabilities required to decode the bitstream. Furthermore, where profiles and levels are also defined for auxiliary information, occupancy map decoding, and point cloud reconstruction (using the decoded results of geometry, texture, auxiliary information, and occupancy maps), these profiles and levels are also signaled in the frame group header 408. In one embodiment, PCC auxiliary information refers to information such as patch information and point local reconstruction information, which is used to reconstruct the point cloud signal from the PCC coded bitstream. In one embodiment, a PCC occupancy map refers to information related to the portion of 3D space occupied by objects, from which texture values ​​and other attributes are sampled.

[0095] As shown in the following syntax, the constraints on the order of different types of data units 402 (also called PCC NAL units) are explicitly specified. In addition, the start of an access unit 410 (which may contain several data units 402, data unit headers 404, etc.) is also explicitly specified.

[0096] In addition, the extraction / construction process of each geometry bitstream or texture bitstream is also clearly indicated in the syntax and / or semantics mentioned below.

[0097] In the second group of methods, different overall syntaxes are used for different video coders. PCC Category 2, which uses HEVC to code geometry and texture, is designated as a modification of HEVC, while PCC Category 2, which uses AVC to code geometry and texture, is designated as a modification of AVC. This group of methods is summarized below.

[0098] For PCC Category 2, which uses HEVC to decode geometry and texture, geometry and texture are treated as three separate layers (e.g., two layers for geometry, d0 and d1, and one layer for texture). SEI messages or new NAL units are used for occupancy maps and auxiliary information. Two new SEI messages are specified, one for occupancy maps and one for auxiliary information. Another sequence-level SEI message is specified to carry the frame group header parameters and other global information. This SEI message is similar to the frame group header 408 in the first group of methods.

[0099] For PCC Category 2, which uses AVC to decode geometry and texture, geometry and texture are treated as three independent layers (e.g., two layers for geometry, d0 and d1, and one layer for texture). SEI messages or new NAL units are used for occupancy maps and auxiliary patch information. The extraction of independently coded non-base layers and the signaling of compliance points (e.g., profile and level) as a single-layer bitstream are specified. Two new SEI message types are specified, one for occupancy maps and one for auxiliary information. Another sequence-level SEI message is specified to carry the frame group header parameters and other global information. This SEI message is similar to the frame group header 408 in the first group of methods.

[0100] The first set of methods can be implemented based on the following definitions, abbreviations, syntax, and semantics: Aspects not specifically mentioned are the same as the latest PCC Cat2 WD.

[0101] The following definitions apply.

[0102] Bitstream: A sequence of bits forming a representation of a coded point cloud frame and associated data forming one or more CPSs.

[0103] Byte: A sequence of 8 bits in which, when written or read as a sequence of bit values, the leftmost and rightmost bits represent the most and least significant bits, respectively.

[0104] Coded PCC sequence (CPS): A sequence of PCC AUs consisting, in decoding order, of a PCC intra random access picture (IRAP) AU, followed by zero or more PCC AUs that are not PCC IRAP AUs, including all subsequent PCC AUs, and ending with but not including any subsequent PCC AU that is a PCC IRAP AU.

[0105] Decoding order: The order in which the decoding process processes syntax elements.

[0106] Decoding process: The process specified in this specification (also known as PCC Cat2 WD) that reads the bitstream and derives a decoded point cloud frame from it.

[0107] Group of frames header NAL unit: PCC NAL unit with PccNalUnitType equal to GOF_HEADER.

[0108] PCC AU: A group of PCC NAL units that are associated with each other according to the specified classification rules, are consecutive in decoding order, and contain all PCC NAL units within a specific presentation time.

[0109] PCC IRAP AU: PCC AU containing the frame group header NAL unit.

[0110] PCC NAL unit: A syntax structure containing an indication of the type of data to follow, and the bytes containing that data in the form of an RBSP intermixed with an emulation prevention byte as required.

[0111] Raw byte sequence payload (RBSP): A syntax structure consisting of an integer number of bytes encapsulated in a PCC NAL unit. The bytes may be empty or in the form of a string of data bits (SODB) containing syntax elements, followed by an RBSP stop bit and a zero bit or multiple subsequent bits equal to zero.

[0112] Raw byte sequence payload (RBSP) stop bit: A bit equal to 1 that is present in the RBSP following the SODB. The end position within the RBSP can be identified by searching for the RBSP stop bit from the end of the RBSP. The stop bit is the last non-zero bit in the RBSP.

[0113] SODB: A sequence of bits representing the syntax elements present in the RBSP preceding the RBSP stop bit, where the leftmost bit is considered to be the first and most significant bit and the rightmost bit is considered to be the last and least significant bit.

[0114] Syntax element: A data element represented in a bitstream.

[0115] Syntax structure: Zero or more syntax elements, appearing consecutively in the bit stream in a specified order.

[0116] Video AU: An access unit for each specific video decoder.

[0117] Video NAL unit: PCC NAL unit with PccNalUnitType equal to GEOMETRY_D0, GEOMETRY_D0, or TEXTURE_NALU.

[0118] The following abbreviations apply:

[0119] AU:access unit

[0120] CPS: coded PCC sequence

[0121] IRAP: intra random access point

[0122] NAL: network abstraction layer network abstraction layer

[0123] PCC: point cloud coding

[0124] RBSP: raw byte sequence payload

[0125] SODB: string of data bits

[0126] The syntax, semantics, and sub-bitstream extraction process are provided below. In this regard, the syntax in clause 7.3 of the latest PCC Cat2 WD is replaced by the following.

[0127] The PCC NAL unit syntax is provided. In particular, the general PCC NAL unit syntax is as follows.

[0128]

[0129] The syntax of the PCC NAL unit header is as follows.

[0130]

[0131] The raw byte sequence payload, trailing bits, and byte alignment syntax are provided. Specifically, the frame group RBSP syntax is as follows.

[0132]

[0133]

[0134] The syntax of the auxiliary information frame RBSP is as follows.

[0135]

[0136] The syntax of the occupancy map frame RBSP is as follows.

[0137]

[0138]

[0139] The RBSP trailing bits syntax in section 7.3.2.11 of the HEVC specification applies. Similarly, the byte alignment syntax in section 7.3.2.12 of the HEVC specification applies. The PCC profile and level syntax is as follows.

[0140]

[0141] The semantics in clause 7.4 of the latest PCC Cat2 WD are replaced by the following and its subclauses.

[0142] In general, the semantics associated with syntax structures and syntax elements within those structures are specified in this subclause. Where a table or set of tables is used to specify syntax element semantics, then, unless otherwise specified, any values ​​not specified in the table shall not appear in the bitstream.

[0143] Regarding PCC NAL unit semantics. For general PCC NAL unit semantics, the general NAL unit semantics in section 7.4.2.1 of the HEVC specification applies. The PCC NAL unit header semantics are as follows.

[0144] forbidden_zero_bit should be equal to 0.

[0145] In bitstreams conforming to this version of the specification, pcc_nuh_reserved_zero_2bits shall be equal to 0. Other values ​​of pcc_nuh_reserved_zero_2bits are reserved for future use by ISO / IEC. Decoders shall ignore the value of pcc_nuh_reserved_zero_2bits.

[0146] pcc_nal_unit_type_plus1 minus 1 specifies the value of the variable PccNalUnitType, which specifies the type of RBSP data structure contained in the PCCNAL unit, as shown in Table 1 (see below). The variable NalUnitType is specified as follows:

[0147] PccNalUnitType=pcc_category2_nal_unit_type_plus1–1 (7-1)

[0148] PCC NAL units with a nal_unit_type in the range UNSPEC25..UNSPEC30 (whose semantics are unspecified), inclusive, shall not affect the decoding process specified in this specification.

[0149] NOTE 1 - PCC NAL unit types in the range UNSPEC25..UNSPEC30 may be used as determined by the application. This specification does not specify the decoding process for PccNalUnitType values. Because different applications may use these PCC NAL unit types for different purposes, caution should be exercised when designing encoders that generate PCC NAL units using these PccNalUnitType values, and when designing decoders that interpret the content of PCC NAL units using these PccNalUnitType values. This specification does not define any management for these values. These PccNalUnitType values ​​may only be used in situations where "conflicts" (e.g., different definitions of the meaning of the PCC NAL unit content for the same PccNalUnitType value) are unimportant or impossible, or are managed (e.g., defined or managed in the controlling application or transport specification, or by the controlling bitstream distribution environment).

[0150] Except for the purpose of determining the amount of data in a bitstream decoding unit, a decoder should ignore (remove from the bitstream and discard) the contents of all PCC NAL units that use the reserved value of PccNalUnitType.

[0151] NOTE 2 – Compatible extensions to this specification may be defined later.

[0152] Table 1 - PCC NAL unit type code

[0153]

[0154] NOTE 3 - The identified video codec (eg, HEVC or AVC) is indicated in the frame group header NAL unit present in the first PCC AU of each CPS.

[0155] Encapsulation of SODB within (informative) RBSP is provided. In this regard, section 7.4.2.3 of the HEVC specification applies.

[0156] The order of PCC NAL units and their association with AUs and CPS is provided. In general, this clause specifies constraints on the order of PCC NAL units in the bitstream.

[0157] Any order of PCC NAL units in a bitstream that adheres to these constraints is referred to herein as the decoding order of the PCC NAL units. Within PCC NAL units that are not video NAL units, the syntax in clause 7.3 specifies the decoding order of the syntax elements. Within video NAL units, the syntax specified in the specification of the identified video decoder specifies the decoding order of the syntax elements. A decoder is capable of receiving PCC NAL units and their syntax elements in decoding order.

[0158] The order of PCC NAL units and their association with PCC AUs is provided.

[0159] This clause specifies the order of PCC NAL units and their association with PCC AUs.

[0160] A PCC AU consists of zero or one frame group header NAL unit, one geometry d0 video AU, one geometry d1 video AU, one auxiliary information frame NAL unit, one occupancy map frame NAL unit, and one texture video AU, in the order listed.

[0161] The association of NAL units with video AUs and the order of NAL units in video AUs are specified in the specifications of the identified video codec (e.g., HEVC or AVC). The identified video codec is indicated in the frame header NAL unit present in the first PCC AU of each CPS.

[0162] The first PCC AU of each CPS starts with a frame group header NAL unit, and each frame group header NAL unit specifies the start of a new PCC AU.

[0163] Other PCC AUs start with a PCC NAL unit that contains the first NAL unit of the geometry d0 video AU. In other words, when the PCC NAL unit containing the first NAL unit of the geometry d0 video AU is not preceded by a frame group header NAL unit, a new PCC AU begins.

[0164] The order of PCC AUs and their association with CPS is provided.

[0165] A bitstream conforming to this specification consists of one or more CPSs.

[0166] A CPS consists of one or more PCC AUs. Clause 7.4.2.4.2 describes the order of PCC NAL units and their association with PCC AUs.

[0167] The first PCC AU of CPS is the PCC IRAP AU.

[0168] The raw byte sequence payload, trailing bits, and byte alignment semantics are provided. The semantics of the RBSP frame group header are as follows.

[0169] identified_codec specifies the identified video codec used to decode the geometry component and the texture component, as shown in Table 2.

[0170] identified_codec identified_codec name Identified video decoder 0 CODEC_HEVC ISO / IEC IS 23008-2 (HEVC) 1 CODEC_AVC ISO / IEC IS 14496-10(AVC) 2..63 CODEC_RSV_2..CODEC_RSV_63 reserve

[0171] frame_width indicates the frame width of geometry video and texture video in pixels. frame_width should be a multiple of occupancyResolution.

[0172] frame_height indicates the frame height of geometry video and texture video in pixels. frame_height should be a multiple of occupancyResolution.

[0173] occupancy_resolution indicates the horizontal and vertical resolution (in pixels) at which patches are packed in the geometry video and texture video. The value of occupancy_resolution should be an even multiple of occupancyPrecision.

[0174] radius_to_smoothing indicates the radius within which adjacent areas are detected for smoothing. The value of radius_to_smoothing should be between 0 and 255 (inclusive).

[0175] neighbor_count_smoothing indicates the maximum number of neighboring areas used for smoothing. The value of neighbor_count_smoothing should be between 0 and 255 (inclusive).

[0176] radius2_boundary_detection indicates the radius of boundary point detection. The value of radius2_boundary_detection should be between 0 and 255 (inclusive).

[0177] threshold_smoothing indicates the smoothing threshold. The value of threshold_smoothing should be between 0 and 255 (inclusive).

[0178] lossless_geometry indicates lossless geometry coding. A lossless_geometry value of 1 indicates that the point cloud geometry information is encoded in a lossless manner. A lossless_geometry value of 0 indicates that the point cloud geometry information is encoded in a lossy manner.

[0179] lossless_texture indicates lossless texture encoding. A lossless_texture value equal to 1 indicates that the point cloud texture information is encoded in a lossless manner. A lossless_texture value equal to 0 indicates that the point cloud texture information is encoded in a lossy manner.

[0180] no_attributes indicates whether attributes are encoded with the geometry data. A value of no_attributes equal to 1 indicates that the encoded point cloud bitstream does not contain any attribute information. A value of no_attributes equal to 0 indicates that the encoded point cloud bitstream contains attribute information.

[0181] lossless_geometry_444 indicates whether the geometry frame uses the 4:2:0 video format or the 4:4:4 video format. A value of lossless_geometry_444 equal to 1 indicates that the geometry video is encoded in the 4:4:4 format. A value of lossless_geometry_444 equal to 0 indicates that the geometry video is encoded in the 4:2:0 format.

[0182] absolute_d1_coding indicates how the geometry layers other than the layer closest to the projection plane are coded. A value of absolute_d1_coding equal to 1 indicates that the actual geometry values ​​of the geometry layers other than the layer closest to the projection plane are coded. absolute_d1_coding equal to 0 indicates that the geometry layers other than the layer closest to the projection plane are differentially coded.

[0183] bin_arithmetic_coding indicates whether binary arithmetic coding is used. A value of bin_arithmetic_coding equal to 1 indicates that all syntax elements use binary arithmetic coding. A value of bin_arithmetic_coding equal to 0 indicates that some syntax elements use non-binary arithmetic coding.

[0184] A gof_header_extension_flag value of 0 specifies that the gof_header_extension_data_flag syntax element is not present in the frame-group header RBSP syntax structure. A gof_header_extension_flag value of 1 specifies that the gof_header_extension_data_flag syntax element is present in the frame-group header RBSP syntax structure. A decoder shall ignore all data following a gof_header_extension_flag value of 1 in a frame-group header NAL unit.

[0185] gof_header_extension_data_flag can take any value. Its presence and value do not affect decoder compliance. Decoders should ignore all syntax elements of gof_header_extension_data_flag.

[0186] Provides the semantics of the auxiliary information frame RBSP.

[0187] patch_count indicates the number of patches in the geometry video and texture video. patch_count should be greater than 0.

[0188] occupancy_precision is the horizontal and vertical resolution of the occupancy map precision (in pixels). It corresponds to the signaled occupancy-related subblock size. To achieve lossless decoding of occupancy maps, occupancy_precision should be set to 1.

[0189] max_candidate_count specifies the maximum number of candidates in the patch candidate list.

[0190] bit_count_u0 specifies the number of bits for the fixed-length decoding of patch_u0.

[0191] bit_count_v0 specifies the number of bits for fixed-length decoding of patch_v0.

[0192] bit_count_u1 specifies the number of bits for fixed-length decoding of patch_u1.

[0193] bit_count_v1 specifies the number of bits for fixed-length decoding of patch_v1.

[0194] bit_count_d1 specifies the number of bits for fixed-length decoding of patch_d1.

[0195] occupancy_aux_stream_size specifies the number of bytes used to decode patch information and occupancy maps.

[0196] The following syntax elements are indicated once per patch.

[0197] patch_u0 specifies the x-coordinate of the top left corner of the patch bounding box of size occupancy_resolution × occupancy_resolution. The value of patch_u0 should be between 0 and frame_width / occupancy_resolution-1 (inclusive).

[0198] patch_v0 specifies the y coordinate of the top left corner of the patch bounding box of size occupancy_resolution × occupancy_resolution. The value of patch_v0 should be between 0 and frame_height / occupancy_resolution-1 (inclusive).

[0199] patch_u1 specifies the minimum x-coordinate of the 3D bounding box of the patch points. The value of patch_u1 should be between 0 and frame_width-1 (inclusive).

[0200] patch_v1 is the minimum y-coordinate of the 3D bounding box of the patch points. The value of patch_v1 should be between 0 and frameHeight-1 (inclusive).

[0201] patch_d1 specifies the minimum depth of the patch. The value of patch_d1 should be between 0 and <255> between (including the end values).

[0202] delta_size_u0 is the difference in patch width between the current patch and the previous patch. The value of delta_size_u0 should be between <-65536> and <65535> between (including the end values).

[0203] delta_size_v0 is the difference in patch height between the current patch and the previous patch. The value of delta_size_v0 should be between <-65536> and <65535> between (including the end values).

[0204] normal_axis specifies the plane projection index. The value of normal_axis should be between 0 and 2 (inclusive). The values ​​of normal_axis 0, 1, and 2 correspond to the X, Y, and Z projection axes respectively.

[0205] The following syntax elements are specified once per block.

[0206] candidate_index is the index into the patch candidate list. The value of candidate_index should be between 0 and max_candidate_count (inclusive).

[0207] patch_index is the index into the sorted list of patches associated with the frame, in descending order of size.

[0208] Provides the semantics of the set of frame occupancy maps.

[0209] The following syntax elements are provided for non-empty blocks.

[0210] is_full specifies whether the current occupied block of size occupancy_resolution × occupancy_resolution blocks is full. is_full equal to 1 specifies that the current block is full. is_full equal to 0 specifies that the current occupied block is not full.

[0211] best_traversal_order_index specifies the scanning order of sub-blocks of size occupancy_precision × occupancy_precision in the current occupancy_resolution × occupancy_resolution block. The value of best_traversal_order_index should be between 0 and 4 (inclusive).

[0212] The run_count_prefix is ​​used in the export of the variable runCountMinusTwo.

[0213] run_count_suffix is ​​used to derive the variable runCountMinusTwo. If run_count_suffix does not exist, the value of run_count_suffix is ​​inferred to be 0.

[0214] When the value of blockToPatch for a particular block is not equal to 0 and the block is not full, the value of runCountMinusTwo plus 2 represents the signaled number of runs of the block. The value of runCountMinusTwo shall be between 0 and (occupancy_resolution * occupancy_resolution) - 1 (inclusive).

[0215] runCountMinusTwo is derived as follows:

[0216] runCountMinusTwo=(1< <run_count_prefix)-1+run_count_suffix (7-85)

[0217] occupancy specifies the occupancy value of the first sub-block (with occupancyPrecision×occupancyPrecision pixels). An occupancy value of 0 specifies that the first sub-block is empty. An occupancy value of 1 specifies that the first sub-block is occupied.

[0218] run_length_idx indicates the run length. The value of runLengthIdx should be between 0 and 14 (inclusive).

[0219] Use Table 3 to derive the variable runLength from run_length_idx.

[0220] Table 3 – Deriving runLength from run_length_idx

[0221]

[0222]

[0223] NOTE - The occupancy map is shared by both the geometry video and the texture video.

[0224] The RBSP trailing bit semantics in section 7.4.3.11 of the HEVC specification apply. The byte alignment semantics in section 7.4.3.12 of the HEVC specification also apply. PCC profile and level semantics are as follows.

[0225] pcc_profile_idc indicates the profile specified in Annex A to which the CPS conforms. The bitstream shall not contain values ​​for pcc_profile_idc other than those specified in Annex A. Other values ​​of pcc_profile_idc are reserved for future use by ISO / IEC.

[0226] In bitstreams conforming to this version of the specification, pcc_pl_reserved_zero_19bits shall be equal to 0. Other values ​​of pcc_pl_reserved_zero_19bits are reserved for future use by ISO / IEC. Decoders shall ignore the value of pcc_pl_reserved_zero_19bits.

[0227] pcc_level_idc indicates the level specified in Annex A to which the CPS conforms. The bitstream shall not contain values ​​of pcc_level_idc other than those specified in Annex A. Other values ​​of pcc_level_idc are reserved for future use by ISO / IEC.

[0228] In the case where a geometry HEVC bitstream extracted as specified in clause 10 is decoded by a conforming HEVC decoder, hevc_ptl_12bytes_geometry shall be equal to the value of the 12 bytes between general_profile_idc and general_level_idc (inclusive) in the activation SPS.

[0229] In the case where a texture HEVC bitstream extracted as specified in clause 10 is decoded by a conforming HEVC decoder, hevc_ptl_12bytes_texture shall be equal to the value of the 12 bytes between general_profile_idc and general_level_idc (inclusive) in the activation SPS.

[0230] In the case where a geometry AVC bitstream extracted as specified in clause 10 is decoded by a conforming AVC decoder, avc_pl_3ytes_geometry shall be equal to the value of the 3 bytes between profile_idc and level_idc (inclusive) in the activation SPS.

[0231] In the case where a geometry AVC bitstream extracted as specified in clause 10 is decoded by a conforming AVC decoder, avc_pl_3ytes_texture shall be equal to the value of the 3 bytes between profile_idc and level_idc (inclusive) in the activation SPS.

[0232] The sub-bitstream extraction process in clause 104 of the latest PCC Cat2 WD is replaced by the following: The input of the sub-bitstream extraction process is the bitstream, geometry d0, geometry d1, or target video component indication of texture components. The output of the process is the sub-bitstream.

[0233] In one embodiment, the bitstream conformance requirement for the input bitstream is that any output sub-bitstream that is the output of the process specified in this clause with a conforming PCC bitstream, and any value of the target video component indication, shall be a conforming video bitstream for each identified video codec.

[0234] The output sub-bitstream is derived in order through the following steps.

[0235] Depending on the value indicated by the target video component, the following applies.

[0236] If the geometry d0 component is indicated, all PCCNAL units with PccNalUnitType not equal to GEOMETRY_D0 are deleted.

[0237] Otherwise, if the geometry d1 component is indicated, delete all PCC NAL units with PccNalUnitType not equal to GEOMETRY_D1.

[0238] Otherwise (texture component is indicated), delete all PCCNAL units whose PccNalUnitType is not equal to TEXTURE_NALU.

[0239] For each PCC NAL unit, the first byte is deleted.

[0240] Another embodiment is provided below.

[0241] In another embodiment of the first group of methods summarized above, the PCC NAL unit header (e.g., Figure 4 The design of the data unit header 404 in the PCC NAL unit allows the decoder used to decode the geometry component and the texture component to be inferred from the type of the PCC NAL unit. For example, the design of the PCC NAL unit header is summarized as follows:

[0242] There is a type indicator (e.g., 7 bits) in the PCC NAL unit header that specifies the type of content carried in the PCC NAL unit payload. For example, the type is determined based on the following:

[0243] 0: Payload contains HEVC NAL units

[0244] 1: Payload contains AVC NAL units

[0245] 2..63: Reserved

[0246] 64: Frame group header NAL unit

[0247] 65: Auxiliary information NAL unit

[0248] 66: Occupancy map NAL unit

[0249] 67..126: Reserved

[0250] A PCC NAL unit having a PCC NAL unit type between 0 and 63 (inclusive) is referred to as a video NAL unit.

[0251] Some reserved PCC NAL unit types can be used for PCC Cat13 data units, thereby unifying PCC Cat2 and PCC Cat13 in one standard specification.

[0252] Figure 5 is an embodiment of a method 500 of point cloud coding implemented by a video decoder, such as video decoder 30. Method 500 may be performed to address one or more of the aforementioned issues associated with point cloud coding.

[0253] In block 502, an encoded bitstream (e.g., data structure 400) is received that includes a frame group header (e.g., frame group header 408) that is not part of or disposed outside of an encoded representation of a texture component or a geometry component, and that identifies a video decoder for decoding the texture component or the geometry component.

[0254] The encoded bitstream is decoded in block 504. The decoded bitstream may be used to generate an image or video for display to a user on a display device.

[0255] In one embodiment, the coded representation is included in the payload of a PCC Network Abstraction Layer (NAL) unit. In one embodiment, the payload of the PCC NAL unit includes a geometry component corresponding to a video decoder. In one embodiment, the payload of the PCC NAL unit includes a texture component corresponding to a video decoder.

[0256] In one embodiment, the video codec is High Efficiency Video Coding (HEVC). In one embodiment, the video codec is Advanced Video Coding (AVC). In one embodiment, the video codec is Versatile Video Coding (VVC). In one embodiment, the video codec is Basic Video Coding (EVC).

[0257] In one embodiment, the geometric component comprises a set of coordinates associated with the point cloud frame. In one embodiment, the set of coordinates is Cartesian coordinates.

[0258] In one embodiment, the texture component comprises a set of luminance sample values ​​of the point cloud frame.

[0259] Figure 6 is an embodiment of a method 600 of point cloud coding implemented by a video encoder, such as video encoder 20. Method 600 may be performed to address one or more of the aforementioned issues associated with point cloud coding.

[0260] In block 602, an encoded bitstream (e.g., data structure 400) is generated that includes a frame group header (e.g., frame group header 408). The frame group header is provided outside of an encoded representation of a texture component or a geometry component and identifies a video decoder for decoding the texture component or the geometry component.

[0261] In block 604, the encoded bitstream is sent to a decoder (eg, video decoder 30). Once the decoder receives the encoded bitstream, it may decode the encoded bitstream to generate an image or video for display to a user on a display device.

[0262] In one embodiment, the coded representation is included in the payload of a PCC Network Abstraction Layer (NAL) unit. In one embodiment, the payload of the PCC NAL unit includes a geometry component corresponding to a video decoder. In one embodiment, the payload of the PCC NAL unit includes a texture component corresponding to a video decoder.

[0263] In one embodiment, the video codec is High Efficiency Video Coding (HEVC). In one embodiment, the video codec is Advanced Video Coding (AVC). In one embodiment, the video codec is Versatile Video Coding (VVC). In one embodiment, the video codec is Basic Video Coding (EVC).

[0264] In one embodiment, the geometric component comprises a set of coordinates associated with the point cloud frame. In one embodiment, the set of coordinates is Cartesian coordinates.

[0265] In one embodiment, the texture component comprises a set of luminance sample values ​​of the point cloud frame.

[0266] Figure 7 is a schematic diagram of a video decoding device 700 (e.g., video encoder 20, video decoder 30, etc.) according to an embodiment of the present disclosure. The video decoding device 700 is suitable for implementing the methods and processes of the present disclosure. The video decoding device 700 includes an inlet port 710 and a receiver unit (Rx) 720 for receiving data; a processor, logic unit, or central processing unit (CPU) 730 for processing data; a transmitter unit (Tx) 740 and an outlet port 750 for sending data; and a memory 760 for storing data. The video decoding device 700 may also include an optical-to-electrical (OE) component and an electro-optical (EO) component connected to the inlet port 710, the receiver unit 720, the transmitter unit 740, and the outlet port 750 for outputting or inputting optical signals or electrical signals.

[0267] The processor 730 is implemented by hardware and software. The processor 730 can be implemented by one or more CPU chips, cores (for example, as a multi-core processor), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 730 communicates with the inlet port 710, the receiver unit 720, the transmitter unit 740, the outlet port 750, and the memory 760. The processor 730 includes a decoding module 770. The decoding module 770 implements the above-mentioned embodiments of the present disclosure. Therefore, the inclusion of the decoding module 770 provides substantial improvements to the functionality of the video decoding device 700 and enables the video decoding device 700 to transition to different states. Optionally, the decoding module 770 is implemented as instructions stored in the memory 760 and executed by the processor 730.

[0268] The video decoding device 700 may further include an input and / or output (I / O) device 780 for communicating data with a user. The I / O device 780 may include an output device, such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 780 may also include an input device, such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with the output devices.

[0269] Memory 760 includes one or more disks, tape drives, and solid-state drives and can be used as an overflow data storage device to store programs when they are selected for execution and to store instructions and data read during program execution. Memory 760 can be volatile or non-volatile and can be read-only memory (ROM), random-access memory (RAM), ternary content-addressable memory (TCAM), and static random-access memory (SRAM).

[0270] Figure 8 is a schematic diagram of an embodiment of a decoding apparatus 800. In one embodiment, decoding apparatus 800 is implemented in a video decoding device 802 (e.g., video encoder 20 or video decoder 30). Video decoding device 802 includes a receiving device 801. Receiving device 801 is configured to receive a picture to be encoded or a bitstream to be decoded. Video decoding device 802 includes a transmitting device 807 connected to receiving device 801. Transmitting device 807 is configured to transmit the bitstream to a decoder or transmit the decoded image to a display device (e.g., one of I / O devices 780).

[0271] Video decoding device 802 includes a storage device 803. Storage device 803 is connected to at least one of receiving device 801 or transmitting device 807. Storage device 803 is used to store instructions. Video decoding device 802 also includes a processing device 805. Processing device 805 is connected to storage device 803. Processing device 805 is used to execute the instructions stored in storage device 803 to perform the method of the present disclosure.

[0272] Although a number of embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered illustrative and not restrictive, and the invention is not to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0273] In addition, without departing from the scope of the present disclosure, the techniques, systems, subsystems, and methods described and shown as discrete or separate in the various embodiments may be combined or integrated with other systems, modules, techniques, or methods. Those skilled in the art may determine and make other examples of changes, substitutions, and modifications without departing from the spirit and scope of the present disclosure.

Claims

1. A point cloud coding method, characterized in that: include: Generate a coded bit stream, the coded bit stream including a point cloud data frame group header unit whose type value is 0, the coded bit stream further including a first point cloud data unit and / or a second point cloud data unit after the point cloud data frame group header unit, the first point cloud data unit including an encoded representation of a texture component of the point cloud data, the second point cloud data unit including an encoded representation of a geometric component of the point cloud data, the point cloud data frame group header unit including a decoder identifier, the decoder identifier being used to identify a video decoder used to decode the texture component and / or the geometric component; as well as The encoded bit stream is transmitted.

2. The method according to claim 1, characterized in that The video coder is one of High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Versatile Video Coding (VVC), and Essential Video Coding (EVC).

3. The method according to claim 1, wherein The point cloud data frame group header unit further includes a configuration file of the texture component and / or a configuration file of the geometry component.

4. The method according to claim 1, wherein The point cloud data frame group header unit further includes the level of the texture component and / or the level of the geometry component.

5. The method according to claim 1, wherein The point cloud data frame group header unit is located in a random access point IRAP access unit AU (PCC IRAP AU) within a point cloud coding (PCC) frame.

6. A point cloud decoding method, characterized in that: include: Receive a coded bitstream, the coded bitstream including a point cloud data frame group header unit whose type value is 0, the coded bitstream further including a first point cloud data unit and / or a second point cloud data unit after the point cloud data frame group header unit, the first point cloud data unit including an encoded representation of a texture component of the point cloud data, the second point cloud data unit including an encoded representation of a geometric component of the point cloud data, the point cloud data frame group header unit including a decoder identifier, the decoder identifier being used to identify a video decoder used to decode the texture component and / or the geometric component; as well as The encoded bitstream is decoded.

7. The method according to claim 6, characterized in that The video coder is one of High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Versatile Video Coding (VVC), and Essential Video Coding (EVC).

8. The method according to claim 6, wherein The point cloud data frame group header unit further includes a configuration file of the texture component and / or a configuration file of the geometry component.

9. The method according to claim 6, wherein The point cloud data frame group header unit further includes the level of the texture component and / or the level of the geometry component.

10. The method according to claim 6, wherein The point cloud data frame group header unit is located in a random access point IRAP access unit AU (PCC IRAP AU) within a point cloud coding (PCC) frame.

11. A coding device, characterized in that: The device comprises: one or more processors; A computer-readable storage medium, coupled to the processor and storing program instructions executed by the processor, wherein the program instructions, when executed by the processor, cause the device to perform the method according to any one of claims 1 to 5.

12. A decoding device, characterized in that: The device comprises: one or more processors; A computer-readable storage medium, coupled to the processor and storing program instructions executed by the processor, wherein the program instructions, when executed by the processor, cause the device to perform the method according to any one of claims 6 to 10.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, wherein when the program instructions are executed by a device or one or more processors, the device executes the method according to any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a coded bit stream obtained by executing the method according to any one of claims 1 to 5 by one or more processors.

15. A method for storing an encoded bit stream, characterized in that The method comprises: Receiving the encoded bit stream generated by the method according to any one of claims 1 to 5; and The encoded bit stream is stored in a storage medium.

16. A system for storing an encoded bit stream, characterized in that include: A receiver for receiving the encoded bit stream generated by the method according to any one of claims 1 to 5; as well as, A storage medium is used to store the encoded bit stream.

17. A method for transmitting an encoded bit stream, characterized in that The method comprises: Acquire an encoded bit stream from a storage medium, wherein the encoded bit stream is generated by the method according to any one of claims 1 to 5 and is stored in the storage medium; and The encoded bit stream is transmitted.

18. A system for transmitting an encoded bit stream, characterized in that The system comprises: a processor, configured to obtain a coded bit stream from a storage medium, wherein the coded bit stream is a coded bit stream generated by the method according to any one of claims 1 to 5 and stored in the storage medium; and A transmitter is configured to transmit the encoded bit stream.

19. A storage device storing a coded bit stream generated by the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and apparatus for encoding and decoding multilayer videos

    US20120044999A1