Method, device and medium for point cloud coding and decoding

By skipping at least one of the prediction or transformation of the transform block, the problem of high complexity of point cloud attribute encoding and decoding in RAHT is solved, and the efficiency of point cloud encoding and decoding is improved.

CN120092264APending Publication Date: 2025-06-03DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071249.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-04
Filing Date
2023-09-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing point cloud encoding and decoding technology has high complexity in regional adaptive hierarchical transformation (RAHT), resulting in inefficient efficiency.

Method used

By skipping at least one of the prediction or transformation of the transform block, a transform block including a predetermined number of sub-blocks, each sub-block containing at least one point of the point cloud sequence, and on this basis the conversion between the point cloud sequence and the bit stream is performed.

Benefits of technology

Reduces the complexity of prediction, transformation and related operations, and improves the efficiency of point cloud encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120092264A_ABST
    Figure CN120092264A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for point cloud coding and decoding. The method comprises: for a conversion between a point cloud sequence comprising a current point cloud (PC) sample associated with a transform block and a bitstream of the point cloud sequence, determining the transform block comprising a predetermined number of sub-blocks, each sub-block comprising at least one point of the point cloud sequence; and performing the transform by skipping at least one of prediction or transform of the transform block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to point cloud encoding and decoding techniques, and more particularly, to improving the encoding and decoding complexity of point cloud attribute in Region Adaptive Hierarchical Transform (RAHT). Background Art

[0002] A point cloud is a collection of individual data points in a three-dimensional (3D) plane, where each point has set coordinates on the X-axis, Y-axis, and Z-axis. Thus, point clouds can be used to represent the physical content of three-dimensional space. For various immersive applications ranging from augmented reality to autonomous vehicles, point clouds have proven to be a promising way to represent 3D visual data.

[0003] Point cloud encoding and decoding standards have mainly evolved through the well-known MPEG organization. MPEG is short for Moving Picture Experts Group, which is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Encoding and Decoding Group (3DG) released a Call for Proposals (CFP) document to start the development of point cloud encoding and decoding standards. The final standard will encompass two categories of solutions. Video-based Point Cloud Compression (V-PCC or VPCC) is applicable to point sets with relatively uniform point distributions. Geometry-based Point Cloud Compression (G-PCC or GPCC) is applicable to more sparse distributions. However, there is generally a desire to further improve the encoding and decoding efficiency of conventional point cloud encoding and decoding techniques. Summary of the Invention

[0004] Embodiments of the present disclosure provide a solution for point cloud encoding and decoding.

[0005] In a first aspect, a method for point cloud encoding and decoding is proposed. The method includes: determining a transform block including a predetermined number of sub-blocks for the conversion between a point cloud sequence including current point cloud (PC) samples associated with the transform block and a bitstream of the point cloud sequence, each sub-block including at least one point of the point cloud sequence; and performing the conversion by skipping at least one of prediction or transformation of the transform block.

[0006] In a second aspect, a device for point cloud encoding and decoding is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a point cloud sequence generated by a method executed by a point cloud processing device. The method includes: determining a transform block including a predetermined number of sub-blocks, each sub-block containing at least one point of the point cloud sequence, the transform block being associated with a current point cloud (PC) sample included in the point cloud sequence; and generating the bitstream by skipping at least one of prediction or transformation of the transform block.

[0009] In a fifth aspect, a method for storing a bitstream of a point cloud sequence is proposed. The method includes: determining a transform block including a predetermined number of sub-blocks, each sub-block containing at least one point of the point cloud sequence, the transform block being associated with a current point cloud (PC) sample included in the point cloud sequence; generating the bitstream by skipping at least one of prediction or transformation of the transform block; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] The present invention content is provided to introduce a selection of concepts further described below in the detailed implementation in a simplified form. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1 is a block diagram showing an exemplary point cloud encoding and decoding system that can utilize the technology of the present disclosure;

[0013] Figure 2 shows a block diagram of an exemplary point cloud encoder according to some embodiments of the present disclosure;

[0014] Figure 3 shows a block diagram of an exemplary point cloud decoder according to some embodiments of the present disclosure;

[0015] Figure 4 shows an exemplary diagram of a parent-level node for each sub-node of a transform unit node according to some embodiments of the present disclosure;

[0016] Figure 5 shows an exemplary diagram of an encoding and decoding process for improving the encoding and decoding complexity of point cloud attributes according to some embodiments of the present disclosure;

[0017] Figure 6A flowchart of a method for point cloud encoding and decoding according to an embodiment of the present disclosure is shown; and

[0018] Figure 7 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0019] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description

[0020] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.

[0021] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0022] References herein to "one embodiment", "an embodiment", "example embodiment", etc., indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of one of ordinary skill in the art in relation to other embodiments.

[0023] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0024] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes", and / or "including" when used herein indicate the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0025] Example environment

[0026] Figure 1 is a block diagram showing an example point cloud encoding / decoding system 100 that can utilize the techniques of the present disclosure. As shown, the point cloud encoding / decoding system 100 can include a source device 110 and a destination device 120. The source device 110 can also be referred to as a point cloud encoding / decoding device, and the destination device 120 can also be referred to as a point cloud decoding device. In operation, the source device 110 can be configured to generate encoded point cloud data, and the destination device 120 can be configured to decode the encoded point cloud data generated by the source device 110. The techniques of the present disclosure generally relate to encoding and decoding (encoding and / or decoding) point cloud data, i.e., supporting point cloud compression. Encoding and decoding can be effective in compressing and / or decompressing point cloud data.

[0027] The source device 100 and the destination device 120 can include any of a variety of devices, including desktop computers, notebooks (i.e., laptops), tablet computers, set-top boxes, telephone handsets (such as smartphones and mobile phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, vehicles (e.g., land or sea vehicles, spacecraft, aircraft, etc.), robots, LIDAR devices, satellites, extended reality devices, and so on. In some cases, the source device 100 and the destination device 120 can be equipped for wireless communication.

[0028] The source device 100 can include a data source 112, a memory 114, a GPCC encoder 116, and an input / output (I / O) interface 118. The destination device 120 can include an input / output (I / O) interface 128, a GPCC decoder 126, a memory 124, and a data consumer 122. According to the present disclosure, the GPCC encoder 116 of the source device 100 and the GPCC decoder 126 of the destination device 120 can be configured to apply the techniques of the present disclosure related to point cloud encoding and decoding. Thus, the source device 100 represents an example of an encoding device, and the destination device 120 represents an example of a decoding device. In other examples, the source device 100 and the destination device 120 can include other components or arrangements. For example, the source device 100 can receive data (e.g., point cloud data) from an internal source or an external source. Similarly, the destination device 120 can interface with an external data consumer instead of including a data consumer in the same device.

[0029] Generally, data source 112 represents a source of point cloud data (i.e., raw, unencoded point cloud data) and can provide a continuous series of “frames” of point cloud data to GPCC encoder 116, which encodes the point cloud data for the frames. In some examples, data source 112 generates the point cloud data. The data source 112 of source device 100 can include a point cloud acquisition device, such as any of a variety of cameras or sensors, e.g., one or more cameras, an archive containing previously acquired point cloud data, a 3D scanner or a light detection and ranging (LIDAR) device, and / or a data feed interface that receives point cloud data from a data content provider. Thus, in some examples, data source 112 can generate point cloud data based on signals from a LIDAR device. Alternatively or additionally, the point cloud data can be generated from a scanner, camera, sensor, or other data by a computer. For example, data source 112 can generate point cloud data, or a combination of real-time point cloud data, archived point cloud data, and computer-generated point cloud data. In each case, GPCC encoder 116 encodes the acquired, pre-acquired, or computer-generated point cloud data. GPCC encoder 116 can rearrange the frames of point cloud data from the received order (sometimes referred to as “display order”) to a codec order for encoding and decoding. GPCC encoder 116 can generate one or more bitstreams including the encoded point cloud data. Source device 100 can then output the encoded point cloud data via I / O interface 118 for reception and / or retrieval by, e.g., the I / O interface 128 of destination device 120. The encoded point cloud data can be directly transmitted to destination device 120 via network 130A through I / O interface 118. The encoded point cloud data can also be stored on storage medium / server 130B for access by destination device 120.

[0030] The memory 114 of the source device 100 and the memory 124 of the destination device 120 may represent general-purpose memories. In some examples, the memories 114 and 124 may store raw point cloud data, e.g., raw point cloud data from the data source 112 and raw, decoded point cloud data from the GPCC decoder 126. Additionally or alternatively, the memories 114 and 124 may store software instructions executable by, for example, the GPCC encoder 116 and the GPCC decoder 126, respectively. Although the memories 114 and 124 are shown separately from the GPCC encoder 116 and the GPCC decoder 126 in this example, it should be understood that the GPCC encoder 116 and the GPCC decoder 126 may also include internal memories for functionally similar or equivalent purposes. Further, the memories 114 and 124 may store encoded point cloud data, e.g., encoded point cloud data output from the GPCC encoder 116 and input to the GPCC decoder 126. In some examples, portions of the memories 114 and 124 may be allocated as one or more caches, e.g., for storing raw point cloud data, decoded point cloud data, and / or encoded point cloud data. For example, the memories 114 and 124 may store point cloud data.

[0031] The I / O interfaces 118 and 128 may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where the I / O interfaces 118 and 128 include wireless components, the I / O interfaces 118 and 128 may be configured to transfer data, such as encoded point cloud data, according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc. In some examples where the I / O interface 118 includes a wireless transmitter, the I / O interfaces 118 and 128 may be configured to transfer data, such as encoded point cloud data, according to other wireless standards such as the IEEE 802.11 specifications. In some examples, the source device 100 and / or the destination device 120 may include respective system-on-a-chip (SoC) devices. For example, the source device 100 may include an SoC device for performing the functions attributed to the GPCC encoder 116 and / or the I / O interface 118, and the destination device 120 may include an SoC device for performing the functions attributed to the GPCC decoder 126 and / or the I / O interface 128.

[0032] The techniques of the present disclosure can be applied to encoding and decoding to support any one of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors, and processing devices (e.g., local or remote servers), geographic mapping, or other applications.

[0033] The I / O interface 128 of the destination device 120 receives the encoded bitstream from the source device 110. The encoded bitstream can include signaling information defined by the GPCC encoder 116, which is also used by the GPCC decoder 126, such as syntax elements having values representing point clouds. The data consumer 122 uses the decoded data. For example, the data consumer 122 can use the decoded point cloud data to determine the location of a physical object. In some examples, the data consumer 122 can include a display for presenting an image based on the point cloud data.

[0034] Each of the GPCC encoder 116 and the GPCC decoder 126 can be implemented as any one of a variety of suitable encoder circuit systems and / or decoder circuit systems, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the GPCC encoder 116 and the GPCC decoder 126 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device. Devices including the GPCC encoder 116 and / or the GPCC decoder 126 can include one or more integrated circuits, microprocessors, and / or other types of devices.

[0035] The GPCC encoder 116 and the GPCC decoder 126 can operate according to encoding and decoding standards, such as the Video Point Cloud Compression (VPCC) standard or the Geometry Point Cloud Compression (GPCC) standard. Generally, the present disclosure can refer to the encoding and decoding (e.g., encoding and decoding) of frames to include the process of encoding data or decoding data. The encoded bitstream typically includes a series of values for syntax elements representing encoding and decoding decisions (e.g., encoding and decoding modes).

[0036] A point cloud can include a set of points in 3D space and can have attributes associated with the points. The attributes can be color information, such as R, G, B or Y, Cb, Cr or reflectivity information, or other attributes. The point cloud can be acquired by various cameras or sensors (such as LIDAR sensors and 3D scanners) and can also be computer-generated. Point cloud data is used in various applications, including but not limited to construction (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors for assisting navigation).

[0037] Figure 2 is a block diagram showing an example of a GPCC encoder 200 according to some embodiments of the present disclosure. The GPCC encoder 200 can be Figure 1 an example of the GPCC encoder 116 in the system 100 shown. Figure 3 is a block diagram showing an example of a GPCC decoder 300 according to some embodiments of the present disclosure. The GPCC decoder 300 can be Figure 1 an example of the GPCC decoder 126 in the system 100 shown.

[0038] In both the GPCC encoder 200 and the GPCC decoder 300, the point cloud positions are first encoded and decoded. The attribute encoding and decoding depend on the decoded geometry. In Figure 2 and Figure 3 , the region adaptive hierarchical transform (RAHT) unit 218, the surface approximation analysis unit 212, the RAHT unit 314, and the surface approximation synthesis unit 310 are options typically used for category 1 data. The level of detail (LOD) generation unit 220, the lifting unit 222, the LOD generation unit 316, and the inverse lifting unit 318 are options typically used for category 3 data. All other units are common between category 1 and category 3.

[0039] For category 3 data, the compressed geometry is typically represented as an octree from the root all the way down to the leaf level of individual voxels. For category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to the leaf level of blocks larger than voxels) plus a model for approximating the surface within each leaf node of the pruned octree. In this way, both category 1 data and category 3 data share the octree encoding and decoding mechanism, while category 1 data can additionally utilize the surface model to approximate the voxels within each leaf node. The surface model used is a triangulation including 1 to 10 triangles per block, resulting in a triangle soup. Therefore, the category 1 geometry codec is referred to as a triangle soup (Trisoup) geometry codec, while the category 3 geometry codec is referred to as an octree geometry codec.

[0040] InFigure 2 In the example of, the GPCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometry reconstruction unit 216, a RAHT unit 218, a LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.

[0041] As Figure 2 shown in the example of, the GPCC encoder 200 may receive a set of positions and a set of attributes. The positions may include the coordinates of the points in the point cloud. The attributes may include information about the points in the point cloud, such as the color associated with the points in the point cloud.

[0042] The coordinate transformation unit 202 may apply a transformation to the coordinates of the points to transform the coordinates from an initial domain to a transformed domain. The present disclosure may refer to the transformed coordinates as the transformation coordinates. The color transformation unit 204 may apply a transformation to convert the color information of the attributes to a different domain. For example, the color transformation unit 204 may convert the color information from the RGB color space to the YCbCr color space.

[0043] In addition, in Figure 2 the example of, the voxelization unit 206 may voxelize the transformation coordinates. The voxelization of the transformation coordinates may include quantizing and removing some of the points of the point cloud. In other words, multiple points of the point cloud may be grouped into a single "voxel", which may thereafter be regarded as a point in some aspects. In addition, the octree analysis unit 210 may generate an octree based on the voxelized transformation coordinates. Additionally, in Figure 2 the example of, the surface approximation analysis unit 212 may analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 may perform arithmetic coding on the syntax elements representing the information of the octree and / or the information of the surface determined by the surface approximation analysis unit 212. The GPCC encoder 200 may output these syntax elements in a geometry bitstream.

[0044] The geometry reconstruction unit 216 may reconstruct the transformation coordinates of the points in the point cloud based on the octree, the data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transformation coordinates reconstructed by the geometry reconstruction unit 216 may be different from the original number of points in the point cloud. The present disclosure may refer to the resulting points as the reconstructed points. The attribute transfer unit 208 may transfer the attributes of the original points of the point cloud to the reconstructed points of the point cloud data.

[0045] In addition, the RAHT unit 218 may apply RAHT coding to the attributes of the reconstructed points. Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 may apply LOD processing and lifting to the attributes of the reconstructed points, respectively. The RAHT unit 218 and the lifting unit 222 may generate coefficients based on the attributes. The coefficient quantization unit 224 may quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 may apply arithmetic coding to the syntax elements representing the quantized coefficients. The GPCC encoder 200 may output these syntax elements in the attribute bitstream.

[0046] In Figure 3 the example of, the GPCC decoder 300 may include a geometric arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometric reconstruction unit 312, a RAHT unit 314, an LOD generation unit 316, an inverse lifting unit 318, a coordinate inverse transformation unit 320, and a color inverse transformation unit 322.

[0047] The GPCC decoder 300 may obtain a geometric bitstream and an attribute bitstream. The geometric arithmetic decoding unit 302 of the decoder 300 may apply arithmetic decoding (e.g., CABAC or other types of arithmetic decoding) to the syntax elements in the geometric bitstream. Similarly, the attribute arithmetic decoding unit 304 may apply arithmetic decoding to the syntax elements in the attribute bitstream.

[0048] The octree synthesis unit 306 may synthesize an octree based on the syntax elements parsed from the geometric bitstream. In the case where surface approximation is used in the geometric bitstream, the surface approximation synthesis unit 310 may determine a surface model based on the syntax elements parsed from the geometric bitstream and based on the octree.

[0049] In addition, the geometric reconstruction unit 312 may perform reconstruction to determine the coordinates of the points in the point cloud. The coordinate inverse transformation unit 320 may apply an inverse transformation to the reconstructed coordinates to transform the reconstructed coordinates (positions) of the points in the point cloud from the transform domain back to the initial domain.

[0050] Additionally, in Figure 3 the example of, the inverse quantization unit 308 may perform inverse quantization on the attribute values. The attribute values may be based on the syntax elements obtained from the attribute bitstream (e.g., including the syntax elements decoded by the attribute arithmetic decoding unit 304).

[0051] Depending on how the attribute values are encoded, the RAHT unit 314 may perform RAHT decoding to determine color values for points in the point cloud based on the de-quantized attribute values. Alternatively, the LOD generation unit 316 and the inverse lifting unit 318 may use a level-of-detail based technique to determine color values for points in the point cloud.

[0052] In addition, in Figure 3 the example of, the color inverse transform unit 322 may apply an inverse color transform to the color values. The inverse color transform may be the inverse of the color transform applied by the color transform unit 204 of the encoder 200. For example, the color transform unit 204 may transform color information from the RGB color space to the YCbCr color space. Correspondingly, the color inverse transform unit 322 may transform color information from the YCbCr color space to the RGB color space.

[0053] Figure 2 and Figure 3 The various units of are shown to assist in understanding the operations performed by the encoder 200 and the decoder 300. These units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function and is preset with respect to the operations that can be performed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations that can be performed. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit may execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuit are generally immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.

[0054] Some exemplary embodiments of the present disclosure will be described in detail below. It should be understood that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to GPCC or other specific point cloud codecs, the disclosed techniques are also applicable to other point cloud coding and decoding techniques. In addition, although some embodiments describe the point cloud coding and decoding steps in detail, it should be understood that the corresponding decoding steps of the decoding will be implemented by the decoder.

[0055] 1. Brief Overview

[0056] This disclosure relates to point cloud encoding and decoding techniques. Specifically, it relates to improving the encoding and decoding complexity of point cloud attribute in the Region Adaptive Hierarchical Transform (RAHT). These concepts can be applied individually or in various combinations to any point cloud encoding and decoding standard or non-standard point cloud codec, such as the Geometry-based Point Cloud Compression (G-PCC) and the Low Latency Low Complexity Codec (L3C2) being developed.

[0057] 2. Abbreviations

[0058] G-PCC Geometry-based Point Cloud Compression

[0059] L3C2 Low Latency Low Complexity Codec

[0060] MPEG Moving Picture Experts Group

[0061] 3DG 3D Graphics Coding Group

[0062] V-PCC Video-based Point Cloud Compression

[0063] RAHT Region Adaptive Hierarchical Transform

[0064] DC Direct Current

[0065] AC Alternating Current

[0066] 3. Introduction

[0067] MPEG is short for the Moving Picture Experts Group, which is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Coding Group (3DG) released a Call for Proposals (CFP) document to start developing a point cloud encoding and decoding standard. The final standard will encompass two categories of solutions. Video-based Point Cloud Compression (V-PCC) is applicable to point sets with relatively uniform point distributions. Geometry-based Point Cloud Compression (G-PCC) is applicable to sparser distributions. Both V-PCC and G-PCC support encoding and decoding for individual point clouds and point cloud sequences.

[0068] In a point cloud, there can be geometric information and attribute information. Geometric information is used to describe the geometric positions of data points. Attribute information is used to record some details of data points, such as texture, normal vector, reflection, etc.

[0069] 3.1 Octree Geometry Compression

[0070] Point cloud codecs can process various information in different ways. Generally, there are many optional tools in the codec to separately support the encoding and decoding of geometric information and attribute information. Among the geometric encoding and decoding tools in G-PCC, octree geometry compression has an important impact on point cloud geometry encoding and decoding performance.

[0071] In G-PCC, one of the important point cloud geometry encoding and decoding tools is octree geometry compression, which utilizes the spatial correlation of point cloud geometry. If the geometry encoding and decoding tool is enabled, the axis-aligned bounding box associated with the octree root node will be determined based on the point cloud geometry information. Then, this bounding box will be subdivided into 8 sub-cubes, which are associated with the 8 child nodes of the root node (the cube is equivalent to the node in the following text). Then, an 8-bit code is generated in a specific order to individually indicate whether each of the 8 child nodes contains points, where one bit is associated with one child node. The bit associated with a child node is named the occupancy bit, and the generated 8-bit code is named the occupancy code. The generated occupancy code will be signaled based on the occupancy information of neighboring nodes. Then, only the nodes that contain points will be further subdivided into 8 child nodes. This process will be recursively executed until the node size is 1. Thus, the point cloud geometry information is converted into a sequence of occupancy codes.

[0072] On the decoder side, the sequence of occupancy codes will be decoded, and the point cloud geometry information can be reconstructed based on the sequence of occupancy codes.

[0073] The breadth-first scanning order will be used for the octree. In one level of the octree, the octree nodes will be scanned in Morton order. If the coordinates of a node are represented by N bits, the coordinates (X, Y, Z) of this node can be represented as follows.

[0074] X = (x N-1 x N-2 …x 1 x 0 )

[0075] Y = (y N-1 y N-2 …y 1 y 0 )

[0076] Z = (z N-1 z N-2 …z 1 z 0 )

[0077] Its Morton code can be represented as follows.

[0078] M = (x N-1 y N-1 z N-1 x N-2 y N-2 z N-2 …x1 y 1 z 1 x 0 y 0 z 0 )

[0079] The Morton order is in ascending or descending order according to the Morton code.

[0080] 3.2 Region Adaptive Hierarchical Transform

[0081] In G-PCC, one of the important point cloud attribute encoding and decoding tools is RAHT. It is a transform as follows that uses the attributes associated with the nodes in the lower levels of the octree to predict the attributes of the nodes in the next level. Assume the positions of the points are given at both the encoder and the decoder. RAHT follows the reverse scan of the octree, from the leaf nodes to the root node. At each step, the nodes are recombined into larger nodes until the root node is reached. At each level of the octree, the nodes are processed in Morton order. At each decomposition, RAHT does not combine eight nodes at once but does so in three steps along each dimension (e.g., along z, then y, then x). If there are L levels in the octree, RAHT takes 3L levels to traverse the tree in reverse.

[0082] Let the node at level l be g l,x,y,z , where x, y, z are integers. g l,x,y,z is obtained by combining g l+1,2x,y,z and g l+1,2x+1,y,z , where the combination along the first dimension is an example. RAHT only processes occupied nodes. If one of the nodes in the pair is not occupied, the other node is promoted to the next level (not processed), i.e., g l-1,x,y,z = g l,2x,y,z , if the latter is the occupied node in the pair. The combination process is repeated until the root is reached. Note that the combination process generates nodes at lower levels that are the result of combining different numbers of voxels along the path. The number of nodes combined to generate the node g l,x,y,z is the weight ω l,x,y,z .

[0083] At each combination of two nodes, say g l,2x,y,z and g l,2x+1,y,z , using their respective weights ω l,2x,y,z and ω l,2x+1,y,z , RAHT applies the following transform:

[0084]

[0085] where ω 1 = ω l,2x,y,z and ω2 = ω l,2x+1,y,z , and

[0086]

[0087] Note that the transformation matrix always changes in accordance with the weights, i.e., in accordance with each g l,x,y,z The actual number of leaf nodes represented. The quantity g l,x,y,z is used to combine and form subsequent nodes at lower levels. h l,x,y,z is the actual high-pass coefficient generated by the transformation and is to be encoded and transmitted. In addition, the weights are accumulated for the upper levels. In the above example,

[0088] ω l-1,2,y,z = ω l,2,x,y,z + ω l,2x+1,y,z .

[0089] In the final stage, the root of the tree, the remaining two voxels g 1,0,0,0 and g 1,1,0,0 are transformed into the final two coefficients as follows:

[0090]

[0091] where g DC = g 0,0,0,0 .

[0092] 3.2.1 Upsampling Transform Domain Prediction

[0093] Introduce transform domain prediction to improve the encoding and decoding efficiency on RAHT. It is formed by two parts.

[0094] First, change the traversal of the RAHT tree from ascending to descending based on the previous ascending method, i.e., for both the encoder and the decoder, construct a tree of the sum of attributes and weights from the root level to the leaf level of the tree, and then perform RAHT. At each level, the nodes are accessed in Morton order. The transformation is performed on the nodes with 2×2×2 child nodes (in the next level). The nodes where the transformation is performed can be called transformation nodes.

[0095] Second, for each child node of the transformation node, generate the corresponding predicted attribute by upsampling the attributes of the previous transformation level. In fact, only the child nodes containing at least one point will generate the corresponding predicted attributes. The transformation nodes containing the predicted attributes are transformed on the encoder side and subtracted from the transformed attributes. The residual of the alternating current (AC) coefficients will be transmitted through the signal. Note that the prediction does not affect the direct current (DC) coefficients.

[0096] Each child node of the transform node is predicted by seven parent-level nodes, including three collinear parent-level neighboring nodes, three coplanar parent-level neighboring nodes, and one parent node. The coplanar neighbors and collinear neighbors are the neighbors that share a face and an edge with the current transform node, respectively. A binary search algorithm is used to find the coplanar parent-level neighboring nodes and the collinear parent-level neighboring nodes. Figure 4 The parent-level nodes for each child node of the transform unit node are shown. Figure 4 Seven parent-level nodes for each child node of the transform node are shown.

[0097] Attribute a of each child node up is predicted as follows according to the distance between it and its parent-level nodes.

[0098] a up = Σω k a k / ∑ω k

[0099] where a k is the attribute of one of its parent-level nodes, and ω k is the distance-dependent weight. In G-PCC, ω parent : ω coplane : ω coline = 4:2:1.

[0100] 3.2.2 Early termination for transform-domain prediction

[0101] Early termination is introduced to reduce complexity. In upsampled transform-domain prediction, seven parent-level neighboring nodes are used to create prediction values for each encoded target node (child node) of the transform node. And there are a total of 19 parent-level neighboring nodes (including the parent node, i.e., the transform unit node) that are used to create prediction values for all eight encoded target nodes of the transform unit node. Since the number of valid neighboring parent nodes is large, the prediction accuracy will be better in a denser point cloud. On the contrary, the prediction accuracy will be poor in a sparse point cloud. Based on this feature, early termination for upsampled transform-domain prediction is introduced to reduce the encoding time. In early termination, the following two parameters are calculated for every eight child nodes of the transform unit node.

[0102] · NumValidP: The total number of valid parent-level neighboring nodes (including the parent node).

[0103] · NumValidGP: The total number of valid grandparent-level neighboring nodes (including the grandparent node).

[0104] Then, the prediction is disabled if NumValidP or NumValidGP is less than the threshold. This means that the prediction is terminated when the number of valid neighboring nodes becomes small.

[0105] 3.3 Problems

[0106] The existing designs for point cloud attribute transform domain prediction in region adaptive hierarchical transformation have the following problems:

[0107] In the current design, a nearest neighbor search process is performed for each transformation node to determine whether early termination is enabled, which results in high complexity. However, when only one child node contains points, there are no AC coefficients in the transformation node, and RAHT does not need to be performed. Therefore, in this case, the nearest neighbor search process can be skipped to reduce complexity.

[0108] 4. Detailed Solutions

[0109] To solve the above problems and some unmentioned problems, the methods outlined below are disclosed. The embodiments should be considered as examples to explain the general concepts and should not be construed in a narrow manner. In addition, these embodiments can be applied individually or in combination in any way.

[0110] 1) The prediction operation of a transformation node can be skipped.

[0111] a. In one example, a transformation node can have only one child node that contains at least one point.

[0112] b. In one example, for a transformation node, the number of child nodes that contain at least one point can be less than or equal to n, where n is an integer and n can be in a range such as from 1 to 8.

[0113] c. In one example, the predicted value can be derived from at least one nearest neighbor through weighted prediction and any other method.

[0114] i. In one example, a nearest neighbor can be a node that shares at least a face, or an edge, or a vertex with the transformation node.

[0115] ii. In one example, a nearest neighbor can be a node that is closest to the transformation node in spatial distance.

[0116] 1. In one example, the distance can be Euclidean distance, Manhattan distance,

[0117] Chebyshev distance, etc.

[0118] iii. In one example, the nearest neighbor and the transformation node can share the same octree depth.

[0119] iv. In one example, the nearest neighbor and the transformation node can have different octree depths.

[0120] d. In one example, the transformation can be a region adaptive hierarchical transformation, a wavelet transformation, a cosine transformation, etc.

[0121] e. In one example, the prediction information can come from the current frame or other frames.

[0122] f. In one example, the prediction information can come from at least one frame, which can be the (multiple) reference frames of the current frame.

[0123] i. In one example, the prediction information can come from one frame.

[0124] ii. In one example, the prediction information can come from two frames.

[0125] iii. In one example, the prediction information can come from n frames, where n is a positive integer.

[0126] 1. In one example, n can be transmitted to the decoder through a signal.

[0127] g. In one example, when a transformation node has only one child node containing points, the prediction operation of the transformation node can be skipped.

[0128] 2) When the prediction operation is skipped, the early termination of the prediction for a transformation node can be disabled.

[0129] a. In one example, if the prediction operation of the parent node is skipped, the early termination of the transformation node can be disabled.

[0130] b. In one example, if the prediction operation of the transformation node is skipped, the early termination of the transformation node can be disabled.

[0131] 3) Neighbor information can be used to determine whether the early termination for a transformation node is disabled.

[0132] a. In one example, a neighbor can be a node that shares at least one face, or edge or vertex with the parent node.

[0133] b. In one example, a neighbor can be a node that is close to the parent node in terms of spatial distance.

[0134] i. In one example, the distance can be the Euclidean distance, the Manhattan distance, the Chebyshev distance, etc.

[0135] c. In one example, the neighbor and the parent node can share the same octree depth.

[0136] i. Alternatively, the neighbor and the parent node can have different octree depths.

[0137] d. In one example, if the neighbor search of the parent node is skipped, the early termination of the transform node can be disabled.

[0138] e. In one example, if the neighbor search of the transform node is skipped, the early termination of the transform node can be disabled.

[0139] f. In one example, if the number of neighbors of the parent node cannot be significantly obtained, the early termination of the transform node can be disabled.

[0140] g. In one example, if the number of neighbors of the transform node cannot be significantly obtained, the early termination of the transform node can be disabled.

[0141] 4) In one example, prediction can be performed in different domains.

[0142] a. In one example, prediction can be performed in the sample domain.

[0143] b. In one example, prediction can be performed in the transform domain.

[0144] 5) Whether to apply the disclosed method and / or how to apply the disclosed method can depend on the encoding / decoding information.

[0145] a. In one example, if a transform node can only have one child node containing at least one point, the prediction operation of a transform node can be skipped.

[0146] b. In one example, if the prediction operation of the parent node is skipped, the early termination of the transform node can be disabled.

[0147] c. In one example, if the neighbor search of the parent node is skipped, the early termination of the transform node can be disabled.

[0148] i. In one example, a neighbor can be a node that shares at least a face, or an edge, or a vertex with the parent node.

[0149] d. In one example, prediction can be performed in the transform domain.

[0150] 6) Whether to apply the disclosed method and / or how to apply the disclosed method can be signaled from the encoder to the decoder.

[0151] a. In one example, whether prediction is performed in the transform domain or the sample domain can be signaled from the encoder to the decoder.

[0152] b. In one example, from which frame the prediction information comes can be signaled from the encoder to the decoder.

[0153] 5. Embodiment

[0154] Figure 5 Examples of encoding and decoding processes for improving the encoding and decoding complexity of point cloud attributes are depicted. At 510, the number of child nodes of a transform node that contain points (denoted as n subnode ) is obtained. That is, first, it is determined how many points are included in the transform node. At 520, it is determined whether the number n subnode is equal to 1. This is to find transform nodes that include only a single point. If n subnode is equal to 1, the process enters the "yes" branch, and at 530, the prediction operation for the transform node is skipped. Then, at 540, early termination for the child nodes of the transform node in the next octree level is disabled. On the other hand, if the number n subnode is not equal to 1, the process enters the "no" branch, and at 550, the prediction operation is performed for the transform node.

[0155] More details will be discussed further below. Figure 6 A flowchart of a method 600 for point cloud encoding and decoding according to an embodiment of the present disclosure is shown.

[0156] It should be understood that in the following embodiments, the term "transform block" refers to a space in which there may or may not be (a plurality of) points of a point cloud, and the (plurality of) points in the transform block will be transformed during the conversion between the point cloud sequence and its bitstream. A transform block may include one or more sub-blocks, and each sub-block may or may not include (a plurality of) points. A transform block may have one or more neighboring blocks adjacent to the transform block.

[0157] Regarding the term "transform node", it refers to a node (e.g., in a tree level) that represents or corresponds to a transform block. A transform node may have child nodes corresponding to sub-blocks of the transform block. A transform node may also have a parent node, e.g., in a tree level, that corresponds to a parent block that includes the transform block.

[0158] At block 610, for the conversion between a point cloud sequence including a current point cloud (PC) sample associated with a transform block and the bitstream of the point cloud sequence, a transform block including a predetermined number of sub-blocks is determined. Each sub-block contains at least one point of the point cloud sequence.

[0159] In some embodiments, the predetermined number is 1. Alternatively, the predetermined number is less than or equal to a positive integer. Alternatively, the predetermined number is less than or equal to a positive integer in the range from 1 to 8.

[0160] At block 620, the conversion is performed by skipping at least one of prediction or transformation of the transform block.

[0161] Method 600 enables reducing the complexity of prediction, transformation, and / or their related operations.

[0162] In some embodiments, the predicted value can be derived from at least one neighboring block of the transform block.

[0163] In some embodiments, a neighboring block can be a block that shares at least one of a face, an edge, or a vertex with the transform block, or a neighboring block can be the block that is closest in terms of distance to the transform block, or the neighboring block and the transform block share the same tree level or have different tree levels.

[0164] In some embodiments, the distance can be one of Euclidean distance, Manhattan distance, or Chebyshev distance, or the tree level includes the octree depth.

[0165] In some embodiments, the transformation can be one of a region adaptive hierarchical transformation, a wavelet transformation, or a cosine transformation.

[0166] In some embodiments, the prediction information for prediction can come from the current frame of the transform block, or the prediction information for prediction can come from at least one frame.

[0167] In some embodiments, the number of frames in the at least one frame can be a positive integer and / or can be indicated in the bitstream, or the at least one frame includes a reference frame of the current frame.

[0168] In some embodiments, if the prediction of the transform block is skipped, the early termination of the prediction for the transform block can be disabled. Alternatively or additionally, the early termination of the prediction for a sub-block of the transform block can be disabled.

[0169] In some embodiments, whether the early termination of the prediction for the transform block can be disabled is based on information about one or more neighboring blocks of the transform block.

[0170] In some embodiments, one of the neighboring blocks can be a block that shares at least one of a face, an edge, or a vertex with the parent block, and the parent block corresponds to the parent node of the transform block. The neighboring block can be the block that is closest in terms of distance to the parent block, or the neighboring block and the transform block can share the same tree level or have different tree levels.

[0171] In some embodiments, the distance can be one of Euclidean distance, Manhattan distance, or Chebyshev distance, or the tree level includes the octree depth.

[0172] In some embodiments, early termination of a transform block is disabled if the neighboring search of the parent node corresponding to the transform block's parent node is skipped, or if the neighboring search of the transform block is skipped, or if the number of neighboring blocks of the parent block cannot be obtained, or if the number of neighboring blocks of the transform block cannot be obtained.

[0173] In some embodiments, prediction is performed in different domains.

[0174] In some embodiments, prediction is performed in the sample domain or the transform domain.

[0175] In some embodiments, the information on whether to apply the method and / or how to apply the method depends on the encoding / decoding information.

[0176] In some embodiments, prediction of a transform block is skipped if the transform block has only one sub-block containing at least one point, or early termination of the transform block may be disabled if prediction of the parent block corresponding to the transform block's parent node is skipped. Alternatively, early termination of the transform block may be disabled if the neighboring search of the parent block is skipped, or prediction is performed in the transform domain.

[0177] In some embodiments, the neighboring blocks of the parent block share at least a face, or an edge, or a vertex with the parent block.

[0178] In some embodiments, first information on whether the method is to be applied and / or how to apply the method is indicated in the bitstream.

[0179] In some embodiments, second information on whether prediction is performed in the transform domain or the sample domain is indicated in the bitstream, or third information on which frame the prediction information is from is indicated in the bitstream.

[0180] In some embodiments, at least one of the first information, the second information, or the third information is indicated from the encoder to the decoder.

[0181] In some embodiments, the current PC sample is one of the following: a frame, a picture, a slice, a sub-frame, a sub-picture, a tile, or a segment.

[0182] In some embodiments, the transformation includes encoding the current PC sample into the bitstream.

[0183] In some embodiments, the transformation includes decoding the current PC sample from the bitstream.

[0184] According to a further embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a point cloud sequence, the bitstream being generated by a method executed by a device for point cloud encoding and decoding. The method includes: determining a transform block including a predetermined number of sub-blocks, each sub-block including at least one point of the point cloud sequence, the transform block being associated with a current point cloud (PC) sample included in the point cloud sequence; and generating the bitstream by skipping at least one of prediction or transformation of the transform block.

[0185] According to still further embodiments of the present disclosure, a method for storing a bitstream of a point cloud sequence is provided. The method includes: determining a transform block including a predetermined number of sub-blocks, each sub-block including at least one point of the point cloud sequence, the transform block being associated with a current point cloud (PC) sample included in the point cloud sequence; generating the bitstream by skipping at least one of prediction or transformation of the transform block; and storing the bitstream in a non-transitory computer-readable recording medium.

[0186] Implementations of the present disclosure may be described according to the following items, and the features may be combined in any reasonable manner.

[0187] Item 1. A method for point cloud encoding and decoding, including: for the conversion between a point cloud sequence including a current point cloud (PC) sample associated with a transform block and a bitstream of the point cloud sequence, determining the transform block including a predetermined number of sub-blocks, each sub-block including at least one point of the point cloud sequence; and performing the conversion by skipping at least one of prediction or transformation of the transform block.

[0188] Item 2. The method according to Item 1, wherein the predetermined number is 1, or wherein the predetermined number is less than or equal to a positive integer, or wherein the predetermined number is less than or equal to a positive integer in the range from 1 to 8.

[0189] Item 3. The method according to Item 1, wherein the predicted value is derived from at least one neighboring block of the transform block.

[0190] Item 4. The method according to Item 3, wherein the neighboring block is a block sharing at least one of a face, an edge, or a vertex with the transform block, or wherein the neighboring block is the block closest to the transform block in terms of distance, or wherein the neighboring block and the transform block share the same tree level or have different tree levels.

[0191] Item 5. The method according to Item 4, wherein the distance is one of an Euclidean distance, a Manhattan distance, or a Chebyshev distance, or wherein the tree level includes an octree depth.

[0192] Item 6. The method according to Item 1, wherein the transformation is one of a region adaptive hierarchical transformation, a wavelet transformation, and a cosine transformation.

[0193] Item 7. The method according to Item 1, wherein the prediction information for the prediction is from the current frame of the transform block, or wherein the prediction information for the prediction is from at least one frame.

[0194] Item 8. The method according to Item 7, wherein the number of frames in the at least one frame is a positive integer and / or is indicated in the bitstream, or wherein the at least one frame includes a reference frame of the current frame.

[0195] Item 9. The method according to Item 1, wherein if the prediction of the transform block is skipped, at least one of the following is disabled: early termination of the prediction for the transform block, or early termination of the prediction for a sub-block of the transform block.

[0196] Item 10. The method according to Item 9, wherein whether the early termination of the prediction for the transform block is disabled is based on information about one or more neighboring blocks of the transform block.

[0197] Item 11. The method according to Item 10, wherein one of the neighboring blocks is a block that shares at least one of a face, an edge, or a vertex with a parent block, the parent block corresponding to the parent node of the transform block, wherein the neighboring block is the block closest to the parent block in terms of distance, or wherein the neighboring block and the transform block share the same tree level or have different tree levels.

[0198] Item 12. The method according to Item 11, wherein the distance is one of an Euclidean distance, a Manhattan distance, and a Chebyshev distance, or wherein the tree level includes an octree depth.

[0199] Item 13. The method according to Item 9, wherein if the neighbor search of the parent node corresponding to the parent node of the transform block is skipped, the early termination of the transform block is disabled, or if the neighbor search of the transform block is skipped, the early termination of the transform block is disabled, or if the number of neighboring blocks of the parent block cannot be obtained, the early termination of the transform block is disabled, or if the number of neighboring blocks of the transform block cannot be obtained, the early termination of the transform block is disabled.

[0200] Item 14. The method according to Item 1, wherein the prediction is performed in different domains.

[0201] Item 15. The method according to Item 14, wherein the prediction is performed in the sample domain or the transform domain.

[0202] Item 16. The method according to any one of Items 1 to 15, wherein information regarding whether to apply the method and / or how to apply the method depends on coding / decoding information.

[0203] Item 17. The method according to Item 16, wherein if the transform block has only one sub-block containing at least one point, the prediction of the transform block is skipped, or wherein if the prediction of the parent block corresponding to the parent node of the transform block is skipped, the early termination of the transform block is disabled, or wherein if the neighboring search of the parent block is skipped, the early termination of the transform block is disabled, or wherein the prediction is performed in the transform domain.

[0204] Item 18. The method according to Item 17, wherein the neighboring block of the parent block shares at least a face, or an edge, or a vertex with the parent block.

[0205] Item 19. The method according to Item 1, wherein first information regarding whether the method is to be applied and / or how to apply the method is indicated in the bitstream.

[0206] Item 20. The method according to Item 19, wherein second information regarding whether the prediction is performed in the transform domain or in the sample domain is indicated in the bitstream, or wherein third information regarding from which frame the prediction information is derived is indicated in the bitstream.

[0207] Item 21. The method according to Item 20, wherein at least one of the first information, the second information, or the third information is indicated from the encoder to the decoder.

[0208] Item 22. The method according to any one of Items 1 to 21, wherein the current PC sample is one of the following: a frame, a picture, a slice, a sub-frame, a sub-picture, a tile, or a segment.

[0209] Item 23. The method according to any one of Items 1 to 22, wherein the transformation includes encoding the current PC sample into the bitstream.

[0210] Item 24. The method according to any one of Items 1 to 22, wherein the transformation includes decoding the current PC sample from the bitstream.

[0211] Item 25. An apparatus for point cloud coding and decoding, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 24.

[0212] Item 26. A non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to any one of Items 1 to 24.

[0213] Item 27. A non-transitory computer-readable recording medium stores a bitstream of a point cloud sequence, the bitstream of the point cloud sequence being generated by a method executed by a point cloud processing device, the method including: determining a transform block including a predetermined number of sub-blocks, each sub-block including at least one point of the point cloud sequence, the transform block being associated with a current point cloud (PC) sample included in the point cloud sequence; and generating the bitstream by skipping at least one of prediction or transformation of the transform block.

[0214] Item 28. A method for storing a bitstream of a point cloud sequence includes: determining a transform block including a predetermined number of sub-blocks, each sub-block including at least one point of the point cloud sequence, the transform block being associated with a current point cloud (PC) sample included in the point cloud sequence; generating the bitstream by skipping at least one of prediction or transformation of the transform block; and storing the bitstream in a non-transitory computer-readable recording medium.

[0215] Example device

[0216] Figure 7 A block diagram of a computing device 700 in which various embodiments of the present disclosure can be implemented is shown. The computing device 700 can be implemented as the source device 110 (or the GPCC encoder 116 or 200) or the destination device 120 (or the GPCC decoder 126 or 300), or can be included in the source device 110 (or the GPCC encoder 116 or 200) or the destination device 120 (or the GPCC decoder 126 or 300).

[0217] It should be understood that Figure 7 the computing device 700 shown is for illustrative purposes only and does not imply any limitation to the functions and scopes of the embodiments of the present disclosure in any way.

[0218] As Figure 7 shown, the computing device 700 includes a general computing device 700. The computing device 700 can include at least one or more processors or processing units 710, a memory 720, a storage unit 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760.

[0219] In some embodiments, computing device 700 may be implemented as any user terminal or server terminal with computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that computing device 700 may support any type of interface to the user (such as "wearable" circuitry, etc.).

[0220] Processing unit 710 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of computing device 700. Processing unit 710 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0221] Computing device 700 generally includes various computer storage media. Such media may be any media accessible by computing device 700, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. Memory 720 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. Storage unit 730 may be any removable or non-removable media, and may include machine-readable media, such as memory, flash drive, disk, or other media that can be used to store information and / or data and can be accessed in computing device 700.

[0222] Computing device 700 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 7 a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0223] The communication unit 740 communicates with another computing device via a communication medium. Additionally, the functionality of the components in the computing device 700 can be implemented by a single computing cluster or multiple computing machines, which can communicate via a communication connection. Thus, the computing device 700 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.

[0224] The input device 750 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and so on. The output device 760 can be one or more of various output devices, such as a display, speaker, printer, and so on. With the aid of the communication unit 740, the computing device 700 can also communicate with one or more external devices (not shown), such as storage devices and display devices, the computing device 700 can also communicate with one or more devices that enable a user to interact with the computing device 700, or if needed, the computing device 700 can also communicate with any device (such as a network card, modem, etc.) that enables the computing device 700 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0225] In some embodiments, some or all of the components of the computing device 700 can also be arranged in a cloud computing architecture instead of being integrated in a single device. In a cloud computing architecture, the components can be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require the end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses suitable protocols to provide services via a wide area network (such as the Internet). For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure can provide services through a shared data center, although to the user, they appear as a single access point. Thus, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein can be provided by a conventional server or directly or otherwise installed on a client device.

[0226] In an embodiment of the present disclosure, a computing device 700 may be used to implement point cloud encoding / decoding. The memory 720 may include one or more point cloud encoding / decoding modules 725 having one or more program instructions. These modules are accessible and executable by the processing unit 710 to perform the functions of the various embodiments described herein.

[0227] In an example embodiment of performing point cloud encoding, an input device 750 may receive point cloud data as an input 770 to be encoded. The point cloud data may be processed, for example, by the point cloud encoding / decoding module 725 to generate an encoded bitstream. The encoded bitstream may be provided as an output 780 via an output device 760.

[0228] In an example embodiment of performing point cloud decoding, an input device 750 may receive the encoded bitstream as an input 770. The encoded bitstream may be processed, for example, by the point cloud encoding / decoding module 725 to generate decoded point cloud data. The decoded point cloud data may be provided as an output 780 via an output device 760.

[0229] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for point cloud encoding and decoding, comprising: determining a transform block including a predetermined number of sub - blocks for the conversion between a point cloud sequence including current point cloud (PC) samples associated with the transform block and a bitstream of the point cloud sequence, each sub - block containing at least one point of the point cloud sequence; and performing the conversion by skipping at least one of prediction or transformation of the transform block.

2. The method according to claim 1, wherein the predetermined number is 1, or wherein the predetermined number is less than or equal to a positive integer, or wherein the predetermined number is less than or equal to a positive integer in the range from 1 to 8.

3. The method according to claim 1, wherein the predicted value is derived from at least one neighboring block of the transform block.

4. The method according to claim 3, wherein the neighboring block is a block that shares at least one of a face, an edge, or a vertex with the transform block, or wherein the neighboring block is the block closest to the transform block in terms of distance, or wherein the neighboring block and the transform block share the same tree level or have different tree levels.

5. The method according to claim 4, wherein the distance is one of Euclidean distance, Manhattan distance, or Chebyshev distance, or wherein the tree level includes octree depth.

6. The method according to claim 1, wherein the transformation is one of a region - adaptive hierarchical transformation, a wavelet transformation, or a cosine transformation.

7. The method according to claim 1, wherein the prediction information for the prediction is from the current frame of the transform block, or wherein the prediction information for the prediction is from at least one frame.

8. The method according to claim 7, wherein the number of frames in the at least one frame is a positive integer and / or is indicated in the bitstream, or wherein the at least one frame includes a reference frame of the current frame.

9. The method according to claim 1, wherein if the prediction of the transform block is skipped, at least one of the following is disabled: early termination of the prediction for the transform block, or early termination of the prediction for sub - blocks of the transform block.

10. The method according to claim 9, wherein whether the early termination of the prediction for the transform block is disabled is based on information about one or more neighboring blocks of the transform block.

11. The method according to claim 10, wherein one of the neighboring blocks is a block that shares at least one of a face, an edge, or a vertex with a parent block, the parent block corresponding to the parent node of the transform block, wherein the neighboring block is the block closest to the parent block in terms of distance, or wherein the neighboring block and the transform block share the same tree level or have different tree levels.

12. The method according to claim 11, wherein the distance is one of Euclidean distance, Manhattan distance, or Chebyshev distance, or wherein the tree level includes octree depth.

13. The method according to claim 9, wherein if the neighbor search of the parent node corresponding to the parent node of the transform block is skipped, the early termination of the transform block is disabled, or If the neighborhood search of the transform block is skipped, the early termination of the transform block is disabled, or If the number of neighboring blocks of the parent block cannot be obtained, the early termination of the transform block is disabled, or If the number of neighboring blocks of the transform block cannot be obtained, the early termination of the transform block is disabled.

14. The method according to claim 1, wherein the prediction is performed in different domains.

15. The method according to claim 14, wherein the prediction is performed in the sample domain or the transform domain.

16. The method according to any one of claims 1 to 15, wherein the information on whether to apply the method and / or how to apply the method depends on the encoding / decoding information.

17. The method according to claim 16, wherein if the transform block has only one sub-block containing at least one point, the prediction of the transform block is skipped, or wherein if the prediction of the parent block corresponding to the parent node of the transform block is skipped, the early termination of the transform block is disabled, or wherein if the neighborhood search of the parent block is skipped, the early termination of the transform block is disabled, or wherein the prediction is performed in the transform domain.

18. The method according to claim 17, wherein the neighboring blocks of the parent block share at least a face, or an edge, or a vertex with the parent block.

19. The method according to claim 1, wherein first information on whether the method is to be applied and / or how to apply the method is indicated in the bitstream.

20. The method according to claim 19, wherein second information on whether the prediction is performed in the transform domain or the sample domain is indicated in the bitstream, or wherein third information on which frame the prediction information is from is indicated in the bitstream.

21. The method according to claim 20, wherein at least one of the first information, the second information, or the third information is indicated from the encoder to the decoder.

22. The method according to any one of claims 1 to 21, wherein the current PC sample is one of the following: frame, picture, slice, sub-frame, sub-picture, tile, or segment.

23. The method according to any one of claims 1 to 22, wherein the transformation includes encoding the current PC sample into the bitstream.

24. The method according to any one of claims 1 to 22, wherein the transformation includes decoding the current PC sample from the bitstream.

25. An apparatus for point cloud encoding and decoding, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 24.

26. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 24.

27. A non-transitory computer-readable recording medium storing a bitstream of a point cloud sequence, the bitstream of the point cloud sequence being generated by a method performed by a point cloud processing apparatus, the method comprising: Determine a transform block including a predetermined number of sub-blocks, each sub-block containing at least one point of the point cloud sequence, the transform block being associated with a current point cloud (PC) sample included in the point cloud sequence; and Generate the bitstream by skipping at least one of prediction or transformation of the transform block.

28. A method for storing a bitstream of a point cloud sequence, comprising:[[]] Determine a transform block including a predetermined number of sub-blocks, each sub-block containing at least one point of the point cloud sequence, the transform block being associated with a current point cloud (PC) sample included in the point cloud sequence; Generate the bitstream by skipping at least one of prediction or transformation of the transform block; and Store the bitstream in a non-transitory computer-readable recording medium.