Method, device and medium for point cloud coding and decoding

By performing attribute prediction in the sample domain of point cloud codec and transforming only the prediction residuals, the complexity problem caused by the execution of attribute prediction in the transform domain in the prior art is solved, and a more efficient encoding and decoding process is achieved.

CN120188488APending Publication Date: 2025-06-20DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071195.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-04
Filing Date
2023-09-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, attribute prediction is performed in the transformation domain during point cloud encoding and decoding, resulting in high complexity and it is difficult to reduce the complexity caused by transformation.

Method used

Perform attribute prediction in the sample field, transforming only the prediction residuals, thereby reducing the complexity of the transformation.

Benefits of technology

By performing attribute prediction in the sample field, only the prediction residuals are transformed, which significantly reduces the complexity in the encoding and decoding process and improves efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120188488A_ABST
    Figure CN120188488A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for point cloud coding and decoding. The method comprises: for a conversion between a point cloud sequence comprising a current point cloud (PC) sample associated with a transform block and a bitstream of the point cloud sequence, determining a transform result of an attribute residual between a neighbor attribute of at least one sub-block of the transform block and a prediction attribute of the at least one sub-block of the transform block, the neighbor attribute is predicted based on an attribute of at least one neighbor block of the transform block; and performing conversion at least based on a conversion result of the attribute residual.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to point cloud encoding and decoding techniques, and more particularly, to sample domain prediction for region adaptive hierarchical transform (RAHT). Background Art

[0002] A point cloud is a collection of individual data points in a three-dimensional (3D) plane, where each point has set coordinates on the X-axis, Y-axis, and Z-axis. Thus, a point cloud can be used to represent the physical content of a three-dimensional space. For various immersive applications ranging from augmented reality to autonomous vehicles, point clouds have proven to be a promising way to represent 3D visual data.

[0003] Point cloud encoding and decoding standards have mainly evolved through the well-known MPEG organization. MPEG is short for the Moving Picture Experts Group, which is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Encoding and Decoding Group (3DG) released a Call for Proposals (CFP) document to initiate the development of point cloud encoding and decoding standards. The final standard will encompass two categories of solutions. Video-based Point Cloud Compression (V-PCC or VPCC) is applicable to point sets with relatively uniform point distributions. Geometry-based Point Cloud Compression (G-PCC or GPCC) is applicable to more sparse distributions. However, there is generally a desire to further improve the encoding and decoding efficiency of conventional point cloud encoding and decoding techniques. Summary of the Invention

[0004] Embodiments of the present disclosure provide a solution for point cloud encoding and decoding.

[0005] In a first aspect, a method for point cloud encoding and decoding is proposed. The method includes: determining a transformed result of an attribute residual between a neighboring attribute of at least one sub-block of a transform block and a prediction attribute of at least one sub-block of the transform block for a conversion between a point cloud sequence including current point cloud (PC) samples associated with the transform block and a bitstream of the point cloud sequence, the neighboring attribute being predicted based on attributes of at least one neighboring block of the transform block; and performing the conversion based at least on the transformed result of the attribute residual.

[0006] In a second aspect, an apparatus for point cloud encoding and decoding is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a point cloud processing device for a point cloud sequence. The method includes: determining a transformed result of an attribute residual between a neighboring attribute of at least one sub-block of a transform block and a predicted attribute of at least one sub-block of the transform block, where the neighboring attribute is predicted based on attributes of at least one neighboring block of the transform block; and generating a bitstream based at least on the transformed result of the attribute residual.

[0009] In a fifth aspect, a method for storing a bitstream of a point cloud sequence is proposed. The method includes: determining a transformed result of an attribute residual between a neighboring attribute of at least one sub-block of a transform block and a predicted attribute of at least one sub-block of the transform block, where the neighboring attribute is predicted based on attributes of at least one neighboring block of the transform block; generating a bitstream based at least on the transformed result of the attribute residual; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] The present invention content is provided to introduce in a simplified form a selection of concepts further described below in the detailed description. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1 is a block diagram showing an exemplary point cloud encoding and decoding system that can utilize the technology of the present disclosure;

[0013] Figure 2 shows a block diagram of an exemplary point cloud encoder according to some embodiments of the present disclosure;

[0014] Figure 3 shows a block diagram of an exemplary point cloud decoder according to some embodiments of the present disclosure;

[0015] Figure 4 shows an exemplary diagram of a parent-level node for each sub-node of a transform unit node according to some embodiments of the present disclosure;

[0016] Figure 5 shows an exemplary diagram of an encoding and decoding process for sample-domain prediction of region-adaptive hierarchical transform according to some embodiments of the present disclosure;

[0017] Figure 6A flowchart of a method for point cloud encoding and decoding according to an embodiment of the present disclosure is shown; and

[0018] Figure 7 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0019] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed implementation manners

[0020] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is only for illustration and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.

[0021] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.

[0022] As used herein, the terms "one embodiment", "embodiment", "example embodiment", etc. indicate that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment must include that specific feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a specific feature, structure, or characteristic is described in connection with an example embodiment, it is contended that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.

[0023] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0024] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes", and / or "including" when used herein indicate the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Example environment

[0025] Figure 1 is a block diagram showing an example point cloud encoding / decoding system 100 that can utilize the techniques of the present disclosure. As shown, the point cloud encoding / decoding system 100 can include a source device 110 and a destination device 120. The source device 110 can also be referred to as a point cloud encoding device, and the destination device 120 can also be referred to as a point cloud decoding device. In operation, the source device 110 can be configured to generate encoded point cloud data, and the destination device 120 can be configured to decode the encoded point cloud data generated by the source device 110. The techniques of the present disclosure generally aim at encoding / decoding (encoding and / or decoding) point cloud data, that is, supporting point cloud compression. Encoding / decoding can be effective in compressing and / or decompressing point cloud data.

[0026] The source device 100 and the destination device 120 can include any of a variety of devices, including desktop computers, notebooks (i.e., laptops) computers, tablet computers, set-top boxes, telephone handsets (such as smart phones and mobile phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, vehicles (e.g., land or sea vehicles, spacecraft, aircraft, etc.), robots, LIDAR devices, satellites, extended reality devices, and so on. In some cases, the source device 100 and the destination device 120 can be equipped for wireless communication.

[0027] The source device 100 can include a data source 112, a memory 114, a GPCC encoder 116, and an input / output (I / O) interface 118. The destination device 120 can include an input / output (I / O) interface 128, a GPCC decoder 126, a memory 124, and a data consumer 122. According to the present disclosure, the GPCC encoder 116 of the source device 100 and the GPCC decoder 126 of the destination device 120 can be configured to apply the techniques of the present disclosure related to point cloud encoding / decoding. Thus, the source device 100 represents an example of an encoding device, and the destination device 120 represents an example of a decoding device. In other examples, the source device 100 and the destination device 120 can include other components or arrangements. For example, the source device 100 can receive data (e.g., point cloud data) from an internal source or an external source. Similarly, the destination device 120 can interface with an external data consumer instead of including a data consumer in the same device.

[0028] In general, data source 112 represents the source of point cloud data (i.e., raw, unencoded point cloud data) and can provide a continuous series of “frames” of point cloud data to GPCC encoder 116, which encodes the point cloud data for the frames. In some examples, data source 112 generates the point cloud data. The data source 112 of source device 100 can include a point cloud acquisition device, such as any of a variety of cameras or sensors, for example, one or more cameras, an archive containing previously acquired point cloud data, a 3D scanner or a light detection and ranging (LIDAR) device, and / or a data feed interface that receives point cloud data from a data content provider. Thus, in some examples, data source 112 can generate point cloud data based on signals from a LIDAR device. Alternatively or additionally, the point cloud data can be generated from a scanner, camera, sensor, or other data by a computer. For example, data source 112 can generate point cloud data, or a combination of real-time point cloud data, archived point cloud data, and computer-generated point cloud data. In each case, GPCC encoder 116 encodes the acquired, pre-acquired, or computer-generated point cloud data. GPCC encoder 116 can rearrange the frames of point cloud data from the received order (sometimes referred to as “display order”) into a codec order for encoding and decoding. GPCC encoder 116 can generate one or more bitstreams including the encoded point cloud data. Source device 100 can then output the encoded point cloud data via I / O interface 118 for reception and / or retrieval by, for example, the I / O interface 128 of destination device 120. The encoded point cloud data can be directly transmitted to destination device 120 via network 130A through I / O interface 118. The encoded point cloud data can also be stored on storage medium / server 130B for access by destination device 120.

[0029] The memory 114 of the source device 100 and the memory 124 of the destination device 120 may represent general-purpose memories. In some examples, the memories 114 and 124 may store raw point cloud data, e.g., raw point cloud data from the data source 112 and raw, decoded point cloud data from the GPCC decoder 126. Additionally or alternatively, the memories 114 and 124 may store software instructions executable by, e.g., the GPCC encoder 116 and the GPCC decoder 126, respectively. Although the memories 114 and 124 are shown separately from the GPCC encoder 116 and the GPCC decoder 126 in this example, it should be understood that the GPCC encoder 116 and the GPCC decoder 126 may also include internal memories for functionally similar or equivalent purposes. Further, the memories 114 and 124 may store encoded point cloud data, e.g., encoded point cloud data output from the GPCC encoder 116 and input to the GPCC decoder 126. In some examples, portions of the memories 114 and 124 may be allocated as one or more caches, e.g., for storing raw point cloud data, decoded point cloud data, and / or encoded point cloud data. For example, the memories 114 and 124 may store point cloud data.

[0030] The I / O interfaces 118 and 128 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where the I / O interfaces 118 and 128 include wireless components, the I / O interfaces 118 and 128 may be configured to transmit data, such as encoded point cloud data, according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc. In some examples where the I / O interface 118 includes a wireless transmitter, the I / O interfaces 118 and 128 may be configured to transmit data, such as encoded point cloud data, according to other wireless standards such as the IEEE 802.11 specifications. In some examples, the source device 100 and / or the destination device 120 may include respective system-on-chip (SoC) devices. For example, the source device 100 may include an SoC device for performing functions attributed to the GPCC encoder 116 and / or the I / O interface 118, and the destination device 120 may include an SoC device for performing functions attributed to the GPCC decoder 126 and / or the I / O interface 128.

[0031] The techniques of the present disclosure can be applied to encoding and decoding to support any one of a variety of applications, such as communication between autonomous vehicles, communication between scanners, cameras, sensors, and processing devices (e.g., local or remote servers), geographical mapping, or other applications.

[0032] The I / O interface 128 of the destination device 120 receives the encoded bitstream from the source device 110. The encoded bitstream can include signaling information defined by the GPCC encoder 116, which is also used by the GPCC decoder 126, such as syntax elements having values representing point clouds. The data consumer 122 uses the decoded data. For example, the data consumer 122 can use the decoded point cloud data to determine the location of a physical object. In some examples, the data consumer 122 can include a display for presenting an image based on the point cloud data.

[0033] The GPCC encoder 116 and the GPCC decoder 126 can each be implemented as any one of a variety of suitable encoder circuit systems and / or decoder circuit systems, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the GPCC encoder 116 and the GPCC decoder 126 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device. Devices including the GPCC encoder 116 and / or the GPCC decoder 126 can include one or more integrated circuits, microprocessors, and / or other types of devices.

[0034] The GPCC encoder 116 and the GPCC decoder 126 can operate according to an encoding / decoding standard, such as the Video Point Cloud Compression (VPCC) standard or the Geometric Point Cloud Compression (GPCC) standard. Generally, the present disclosure can refer to the encoding / decoding (e.g., encoding and decoding) of frames to include the process of encoding data or decoding data. The encoded bitstream typically includes a series of values for syntax elements representing encoding / decoding decisions (e.g., encoding / decoding modes).

[0035] A point cloud can include a set of points in 3D space and can have attributes associated with the points. The attributes can be color information, such as R, G, B or Y, Cb, Cr or reflectivity information, or other attributes. The point cloud can be acquired by various cameras or sensors (such as LIDAR sensors and 3D scanners) and can also be computer-generated. Point cloud data is used in various applications, including but not limited to construction (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors for assisting navigation).

[0036] Figure 2 is a block diagram showing an example of the GPCC encoder 200 according to some embodiments of the present disclosure. The GPCC encoder 200 can be Figure 1 an example of the GPCC encoder 116 in the system 100 shown. Figure 3 is a block diagram showing an example of the GPCC decoder 300 according to some embodiments of the present disclosure. The GPCC decoder 300 can be Figure 1 an example of the GPCC decoder 126 in the system 100 shown.

[0037] In both the GPCC encoder 200 and the GPCC decoder 300, the point cloud positions are first encoded and decoded. Attribute encoding and decoding depend on the decoded geometry. In Figure 2 and Figure 3 , the Region Adaptive Hierarchical Transform (RAHT) unit 218, the Surface Approximation Analysis unit 212, the RAHT unit 314, and the Surface Approximation Synthesis unit 310 are options typically used for Category 1 data. The Level of Detail (LOD) Generation unit 220, the Lift unit 222, the LOD Generation unit 316, and the Inverse Lift unit 318 are options typically used for Category 3 data. All other units are common between Category 1 and Category 3.

[0038] For Category 3 data, the compressed geometry is typically represented as an octree from the root all the way down to the leaf level of individual voxels. For Category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to the leaf level of blocks larger than voxels) plus a model for approximating the surface within each leaf node of the pruned octree. In this way, both Category 1 data and Category 3 data share the octree encoding and decoding mechanism, while Category 1 data can additionally utilize the surface model to approximate the voxels within each leaf node. The surface model used is a triangulation including 1 to 10 triangles per block, resulting in a triangle soup. Therefore, the Category 1 geometry codec is referred to as the Triangle Soup (Trisoup) geometry codec, while the Category 3 geometry codec is referred to as the octree geometry codec.

[0039] InFigure 2 In the example of, the GPCC encoder 200 may include a coordinate transformation unit 202, a color transformation unit 204, a voxelization unit 206, an attribute transfer unit 208, an octree analysis unit 210, a surface approximation analysis unit 212, an arithmetic coding unit 214, a geometric reconstruction unit 216, a RAHT unit 218, a LOD generation unit 220, a lifting unit 222, a coefficient quantization unit 224, and an arithmetic coding unit 226.

[0040] As Figure 2 shown in the example of, the GPCC encoder 200 may receive a set of positions and a set of attributes. The positions may include the coordinates of the points in the point cloud. The attributes may include information about the points in the point cloud, such as the color associated with the points in the point cloud.

[0041] The coordinate transformation unit 202 may apply a transformation to the coordinates of the points to transform the coordinates from an initial domain to a transformed domain. The present disclosure may refer to the transformed coordinates as transformation coordinates. The color transformation unit 204 may apply a transformation to convert the color information of the attributes to a different domain. For example, the color transformation unit 204 may convert the color information from the RGB color space to the YCbCr color space.

[0042] In addition, in Figure 2 the example of, the voxelization unit 206 may voxelize the transformation coordinates. The voxelization of the transformation coordinates may include quantizing and removing some of the points of the point cloud. In other words, multiple points of the point cloud may be grouped into a single "voxel", which may thereafter be considered as a point in some aspects. In addition, the octree analysis unit 210 may generate an octree based on the voxelized transformation coordinates. Additionally, in Figure 2 the example of, the surface approximation analysis unit 212 may analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 214 may perform arithmetic coding on the syntax elements representing the information of the octree and / or the information of the surface determined by the surface approximation analysis unit 212. The GPCC encoder 200 may output these syntax elements in a geometric bitstream.

[0043] The geometric reconstruction unit 216 may reconstruct the transformation coordinates of the points in the point cloud based on the octree, the data indicating the surface determined by the surface approximation analysis unit 212, and / or other information. Due to voxelization and surface approximation, the number of transformation coordinates reconstructed by the geometric reconstruction unit 216 may be different from the original number of points in the point cloud. The present disclosure may refer to the resulting points as reconstructed points. The attribute transfer unit 208 may transfer the attributes of the original points of the point cloud to the reconstructed points of the point cloud data.

[0044] In addition, the RAHT unit 218 may apply RAHT coding to the attributes of the reconstructed points. Alternatively or additionally, the LOD generation unit 220 and the lifting unit 222 may apply LOD processing and lifting to the attributes of the reconstructed points, respectively. The RAHT unit 218 and the lifting unit 222 may generate coefficients based on the attributes. The coefficient quantization unit 224 may quantize the coefficients generated by the RAHT unit 218 or the lifting unit 222. The arithmetic coding unit 226 may apply arithmetic coding to the syntax elements representing the quantized coefficients. The GPCC encoder 200 may output these syntax elements in the attribute bitstream.

[0045] In Figure 3 the example of, the GPCC decoder 300 may include a geometric arithmetic decoding unit 302, an attribute arithmetic decoding unit 304, an octree synthesis unit 306, an inverse quantization unit 308, a surface approximation synthesis unit 310, a geometric reconstruction unit 312, a RAHT unit 314, a LOD generation unit 316, an inverse lifting unit 318, a coordinate inverse transformation unit 320, and a color inverse transformation unit 322.

[0046] The GPCC decoder 300 may obtain a geometric bitstream and an attribute bitstream. The geometric arithmetic decoding unit 302 of the decoder 300 may apply arithmetic decoding (e.g., CABAC or other types of arithmetic decoding) to the syntax elements in the geometric bitstream. Similarly, the attribute arithmetic decoding unit 304 may apply arithmetic decoding to the syntax elements in the attribute bitstream.

[0047] The octree synthesis unit 306 may synthesize an octree based on the syntax elements parsed from the geometric bitstream. In the case where surface approximation is used in the geometric bitstream, the surface approximation synthesis unit 310 may determine a surface model based on the syntax elements parsed from the geometric bitstream and based on the octree.

[0048] In addition, the geometric reconstruction unit 312 may perform reconstruction to determine the coordinates of the points in the point cloud. The coordinate inverse transformation unit 320 may apply an inverse transformation to the reconstructed coordinates to transform the reconstructed coordinates (positions) of the points in the point cloud from the transform domain back to the initial domain.

[0049] Additionally, in Figure 3 the example of, the inverse quantization unit 308 inverse quantizes the attribute values. The attribute values may be based on the syntax elements obtained from the attribute bitstream (e.g., including the syntax elements decoded by the attribute arithmetic decoding unit 304).

[0050] Depending on how the attribute values are encoded, the RAHT unit 314 may perform RAHT decoding to determine color values for points in the point cloud based on the dequantized attribute values. Alternatively, the LOD generation unit 316 and the inverse lifting unit 318 may use a level-of-detail-based technique to determine color values for points in the point cloud.

[0051] In addition, in Figure 3 the example of, the color inverse transform unit 322 may apply an inverse color transform to the color values. The inverse color transform may be the inverse of the color transform applied by the color transform unit 204 of the encoder 200. For example, the color transform unit 204 may transform color information from the RGB color space to the YCbCr color space. Correspondingly, the color inverse transform unit 322 may transform color information from the YCbCr color space to the RGB color space.

[0052] Figure 2 and Figure 3 The various units of are shown to assist in understanding the operations performed by the encoder 200 and the decoder 300. These units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function and is preset with respect to the operations that can be performed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations that can be performed. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit may execute software instructions (e.g., to receive parameters or output parameters), but the type of operations performed by the fixed-function circuit is generally immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.

[0053] Some exemplary embodiments of the present disclosure will be described in detail below. It should be understood that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to GPCC or other specific point cloud codecs, the disclosed techniques are also applicable to other point cloud coding and decoding techniques. In addition, although some embodiments describe the point cloud coding and decoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. 1. Brief Overview This disclosure relates to point cloud encoding and decoding techniques. Specifically, this disclosure relates to sample domain prediction for Region Adaptive Hierarchical Transform (RAHT). These concepts can be applied, either alone or in various combinations, to any point cloud encoding and decoding standard or non-standard point cloud codec, such as the Geometry-based Point Cloud Compression (G-PCC) and Low Latency Low Complexity Codec (L3C2) that are under development. 2. Abbreviations G-PCC Geometry-based Point Cloud Compression L3C2 Low Latency Low Complexity Codec MPEG Moving Picture Experts Group 3DG 3D Graphics Coding Team CFP Call for Proposals V-PCC Video-based Point Cloud Compression RAHT Region Adaptive Hierarchical Transform DC Direct Current AC Alternating Current 3. Introduction MPEG is short for Moving Picture Experts Group, which is one of the main standardization groups dealing with multimedia. In 2017, the MPEG 3D Graphics Coding Team (3DG) released a Call for Proposals (CFP) document to start developing a point cloud encoding and decoding standard. The final standard will encompass two categories of solutions. Video-based Point Cloud Compression (V-PCC or VPCC) is applicable to point sets with relatively uniform point distributions. Geometry-based Point Cloud Compression (G-PCC or GPCC) is applicable to sparser distributions. Both V-PCC and G-PCC support encoding and decoding for individual point clouds and point cloud sequences. In a point cloud, there can be geometric information and attribute information. Geometric information is used to describe the geometric positions of data points. Attribute information is used to record some details of data points, such as texture, normal vector, reflection, etc. 3.1 Octree Geometry Compression Point cloud codecs can process various information in different ways. Generally, there are many optional tools in the codec to separately support the encoding and decoding of geometric information and attribute information. Among the geometric encoding and decoding tools in G-PCC, octree geometry compression has an important impact on point cloud geometry encoding and decoding performance. In G-PCC, one of the important point cloud geometry encoding and decoding tools is octree geometry compression, which utilizes the spatial correlation of point cloud geometry. If the geometry encoding and decoding tool is enabled, a cube axis-aligned bounding box associated with the octree root node will be determined based on the point cloud geometry information. Then, the bounding box will be subdivided into 8 sub-cubes, and the 8 sub-cubes are associated with the 8 child nodes of the root node (the cube is equivalent to the node in the following text). Then, 8-bit codes are generated in a specific order to indicate whether the 8 child nodes contain points respectively, where one bit is associated with one child node. The bit associated with a child node is named the occupancy bit, and the generated 8-bit code is named the occupancy code. The generated occupancy code will be signaled based on the occupancy information of neighboring nodes. Then, only the nodes containing points will be further subdivided into 8 child nodes. This process will be executed recursively until the node size is 1. Thus, the point cloud geometry information is converted into a sequence of occupancy codes. On the decoder side, the sequence of occupancy codes will be decoded, and the point cloud geometry information can be reconstructed based on the sequence of occupancy codes. The breadth-first scan order will be used for the octree. In one level of the octree, the octree nodes will be scanned in Morton order. If the coordinates of a node are represented by N bits, the coordinates (X, Y, Z) of the node can be represented as follows. X = (x N-1 x N-2 …x1x0) Y = (y N-1 y N-2 …y1y0) Z = (z N-1 Z N-2 …z1z0) The Morton code of the coordinates (X, Y, Z) can be represented as follows. M = (x N-1 y N-1 z N-1 x N-2 y N-2 z N-2 …x1y1z1x0y0z0) The Morton order is based on the ascending or descending order of the Morton code. 3.2 Region Adaptive Hierarchical Transform In G-PCC, one of the important point cloud attribute encoding and decoding tools is RAHT. RAHT is a transform that uses the attributes associated with nodes in the lower levels of an octree to predict the attributes of nodes in the next level. Assume that the positions of the points are given at both the encoder and the decoder. RAHT traverses the octree in reverse, from the leaf nodes to the root node, recombining the nodes into larger nodes at each step until the root node is reached. At each level of the octree, the nodes are processed in Morton order. At each decomposition, instead of aggregating eight nodes at once, RAHT performs the aggregation in three steps along each dimension (e.g., along z, then y, then x). If there are L levels in the octree, RAHT uses 3L levels to traverse the tree in reverse. For integers x, y, z, let the node be g at level l l,x,y,z . g l,x,y,z is obtained by aggregating g l+1,2x,y,z and g l+1,2x+1,y,z , where the aggregation along the first dimension is an example. RAHT only processes occupied nodes. If one of the nodes in the pair is not occupied, the other node is promoted to the next level without being processed, i.e., if g l,2x,y,z is the occupied node in the pair, then g l-1,x,y,z = g l,2x,y,x . The aggregation process is repeated until the root is reached. Note that the aggregation process generates nodes at the lower levels that are the result of aggregating different numbers of voxels along the path. The number of nodes that are aggregated to generate the node g l,x,y,z is the weight ω l,x,y,z of that node. At each aggregation of two nodes (say g l,2x,y,z and g l,2x+1,y,z ) using their respective weights (ω l,2x,y,z and ω l,2x+1,y,z ), RAHT applies the following transform: where ω1 = ω l,2x,y,z and ω2 = ω l,2x+1,y,z , and Note that the transformation matrix always changes, adapting to the weights, i.e., adapting to the number of leaf nodes that each g l,x,y,z actually represents. The quantity g l,x,y,z is used for aggregating and composing other nodes at the lower levels. h l,x,y,z is the actual high-pass coefficient generated by the transform and to be encoded and transmitted. In addition, the weights are accumulated for the above levels. In the above example, ω l-1,2,y,z = ω l,2x,y,z + ω l,2x+1,y,z In the last level (the root of the tree), the remaining two voxels g 1,0,0,0 and g 1,1,0,0 are transformed into the final two coefficients as follows: where g DC = g 0,0,0,0 . 3.2.1 Upsampling Transform Domain Prediction Transform domain prediction is introduced to improve the coding and decoding efficiency of RAHT. The transform domain prediction consists of two parts. First, the RAHT tree traversal is changed from the previous ascending order method to a descending order based one, i.e., a tree of the sum of attributes and weights is constructed, and then RAHT is performed for both the encoder and the decoder from the root level to the leaf level of the tree. At each level, the nodes are accessed in Morton order. The transformation is performed in the nodes with 2×2×2 child nodes in the next level. The nodes in which the transformation is performed can be called transform nodes. Second, for each child node of the transform node, the corresponding predicted attribute is generated by upsampling the attributes of the previous transform level. In fact, only the child nodes containing at least one point will generate the corresponding predicted attributes. On the encoder side, the transform nodes containing the predicted attributes are transformed and subtracted from the transformed attributes. The residuals of the alternating current (AC) coefficients will be signaled. Note that the prediction does not affect the direct current (DC) coefficients. Each child node of the transform node is predicted by 7 parent level nodes, where 3 collinear parent level neighboring nodes, 3 coplanar parent level neighboring nodes and 1 parent node. The coplanar neighbors and collinear neighbors are the neighbors sharing a face and an edge with the current transform node respectively. The binary search algorithm is used to find the coplanar parent level neighbors and collinear parent level neighbors. Figure 4 Shows the parent level nodes of each child node of the transform unit node. Figure 4 Shows the 7 parent level nodes for each child node of the transform node. The attribute a up of each child node is predicted as follows depending on the distance between the child node and its parent level nodes. a up = ∑ω k a k / ∑ω k where a k is the attribute of a parent level node of the child node, and ω k is the weight depending on the distance. In G-PCC, ω parent : ω coplane : ω coline = 4∶2∶1. 3.2.2 Early Termination of Transform Domain Prediction Early termination is introduced to reduce complexity. In upsampled transform domain prediction, 7 parent-level neighboring nodes are used to create a predicted value for each encoded target node (child node) of the transform node. And a total of 19 parent-level neighboring nodes (including the parent node, i.e., the transform unit node) are used to create predicted values for all 8 encoded target nodes of the transform unit node. Since the number of valid neighboring parent nodes is larger, the prediction accuracy will be better in denser point clouds. On the contrary, the prediction accuracy will be worse in sparser point clouds. Based on this feature, early termination for upsampling transform domain prediction is introduced to reduce the encoding and decoding time. At early termination, the following two parameters are calculated for each of the 8 child nodes of the transform unit node. · NumValidP: The total number of valid parent-level neighboring nodes (including the parent node). · NumValidGP: The total number of valid grandparent-level neighboring nodes (including the grandparent node). Then, the prediction will be disabled if NumValidP or NumValidGP is less than the threshold. This means that the prediction is terminated when the number of valid neighboring nodes becomes small. 3.3 Problems The existing designs for point cloud attribute transform domain prediction in region adaptive hierarchical transform have the following problems: Attribute prediction is performed in the transform domain, and the attribute prediction is sub-optimal in terms of complexity. In transform domain prediction, transforms need to be performed on the transform node and the predicted transform node. If the attribute prediction can be performed in the sample domain, only the transform on the predicted residual transform node is required. The complexity from the transform can be reduced in sample domain prediction. At the same time, sample domain prediction and transform domain prediction are mathematically equivalent. 4. Detailed Solutions To solve the above problems and some unmentioned problems, the method outlined below is disclosed. The embodiments should be regarded as examples for explaining general concepts and should not be interpreted in a narrow way. In addition, these embodiments can be applied individually or in any combination. 1) The attributes of at least one neighbor can be used to predict the attributes of at least one child node of the transform node. a. In one example, a neighbor can be a node that shares at least a face, or an edge, or a vertex with the transform node. i. The neighbor and the transform node can share the same octree depth. ii. The neighbor and the transform node can have different octree depths. b. In one example, a neighbor can be a node that shares at least a face, or an edge, or a vertex with at least one child node of the transformation node. i. The neighbor and at least one child node of the transformation node can share the same octree depth. ii. The neighbor and at least one child node of the transformation node can have different octree depths. c. In one example, a neighbor can be a node that is spatially close to the transformation node or at least one child node of the transformation node. i. In one example, the distance can be Euclidean distance, Manhattan distance, Chebyshev distance, etc. ii. The neighbor and at least one child node of the transformation node can share the same octree depth. iii. The neighbor and at least one child node of the transformation node can have different octree depths. iv. The neighbor and the transformation node can share the same octree depth. v. The neighbor and the transformation node can have different octree depths. d. In one example, for a neighbor, whether it is used in the prediction and / or how it is used in the prediction can be signaled from the encoder to the decoder. e. In one example, for a neighbor, whether it is used in the prediction and / or how it is used in the prediction can be derived by the decoder. f. In one example, which neighbors to be used in the prediction can be signaled from the encoder to the decoder. g. In one example, which neighbors to be used in the prediction can be derived by the decoder. 2) There can be prediction weights for at least one. a. In one example, the prediction result from at least one neighbor can be a weighted average of the attributes of the neighbor. b. In one example, there can be prediction weights for each neighbor. c. In one example, the prediction weight of a neighbor can be derived based on the position and / or the distance. i. In one example, the distance can be negatively correlated with the distance. 1. In one example, the distance can be Euclidean distance, Manhattan distance, Chebyshev distance, etc. ii. In one example, the distance can be the distance between the neighbor and the transformation node. iii. In one example, the distance can be the distance between the neighbor and a child node of the transformation node. 3) There can be a predicted attribute for a child node of the transformation node. a. In one example, if a child node includes points, this child node may have a prediction attribute. i. In one example, the prediction attribute may be a prediction result. b. The prediction attribute may be a fusion of neighboring attributes. i. In one example, the fusion may be a weighted average of neighboring attributes. 1. In one example, the weighted average may be linear. 2. In one example, the weighted average may be non - linear. 4) A prediction residual for a child node of a transform node can be derived. a. In one example, if a child node includes points, this child node may have a prediction residual. b. In one example, the prediction residual may be the difference between the prediction attribute and the attribute of a child node of the transform node. 5) The prediction residual can be transformed to obtain transform coefficients at the encoder and / or decoder. a. In one example, the transformation may be a region - adaptive hierarchical transformation, a wavelet transformation, a cosine transformation, etc. b. In one example, the transform coefficients of the prediction residual can be transmitted via a signal. i. In one example, only the AC coefficients may be transmitted via a signal. ii. In one example, the DC coefficients can be inherited from a previous transformation process. iii. In one example, the transform coefficients can be further processed multiple times before signaling. 1. In one example, the processing may be quantization. 2. In one example, the processing may be binarization using fixed - length coding / decoding, binarization using EG coding / decoding, binarization using (rounding) unary coding / decoding, etc. 3. Information about the quantization step size can be transmitted via a signal or derived. iv. In one example, the transform coefficients can be coded / decoded using at least one context in arithmetic coding. v. In one example, the transform coefficients can be bypass - coded / decoded. vi. In one example, the transform coefficients can be coded / decoded by run - length coding. 6) The transform coefficients can be inverse - transformed to obtain the prediction residual at the encoder and / or decoder. a. The coefficients can be de - quantized before inverse - transformation. b. In one example, the transformation may be a region - adaptive hierarchical transformation, a wavelet transformation, a cosine transformation, etc. c. In one example, the AC coefficients can be reconstructed from the bitstream. d. In one example, the DC coefficients can be estimated. i. In one example, the DC coefficient of the transform of a transform node can be inherited from a previous transform process. ii. In one example, the DC coefficient of the transform of a prediction residual can be the difference between the mean of the inherited value and the mean of the prediction attributes of the children nodes of the transform node. e. In one example, the AC and DC coefficients of the transform of a prediction residual can be inverse-transformed to obtain the prediction residual. 7) Whether to apply the transform / inverse transform can be signaled from the encoder or deduced at the decoder. 8) Which transform / inverse transform is used can be signaled from the encoder or deduced at the decoder. 5. Embodiments Figure 5 An example of an encoding / decoding process 500 for sample-domain prediction of region-adaptive hierarchical transform is shown. At 510, the neighboring attributes of a transform node are obtained. Then, at 520, prediction attributes are calculated for each child node of the transform node according to the neighboring attributes of the transform node and corresponding weights. At 530, an attribute residual for each child node of the transform node is determined by calculating the difference between the prediction attribute and the child node attribute of the transform node. At 540, a transform is performed on the attribute residual determined at 530. At 550, the attribute residual on which the transform has been performed is quantized. At 560, the quantized attribute residual is signaled.

[0054] More details will be discussed further below. Figure 6 A flowchart of a method 600 for point cloud encoding / decoding according to an embodiment of the present disclosure is shown.

[0055] At block 610, a transform result of an attribute residual between neighboring attributes of at least one sub-block of a transform block and prediction attributes of at least one sub-block of the transform block is determined for a conversion between a point cloud sequence including a current point cloud (PC) sample associated with the transform block and a bitstream of the point cloud sequence. The neighboring attributes are predicted based on attributes of at least one neighboring block of the transform block.

[0056] At block 620, the conversion is performed based at least on the transform result of the attribute residual.

[0057] According to method 600, attribute prediction can be performed in the sample domain, so that the transform only needs to be performed on prediction residual transform nodes. In this way, the complexity caused by the transform can be reduced.

[0058] In some embodiments, at least one neighboring block includes a block that shares at least a face, or an edge, or a vertex with a transform block.

[0059] In some embodiments, at least one neighboring block and a transform block share the same tree level or have different tree levels.

[0060] In some embodiments, the tree level includes an octree depth.

[0061] In some embodiments, at least one neighboring block includes a block that shares at least a face, or an edge, or a vertex with at least one sub-block of a transform block.

[0062] In some embodiments, at least one neighboring block and at least one sub-block share the same tree level or have different tree levels.

[0063] In some embodiments, the tree level includes an octree depth.

[0064] In some embodiments, at least one neighboring block includes a block that is closest in terms of distance to a transform block or at least one sub-block.

[0065] In some embodiments, the distance is one of the following: Euclidean distance, Manhattan distance, or Chebyshev distance.

[0066] In some embodiments, at least one neighboring block and at least one sub-block share the same tree level or have different tree levels.

[0067] In some embodiments, at least one neighboring block and a transform block share the same tree level or have different tree levels.

[0068] In some embodiments, the tree level includes an octree depth.

[0069] In some embodiments, for one neighboring block among at least one neighboring block, whether it is used in prediction and / or how it is used in prediction is indicated from an encoder to a decoder or is deduced by the decoder. Alternatively or additionally, in some embodiments, information about at least one neighboring block to be used in prediction is indicated from the encoder to the decoder or is deduced by the decoder.

[0070] In some embodiments, there is a prediction weight for each neighboring block among at least one neighboring block.

[0071] In some embodiments, the result of the neighboring attribute is a weighted average obtained based on the attributes of the neighboring block and the corresponding prediction weights of the neighboring block.

[0072] In some embodiments, the prediction weight of the neighboring block is deduced based on a distance including at least one of the following: the distance between the neighboring block and the transform block, the distance between the neighboring block and a sub-block of the transform block.

[0073] In some embodiments, the prediction weight is negatively correlated with the distance, or the distance is one of the following: Euclidean distance, Manhattan distance, or Chebyshev distance.

[0074] In some embodiments, there is a prediction attribute for a sub-block of a transform block.

[0075] In some embodiments, if a sub-block includes at least one point, the sub-block has a prediction attribute.

[0076] In some embodiments, the prediction attribute is a prediction result, or the prediction attribute is a fusion of multiple neighbor attributes.

[0077] In some embodiments, the fusion is a weighted average of multiple neighbor attributes.

[0078] In some embodiments, the weighted average can be linear or can be non-linear.

[0079] In some embodiments, an attribute residual is determined based on neighbor attributes of at least one sub-block of a transform block and prediction attributes of at least one sub-block of the transform block.

[0080] In some embodiments, if a sub-block includes at least one point, the sub-block has an attribute residual.

[0081] In some embodiments, the attribute residual is the difference between a neighbor attribute and a prediction attribute of a sub-block of a transform block.

[0082] In some embodiments, a transform result is obtained by transforming an attribute residual at an encoder and / or a decoder.

[0083] In some embodiments, the transform is one of the following: region adaptive hierarchical transform, wavelet transform, or cosine transform.

[0084] In some embodiments, the transform result includes transform coefficients, and the transform coefficients are indicated in a bitstream.

[0085] In some embodiments, the transform coefficients include alternating current (AC) coefficients and direct current (DC) coefficients, and only the AC coefficients are indicated in the bitstream.

[0086] In some embodiments, the DC coefficients are inherited from a previous transform process at a decoder.

[0087] In some embodiments, the transform coefficients are further processed multiple times before being indicated in the bitstream.

[0088] In some embodiments, the processing includes at least one of the following: quantization, or binarization using fixed-length coding / decoding, binarization using EG coding / decoding, binarization using unary coding / decoding, or binarization using truncated unary coding / decoding.

[0089] In some embodiments, information about the quantization step size is indicated in the bitstream or derived at the decoder.

[0090] In some embodiments, the transform coefficients are coded using at least one context in arithmetic coding. Alternatively, the transform coefficients are bypass-coded. As another alternative, the transform coefficients are run-length coded.

[0091] In some embodiments, the transform coefficients are inverse-transformed to obtain the attribute residuals at the encoder and / or decoder.

[0092] In some embodiments, the coefficients are dequantized before the inverse transform. Alternatively, the transform is one of the following: region-adaptive hierarchical transform, wavelet transform, cosine transform. Alternatively, the AC coefficients are reconstructed from the bitstream. As another alternative, the DC coefficients are estimated.

[0093] In some embodiments, the DC coefficient of the transform of the transform block is inherited from a previous transform process. Alternatively, the DC coefficient of the transform of the attribute residuals is the difference between the mean of the inherited values and the mean of the predicted attributes of the sub-blocks of the transform block.

[0094] In some embodiments, the AC and DC coefficients of the transform of the attribute residuals are inverse-transformed to obtain the attribute residuals.

[0095] In some embodiments, whether to apply the transform / inverse transform is indicated from the encoder to the decoder. Alternatively, whether to apply the transform / inverse transform is derived at the decoder.

[0096] In some embodiments, information about the transform / inverse transform being used is indicated from the encoder to the decoder. Alternatively, this information can be derived at the decoder.

[0097] In some embodiments, the current PC sample is one of the following: frame, picture, slice, sub-frame, sub-picture, tile, or segment.

[0098] In some embodiments, the conversion includes encoding the current PC sample into the bitstream.

[0099] In some embodiments, the conversion includes decoding the current PC sample from the bitstream.

[0100] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for point cloud encoding and decoding. According to the method, a transformed result of an attribute residual between a neighboring attribute of at least one sub-block of a transform block and a predicted attribute of at least one sub-block of the transform block is determined. The neighboring attribute is predicted based on an attribute of at least one neighboring block of the transform block. Then, a bitstream is generated based at least on the transformed result of the attribute residual.

[0101] According to still some other embodiments of the present disclosure, a method for storing a bitstream of a point cloud sequence is provided. According to the method, a transformed result of an attribute residual between a neighboring attribute of at least one sub-block of a transform block and a predicted attribute of at least one sub-block of the transform block is determined. The neighboring attribute is predicted based on an attribute of at least one neighboring block of the transform block. Then, a bitstream is generated based at least on the transformed result of the attribute residual. The generated bitstream is stored in a non-transitory computer-readable recording medium.

[0102] Embodiments of the present disclosure may be described according to the following items, and the features of these items may be combined in any reasonable manner.

[0103] Item 1. A method for point cloud encoding and decoding, comprising: for a conversion between a point cloud sequence including a current point cloud (PC) sample associated with a transform block and a bitstream of the point cloud sequence, determining a transformed result of an attribute residual between a neighboring attribute of at least one sub-block of the transform block and the predicted attribute of the at least one sub-block of the transform block, the neighboring attribute being predicted based on an attribute of at least one neighboring block of the transform block; and performing the conversion based at least on the transformed result of the attribute residual.

[0104] Item 2. The method according to Item 1, wherein the at least one neighboring block includes a block that shares at least a face, or an edge, or a vertex with the transform block.

[0105] Item 3. The method according to Item 1 or 2, wherein the at least one neighboring block and the transform block share the same tree level or have different tree levels.

[0106] Item 4. The method according to Item 3, wherein the tree level includes an octree depth.

[0107] Item 5. The method according to Item 1, wherein the at least one neighboring block includes a block that shares at least a face, or an edge, or a vertex with the at least one sub-block of the transform block.

[0108] Item 6. The method according to Item 1 or 5, wherein the at least one neighboring block and the at least one sub-block share the same tree level or have different tree levels.

[0109] Item 7. The method according to Item 6, wherein the tree level includes the octree depth.

[0110] Item 8. The method according to Item 1, wherein the at least one neighboring block includes the block that is closest in terms of distance to the transform block or the at least one sub-block.

[0111] Item 9. The method according to Item 8, wherein the distance is one of the following: Euclidean distance, Manhattan distance, or Chebyshev distance.

[0112] Item 10. The method according to Item 8, wherein the at least one neighboring block and the at least one sub-block share the same tree level or have different tree levels.

[0113] Item 11. The method according to Item 8, wherein the at least one neighboring block and the transform block share the same tree level or have different tree levels.

[0114] Item 12. The method according to Item 10 or 11, wherein the tree level includes the octree depth.

[0115] Item 13. The method according to Item 1, wherein for one neighboring block among the at least one neighboring block, whether it is used in the prediction and / or how it is used in the prediction is indicated from the encoder to the decoder or is deduced by the decoder, or wherein the information of the at least one neighboring block to be used in the prediction is indicated from the encoder to the decoder or is deduced by the decoder.

[0116] Item 14. The method according to Item 1, wherein there are prediction weights for each neighboring block among the at least one neighboring block.

[0117] Item 15. The method according to Item 14, wherein the result of the neighboring property is a weighted average obtained based on the properties of the neighboring blocks and the corresponding prediction weights of the neighboring blocks.

[0118] Item 16. The method according to Item 14 or 15, wherein the prediction weights of the neighboring blocks are deduced based on a distance including at least one of the following: the distance between the neighboring block and the transform block, the distance between the neighboring block and the sub-block of the transform block.

[0119] Item 17. The method according to Item 16, wherein the prediction weight is negatively correlated with the distance, or wherein the distance is one of the following: Euclidean distance, Manhattan distance, or Chebyshev distance.

[0120] Item 18. The method according to Item 1, wherein there is a prediction attribute for a sub-block of the transform block.

[0121] Item 19. The method according to Item 18, wherein if the sub-block includes at least one point, the sub-block has a prediction attribute.

[0122] Item 20. The method according to Item 18 or 19, wherein the prediction attribute is the prediction result, or wherein the prediction attribute is a fusion of multiple neighbor attributes.

[0123] Item 21. The method according to Item 22, wherein the fusion is a weighted average of the multiple neighbor attributes.

[0124] Item 22. The method according to Item 21, wherein the weighted average is linear or non-linear.

[0125] Item 23. The method according to Item 21, wherein the attribute residual is determined based on the neighbor attributes of at least one sub-block of the transform block and the prediction attributes of at least one sub-block of the transform block.

[0126] Item 24. The method according to Item 23, wherein if a sub-block includes at least one point, the sub-block has an attribute residual.

[0127] Item 25. The method according to Item 23 or 24, wherein the attribute residual is the difference between the neighbor attribute and the prediction attribute of a sub-block of the transform block.

[0128] Item 26. The method according to any one of Items 1 to 25, wherein the transform result is obtained by transforming the attribute residual at the encoder and / or decoder.

[0129] Item 27. The method according to any one of Items 1 to 26, wherein the transform is one of the following: region adaptive hierarchical transform, wavelet transform, or cosine transform.

[0130] Item 28. The method according to any one of Items 1 to 27, wherein the transform result includes transform coefficients, and the transform coefficients are indicated in the bitstream.

[0131] Item 29. The method according to Item 28, wherein the transform coefficients include alternating current (AC) coefficients and direct current (DC) coefficients, and wherein only the AC coefficients are indicated in the bitstream.

[0132] Item 30. The method according to Item 29, wherein at the decoder the DC coefficients are inherited from a previous transform process.

[0133] Item 31. The method according to any one of Items 28 to 30, wherein the transform coefficients are further processed multiple times before being indicated in the bitstream.

[0134] Item 32. The method according to Item 31, wherein the processing includes at least one of the following: quantization, or binarization using fixed-length coding and decoding, binarization using exponential-Golomb coding, binarization using unary coding, or binarization using truncated unary coding.

[0135] Item 33. The method according to Item 32, wherein information about the quantization step size is indicated in the bitstream or is derived at the decoder.

[0136] Item 34. The method according to Item 28, wherein the transform coefficients are coded using at least one context in arithmetic coding, or wherein the transform coefficients are bypass-coded, or wherein the transform coefficients are run-length coded.

[0137] Item 35. The method according to Item 28, wherein the transform coefficients are inverse-transformed to obtain the attribute residual at the encoder and / or decoder.

[0138] Item 36. The method according to Item 35, wherein the coefficients are dequantized before the inverse transform, or wherein the transform is one of the following: region-adaptive hierarchical transform, wavelet transform, cosine transform, or wherein the AC coefficients are reconstructed from the bitstream, or wherein the DC coefficients are estimated.

[0139] Item 37. The method according to Item 36, wherein the DC coefficients of the transform of the transform block are inherited from a previous transform process, or wherein the DC coefficients of the transform of the attribute residual are the difference between the mean of the inherited values and the mean of the predicted attributes of the sub-blocks of the transform block.

[0140] Item 38. The method according to Item 35, wherein the AC coefficients and the DC coefficients of the transform of the attribute residual are inverse-transformed to obtain the attribute residual.

[0141] Item 39. The method according to any one of Items 1 to 38, wherein whether to apply the transformation / inverse transformation is indicated from the encoder or deduced at the decoder.

[0142] Item 40. The method according to any one of Items 1 to 39, wherein information about the transformation / inverse transformation being used is indicated from the encoder or deduced at the decoder.

[0143] Item 41. The method according to any one of Items 1 to 40, wherein the current PC sample is one of the following: a frame, a picture, a slice, a sub-frame, a sub-picture, a tile, or a segment.

[0144] Item 42. The method according to any one of Items 1 to 41, wherein the transformation includes encoding the current PC sample into the bitstream.

[0145] Item 43. The method according to any one of Items 1 to 41, wherein the transformation includes decoding the current PC sample from the bitstream.

[0146] Item 44. An apparatus for point cloud encoding and decoding, including a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 43.

[0147] Item 45. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 43.

[0148] Item 46. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by a point cloud processing device for a point cloud sequence, wherein the method includes: determining a transformation result of an attribute residual between a neighboring attribute of at least one sub-block of a transformation block and a prediction attribute of the at least one sub-block of the transformation block, the neighboring attribute being predicted based on an attribute of at least one neighboring block of the transformation block; and generating the bitstream based at least on the transformation result of the attribute residual.

[0149] Item 47. A method for storing a bitstream of a point cloud sequence, including: determining a transformation result of an attribute residual between a neighboring attribute of at least one sub-block of a transformation block and a prediction attribute of the at least one sub-block of the transformation block, the neighboring attribute being predicted based on an attribute of at least one neighboring block of the transformation block; generating the bitstream based at least on the transformation result of the attribute residual; and storing the bitstream in a non-transitory computer-readable recording medium. Example device

[0150] Figure 7FIG. shows a block diagram of a computing device 700 in which various embodiments of the present disclosure may be implemented. The computing device 700 may be implemented as the source device 110 (or the GPCC encoder 116 or 200) or the destination device 120 (or the GPCC decoder 126 or 300), or may be included in the source device 110 (or the GPCC encoder 116 or 200) or the destination device 120 (or the GPCC decoder 126 or 300).

[0151] It should be understood that Figure 7 the computing device 700 shown in is for illustrative purposes only and does not imply any limitation on the functionality and scope of the embodiments of the present disclosure in any way.

[0152] As Figure 7 shown, the computing device 700 includes a general computing device 700. The computing device 700 may include at least one or more processors or processing units 710, a memory 720, a storage unit 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760.

[0153] In some embodiments, the computing device 700 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is conceivable that the computing device 700 may support any type of interface to the user (such as "wearable" circuitry, etc.).

[0154] The processing unit 710 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in the memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 700. The processing unit 710 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0155] Computing device 700 generally includes various computer storage media. Such media can be any media accessible by computing device 700, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. Memory 720 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. Storage unit 730 can be any removable or non-removable media and can include machine-readable media such as memory, flash drives, magnetic disks, or other media that can be used to store information and / or data and can be accessed in computing device 700.

[0156] Computing device 700 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 7 , a disk drive for reading from and / or writing to a removable non-volatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In such a case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0157] Communication unit 740 communicates with another computing device via a communication medium. Additionally, the functions of the components in computing device 700 may be implemented by a single computing cluster or multiple computer machines that may communicate via a communication connection. Thus, computing device 700 may operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.

[0158] Input device 750 can be one or more of various input devices such as a mouse, keyboard, trackball, voice input device, etc. Output device 760 can be one or more of various output devices such as a display, speaker, printer, etc. With the aid of communication unit 740, computing device 700 can also communicate with one or more external devices (not shown), such as storage devices and display devices, computing device 700 can also communicate with one or more devices that enable a user to interact with computing device 700, or if needed, computing device 700 can also communicate with any device that enables computing device 700 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be carried out via an input / output (I / O) interface (not shown).

[0159] In some embodiments, some or all components of computing device 700 may also be arranged in a cloud computing architecture rather than integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses suitable protocols to provide services via a wide area network, such as the Internet. For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. The computing resources in a cloud computing environment may be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure may provide services through shared data centers, although to a user, they appear as a single access point. Thus, the cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server or installed directly or otherwise on a client device.

[0160] In an embodiment of the present disclosure, computing device 700 may be used to implement point cloud encoding / decoding. Memory 720 may include one or more point cloud encoding / decoding modules 725 having one or more program instructions. These modules are accessible and executable by processing unit 710 to perform the functions of the various embodiments described herein.

[0161] In an example embodiment of performing point cloud encoding, input device 750 may receive point cloud data as input 770 to be encoded. The point cloud data may be processed, for example, by point cloud encoding / decoding module 725 to generate an encoded bitstream. The encoded bitstream may be provided as output 780 via output device 760.

[0162] In an example embodiment of performing point cloud decoding, input device 750 may receive the encoded bitstream as input 770. The encoded bitstream may be processed, for example, by point cloud encoding / decoding module 725 to generate decoded point cloud data. The decoded point cloud data may be provided as output 780 via output device 760.

[0163] Although the present disclosure has been specifically shown and described with reference to preferred embodiments of the present disclosure, those skilled in the art will understand that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for point cloud encoding and decoding, comprising: For the conversion between a point cloud sequence including a current point cloud (PC) sample associated with a transform block and a bitstream of the point cloud sequence, determine a transform result of an attribute residual between a neighboring attribute of at least one sub-block of the transform block and a prediction attribute of the at least one sub-block of the transform block, the neighboring attribute being predicted based on an attribute of at least one neighboring block of the transform block; and perform the conversion based at least on the transform result of the attribute residual.

2. The method according to claim 1, wherein the at least one neighboring block includes a block that shares at least a face, or an edge, or a vertex with the transform block.

3. The method according to claim 1 or 2, wherein the at least one neighboring block and the transform block share the same tree level or have different tree levels.

4. The method according to claim 3, wherein the tree level includes the octree depth.

5. The method according to claim 1, wherein the at least one neighboring block includes a block that shares at least a face, or an edge, or a vertex with at least one sub-block of the transform block.

6. The method according to claim 1 or 5, wherein the at least one neighboring block and the at least one sub-block share the same tree level or have different tree levels.

7. The method according to claim 6, wherein the tree level includes the octree depth.

8. The method according to claim 1, wherein the at least one neighboring block includes a block that is closest to the transform block or the at least one sub-block in terms of distance.

9. The method according to claim 8, wherein the distance is one of the following: Euclidean distance, Manhattan distance, or Chebyshev distance.

10. The method according to claim 8, wherein the at least one neighboring block and the at least one sub-block share the same tree level or have different tree levels.

11. The method according to claim 8, wherein the at least one neighboring block and the transform block share the same tree level or have different tree levels.

12. The method according to claim 10 or 11, wherein the tree level includes the octree depth.

13. The method according to claim 1, wherein for one neighboring block among the at least one neighboring block, whether it is used in the prediction and / or how it is used in the prediction is indicated from the encoder to the decoder or is deduced by the decoder, or wherein the information of the at least one neighboring block to be used in the prediction is indicated from the encoder to the decoder or is deduced by the decoder.

14. The method according to claim 1, wherein there are prediction weights for each neighboring block among the at least one neighboring block.

15. The method according to claim 14, wherein the result of the neighboring property is a weighted average obtained based on the properties of neighboring blocks and the corresponding prediction weights of the neighboring blocks.

16. The method according to claim 14 or 15, wherein the prediction weights of the neighboring blocks are derived based on a distance including at least one of the following: The distance between the neighboring block and the transform block, The distance between the neighboring block and a sub-block of the transform block.

17. The method according to claim 16, wherein the prediction weight is negatively correlated with the distance, or wherein the distance is one of the following: Euclidean distance, Manhattan distance, or Chebyshev distance.

18. The method according to claim 1, wherein there is a prediction property for a sub-block of the transform block.

19. The method according to claim 18, wherein if the sub-block includes at least one point, the sub-block has a prediction property.

20. The method according to claim 18 or 19, wherein the prediction property is the prediction result, or wherein the prediction property is a fusion of multiple neighboring properties.

21. The method according to claim 22, wherein the fusion is a weighted average of the multiple neighboring properties.

22. The method according to claim 21, wherein the weighted average is linear or non-linear.

23. The method according to claim 21, wherein the property residual is determined based on the neighboring property of the at least one sub-block of the transform block and the prediction property of the at least one sub-block of the transform block.

24. The method according to claim 23, wherein if a sub-block includes at least one point, the sub-block has a property residual.

25. The method according to claim 23 or 24, wherein the property residual is the difference between the neighboring property and the prediction property of a sub-block of the transform block.

26. The method according to any one of claims 1 to 25, wherein the transform result is obtained by transforming the property residual at the encoder and / or decoder.

27. The method according to any one of claims 1 to 26, wherein the transform is one of the following: region adaptive hierarchical transform, wavelet transform, or cosine transform.

28. The method according to any one of claims 1 to 27, wherein the transform result includes transform coefficients, and the transform coefficients are indicated in the bitstream.

29. The method according to claim 28, wherein the transform coefficients include alternating current (AC) coefficients and direct current (DC) coefficients, and wherein only the AC coefficients are indicated in the bitstream.

30. The method according to claim 29, wherein at the decoder the DC coefficients are inherited from a previous transform process.

31. The method according to any one of claims 28 to 30, wherein the transform coefficients are further processed multiple times before being indicated in the bitstream.

32. The method according to claim 31, wherein the processing comprises at least one of the following: quantization, or binarization using fixed - length coding / decoding, binarization using exponential - Golomb (EG) coding / decoding, binarization using unary coding / decoding, or binarization using truncated unary coding / decoding.

33. The method according to claim 32, wherein information about the quantization step size is indicated in the bitstream or is derived at the decoder.

34. The method according to claim 28, wherein the transform coefficients are coded using at least one context in arithmetic coding, or wherein the transform coefficients are bypass - coded, or wherein the transform coefficients are coded by run - length coding.

35. The method according to claim 28, wherein the transform coefficients are inverse - transformed to obtain the attribute residual at the encoder and / or decoder.

36. The method according to claim 35, wherein the coefficients are de - quantized before the inverse transform, or wherein the transform is one of: region - adaptive hierarchical transform, wavelet transform, cosine transform, or wherein the AC coefficients are reconstructed from the bitstream, or wherein the DC coefficients are estimated.

37. The method according to claim 36, wherein the DC coefficients of the transform of the transform block are inherited from a previous transform process, or wherein the DC coefficients of the transform of the attribute residual are the difference between the mean of the inherited values and the mean of the predicted attributes of the sub - blocks of the transform block.

38. The method according to claim 35, wherein the AC coefficients and the DC coefficients of the transform of the attribute residual are inverse - transformed to obtain the attribute residual.

39. The method according to any one of claims 1 to 38, wherein whether to apply the transform / inverse transform is indicated from the encoder or is derived at the decoder.

40. The method according to any one of claims 1 to 39, wherein information about which transform / inverse transform is used is indicated from the encoder or derived at the decoder.

41. The method according to any one of claims 1 to 40, wherein the current PC sample is one of the following: frame, picture, slice, sub - frame, sub - picture, tile, or segment.

42. The method according to any one of claims 1 to 41, wherein the transformation includes encoding the current PC sample into the bitstream.

43. The method according to any one of claims 1 to 41, wherein the transformation includes decoding the current PC sample from the bitstream.

44. An apparatus for point cloud encoding and decoding, comprising a processor and a non - transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 43.

45. A non - transitory computer - readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 43.

46. A non - transitory computer - readable recording medium storing a bitstream generated by a method executed by a point cloud processing apparatus for a point cloud sequence, wherein the method includes: Determine a transform result of an attribute residual between a neighboring attribute of at least one sub-block of a transform block and a prediction attribute of the at least one sub-block of the transform block, the neighboring attribute being predicted based on an attribute of at least one neighboring block of the transform block; and generate the bitstream based at least on the transform result of the attribute residual.

47. A method for storing a bitstream of a point cloud sequence, comprising: Determine a transform result of an attribute residual between a neighboring attribute of at least one sub-block of a transform block and a prediction attribute of the at least one sub-block of the transform block, the neighboring attribute being predicted based on an attribute of at least one neighboring block of the transform block; generate the bitstream based at least on the transform result of the attribute residual; and store the bitstream in a non-transitory computer-readable recording medium.