Point cloud encoding method, point cloud decoding method, bitstream, encoder, decoder and storage medium

By parsing the bitstream and determining the decoding method within the G-PCC framework, the problem of low RAHT encoding efficiency for point cloud attributes is solved, achieving efficient encoding and decoding of point cloud attributes.

WO2026007149A1PCT designated stage Publication Date: 2026-01-08GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/104102
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-05
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

In the point cloud compression encoding and decoding framework of the Moving Picture Experts Group (G-PCC), existing technologies have failed to fully consider the attribute distribution characteristics of the encoded nodes in the slice, resulting in low RAHT encoding efficiency for point cloud attributes.

Method used

By parsing the bitstream, the value of the first syntax element of the unit to be decoded is determined, the decoding method is determined based on the value, and the attributes of the nodes of the unit to be decoded are decoded, thereby improving the encoding and decoding efficiency of point cloud attributes.

Benefits of technology

The efficiency of point cloud attribute encoding and decoding has been improved by taking into account the distribution characteristics of node attributes and optimizing the point cloud decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024104102_08012026_PF_FP_ABST
    Figure CN2024104102_08012026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a point cloud encoding method, a point cloud decoding method, a bitstream, an encoder, a decoder and a storage medium. Taking a decoding end as an example, the point cloud decoding method comprises: parsing a bitstream, and determining a value of a first syntax element of a unit to be decoded; and on the basis of the value of the first syntax element of said unit, performing attribute decoding on a node of said unit to determine an attribute reconstructed value of the node of said unit. In this way, the attribute distribution characteristics of the node of said unit are taken into full consideration, the first syntax element is introduced for said unit to indicate a decoding mode for said unit, and the decoding mode is used to perform point cloud attribute reconstruction on the node of said unit, thereby improving the efficiency of point cloud attribute decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Point cloud coding method, code stream, encoder, decoder and storage medium TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of point cloud coding, and particularly relate to a point cloud coding method, a code stream, an encoder, a decoder and a storage medium. BACKGROUND

[0002] In a Geometry-based Point Cloud Compression (G-PCC) coding framework provided by Moving Picture Experts Group (MPEG), geometry information and attribute information of a point cloud are encoded separately. The attribute encoding of G-PCC can include a Predicting Transform (PT), a Lifting Transform (LT) and a Region Adaptive Hierarchical Transform (RAHT). The first two are point cloud prediction encoding based on the generation order of a Level of Detail (LOD), and the RAHT is an adaptive transform of attribute information from bottom to top based on the construction level of an octree.

[0003] In G-PCC attribute RAHT encoding, a syntax element is set in an attribute parameter set to determine the attribute encoding mode of a current slice. This attribute encoding scheme does not take into account the attribute distribution characteristics of the encoding nodes in the slice, thereby resulting in low attribute RAHT encoding efficiency of the point cloud.

[0004] SUMMARY

[0005] Embodiments of the present application provide a point cloud coding method, a code stream, an encoder, a decoder and a storage medium, which can improve the coding efficiency of point cloud attributes.

[0006] The technical solutions of embodiments of the present application can be implemented as follows:

[0007] In a first aspect, the embodiments of the present application provide a point cloud decoding method applied to a decoder, and the method comprises:

[0008] parsing a code stream to determine the value of a first syntax element of a to-be-decoded unit;

[0009] determining a decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit;

[0010] Attribute decoding is performed on the node of the to-be-decoded unit based on the determined decoding mode, and an attribute reconstruction value of the node of the to-be-decoded unit is determined.

[0011] In a second aspect, an embodiment of the present application provides a point cloud encoding method applied to an encoder, the method comprising:

[0012] Attribute encoding is performed on the node of the to-be-encoded unit based on the candidate encoding mode of the to-be-encoded unit, and an attribute reconstruction value of the node of the to-be-encoded unit is determined.

[0013] Encoding decision is performed on the candidate encoding mode based on the attribute reconstruction value of the node of the to-be-encoded unit, and an encoding mode of the to-be-encoded unit is determined.

[0014] Based on the encoding mode of the to-be-encoded unit, a value of a first syntax element of the to-be-encoded unit is determined.

[0015] The first syntax element of the to-be-encoded unit is encoded, and the obtained encoding bits are written into a bitstream.

[0016] In a third aspect, an embodiment of the present application provides a bitstream, wherein the bitstream is generated by bit encoding to-be-encoded information; wherein the to-be-encoded information comprises at least one of the following: a first syntax element, a second syntax element, a third syntax element, a fourth syntax element, a fifth syntax element and a sixth syntax element.

[0017] The value of the first syntax element is used to indicate a decoding mode of a to-be-decoded unit.

[0018] The value of the second syntax element is used to indicate whether the to-be-decoded unit is allowed to enable region adaptive hierarchical transform prediction.

[0019] The third syntax element is used to indicate whether the to-be-decoded unit is allowed to refer to a first inter-frame reference unit.

[0020] The fourth syntax element is used to indicate whether the to-be-decoded unit is allowed to refer to a second inter-frame reference unit.

[0021] The fifth syntax element is used to indicate a number of to-be-decoded units.

[0022] The sixth syntax element is used to indicate a node number N of a to-be-decoded group.

[0023] In a fourth aspect, an embodiment of the present application provides an encoder, which comprises a first prediction unit, a first determination unit and an encoding unit; wherein:

[0024] The first prediction unit is configured to perform attribute coding on the node of the to-be-encoded unit based on the candidate coding mode of the to-be-encoded unit, and determine an attribute reconstruction value of the node of the to-be-encoded unit.

[0025] The first determination unit is configured to perform coding decision on the candidate coding mode based on the attribute reconstruction value of the node of the to-be-encoded unit, determine the coding mode of the to-be-encoded unit, and determine a value of a first syntax element of the to-be-encoded unit based on the coding mode of the to-be-encoded unit.

[0026] The encoding unit is configured to perform encoding processing on the first syntax element of the to-be-encoded unit, and write obtained coding bits into a bitstream.

[0027] In a fifth aspect, an embodiment of the present application provides an encoder, which comprises a first memory and a first processor; wherein,

[0028] The first memory is configured to store a computer program capable of running on the first processor.

[0029] The first processor is configured to execute the method in the second aspect when running the computer program.

[0030] In a sixth aspect, an embodiment of the present application provides a decoder, which comprises a decoding unit, a second determination unit and a second prediction unit; wherein,

[0031] The decoding unit is configured to parse a bitstream, and determine a value of a first syntax element of a to-be-decoded unit.

[0032] The second determination unit is configured to determine a decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit.

[0033] The second prediction unit is configured to perform attribute decoding on a node of the to-be-decoded unit based on the determined decoding mode, and determine an attribute reconstruction value of the node of the to-be-decoded unit.

[0034] In a seventh aspect, an embodiment of the present application provides a decoder, which comprises a second memory and a second processor; wherein,

[0035] The second memory is configured to store a computer program capable of running on the second processor.

[0036] The second processor is configured to execute the method in the first aspect when running the computer program.

[0037] In an eighth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a bitstream generated by the encoding method.

[0038] In a ninth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program. When the computer program is executed, the method according to the first aspect or the method according to the second aspect is implemented.

[0039] The method provided by the embodiment of the present application, the code stream, the encoder, the decoder and the storage medium, the method comprises: parsing a code stream, determining a value of a first syntax element of a to-be-decoded unit; based on the value of the first syntax element of the to-be-decoded unit, performing attribute decoding on a node of the to-be-decoded unit, and determining a property reconstruction value of the node of the to-be-decoded unit. In this way, the first syntax element is introduced for the to-be-decoded unit to indicate a decoding mode of the to-be-decoded unit by fully considering the node property distribution characteristics of the to-be-decoded unit, and the node of the to-be-decoded unit is reconstructed by using the decoding mode, thereby improving the point cloud attribute decoding efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0040] FIG. 1A is a schematic diagram of a three-dimensional point cloud image;

[0041] FIG. 1B is a partial enlarged view of a three-dimensional point cloud image;

[0042] FIG. 2A is a schematic diagram of six viewing angles of a point cloud image;

[0043] FIG. 2B is a schematic diagram of a data storage format corresponding to a point cloud image;

[0044] FIG. 3 is a schematic diagram of a network architecture of point cloud coding;

[0045] FIG. 4A is a schematic diagram of a composition framework of a G-PCC encoder;

[0046] FIG. 4B is a schematic diagram of a composition framework of a G-PCC decoder;

[0047] FIG. 5A is a schematic diagram of a low-plane position in a Z-axis direction;

[0048] FIG. 5B is a schematic diagram of a high-plane position in a Z-axis direction;

[0049] FIG. 6 is a schematic diagram of a node coding order;

[0050] FIG. 7A is a schematic diagram of plane identification information;

[0051] FIG. 7B is another schematic diagram of plane identification information;

[0052] FIG. 8 is a schematic diagram of IDCM coding;

[0053] FIG. 9A is a schematic diagram of three intersection points included in a sub-block;

[0054] FIG. 9B is a schematic diagram of a triangle patch set fitted with three intersection points;

[0055] FIG. 9C is a schematic diagram of upsampling of a triangle patch set;

[0056] FIG. 10 is a schematic diagram of a distance-based LOD construction process;

[0057] FIG. 11 is a schematic diagram of a fast lookup-based attribute inter-frame prediction;

[0058] FIG. 12 is a schematic diagram of a block-based neighbor lookup structure;

[0059] FIG. 13 is a schematic diagram of a RAHT transform structure;

[0060] FIG. 14 is a schematic diagram of a RAHT transform process along x, y, and z directions;

[0061] FIG. 15 is a schematic diagram of a RAHT forward transform process;

[0062] FIG. 16 is a schematic diagram of a RAHT inverse transform process;

[0063] FIG. 17 is a schematic diagram of an attribute encoding block;

[0064] FIG. 18 is a schematic diagram of a RAHT-based attribute prediction transform encoding principle;

[0065] FIG. 19 is a schematic diagram of coplanar and collinear spatial relationships;

[0066] FIG. 20 is a schematic diagram of a neighbor prediction relationship of attribute prediction;

[0067] FIG. 21 is a schematic diagram of a flow of a point cloud decoding method provided by an embodiment of the present application;

[0068] FIG. 22 is a schematic diagram of a decoding order defined by a decoding mode in an embodiment of the present application;

[0069] FIG. 23 is a schematic diagram of a flow of an attribute decoding method based on a decoding mode in an embodiment of the present application;

[0070] FIG. 24 is a schematic diagram of a RAHT decoding layer;

[0071] FIG. 25 is a schematic diagram of a division result of a decoding group;

[0072] FIG. 26 is a schematic diagram of a flow of a point cloud encoding method provided by an embodiment of the present application;

[0073] FIG. 27 is a schematic diagram of a composition structure of an encoder provided by an embodiment of the present application;

[0074] FIG. 28 is a schematic diagram of a specific hardware structure of an encoder according to an embodiment of the present application;

[0075] FIG. 29 is a schematic diagram of a composition structure of a decoder according to an embodiment of the present application;

[0076] FIG. 30 is a schematic diagram of a specific hardware structure of a decoder according to an embodiment of the present application;

[0077] FIG. 31 is a schematic diagram of a composition structure of a codec system according to an embodiment of the present application. DETAILED DESCRIPTION

[0078] In order to enable a person skilled in the art to more fully understand the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings, which are used only for reference and are not intended to limit the embodiments of the present application.

[0079] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terminology used in the specification is for the purpose of describing the embodiments of the present application only and is not intended to be limiting of the present application.

[0080] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0081] It should also be noted that the terms "first", "second", "third" used in the embodiments of the present application are only used to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first", "second", "third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0082] Point cloud is a three-dimensional representation of the surface of an object, which can be collected by photoelectric radar, laser radar, laser scanner, multi-view camera and other collection devices.

[0083] (1) Point cloud data form

[0084] Point cloud is a discrete point set in three-dimensional space, which contains geometric information for representing spatial position and attribute information for representing the appearance texture of point cloud. FIG. 1A shows a three-dimensional point cloud image and FIG. 1B shows a local enlarged view of the three-dimensional point cloud image. It can be seen that the surface of the point cloud is composed of densely distributed points.

[0085] A two-dimensional image has information expressed at each pixel point, and has the characteristic of regular distribution, so it does not need to record the position information. However, the points in the point cloud have randomness and irregularity in the three-dimensional space, so the position of each point in the space needs to be recorded in order to completely represent the object in the three-dimensional space. Similar to the two-dimensional image, each position has corresponding attribute information in the acquisition process, usually including color information and reflectivity information. The color information reflects the color of the object, usually represented by RGB; the reflectivity information reflects the surface material of the object, usually represented by reflectance. The point cloud data is usually composed of geometric information (x, y, z) representing the position in the three-dimensional space and attribute information such as color information (r, g, b) and reflectivity information (reflectance).

[0086] Therefore, the point cloud data usually includes the position information of the points and the attribute information of the points. The position information of the points can also be referred to as the geometric information of the points. For example, the geometric information of the points can be the three-dimensional coordinate information (x, y, z) of the points. The attribute information of the points can include color information and / or reflectivity, etc. For example, the reflectivity can be one-dimensional reflectivity information (r); the color information can be information on any color space, or the color information can also be three-dimensional color information such as RGB information. Here, R represents red (Red, R), G represents green (Green, G), and B represents blue (Blue, B). For another example, the color information can be luminance chrominance (YCbCr, YUV) information. Among them, Y represents brightness (Luma), Cb (U) represents blue color difference, and Cr (V) represents red color difference.

[0087] As shown in FIGS. 2A and 2B are a point cloud image and its corresponding data storage format. FIG. 2A provides six viewing dimensions of the point cloud image, and FIG. 2B is composed of a file header information part and a data part, wherein the header information includes data format, data representation type, total number of points of the point cloud, and content represented by the point cloud. For example, the point cloud file format in this example is “.ply”, represented by ASCII code, the total number of points is 207242, and each point has three-dimensional position information (x, y, z) and three-dimensional color information (r, g, b).

[0088] The point cloud can be divided into the following types according to the acquisition method:

[0089] Static point cloud: the object is static, and the device for acquiring the point cloud is also static;

[0090] Dynamic point cloud: the object is moving, but the device for acquiring the point cloud is static;

[0091] Dynamic acquisition of point cloud: the device for acquiring the point cloud is moving.

[0092] For example, point clouds can be classified into two categories according to their use:

[0093] Category one: machine perception point clouds, which can be used in autonomous navigation systems, real-time inspection systems, geographic information systems, vision sorting robots, disaster rescue robots, and the like;

[0094] Category two: human eye perception point clouds, which can be used in digital cultural heritage, free-viewpoint broadcasting, three-dimensional immersive communication, three-dimensional immersive interaction, and the like point cloud application scenarios.

[0095] (2) Point cloud compression background

[0096] Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes, and can provide strong realism on the premise of ensuring accuracy because point clouds can be obtained by directly sampling real objects. Therefore, point clouds are widely used in virtual reality games, computer-aided design, geographic information systems, autonomous navigation systems, digital cultural heritage, free-viewpoint broadcasting, three-dimensional immersive remote presentation, and three-dimensional reconstruction of biological tissues and organs.

[0097] Point clouds can be collected in the following ways: computer generation, 3D laser scanning, 3D photogrammetry, and the like. A computer can generate point clouds of virtual three-dimensional objects and scenes; 3D laser scanning can obtain point clouds of static real-world three-dimensional objects or scenes, and can obtain millions of points per second; 3D photogrammetry can obtain point clouds of dynamic real-world three-dimensional objects or scenes, and can obtain tens of millions of points per second. These point cloud collection technologies reduce the cost and time period of obtaining point cloud data, and improve the accuracy of the data, further promoting the practical application of point clouds. Due to the continuous industrialization of point cloud data acquisition methods, it is possible to obtain a large amount of point cloud data. However, with the growth of application demand, the processing of massive 3D point cloud data has encountered a bottleneck in storage space and transmission bandwidth.

[0098] Taking a point cloud video with a frame rate of 30 fps (frames per second) as an example, the number of points of each frame of point cloud is 700,000, each point contains coordinate information xyz (float) and color information RGB (uchar), and the data volume of a 10s point cloud video is about 0.7 million · (4 Byte · 3 + 1 Byte · 3) · 30 fps · 10s = 3.15 GB. Correspondingly, the data volume of a 10s 1280*720 two-dimensional video with a YUV sampling format of 4:2:0 and a frame rate of 30 fps is about 1280*720*12bit*30frames*10s = 0.39 GB, and the data volume of a 10s two-view 3D video is about 0.39*2 = 0.78 GB. As can be seen, the data volume of the point cloud video is much larger than that of the two-dimensional video and the three-dimensional video of the same length. Therefore, in order to better realize data management, save server storage space, and reduce transmission bandwidth and transmission time between the server and the client, point cloud compression has attracted widespread attention in the industry.

[0099] That is, since the point cloud is a collection of a large number of points, storing the point cloud not only consumes a large amount of memory, but is also not conducive to transmission, and there is no such large bandwidth to support the transmission of the point cloud directly on the network layer without compression, therefore, the point cloud needs to be compressed.

[0100] At present, the point cloud encoding framework that can compress the point cloud can be a geometry-based point cloud compression (G-PCC) coding framework or a video-based point cloud compression (V-PCC) coding framework provided by the Moving Picture Experts Group (MPEG), or an AVS-PCC coding framework provided by the AVS. The G-PCC coding framework can be used for compression of the first type of static point cloud and the third type of dynamically acquired point cloud, and can be based on the Test Model Compression 13 (TMC13). The V-PCC coding framework can be used for compression of the second type of dynamic point cloud, and can be based on the Test Model Compression 2 (TMC2). Therefore, the G-PCC coding framework is also called the point cloud codec TMC13, and the V-PCC coding framework is also called the point cloud codec TMC2.

[0101] The embodiment of the present application provides a network architecture of a point cloud coding system comprising a decoding method and an encoding method. FIG. 3 is a schematic diagram of a network architecture of a point cloud coding provided by the embodiment of the present application. As shown in FIG. 3, the network architecture comprises one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01. The electronic devices in the implementation process can be various types of devices with point cloud coding functions, for example, the electronic devices can comprise a mobile phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital telephone, a video telephone, a television, a sensing device, a server and the like, and the embodiment of the present application is not limited thereto. The decoder or the encoder in the embodiment of the present application can be the above-mentioned electronic devices.

[0102] In the embodiment of the present application, the electronic device with the point cloud coding function generally comprises a point cloud encoder (i.e., an encoder) and a point cloud decoder (i.e., a decoder).

[0103] (3) G-PCC coding framework

[0104] In the point cloud G-PCC coding framework, for the point cloud data to be encoded, the point cloud data is first divided into multiple slices through slice division. In each slice, the geometry information of the point cloud and the attribute information corresponding to each point are encoded separately.

[0105] FIG. 4A shows a schematic diagram of a G-PCC encoder. As shown in FIG. 4A, in the geometry encoding process, the geometry information is first converted in coordinates, so that all the point clouds are contained in a bounding box, and then quantized, which mainly plays a role of scaling. Due to the quantization rounding, the geometry information of a part of the point clouds is the same, and then it is determined based on the parameters whether to remove the duplicate points. This process of quantization and removal of duplicate points is also called voxelization process. Then, the bounding box is divided by octree or a prediction tree is constructed. In this process, the points in the divided leaf nodes are arithmetically encoded to generate a binary geometry bitstream, or the vertices generated by the division are arithmetically encoded (surface fitting based on the vertices) to generate a binary geometry bitstream. In the attribute encoding process, after the geometry encoding is completed and the geometry information is reconstructed, color conversion is needed to convert the color information (i.e., attribute information) from the RGB color space to the YUV color space. Then, the point cloud is recolored using the reconstructed geometry information, so that the unencoded attribute information corresponds to the reconstructed geometry information. Attribute encoding is mainly for color information. In the color information encoding process, there are mainly two transformation methods, one is distance-based lifting transformation depending on level of detail (LOD) division, and the other is direct region adaptive hierarchal transform (RAHT). Both of these two methods convert the color information from the spatial domain to the frequency domain, obtain high-frequency coefficients and low-frequency coefficients through transformation, and finally quantize the coefficients, and then arithmetically encode the quantized coefficients to generate a binary attribute bitstream.

[0106] FIG. 4B shows a schematic diagram of a G-PCC decoder. As shown in FIG. 4B, for the obtained binary bitstream, the geometry bitstream and the attribute bitstream in the binary bitstream are first independently decoded. In the decoding of the geometry bitstream, the geometry information of the point cloud is obtained through arithmetic decoding-reconstructing octree / reconstructing prediction tree-reconstructing geometry-coordinate inverse conversion; in the decoding of the attribute bitstream, the attribute information of the point cloud is obtained through arithmetic decoding-inverse quantization-LOD division / RAHT-color inverse conversion, and the point cloud data to be encoded (i.e., the output point cloud) is restored based on the geometry information and the attribute information.

[0107] It should be noted that, as shown in FIG. 4A or FIG. 4B, the geometry encoding and decoding of the current G-PCC can be divided into octree-based geometry encoding and decoding (identified by a dashed line box) and prediction tree-based geometry encoding and decoding (identified by a dotted line box).

[0108] Geometry information encoding:

[0109] For Octree geometry encoding (OctGeomEnc), first, the geometry information is converted to a coordinate system so that all the points are contained in a bounding box. Then, quantization is performed, which mainly serves as a scaling function. Due to the quantization, the geometry information of some points is the same, and whether to remove the duplicate points is determined according to the parameters. The process of quantization and removal of duplicate points is also called voxelization. Next, the bounding box is divided into a multi-tree (such as octree, quadtree, binary tree, etc.) in the order of breadth-first traversal, and the occupancy code of each node is encoded. At the July 2019 MPEG meeting in Gothenburg, Sweden, an implicit geometry division method was proposed. First, the bounding box of the point cloud is calculated Assume d x > d y > d z The bounding box corresponds to a cuboid. When dividing the geometry, first, a binary tree is divided based on the x-axis to obtain two child nodes; until the condition d x = d y > d z is met, a quadtree is divided based on the x and y axes to obtain four child nodes; when the condition d x = d y = d z is met, an octree is divided until the leaf node obtained by the division is a 1x1x1 unit cube, and the points in the leaf node are encoded to generate a binary code stream. In the division process based on the binary tree / quadtree / octree, two parameters K and M are introduced. Parameter K indicates the maximum number of binary tree / quadtree divisions before octree division; parameter M is used to indicate the minimum block length corresponding to the binary tree / quadtree division, which is 2 M . At the same time, K and M must satisfy the condition: assume d max = max(d x , d y , d z ), d min = min(d x , d y , d z ), parameter K satisfies: K ≥ d max -d min ; parameter M satisfies: M ≥ d minThe parameters K and M satisfy the above conditions because, in the current G-PCC, the priority of the partitioning mode is binary tree, quad tree, and octree, and when the node block size does not satisfy the condition of the binary tree / quad tree, the octree partitioning is performed on the node until the leaf node of the minimum unit 1*1*1 is obtained. The octree-based geometry coding mode can effectively encode the geometry information of the point cloud by using the correlation between the adjacent points in the space, and for some relatively flat nodes or nodes with a planar feature, the coding efficiency of the geometry information of the point cloud can be further improved by using the planar coding.

[0110] Exemplarily, FIGS. 5A and 5B provide a schematic diagram of a planar position. FIG. 5A shows a low planar position schematic diagram in the Z-axis direction, and FIG. 5B shows a high planar position schematic diagram in the Z-axis direction. As shown in FIG. 5A, (a), (a0), (a1), (a2), and (a3) all belong to the low planar position in the Z-axis direction. For example, (a), the four occupied sub-nodes in the current node are located in the low planar position of the current node in the Z-axis direction, and thus the current node belongs to a Z plane and is a low planar in the Z-axis direction. Similarly, as shown in FIG. 5B, (b), (b0), (b1), (b2), and (b3) all belong to the high planar position in the Z-axis direction. For example, (b), the four occupied sub-nodes in the current node are located in the high planar position of the current node in the Z-axis direction, and thus the current node belongs to a Z plane and is a high planar in the Z-axis direction.

[0111] Further, taking (a) in FIG. 5A as an example, the efficiency of the octree coding and the planar coding is compared, and FIG. 6 provides a schematic diagram of a node coding order, that is, the node coding is performed in the order of 0, 1, 2, 3, 4, 5, 6, and 7 shown in FIG. 6. Here, if the octree coding mode is used for (a) in FIG. 5A, the occupancy information of the current node is represented as 11001100. If the planar coding mode is used, first, an identifier is coded to represent that the current node is a planar in the Z-axis direction and the planar position of the current node needs to be represented; second, only the occupancy information of the low planar node in the Z-axis direction needs to be coded (that is, the occupancy information of the four sub-nodes 0246). Therefore, 6 bits are needed to code the current node based on the planar coding mode, which can reduce 2 bits compared with the original octree coding mode. Based on this analysis, the planar coding mode has a relatively obvious coding efficiency compared with the octree coding mode. For an occupied node, if the planar coding mode is used for coding in a certain dimension, first, the planar identifier and the planar position information of the current node in the dimension are represented, and then the occupancy information of the current node is coded based on the planar information of the current node.

[0112] For example, Fig. 7A shows a schematic diagram of planar identification information. As shown in Fig. 7A, here the Z-axis direction is a low plane; correspondingly, the value of the planar identification information is true or 1, i.e. planarMode Z = true; the planar position information is low plane, i.e. PlanePosition Z = low. Fig. 7B shows another schematic diagram of planar identification information. As shown in Fig. 7B, here the Z-axis direction is not a plane; correspondingly, the value of the planar identification information is false or 0, i.e. planarMode Z

[0113] It should be noted that PlaneMode i (i = 0, 1, 2): 0 represents that the current node is not a plane in the i-axis direction. If the current node is a plane in the i-axis direction (i.e. PlaneMode i (i = 0, 1, 2): 1), PlanePosition i : 0 represents that the current node is a plane in the i-axis direction, and the plane position is a low plane; PlanePosition i : 1 represents that the current node is a high plane in the i-axis direction. Wherein, i represents the coordinate dimension, which can be the X-axis direction, the Y-axis direction or the Z-axis direction, so i = 0, 1, 2.

[0114] The octree-based geometry information coding mode has high compression efficiency for points with strong correlation in space, but for points in isolated positions in the geometry space, using the DCM coding mode can improve the compression efficiency while also reducing the coding complexity to a certain extent. For all nodes in the octree, the use of DCM is not represented by flag information, but is inferred from the parent node and neighbor information of the current node. There are three ways to determine whether the current node has DCM coding qualifications, as follows:

[0115] (1) The current node has no brother and sister child nodes (i.e. the parent node of the current node has only one child node), and the parent node of the current node has only two occupied child nodes (i.e. the current node has at most one neighbor node).

[0116] (2) The parent node of the current node has only one occupied child node of the current node, and the six neighbor nodes sharing a face with the current node all belong to empty nodes.

[0117] (3) The number of brother and sister nodes of the current node is greater than 1. ​

[0118] Exemplarily, FIG. 8 provides an IDCM encoding schematic diagram. If the current node does not have DCM encoding qualification, it will be octree divided, and if it has DCM encoding qualification, it will further judge the number of points contained in the node: when the number of points is less than a threshold 2 (i.e. indicating that the current node is a real isolated point), the node is DCM encoded; otherwise, the octree division will be continued. When the DCM encoding mode is applied, firstly, an identification bit (i.e. IDCM_flag) needs to be encoded to indicate whether the current node is a real isolated point. When IDCM_flag is true, the current node adopts the DCM encoding mode; otherwise, the current node still adopts the octree encoding mode. When the current node meets the DCM encoding condition, the DCM encoding mode of the current node needs to be encoded. There are two DCM encoding modes at present, which are: 1) only one point exists (or multiple points exist, but belong to repeated points); 2) contains two points. Secondly, the geometric information of each point needs to be encoded. Assuming that the edge length of the node is 2d, when encoding each component of the geometric coordinates of the node, d bits are needed, and the bit information is directly encoded into the code stream. It should be noted that when encoding the laser radar point cloud, by using the laser radar acquisition parameters to predict the encoding of the three-dimensional coordinate information, the encoding efficiency of the geometric information can be further improved.

[0119] It should be noted that when the node is divided into a leaf node, in the case of lossless geometric encoding, the number of repeated points in the leaf node needs to be encoded. Finally, the occupancy information of all nodes is encoded to generate a binary code stream.

[0120] Based on the octree-based geometry decoding mode, the decoding end will use the reconstructed geometry information to determine whether the current node is subjected to plane decoding or IDCM decoding in the order of breadth-first traversal before decoding the occupancy information of each node. If the current node meets the conditions of plane decoding, the plane identifier and plane position information of the current node will be decoded, and the occupancy information of the current node will be decoded based on the plane identifier and the plane position information. If the current node meets the conditions of IDCM decoding, IDCM_flag needs to be further parsed to determine whether the current node is a real IDCM node. If IDCM_flag: 1, it indicates that the current node is a real IDCM node, and the DCM decoding mode of the current node will be parsed, the number of points in the current DCM node can be obtained, and finally the geometry information of each point is decoded. For the node that does not meet the plane decoding mode and the DCM decoding mode, the occupancy information of the current node will be decoded. In this way, the occupancy code of each node is parsed continuously, and the nodes are continuously divided in sequence until the unit cube of 1x1x1 is divided, the number of points contained in each leaf node is parsed, and finally the geometry reconstruction point cloud information is recovered.

[0121] In the geometry information coding framework based on triangle soup (trisoup), geometry division is also performed first, but unlike the geometry information coding based on binary tree / quaternary tree / octree, this method does not need to divide the point cloud to the unit cube with an edge length of 1x1x1, but stops dividing when the edge length of the block is W. Based on the surface formed by the distribution of the point cloud in each block, at most twelve intersection points (vertices) generated by the surface and the twelve edges of the block are obtained. The vertex coordinates of each block are sequentially encoded to generate a binary code stream.

[0122] Point cloud geometry information reconstruction based on trisoup

[0123] When reconstructing the point cloud geometry information at the decoding end, the vertex coordinates are first decoded to complete the triangle patch reconstruction, and the process is shown in FIGS. 9A, 9B and 9C. As shown in FIG. 9A, there are three intersection points (v1, v2, v3) in the block, and the triangle patch set formed by the three intersection points in a certain order is called triangle soup, i.e., trisoup, as shown in FIG. 9B. Then, sampling is performed on the triangle patch set, and the obtained sampling points are used as the reconstructed point cloud in the block, as shown in FIG. 9C.

[0124] The geometry coding based on the predictive geometry coding (PredGeomTree) firstly sorts the input point cloud, and the currently used sorting methods include unordered, Morton order, azimuth angle order and radial distance order. At the encoding end, the predictive tree structure is established by using two different ways, including KD-Tree (high latency and slow mode) and low latency and fast mode (using laser radar calibration information). When using the laser radar calibration information, each point is divided into different Lasers, and the predictive tree structure is established according to different Lasers. Next, based on the structure of the predictive tree, each node in the predictive tree is traversed, the geometry position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometry prediction residual is quantized by using the quantization parameter. Finally, by continuously iterating, the prediction residual of the position information of the node of the predictive tree, the predictive tree structure and the quantization parameter are encoded to generate a binary code stream.

[0125] The geometry decoding based on the predictive tree, at the decoding end, the predictive tree structure is reconstructed by continuously analyzing the code stream, then the geometry position prediction residual information and the quantization parameter of each prediction node are obtained by analysis, the prediction residual is dequantized, and the reconstructed geometry position information of each node is recovered, and finally the geometry reconstruction at the decoding end is completed.

[0126] After the geometry coding is completed, the geometry information needs to be reconstructed. At present, the attribute coding is mainly for color information. First, the color information is converted from the RGB color space to the YUV color space. Then, the point cloud is recolored using the reconstructed geometry information, so that the uncoded attribute information corresponds to the reconstructed geometry information. In color information coding, there are mainly two transformation methods, one is distance-based lifting transformation depending on LOD division, and the other is direct RAHT transformation. Both of these two methods will convert the color information from the spatial domain to the frequency domain, obtain the high frequency coefficient and the low frequency coefficient through transformation, and finally quantize and encode the coefficients to generate a binary code stream, as shown in FIG. 4A and FIG. 4B.

[0127] Further, when using geometry information to predict attribute information, the nearest neighbor search can be performed by using the Morton code. The Morton code corresponding to each point in the point cloud can be obtained from the geometry coordinates of the point. The specific method of calculating the Morton code is described as follows. For a three-dimensional coordinate represented by d-bit binary numbers for each component, the three components can be represented as:

[0128] wherein, are the highest bits of x, y and z respectively are the lowest bits of x, y and z respectively The corresponding binary value. The Morton code M arranges x, y, and z in a crosswise manner starting from the highest bit to the lowest bit. The calculation formula of M is as follows: To the lowest bit, the calculation formula of M is as follows:

[0129] Where Are the highest bit of M To the lowest bit Values. After obtaining the Morton code M of each point in the point cloud, the points in the point cloud are arranged in ascending order of the Morton code, and the weight value w of each point is set to 1.

[0130] Point cloud attribute coding

[0131] As shown in FIG. 4A or FIG. 4B, the current G-PCC coding framework includes three attribute coding methods: Predicting Transform (PT), Lifting Transform (LT), and Region Adaptive Hierarchical Transform (RAHT). Among them, the first two perform predictive coding on the point cloud based on the generation order of LODs, while RAHT adaptively transforms the attribute information from bottom to top according to the construction levels of the octree. The following will specifically introduce these three point cloud attribute coding methods.

[0132] Predictive coding of point cloud attribute information

[0133] Currently, the attribute prediction module of G-PCC adopts a nearest neighbor attribute prediction coding scheme based on a hierarchical (Level-of-details, LoDs) structure. The construction methods of LODs include a distance-based LOD construction scheme, a fixed sampling rate-based LOD construction scheme, and an octree-based LOD construction scheme, etc. In the distance threshold-based LOD construction scheme, before constructing LOD, the point cloud is first sorted by Morton to ensure strong attribute correlation between adjacent points. FIG. 10 is a schematic diagram of a distance-based LOD construction process. As shown in FIG. 10, according to L Manhattan (Manhattan) distances (dl) preset by the user in advance, l = 0, 1,..., L - 1; the point cloud is divided into L different point cloud detail levels (Rl), l = 0, 1,..., L - 1, where (dl)l = 0, 1,..., L - 1 satisfies dl < dl-1. The construction process of LOD is as follows:

[0134] (1) Firstly, all points in the point cloud are marked as unvisited, and a set V is established to store the set of points that have been visited; (2) for each iteration l, by traversing the points in the point cloud, if the current point has been visited, the point is ignored, otherwise the minimum distance D of the current point to the point set V is calculated, if D < dl, the point is ignored; otherwise, the current point is marked as visited and the current point is added to the refinement layer Rl and the point set V; (3) the points in the detail level LODl are composed of the points in the refinement layers R0, R1, R2…Rl; (4) the above steps are repeatedly executed until all points are marked as visited.

[0135] On the basis of the structure of LOD, the attribute value of each point is predicted by linearly weighting the reconstructed attribute values of the points in the same layer or the next higher layer of LOD, wherein the maximum number of reference prediction neighbors is specified by an encoder high-level syntax element. For each attribute of a point, a rate-distortion optimization algorithm is used at the encoding end to select the attribute of the N nearest neighbor points found to be used for weighted prediction or select the attribute of a single nearest neighbor point for prediction, and finally the selected prediction mode and prediction residual are encoded.

[0136] wherein N represents the number of prediction points in the nearest neighbor set of point i, Pi represents the sum of the N nearest neighbor points of point i, Dm represents the spatial geometric distance of the nearest neighbor point m to the current point i, Attrm represents the attribute value of the nearest neighbor point m after reconstruction, Attr i i represents the predicted value of the attribute of the current point i, and the point number N is a preset value.

[0137] Inter-frame nearest neighbor search

[0138] FIG. 11 is a schematic diagram of attribute inter-frame prediction based on fast search. As shown in FIG. 11, when performing attribute inter-frame prediction, first, the Morton code corresponding to the current point is obtained using the geometric coordinates of the current point to be encoded, second, the first reference point (j) greater than the Morton code of the current point is found in the reference image based on the Morton code of the current point, and third, nearest neighbor search is performed within the range of [j-searchRange, j+searchRange].

[0139] Currently, when performing nearest neighbor search for intra-frame and inter-frame, neighborhood search is performed based on blocks. For details, refer to FIG. 12. As shown in FIG. 12, when performing neighborhood search on the current point (Morton code index i), first, the points in the reference image are divided into N (N=3) layers according to the Morton code, and the specific division algorithm is as follows:

[0140] • The first layer: the points in the hypothetical reference image are divided into numPoints, and first, the points in the reference image are divided into a block every M (M=2 5 =32) points;

[0141] • Second layer: On the basis of the first layer, the blocks of the first layer are divided into a block every M (M = 2 5 = 32) blocks in the order of the Morden code;

[0142] • Third layer: On the basis of the second layer, the blocks of the second layer are divided into a block every M (M = 2 5 = 32) blocks in the order of the Morden code;

[0143] Finally, the prediction structure as shown in FIG. 12 is obtained.

[0144] In the attribute prediction based on the prediction structure as shown in FIG. 12, assuming that the Morden code index of the current point to be coded is i, first, the point whose index is j in the reference image and which is greater than or equal to the Morden code of the current point is obtained. Secondly, the block index of the reference point is calculated based on j, and the specific calculation method is as follows:

[0145] • First layer: BucketSize_0 = 2 5 = 32;

[0146] • Second layer: BucketSize_1 = 2 5 = 32 x BucketSize_0 = 1024;

[0147] • Third layer: BucketSize_2 = 2 5 = 32 x BucketSize_1 = 32768.

[0148] Assuming that the reference range in the prediction frame of the current point is [j-searchRange, j+searchRange], the starting index of the third layer is calculated by using j-searchRange, and the terminal index of the third layer is calculated by using j+searchRange; secondly, first, it is judged in the blocks of the third layer whether some blocks of the second layer need to perform the nearest neighbor search, and then to the second layer, it is judged for each block in the first layer whether the search needs to be performed, if some blocks of the first layer need to perform the nearest neighbor search, then the points in some blocks of the first layer are judged point by point to update the nearest neighbor.

[0149] Next, the algorithm for calculating the block based on the index is introduced, assuming that the Morden code index corresponding to the current point is index, then the index of the third layer block corresponding thereto is: idx_2 = index / BucketSize_2 (4)

[0150] After obtaining the third layer's block index idx_2, the start index and end index of the current block in the second layer's corresponding block can be obtained using idx_2: startIdx1 = idx_2 x BucketSize_1 (5) endIdx = idx_2 x BucketSize_1 + BucketSize_1 - 1 (6)

[0151] Similarly, the first layer's block index is obtained based on the same algorithm based on the second layer's block index.

[0152] When performing the nearest neighbor search based on the block, it is first determined whether the current block needs to perform the nearest neighbor search, that is, the nearest neighbor search of the screening block. Each spatial block can be obtained by two variables minPos and maxPos, minPos representing the minimum value of the block, and maxPos representing the maximum value of the block.

[0153] Suppose the distance of the farthest point in the N nearest neighbors of the current point is Dist, the coordinates of the to-be-encoded point are (x, y, z), and the current block is represented as (minPos, maxPos), wherein minPos is the minimum value of the bounding box in three dimensions, and maxPos is the maximum value of the bounding box in three dimensions. The distance D between the current point and the bounding box is calculated as follows: int dx = int(std::max(std::max(minPos[0] - point[0], 0), point[0] - maxPos[0])); int dy = int(std::max(std::max(minPos[1] - point[1], 0), point[1] - maxPos[1])); int dz = int(std::max(std::max(minPos[2] - point[2], 0), point[2] - maxPos[2])); D = dx + dy + dz;

[0154] When D is less than or equal to Dist, the points in the current block are traversed.

[0155] Region adaptive hierarchical transform

[0156] Region Adaptive Hierarchical Transform (RAHT) is a kind of Haar wavelet transform, which can transform the point cloud attribute information from spatial domain to frequency domain, further reducing the correlation between the point cloud attributes. The main idea is to transform the nodes in each layer from X, Y, Z three dimensions according to the octree structure in a bottom-up manner (as shown in FIG. 14), and iterate until the root node of the octree. As shown in FIG. 13, the basic idea is to perform wavelet transform based on the hierarchical structure of the octree, associate the attribute information with the octree nodes, recursively transform the attributes of the occupied nodes in the same parent node in a bottom-up manner, transform the nodes in each layer from X, Y, Z three dimensions, and continue to transform until the root node of the octree is reached. During the hierarchical transform, the low pass / low frequency (DC) coefficients obtained after the transform of the nodes in the same layer are passed to the nodes in the next layer for further transform, and all the high pass / high frequency (AC) coefficients can be encoded by an arithmetic encoder.

[0157] During the transform process, the DC coefficients (direct current component) after the transform of the nodes in the same layer will be passed to the previous layer for further transform, and the AC coefficients (alternating current component) after the transform of each layer will be quantized and encoded. The main transform process will be introduced below.

[0158] FIG. 15 is a schematic diagram of the process of RAHT forward transform, and FIG. 16 is a schematic diagram of the process of RAHT inverse transform. For the transform and inverse transform process corresponding to RAHT, it is assumed that g′ L,2x,y,z and g′L,2x+1,y,z are two attribute DC coefficients of the nodes that are adjacent to each other in the L layer. After linear transform, the information of the L-1 layer is the AC coefficient f′ L-1,x,y,z and the DC coefficient g′ L-1,x,y,z ; then, f′ L-1,x,y,z will no longer be transformed and will be directly quantized and encoded, and g′ L-1,x,y,z will continue to find neighbors for transform, and if no neighbor is found, it will be directly passed to the L-2 layer, i.e. the RAHT transform is only valid for the nodes with neighbors, and the nodes without neighbors will be directly passed to the previous layer. In the above transform process, the weights corresponding to g′ L,2x,y,z and g′L,2x+2,y,z are w′ L,2x,y,z and w′L,2x+1,y,z (abbreviated as w′0 and w′1), and the weight of g′ L-1,x,y,z is w′ L-1,x,y,z , and the general transform formula is:

[0159] where T w0,w1 is the transform matrix:

[0160] The transformation matrix is updated adaptively according to the corresponding weight of each point. The above process is iteratively updated according to the division structure of the octree until the root node of the octree.

[0161] Region adaptive hierarchical intra prediction transform coding

[0162] Region adaptive hierarchical prediction transform coding is based on the prediction of RAHT transform coding. As shown in FIG. 13, the RAHT attribute transform is based on the order of the octree hierarchy, and the transformation is continuously performed from the voxel level until the root node is obtained, thereby completing the hierarchical transform coding of the entire attribute. In the prediction transform coding, the attribute prediction transform coding is also based on the order of the octree hierarchy, but is continuously transformed from the root node until the voxel level. In the process of each RAHT attribute transform, the attribute prediction transform coding is based on the 2x2x2 block. FIG. 17 is a schematic diagram of an attribute coding block, as shown in FIG. 17, it can be seen that the dark filled block is the current block to be coded, and the light filled block is some neighbor block coplanar and collinear with the current block to be coded.

[0163] FIG. 18 is a schematic diagram of an attribute prediction transform coding based on RAHT, first, the attribute A of the current block can be obtained from the attribute of the point contained in the current block node As shown in FIG. 18(a), specifically as follows: A node =∑ p∈node attribute(p) (9)

[0164] Secondly, the number of points in the current block is normalized to obtain the mean value a of the attribute of the current block node As shown in FIG. 18(b), the specific normalization process is as follows: w node =∑ p∈node 1=#{p∈node} (10) a node =A node / w node (11)

[0165] The attribute transform coding is performed using the mean value of the attribute of the current block.

[0166] Further, the attribute of each sub-block is linearly weighted predicted by using the spatial geometric distance between the neighbor blocks of the current block and each sub-block of the current block, to obtain the predicted attribute of the sub-block of the current block, also called predicted attribute block, as shown in Fig. 18(c). The predicted attribute of the current block is inverse normalized to obtain the final predicted attribute, as shown in Fig. 18(e). Fig. 18(d) is the original attribute of the current block, also called attribute original block. Finally, the predicted attribute and the original attribute of the sub-block are respectively subjected to attribute transformation to obtain the original AC coefficient and the predicted AC coefficient, and the AC coefficient residual is obtained according to the original AC coefficient and the predicted AC coefficient, and Fig. 18(f) is the AC coefficient residual of the current block, and the AC coefficient parameter is encoded.

[0167] Fig. 19 is a schematic diagram of a coplanar and collinear spatial relationship, and Fig. 20 is a schematic diagram of a neighbor prediction relationship of attribute prediction. Firstly, 19 neighbor blocks of the current block are determined as shown in Fig. 19, and secondly, the attribute of each sub-block is linearly weighted predicted by using the spatial geometric distance between the neighbor blocks and each sub-block a up

[0168] Region adaptive hierarchical inter prediction transform coding

[0169] In the existing G-PCC attribute inter prediction coding, if the inter prediction coding is started, the RAHT attribute transform coding structure is constructed based on the geometric information of the current to-be-coded node, that is, the node merging is continuously performed at the voxel level until the root node of the entire RAHT transform tree is obtained, so that the transform coding hierarchical structure of the attribute is obtained. According to the RAHT transform structure, each node is divided from the root node to obtain N sub-nodes (N is less than or equal to 8) of each node. In the inter prediction coding scheme, the attributes of the N sub-nodes are independently and orthogonally transformed by using the RAHT transform to obtain DC and AC coefficients. The AC coefficients of the N sub-nodes are inter predicted in attributes in the following manner:

[0170] The inter prediction node of the current node is valid: that is, the homochromatic node exists, and the attribute of the prediction node is directly taken as the attribute prediction value of the current to-be-coded node.

[0171] The current node can find a node with the same position as the current node in the cache of the reference image: that is, the homochromatic node exists, and the AC coefficients of the M sub-nodes contained in the homochromatic node are directly taken as the AC coefficient attribute prediction value of the N sub-nodes of the current node.

[0172] 1. If the AC coefficient of the prediction node is not zero: the AC coefficient of the prediction node is directly taken as the prediction value;​

[0173] 2. If the AC coefficients of the prediction node are zero, the AC coefficients of the corresponding child node are used as the prediction value of the intra prediction

[0174] The inter prediction node of the current node is invalid: that is, the sibling node does not exist, and the attribute prediction value of the intra adjacent node is used as the attribute prediction value of the node to be encoded

[0175] And on this basis, the existing RAHT inter coding will select the best RAHT coding mode for each layer: intra prediction coding or inter prediction coding, when the cost of the intra prediction coding mode is less than the cost of the inter prediction coding mode, the current layer is subjected to RAHT intra prediction, otherwise, RAHT inter prediction is performed.

[0176] General test conditions of GPCC

[0177] 1) There are four test conditions in total:

[0178] Condition 1: limited loss of geometric position and loss of attribute

[0179] Condition 2: lossless geometric position and loss of attribute

[0180] Condition 3: lossless geometric position and limited loss of attribute

[0181] Condition 4: lossless geometric position and lossless attribute

[0182] 2) The general test sequence includes Cat1A, Cat1B, Cat3-fused, and Cat3-frame, wherein Cat2-frame point cloud only contains reflectance attribute information, Cat1A and Cat1B point cloud only contains color attribute information, and Cat3-fused point cloud contains color and reflectance attribute information.

[0183] 3) Technical route: there are two kinds, which are distinguished by the algorithm used for geometric compression.

[0184] Technical route 1: octree coding branch:

[0185] At the encoding end, the bounding box is divided into sub-cubes in sequence, and the non-empty (containing points in the point cloud) sub-cubes are further divided until the leaf nodes obtained by the division are 1x1x1 unit cubes, at which time the division is stopped. In the case of lossless geometric coding, the number of points contained in the leaf nodes needs to be encoded, and finally the geometric octree coding is completed to generate a binary code stream.

[0186] At the decoding end, the decoding end parses the placeholder code of each node in the breadth-first traversal order, and continuously divides the nodes one by one until the 1x1x1 unit cube is obtained, and stops dividing. In the case of geometric lossless decoding, the number of points contained in each leaf node needs to be parsed, and finally the geometric reconstruction point cloud information is recovered.

[0187] Technical route 2: prediction tree coding branch:

[0188] At the encoding end, two different ways are used to establish the prediction tree structure, including KD-Tree (high latency slow mode) and using laser radar calibration information to divide each point into different lasers, and establishing a prediction structure according to different lasers (low latency fast mode). Next, based on the structure of the prediction tree, each node in the prediction tree is traversed, the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter. Finally, by continuously iterating, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameter are encoded to generate a binary code stream;

[0189] At the decoding end, the decoding end reconstructs the prediction tree structure by continuously parsing the code stream, and then parses the geometric position prediction residual information and the quantization parameter of each prediction node, and dequantizes the prediction residual to recover the reconstructed geometric position information of each node. Finally, the geometric reconstruction at the decoding end is completed.

[0190] In the existing G-PCC attribute RAHT inter-frame prediction coding scheme, when the attribute inter-frame prediction coding is enabled (attrInterPredictionEnabled is True), the attribute information of the current node to be coded is predicted and coded by using the reconstructed geometry information and attribute information of the reference image homologous node. Specifically, the spatial position of the node to be coded is used to obtain the homologous node (i.e. the spatial position is exactly the same) in the reference image. If the homologous node exists, the reconstructed attribute of the homologous node is used to predict and code the attribute of the representative node, otherwise, the attribute information of the current node is predicted and coded, and the inter-frame prediction coding of the attribute information is completed based on this coding scheme.

[0191] In the existing RAHT attribute inter-frame prediction coding, if there are two reference units for the current frame to be coded, the attribute of the current slice can be inter-frame prediction coded by means of the reconstructed attribute information of the two reference units. However, in the existing coding scheme, whether to start the unidirectional inter-frame prediction or the bidirectional inter-frame prediction coding of the current slice is adaptively determined by using the geometric information of the reference unit and the unit to be coded at the encoding end. This coding scheme can adaptively determine whether to perform bidirectional prediction for each slice unit to some extent. However, this attribute coding scheme does not consider the attribute distribution characteristics of the coding nodes in the slice, thereby resulting in a low attribute RAHT coding efficiency of the point cloud.

[0192] Based on this, the embodiment of the present application provides a point cloud coding method. At the encoding end, the node attribute distribution characteristics of the unit to be coded are fully considered, the optimal coding mode is adaptively selected for the unit to be coded, a first syntax element indicating the optimal coding mode is written into a bitstream, at the decoding end, the first syntax element is determined from the bitstream, thereby determining the decoding mode of the unit to be decoded, and at the encoding and decoding end, the point cloud attribute reconstruction is performed by using the decoding and coding mode of the unit, thereby improving the point cloud attribute decoding efficiency.

[0193] To facilitate understanding of the technical solutions of the embodiments of the present application, the technical solutions of the present application are described in detail below through specific embodiments. The above related technologies can be combined with the technical solutions of the embodiments of the present application in any manner, and all of them belong to the protection scope of the embodiments of the present application. The embodiments of the present application include at least part of the following contents. The present application provides a point cloud coding method, and more specifically provides a point cloud attribute coding method. In some embodiments, the method can include: a layer-based point cloud attribute coding method, a coding group-based point cloud attribute coding method, and a node-based point cloud attribute coding method.

[0194] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0195] In an embodiment of the present application, referring to FIG. 21, a flowchart of a point cloud decoding method provided by the embodiment of the present application is shown. As shown in FIG. 21, the method can include:

[0196] S101: parsing a bitstream to determine the value of the first syntax element of the unit to be decoded;

[0197] In some embodiments, the unit to be decoded can include at least one of the following: a layer to be decoded, a decoding group, and a node to be decoded.

[0198] In some embodiments, the to-be-decoded unit can also include at least one of a to-be-decoded sequence, a to-be-decoded frame, and a to-be-decoded slice.

[0199] The first syntax element is used to indicate the decoding manner of the to-be-decoded unit. The first syntax element can include at least one of a sequence-level syntax element, a frame-level syntax element, a slice-level syntax element, a layer-level syntax element, a decoding group-level syntax element, and a node-level syntax element.

[0200] In some embodiments, the to-be-decoded unit is a to-be-decoded layer, and the first syntax element can be added to an attribute header (ABH) parameter set. The to-be-decoded unit is a to-be-decoded group, and the first syntax element can be added to the ABH. The to-be-decoded unit is a to-be-decoded node, and the first syntax element can be added to the ABH. The decoding end obtains the value of the first syntax element through the ABH.

[0201] In some embodiments, the value of the first syntax element of the current block is determined when the parsing condition of the first syntax element of the current block is met. The decoder determines whether to decode the first syntax element when decoding the current block by judging whether the parsing condition is met, and the optimization of the parsing condition of the first syntax element can reduce unnecessary transmission of the first syntax element and effectively control the code stream overhead.

[0202] In some embodiments, the decoding manner includes a decoding order of one or more decoding modes, and if the decoding mode includes an attribute prediction mode, the method can further include: determining the value of the first syntax element of the to-be-decoded unit when it is determined that the to-be-decoded unit is allowed to enable attribute prediction. That is, the parsing condition of the first syntax element at least includes allowing the to-be-decoded unit to enable attribute prediction, and only when attribute prediction is allowed to be enabled, the to-be-decoded unit can use the decoding manner indicated by the first syntax element, and the first syntax element is parsed.

[0203] In some embodiments, the method further includes: determining the value of the second syntax element; and determining that the to-be-decoded unit is allowed to enable attribute prediction when it is determined that the to-be-decoded unit is allowed to enable region adaptive hierarchical transform prediction decoding based on the value of the second syntax element.

[0204] The second syntax element is used to indicate whether the to-be-decoded unit is allowed to enable region adaptive hierarchical transform (RAHT) prediction decoding. When the second syntax element indicates that the to-be-decoded unit is enabled to perform RAHT prediction decoding, it can be understood that the to-be-decoded unit is allowed to enable attribute prediction. The second syntax element can be an enabling switch for controlling RAHT prediction decoding. The second syntax element can include a sequence level syntax element. The second syntax element can also include at least one of the following: a frame level syntax element, a slice level syntax element, a layer level syntax element, and a group level syntax element.

[0205] In some embodiments, when the attribute prediction mode includes attribute inter-frame prediction, the method can further include: determining a value of a first syntax element of the to-be-decoded unit in a case where it is determined that the to-be-decoded unit is allowed to enable attribute inter-frame prediction. That is, the parsing condition of the first syntax element at least includes that the to-be-decoded unit is allowed to enable attribute inter-frame prediction. Only when attribute inter-frame prediction is enabled, the to-be-decoded unit is allowed to use the decoding mode indicated based on the first syntax element to parse the first syntax element.

[0206] The related syntax element is used to indicate whether the to-be-decoded unit is enabled to perform attribute inter-frame prediction.

[0207] In some embodiments, the related syntax element is used to indicate a type of the to-be-decoded frame or the to-be-decoded slice. If the to-be-decoded frame is a P frame or a B frame, it is determined that the to-be-decoded frame and the to-be-decoded slice, the to-be-decoded layer, the to-be-decoded group, and the to-be-decoded point in the to-be-decoded frame are allowed to enable attribute inter-frame prediction. If the to-be-decoded slice is a P-slice or a B-slice, it is determined that the to-be-decoded slice, and the to-be-decoded layer, the to-be-decoded group, and the to-be-decoded point in the to-be-decoded slice are allowed to enable attribute inter-frame prediction.

[0208] In some embodiments, the related syntax element includes a third syntax element and / or a fourth syntax element. For example, the method further includes: determining a value of the third syntax element and / or a value of the fourth syntax element; determining that the to-be-decoded unit is allowed to enable attribute inter-frame prediction based on a determination that the to-be-decoded unit is allowed to refer to a first inter-frame reference unit based on the value of the third syntax element and / or a determination that the to-be-decoded unit is allowed to refer to a second inter-frame reference unit based on the value of the fourth syntax element.

[0209] The third syntax element is used to indicate whether the to-be-decoded unit is allowed to refer to the first inter-frame reference unit. The fourth syntax element is used to indicate whether the to-be-decoded unit is allowed to refer to the second inter-frame reference unit. For example, a value of the third syntax element being a first numerical value determines that the to-be-decoded unit is allowed to refer to the first inter-frame reference unit. A value of the fourth syntax element being the first numerical value determines that the to-be-decoded unit is allowed to refer to the second inter-frame reference unit. For example, the first numerical value can be 1.

[0210] The third syntax element comprises at least one of a frame-level syntax element, a slice-level syntax element, or another decoding unit-level syntax element. The fourth syntax element comprises at least one of a frame-level syntax element, a slice-level syntax element, or another decoding unit-level syntax element.

[0211] In some embodiments, when the to-be-decoded unit is a to-be-decoded layer or a to-be-decoded group, the method further comprises: parsing the bitstream to determine a value of a fifth syntax element, the fifth syntax element being used to indicate a number of the to-be-decoded units; and determining the value of the first syntax element of the to-be-decoded unit based on the value of the fifth syntax element.

[0212] It should be noted that the to-be-decoded unit can be understood as a unit that is subjected to attribute decoding based on the decoding mode indicated by the first syntax element, and thus the fifth syntax element can be understood as being used to indicate a number of the first syntax elements or a number of the decoding modes transmitted in the bitstream.

[0213] In some embodiments, the to-be-decoded layer can be a to-be-decoded RAHT layer. The to-be-decoded layer can also be a to-be-decoded LOD layer.

[0214] In some embodiments, when the to-be-decoded unit is a to-be-decoded group, the method further comprises: parsing the bitstream to determine a value of a sixth syntax element of the to-be-decoded group; determining a number N of nodes of the to-be-decoded group based on the value of the sixth syntax element, N being a positive integer; and performing node division on the to-be-decoded layer based on the number N of nodes of the to-be-decoded group to determine the to-be-decoded group.

[0215] In some embodiments, performing node division on the to-be-decoded layer based on the number N of nodes of the to-be-decoded group to determine the to-be-decoded group comprises: dividing every N nodes into one to-be-decoded group based on the number N of nodes of the to-be-decoded group and a node order of the to-be-decoded layer, starting from a first node; and if there are less than N nodes remaining in the to-be-decoded layer, dividing the less than N nodes remaining into one to-be-decoded group.

[0216] In some embodiments, if a hardware condition permits, the to-be-decoded unit can also be a to-be-decoded node.

[0217] S102: determining a decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit;

[0218] In some embodiments, the decoding mode comprises a decoding order of one or more decoding modes.

[0219] In some embodiments, the one or more decoding modes comprise at least one of an inter-intra fusion prediction transform mode, a bi-directional inter prediction transform mode, a forward inter prediction transform mode, a backward inter prediction transform mode, an intra prediction transform mode, and a transform mode.

[0220] The inter-intra fusion prediction transform mode, the bi-directional inter prediction transform mode, the forward inter prediction transform mode, the backward inter prediction transform mode and the intra prediction transform mode can be understood as an attribute prediction transform mode. If the neighborhood attribute distribution characteristics of the to-be-decoded node are relatively smooth, the attribute prediction transform mode can be used based on the distribution characteristics. If the neighborhood attribute distribution of the to-be-decoded node is relatively jittered, only the attribute transform mode is used.

[0221] The inter-intra fusion prediction transform mode can be understood as inter-intra fusion prediction + transform. The attribute prediction value is obtained by merging the inter prediction value and the intra prediction value of the to-be-decoded node, and the attribute prediction value is inversely transformed to determine the attribute reconstruction value of the to-be-decoded node. For example, the inter prediction value and the intra prediction value can be merged according to different weights to obtain the final prediction value, so as to further improve the point cloud attribute decoding efficiency. Specifically, assuming that the intra prediction value of the current node is predIntraVal and the inter prediction value is predInterVal, the final prediction value predVal is predVal = w1*predIntraVal + w2*predInterVal.

[0222] The bi-directional inter prediction transform mode can be understood as bi-directional inter prediction + transform. The attribute prediction value is obtained by merging the forward inter prediction value and the backward inter prediction value of the to-be-decoded node, and the attribute prediction value is inversely transformed to determine the attribute reconstruction value of the to-be-decoded node.

[0223] The forward inter prediction transform mode can be understood as forward inter prediction + transform. The forward inter prediction value of the to-be-decoded node is determined, and the forward inter prediction value is inversely transformed to determine the attribute reconstruction value of the to-be-decoded node.

[0224] The backward inter prediction transform mode can be understood as backward inter prediction + transform. The backward inter prediction value of the to-be-decoded node is determined, and the backward inter prediction value is inversely transformed to determine the attribute reconstruction value of the to-be-decoded node.

[0225] The intra prediction transform mode can be understood as intra prediction + transform. The intra prediction value of the to-be-decoded node is determined, and the intra prediction value is inversely transformed to determine the attribute reconstruction value of the to-be-decoded node.

[0226] The transform mode can be understood as not using attribute prediction, but only using transform to determine the attribute reconstruction value of the to-be-decoded node.

[0227] In some embodiments, in a case where the first syntax element of the to-be-decoded unit has a first value, then the decoding order corresponding to the decoding manner of the to-be-decoded unit is determined to be in turn the first inter prediction transform mode, the second inter prediction transform mode, the third inter prediction transform mode, the intra prediction transform mode, and the transform mode. The first value can be any value used to identify such decoding order.

[0228] In some embodiments, in a case where the first syntax element of the to-be-decoded unit has a second value, then the decoding order corresponding to the decoding manner of the to-be-decoded unit is determined to be in turn the second inter prediction transform mode, the intra prediction transform mode, and the transform mode. The second value can be any value used to identify such decoding order.

[0229] In some embodiments, in a case where the first syntax element of the to-be-decoded unit has a third value, then the decoding order corresponding to the decoding manner of the to-be-decoded unit is determined to be in turn the third inter prediction transform mode, the intra prediction transform mode, and the transform mode. The third value can be any value used to identify such decoding order.

[0230] In some embodiments, in a case where the first syntax element of the to-be-decoded unit has a fourth value, then the decoding order corresponding to the decoding manner of the to-be-decoded unit is determined to be in turn the intra prediction transform mode and the transform mode. The fourth value can be any value used to identify such decoding order.

[0231] In some embodiments, in a case where the first syntax element of the to-be-decoded unit has a fifth value, then the decoding order corresponding to the decoding manner of the to-be-decoded unit is determined to be in turn the transform mode. The fifth value can be any value used to identify such decoding order.

[0232] In some embodiments, in a case where the first syntax element of the to-be-decoded unit has a sixth value, then the decoding order corresponding to the decoding manner of the to-be-decoded unit is determined to be in turn the inter-intra prediction transform mode, the first inter prediction transform mode, the second inter prediction transform mode, the third inter prediction transform mode, the intra prediction transform mode, and the transform mode. The sixth value can be any value used to identify such decoding order.

[0233] It should be noted that the above six decoding manners are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application. Any combination of the above six decoding manners can be made to constitute one or more optional decoding manners of the to-be-decoded unit.

[0234] In some embodiments, the first inter-prediction transform mode is a bi-directional inter-prediction transform mode, the second inter-prediction transform mode is a forward inter-prediction transform mode, and the third inter-prediction transform mode is a backward inter-prediction transform mode. That is, if the decoding manner simultaneously exists two or three inter-prediction modes, the decoding order of the bi-directional inter-prediction transform mode is before the decoding order of the forward inter-prediction transform mode, the decoding order of the bi-directional inter-prediction transform mode is before the decoding order of the backward inter-prediction transform mode, and the decoding order of the forward inter-prediction transform mode is before the decoding order of the backward inter-prediction transform mode.

[0235] In some embodiments, the decoding order of the plurality of decoding modes is preset, and the value of the first syntax element is used to indicate a first decoding mode of the decoding manner, and the first decoding mode to the last decoding mode constitute the decoding manner.

[0236] As shown in FIG. 22, assuming that the decoding order of the plurality of decoding modes is: inter-intra fusion prediction transform mode, bi-directional inter-prediction transform mode, forward inter-prediction transform mode, backward inter-prediction transform mode, intra-prediction transform mode, and transform mode in turn. The first syntax element is denoted as attr_code_mode, which is used to indicate the best decoding mode of the to-be-decoded unit. For example, when the value of attr_code_mode is Idxl, the best decoding mode of the to-be-decoded unit is the bi-directional inter-prediction transform mode. Based on the best decoding mode indicated by the first syntax element and the preset decoding order, a decoding manner is determined, and the decoding order corresponding to the decoding manner can be: bi-directional inter-prediction transform mode, forward inter-prediction transform mode, backward inter-prediction transform mode, intra-prediction transform mode, and transform mode.

[0237] The bi-directional inter-prediction transform mode is further attempted to decode one or more decoding modes after the bi-directional inter-prediction transform mode in the case that the bi-directional inter-prediction transform mode is unavailable for the node of the to-be-decoded unit.

[0238] For example, taking the to-be-decoded layer as an example, if the bi-directional inter-prediction decoding is enabled for the current slice, a possible index mode is shown in Table 1.

[0239] Table 1

[0240] attr_code_mode[i] represents the decoding manner of the to-be-decoded layer i, and non-prediction represents only using the transform mode.

[0241] A possible index mode is shown in Table 2.

[0242] Table 2

[0243] A possible index mode is shown in Table 3.

[0244] Table 3

[0245] S103: Perform attribute decoding on the node of the to-be-decoded unit based on the determined decoding manner, and determine the attribute reconstruction value of the node of the to-be-decoded unit.

[0246] In some embodiments, the performing attribute decoding on the node of the to-be-decoded unit based on the determined decoding manner, and determining the attribute reconstruction value of the node of the to-be-decoded unit comprises: determining a current decoding mode based on a decoding order of the determined decoding manner; judging whether the current decoding mode is available for the node of the to-be-decoded unit; in a case where the current decoding mode is available for the node of the to-be-decoded unit, performing attribute decoding on the node of the to-be-decoded unit based on the current decoding mode, and determining the attribute reconstruction value of the node of the to-be-decoded unit; in a case where the current decoding mode is unavailable for the node of the to-be-decoded unit, determining a next decoding mode based on the determined decoding manner, and continuing to judge whether the next decoding mode is available for the node of the to-be-decoded unit.

[0247] That is, based on the decoding order defined by the decoding manner, the decoding modes in the decoding manner are traversed, and an available decoding mode is selected to perform attribute decoding on the to-be-decoded unit.

[0248] In some embodiments, the method further comprises: in a case where the current decoding mode is the last decoding mode, performing attribute decoding on the node of the to-be-decoded unit based on the current decoding mode, and determining the attribute reconstruction value of the node of the to-be-decoded unit. That is, if the last decoding mode in the decoding manner is reached, the last decoding mode is directly used for attribute decoding, and the last decoding mode is available for the node of the to-be-decoded unit when attribute decoding is performed based on the decoding manner indicated by the first syntax element.

[0249] FIG. 23 is a flow diagram of a method of attribute decoding based on a decoding manner according to an embodiment of the present application. As shown in FIG. 23, the method of attribute decoding based on a decoding manner according to an embodiment of the present application can specifically comprise:

[0250] S201: Determine a current decoding mode of a to-be-decoded unit based on a determined decoding manner;

[0251] S202: Judge whether the determined current decoding mode is the last decoding mode in the to-be-decoded manner; if not, perform step S203; if yes, perform step S204;

[0252] S203: Judge whether the determined current decoding mode is available for the node of the to-be-decoded unit; if yes, perform step S204; if not, return to step S201, and continue to determine a next decoding mode based on the determined decoding manner;

[0253] S204: In a case that the current decoding mode is available for the node of the to-be-decoded unit, attribute decoding is performed on the node of the to-be-decoded unit based on the current decoding mode, and an attribute reconstruction value of the node of the to-be-decoded unit is determined.

[0254] In a case that the current decoding mode is the inter-prediction transform mode, determining whether the determined current decoding mode is available for the node of the to-be-decoded unit includes: in a case that the node of the to-be-decoded unit has an inter-reference node, determining that the inter-prediction transform mode is available for the node of the to-be-decoded unit; and in a case that the node of the to-be-decoded unit does not have an inter-reference node, determining that the inter-prediction transform mode is unavailable for the node of the to-be-decoded unit.

[0255] Specifically, in a case that the current decoding mode is the bi-directional inter-prediction transform mode, the bi-directional inter-prediction transform mode is determined to be available for the node of the to-be-decoded unit in a case that the to-be-decoded unit has a forward inter-reference node and a backward inter-reference node; otherwise, the bi-directional inter-prediction transform mode is determined to be unavailable.

[0256] In a case that the current decoding mode is the forward inter-prediction transform mode, the forward inter-prediction transform mode is determined to be available for the node of the to-be-decoded unit in a case that the to-be-decoded unit has a forward inter-reference node; otherwise, the forward inter-prediction transform mode is determined to be unavailable.

[0257] In a case that the current decoding mode is the backward inter-prediction transform mode, the forward inter-prediction transform mode is determined to be available for the node of the to-be-decoded unit in a case that the to-be-decoded unit has a backward inter-reference node; otherwise, the forward inter-prediction transform mode is determined to be unavailable.

[0258] It should be noted that the inter-reference node can include at least one of the following: an inter-sibling node of the to-be-encoded node, an inter-sibling parent node, or a motion-compensated node.

[0259] In a case that the current decoding mode is the intra-prediction transform mode, determining whether the determined current decoding mode is available for the node of the to-be-decoded unit includes: determining a number of neighbor nodes of the node of the to-be-decoded unit, and determining that the intra-prediction transform mode is available for the node of the to-be-decoded unit in a case that the number of neighbor nodes is greater than or equal to a first threshold; otherwise, determining that the intra-prediction transform mode is unavailable for the node of the to-be-decoded unit.

[0260] In a case that the current decoding mode is the intra-prediction transform mode, determining whether the determined current decoding mode is available for the node of the to-be-decoded unit includes: determining a number of neighbor nodes of the node of the to-be-decoded unit, and determining that the intra-prediction transform mode is available for the node of the to-be-decoded unit in a case that the number of neighbor nodes is greater than or equal to a first threshold; otherwise, determining that the intra-prediction transform mode is unavailable for the node of the to-be-decoded unit.

[0261] In some embodiments, based on the determined decoding manner, attribute decoding is performed on the node of the to-be-decoded unit to determine the attribute reconstruction value of the node of the to-be-decoded unit, including: based on the determined decoding manner, a decoding mode of the to-be-decoded unit is determined; and based on the determined decoding mode, attribute decoding is performed on the node of the to-be-decoded unit to determine the attribute reconstruction value of the node of the to-be-decoded unit.

[0262] If the current decoding mode is the inter-intra fusion prediction transform mode, it is determined whether the determined current decoding mode is available for the node of the to-be-decoded unit, including: if the inter prediction of the node of the to-be-decoded unit is available and the intra prediction is available, it is determined that the inter-intra fusion prediction transform mode is available for the node of the to-be-decoded unit, otherwise, it is unavailable.

[0263] It can be understood that if the decoding manner includes one decoding mode, the to-be-decoded unit is attribute decoded using the decoding mode. If the decoding manner includes multiple decoding modes, the to-be-decoded unit is attribute decoded based on a decoding order of the multiple decoding modes, and a suitable decoding mode is selected considering the attribute distribution characteristics of different coding nodes to improve the attribute coding efficiency. For example, for a to-be-decoded unit using the bidirectional inter prediction transform mode, if the to-be-decoded node does not exist a bidirectional parity node, a unidirectional inter prediction transform mode can be selected, or an intra prediction transform mode can be selected, or only a transform mode can be used.

[0264] In some embodiments, if the current decoding mode is a prediction transform mode (including inter prediction and intra prediction), attribute decoding is performed on the node of the to-be-decoded unit based on the current decoding mode, including: predicting the node of the to-be-decoded unit (also referred to as a to-be-decoded node) to determine a prediction value of an alternating current coefficient of the to-be-decoded node; parsing a code stream to determine a residual value of the alternating current coefficient of the to-be-decoded node; based on the prediction value of the alternating current coefficient of the to-be-decoded node and the residual value of the alternating current coefficient, a reconstruction value of the alternating current coefficient of the to-be-decoded node is determined; and inverse transform is performed on the reconstruction value of the alternating current coefficient and a direct current coefficient reconstruction value of the to-be-decoded node to determine an attribute reconstruction value of the to-be-decoded node. For bidirectional inter prediction, the prediction value of the alternating current coefficient of the to-be-decoded node is determined based on the reconstruction value of the alternating current coefficient of a bidirectional inter reference node; for unidirectional inter prediction, the prediction value of the alternating current coefficient of the to-be-decoded node is determined based on the reconstruction value of the alternating current coefficient of a unidirectional inter reference node; and for intra prediction, the prediction value of the alternating current coefficient of the to-be-decoded node is determined based on the reconstruction value of the alternating current coefficient of an intra reference node.

[0265] In some embodiments, if the current decoding mode is a transform mode, attribute decoding is performed on the node of the to-be-decoded unit based on the current decoding mode, including: parsing a code stream to determine a reconstruction value of an alternating current coefficient of the to-be-decoded node; and inverse transform is performed on the reconstruction value of the alternating current coefficient and a direct current coefficient reconstruction value of the to-be-decoded node to determine an attribute reconstruction value of the to-be-decoded node.

[0266] By adopting the technical scheme, the decoding end determines the first syntax element from the code stream, thereby determining the decoding mode of the to-be-decoded unit. The decoding mode is the best decoding mode adaptively selected by fully considering the node attribute distribution characteristics of the to-be-decoded unit. The point cloud attribute reconstruction is performed on the to-be-decoded unit by using the decoding mode, and the point cloud attribute decoding efficiency can be improved.

[0267] The attribute decoding process of different to-be-decoded units is further illustrated below.

[0268] The to-be-decoded unit is taken as an example of the RAHT decoding layer.

[0269] First, the RAHT decoding layer, also referred to as the RAHT transform layer or the RAHT layer, is defined. The RAHT attribute transformation is based on the order of the octree hierarchy, and the transformation is continuously performed from the root node until the voxel level (1x1x1), thereby completing the attribute reconstruction of the entire point cloud attribute. Here, we define each layer obtained by performing once down-sampling on the geometric information along the Z direction, the Y direction, and the X direction as a RAHT transform layer, that is, layer, as shown in FIG. 24. The attribute decoding method of the to-be-decoded RAHT layer includes:

[0270] 1. When decoding the attribute information of the current slice, the reconstructed geometric information of the current slice is obtained first. The same down-sampling algorithm as that of the encoding end is adopted to down-sample the points of the entire slice. Specifically, once down-sampling is performed on the geometric information along the Z direction, the Y direction, and the X direction each time, and a RAHT decoding layer is obtained.

[0271] 2. Step 1 is repeatedly performed to down-sample the point geometric information of the current slice until the root node is reached, that is, the down-sampling is stopped when a node is reached.

[0272] 3. The root node is up-sampled. Specifically, once up-sampling is performed on the geometric information along the Z direction, the Y direction, and the X direction each time, and a RAHT decoding layer is obtained.

[0273] 4. After obtaining the RAHT decoding layer, the decoding end parses the first syntax element of the current RAHT layer from the code stream, determines the decoding mode of the current RAHT layer, reconstructs the AC coefficients of the nodes of the current RAHT layer according to the decoding mode, combines the corresponding DC coefficients, performs independent orthogonal inverse transformation by using the inverse transformation matrix of the region adaptive hierarchical transformation, and reconstructs the reconstructed attributes of the sub-nodes contained in each node of the current RAHT decoding layer.

[0274] 5. Steps 3 and 4 are repeatedly performed to reconstruct and recover the attribute information of the current slice based on the RAHT decoding layer.

[0275] When decoding the attribute of the RAHT layer, the syntax element (Attribute data unit header syntax) in the attribute header information is described as shown in Table 4.

[0276] Table 4

[0277] disableAttrInterPred corresponds to the third syntax element, which is used to indicate whether the RAHT layer of the current slice is allowed to refer to the first inter-frame reference unit, and can also be understood as whether the attribute inter prediction is enabled for the current slice. disableAttrInterPred2 corresponds to the fourth syntax element, which is used to indicate whether the RAHT layer of the current slice is allowed to refer to the second inter-frame reference unit, and can also be understood as whether the second inter-frame reference unit is used for the current slice. raht_prediction_enabled is used to indicate whether the RAHT prediction decoding is allowed for the current slice. When it is true, it indicates that the RAHT prediction decoding is allowed, and when it is false, it indicates that the RAHT prediction decoding is not allowed. attr_code_mode_cnt corresponds to the fifth syntax element, which is used to indicate the number of RAHT layers, and can also be understood as the number of decoding modes contained in the current slice stream. attr_code_mode[i] is used to indicate the coding mode of the i-th layer of the current slice. At present, attr_code_mode_cnt and attr_code_mode[i] of the RAHT decoding layer can be stored in the ABH, and the decoding mode of the current RAHT decoding layer is obtained through the ABH at the decoding end. The form of encoding this parameter is not limited here.

[0278] In some embodiments, the attr_code_mode[i] of the current slice can be used to indicate the decoding mode of the part of the RAHT decoding layer (such as the root node).

[0279] In some embodiments, the attr_code_mode[i] determines the decoding mode of each RAHT coding layer when the current slice starts the bidirectional inter-frame prediction decoding. For example, the present application introduces four attribute decoding modes in some embodiments: forward inter-frame prediction + transform, bidirectional inter-frame prediction + transform, backward inter-frame prediction + transform, and transform only.

[0280] attr_code_mode[i] is 0: indicates that if the current node of the RAHT layer to be encoded has a homonym in the forward reference image, forward inter-frame prediction is performed; if the current node does not meet the forward inter-frame prediction, it is determined whether the current node can use intra-frame prediction, if the conditions for intra-frame prediction are met, intra-frame prediction is performed; otherwise, non-predictive decoding is performed on the current node, that is, only transform decoding is used.

[0281] attr_code_mode[i] is 1: indicates that if the current node has a homonym in the backward reference image, backward inter-frame prediction is performed; if the current node does not meet the backward inter-frame prediction, it is determined whether the current node can use intra-frame prediction, if the conditions for intra-frame prediction are met, intra-frame prediction is performed; otherwise, non-predictive decoding is performed on the current node.

[0282] attr_code_mode[i] is 2: indicates that if the current node has homonyms in both the forward and backward reference images, bidirectional inter-frame prediction is performed; if the current node does not meet the bidirectional inter-frame prediction, it is determined whether the current node has a homonym in the forward reference image, if the homonym exists, forward inter-frame prediction is performed; if the current node does not have a homonym in the forward reference image, it is determined whether the current node has a homonym, if the homonym exists, backward inter-frame prediction is performed; if the current node does not meet the backward inter-frame prediction, it is determined whether the current node can use intra-frame prediction, if the conditions for intra-frame prediction are met, intra-frame prediction is performed; otherwise, non-predictive decoding is performed on the current node.

[0283] attr_code_mode[i] is 3: indicates non-predictive decoding.

[0284] In the RAHT attribute decoding of the embodiments of the present application, the attribute distribution characteristics of different RAHT layers in the node slice are fully considered, a first syntax element is introduced for the RAHT layer to indicate the best decoding mode, and the node of the RAHT layer is reconstructed using the decoding mode to improve the point cloud attribute coding and decoding efficiency.

[0285] In some embodiments, the attribute decoding mode can be changed to: inter-frame prediction + transform, intra-frame prediction + transform, and transform coding mode. The inter-frame prediction can be any one of bidirectional inter-frame prediction, forward inter-frame prediction and backward inter-frame prediction.

[0286] In the above scheme, for any decoding mode, firstly, it is judged whether the attribute inter-frame prediction value is equal to zero. If not, the attribute inter-frame prediction value is directly taken as the prediction value of the AC coefficient of the current node. Otherwise, intra-frame prediction is performed to obtain the AC coefficient prediction value of the current node. In other embodiments, inter-intra fusion prediction + transform can also be introduced. By merging the inter-frame prediction value and the intra-frame prediction value of different RAHT transform layers, the best prediction value is finally obtained according to different weights, so that the point cloud attribute RAHT coding efficiency can be further improved. The specific prediction decoding scheme is as follows: assuming that the RAHT intra-frame prediction value of the current node is predIntraVal, and the inter-frame prediction value is predInterVal, the final prediction value predVal is predVal=w1*predIntraVal+w2*predIntraVal

[0287] The to-be-decoded unit is taken as an example of a decoding group (DG).

[0288] The RAHT coding layer unit is divided to obtain different DGs. Similarly, each DG needs to transmit a best prediction coding mode to reconstruct the attribute information of each node in the current DG. Here, the division of the DG group can be adaptive division, for example, the maximum number of nodes of each DG is fixed as N, and each N nodes of each RAHT decoding layer are divided into a DG group, and then the best coding mode is selected for each DG at the encoding end. The attribute decoding method of the to-be-decoded group can include:

[0289] 1. When decoding the attribute information of the current slice, the reconstructed geometry information of the current slice is obtained. The same downsampling algorithm as that of the encoding end is used to downsample the points of the entire slice. Specifically, the geometry information is downsampled once along the Z direction, Y direction and X direction each time, and a RAHT coding layer is obtained.

[0290] 2. Step 1 is repeatedly performed to downsample the point geometry information of the current slice until the root node is reached, i.e., a node is downsampled and the downsample is stopped.

[0291] 3. The root node is upsampled. Specifically, the geometry information is upsampled once along the Z direction, Y direction and X direction each time, and a RAHT decoding layer is obtained.

[0292] 4. After getting the RAHT decoding layer, suppose the number of nodes in the current RAHT decoding layer is Q. Divide every N nodes in the current RAHT decoding layer into a decoding group DG (decoding group), based on such a way, get M decoding group layers. As shown in FIG. 25, divide every 4 nodes in the RAHT decoding layer into a DG group, and divide the last 1 node which is less than 4 into a DG group.

[0293] 5. After dividing the current RAHT decoding layer, get different decoding groups, and parse the attribute prediction mode of different decoding groups.

[0294] 6. After determining the attribute prediction mode of each node in the DG, before reconstructing the attribute of each node in the DG, parse the DC coefficient and the corresponding AC coefficient of each node in the current DG decoding unit from the code stream, and use the inverse transform matrix of the background "region adaptive hierarchical transform" to perform independent orthogonal inverse transform, to reconstruct the reconstructed attribute of the sub-nodes contained in each node in the current DG decoding unit.

[0295] 7. Continuously repeat steps 3 and 4 to reconstruct and recover the attribute information of the current slice based on the RAHT decoding layer.

[0296] When decoding the attribute of the decoding group, the syntax element (Attribute data unit header syntax) in the attribute header information is described as shown in Table 5.

[0297] Table 5

[0298] attr_code_group_size is used to indicate the number of nodes N in each decoding group. When the syntax element does not exist, it is 0 by default. attr_code_mode_cnt corresponds to the fifth syntax element, which is used to indicate the number of RAHT layers, and can also be understood as indicating the number of decoding groups of the current slice, and can also be understood as the number of decoding modes for the current slice. attr_code_mode[i] is used to indicate the coding mode of the i-th decoding group of the current slice.

[0299] Currently, attr_code_group_size, attr_code_mode_cnt and attr_code_mode[i] of the decoding group can be stored in ABH, and the decoding mode of the current decoding group is obtained through ABH at the decoding end. The form of encoding this parameter is not limited here.

[0300] In some embodiments, the attr_code_mode[i] determines the decoding mode for each decoding group when the current slice enables bi-directional inter-frame prediction decoding. For example, the present application introduces four attribute decoding modes in some embodiments: forward inter-frame prediction + transform, bi-directional inter-frame prediction + transform, backward inter-frame prediction + transform.

[0301] attr_code_mode[i] is 0: indicating that if the current node of the to-be-encoded group has a homonym node in the forward reference image, forward inter-frame prediction is performed; if the current node does not meet the forward inter-frame prediction, it is judged whether the current node can adopt intra-frame prediction, if the intra-frame prediction condition is met, intra-frame prediction is performed; otherwise, non-prediction decoding is performed on the current node.

[0302] attr_code_mode[i] is 1: indicating that if the current node has a homonym node in the backward reference image, backward inter-frame prediction is performed; if the current node does not meet the backward inter-frame prediction, it is judged whether the current node can adopt intra-frame prediction, if the intra-frame prediction condition is met, intra-frame prediction is performed; otherwise, non-prediction decoding is performed on the current node.

[0303] attr_code_mode[i] is 2: indicating that if the current node has a homonym node in both the forward and backward reference images, bi-directional inter-frame prediction is performed; if the current node does not meet the bi-directional inter-frame prediction, it is judged whether the current node has a homonym node in the forward reference image, if the homonym node exists, forward inter-frame prediction is performed; if the current node does not have a homonym node in the forward reference image, it is judged whether the current node has a homonym node, if the homonym node exists, backward inter-frame prediction is performed; if the current node does not meet the backward inter-frame prediction, it is judged whether the current node can adopt intra-frame prediction, if the intra-frame prediction condition is met, intra-frame prediction is performed; otherwise, non-prediction decoding is performed on the current node.

[0304] In the RAHT attribute decoding of the embodiments of the present application, the attribute distribution characteristics of different RAHT layers in the node slice are fully considered, and the RAHT layer is further divided into decoding groups, a first syntax element is introduced for the decoding group to indicate the to-be-optimal decoding mode, the point cloud attribute reconstruction of the node of the decoding group is performed by using the decoding mode, and the point cloud attribute decoding efficiency is improved.

[0305] In an embodiment of the present application, referring to FIG. 26, a flowchart of a point cloud encoding method provided by the embodiments of the present application is shown. As shown in FIG. 26, the method can include:

[0306] S301: based on the candidate encoding mode of the to-be-encoded unit, performing attribute encoding on the node of the to-be-encoded unit to determine the attribute reconstruction value of the node of the to-be-encoded unit;

[0307] In some embodiments, the to-be-encoded unit can include at least one of the following: a to-be-encoded layer, a to-be-encoded group, and a to-be-encoded node.

[0308] In some embodiments, the to-be-encoded unit can also include at least one of the following: a to-be-encoded sequence, a to-be-encoded frame, and a to-be-encoded slice.

[0309] In some embodiments, the method further includes: in a case where the coding condition of the first syntax element of the current block is satisfied, performing steps S301 to S304.

[0310] In some embodiments, the method further includes: in a case where it is determined that the to-be-encoded unit is allowed to enable attribute prediction, performing attribute coding on the node of the to-be-encoded unit based on the candidate coding mode of the to-be-encoded unit, and determining the attribute reconstruction value of the node of the to-be-encoded unit.

[0311] In some embodiments, the method further includes: determining the value of the second syntax element; in a case where it is determined that the to-be-encoded unit is allowed to enable region adaptive hierarchical transform prediction coding based on the value of the second syntax element, determining that the to-be-encoded unit is allowed to enable attribute prediction; and performing coding processing on the second syntax element, and writing the obtained coding bits into the bitstream.

[0312] In some embodiments, the method further includes: in a case where it is determined that the to-be-encoded unit is allowed to enable attribute inter prediction, performing attribute coding on the node of the to-be-encoded unit based on the candidate coding mode of the to-be-encoded unit, and determining the attribute reconstruction value of the node of the to-be-encoded unit.

[0313] In some embodiments, the method further includes: determining the value of the third syntax element and / or the value of the fourth syntax element; in a case where it is determined that the to-be-encoded unit is allowed to refer to the first inter reference unit based on the value of the third syntax element, and / or the to-be-encoded unit is allowed to refer to the second inter reference unit based on the value of the fourth syntax element, determining that the to-be-encoded unit is allowed to enable attribute inter prediction; and performing coding processing on the third syntax element and / or the fourth syntax element, and writing the obtained coding bits into the bitstream.

[0314] In some embodiments, when the to-be-encoded unit is a to-be-encoded layer or a to-be-encoded group, the method further includes: determining the value of a fifth syntax element based on the number of the to-be-encoded unit; and performing coding processing on the fifth syntax element, and writing the obtained coding bits into the bitstream.

[0315] In some embodiments, when the to-be-encoded unit is a to-be-encoded group, the method further comprises: performing node division on the to-be-encoded layer based on the node number N of the to-be-encoded group to determine the to-be-encoded group; determining the value of the sixth syntax element of the to-be-encoded group based on the node number N of the to-be-encoded group; and performing encoding processing on the sixth syntax element and writing the obtained encoding bits into the code stream.

[0316] In some embodiments, performing node division on the to-be-encoded layer based on the node number N of the to-be-encoded group to determine the to-be-encoded group comprises: dividing every N nodes into a to-be-encoded group based on the node number N of the to-be-encoded group and the node order of the to-be-encoded layer, starting from the first node; and if there are less than N nodes remaining in the to-be-encoded layer, dividing the less than N nodes remaining into a to-be-encoded group.

[0317] In some embodiments, the candidate encoding mode comprises one or more candidates, and the encoding mode can comprise the encoding order of one or more encoding modes.

[0318] In some embodiments, based on the candidate encoding mode of the to-be-encoded unit, performing attribute encoding on the nodes of the to-be-encoded unit to determine the attribute reconstruction value of the nodes of the to-be-encoded unit comprises: determining the current encoding mode based on the encoding order of the candidate encoding mode; judging whether the current encoding mode is available for the nodes of the to-be-encoded unit; in the case that the current encoding mode is available for the nodes of the to-be-encoded unit, performing attribute encoding on the nodes of the to-be-encoded unit based on the current encoding mode to determine the attribute reconstruction value of the nodes of the to-be-encoded unit; and in the case that the current encoding mode is not available for the nodes of the to-be-encoded unit, determining the next encoding mode based on the determined encoding mode and continuing to judge whether the next encoding mode is available for the nodes of the to-be-encoded unit.

[0319] In some embodiments, based on the candidate encoding mode of the to-be-encoded unit, performing attribute encoding on the nodes of the to-be-encoded unit to determine the attribute reconstruction value of the nodes of the to-be-encoded unit comprises: determining the encoding mode of the to-be-encoded unit based on the candidate encoding mode; and performing attribute encoding on the nodes of the to-be-encoded unit based on the determined encoding mode to determine the attribute reconstruction value of the nodes of the to-be-encoded unit.

[0320] S302: based on the attribute reconstruction value of the nodes of the to-be-encoded unit, performing encoding decision on the candidate encoding mode to determine the encoding mode of the to-be-encoded unit;

[0321] In some embodiments, different attribute reconstruction values can be obtained by cycling through different candidate encoding manners, and an encoding cost is calculated according to the attribute reconstruction values, so as to determine the optimal encoding manner. Exemplarily, the encoding cost includes, but is not limited to, sum of absolute difference (SAD), sum of absolute transformed difference (SATD), sum of squared difference (SSE), mean absolute difference (MAD), mean absolute error (MAE), mean squared error (MSE), rate-distortion cost (RDO), etc.

[0322] Exemplarily, for determining the optimal encoding manner of the to-be-encoded unit by the rate-distortion optimization algorithm, firstly, a distortion D of the attribute reconstruction value and the attribute original value of each encoding manner is calculated, and secondly, a code stream R required for encoding of each encoding manner is obtained, and then the rate-distortion cost is calculated as follows: J=D+λxR

[0323] Wherein, λ can be calculated by the attribute quantization parameter, and the current λ calculation manner is as follows:

[0324] The parameter N is currently set to different values according to the reflectivity and color.

[0325] S303: determining the value of the first syntax element of the to-be-encoded unit based on the encoding manner of the to-be-encoded unit;

[0326] In some embodiments, the encoding order corresponding to the encoding manner of the to-be-encoded unit is first inter-prediction transform mode, second inter-prediction transform mode, third inter-prediction transform mode, intra-prediction transform mode and transform mode in turn, and the value of the first syntax element of the to-be-encoded unit is determined as the first numerical value.

[0327] In some embodiments, the encoding order corresponding to the encoding manner of the to-be-encoded unit is second inter-prediction transform mode, intra-prediction transform mode and transform mode in turn, and the value of the first syntax element of the to-be-encoded unit is determined as the second numerical value.

[0328] In some embodiments, the encoding order corresponding to the encoding manner of the to-be-encoded unit is third inter-prediction transform mode, intra-prediction transform mode and transform mode in turn, and the value of the first syntax element of the to-be-encoded unit is determined as the third numerical value.

[0329] In some embodiments, the encoding order corresponding to the encoding manner of the to-be-encoded unit is intra-prediction transform mode and transform mode in turn, and the value of the first syntax element of the to-be-encoded unit is determined as the fourth numerical value.

[0330] In some embodiments, the encoding order corresponding to the encoding manner of the to-be-encoded unit is transform mode in turn, and the value of the first syntax element of the to-be-encoded unit is determined as the fifth numerical value.

[0331] In some embodiments, the coding order corresponding to the coding manner of the to-be-encoded unit is in the order of inter-frame intra-frame fusion prediction transform mode, first inter-frame prediction transform mode, second inter-frame prediction transform mode, third inter-frame prediction transform mode, intra-frame prediction transform mode, and transform mode, and the value of the first syntax element of the to-be-encoded unit is determined as the sixth numerical value.

[0332] In some embodiments, the first inter-frame prediction transform mode is a bi-directional inter-frame prediction transform mode, the second inter-frame prediction transform mode is a forward inter-frame prediction transform mode, and the third inter-frame prediction transform mode is a backward inter-frame prediction transform mode.

[0333] S304: Perform encoding processing on the first syntax element of the to-be-encoded unit, and write the obtained encoding bits into the code stream.

[0334] By using the above technical solution, the coding end fully considers the node attribute distribution characteristics of the to-be-encoded unit, adaptively selects the best coding manner for the to-be-encoded unit, writes the first syntax element indicating the best coding manner into the code stream, and uses the best coding manner to perform point cloud attribute reconstruction on the to-be-encoded unit, thereby improving the point cloud attribute coding efficiency.

[0335] The attribute coding process of different to-be-encoded units will be further illustrated below.

[0336] Taking the RAHT coding layer as an example.

[0337] First, the RAHT coding layer, also known as the RAHT transform layer or the RAHT layer, is defined. The current attribute RAHT transform coding order is to perform division from the root node until the voxel level (1x1x1) is reached, that is, the node merging is performed from the voxel level, and the root node of the entire RAHT transform tree is obtained, thereby completing the attribute reconstruction of the entire point cloud attribute. Here, we define each layer obtained by performing downsampling on the geometric information along the Z direction, Y direction, and X direction as a RAHT transform layer, i.e., layer, as shown in FIG. 24. The attribute coding method of the to-be-encoded RAHT layer includes:

[0338] 1. When coding the attribute information of the current slice, the geometric information and attribute information of the current slice are obtained. The points of the current slice are downsampled by using the same algorithm principle as the octree. Specifically, the geometric information is downsampled once along the Z direction, Y direction, and X direction each time, and a RAHT coding layer is obtained.

[0339] 2. After getting the RAHT coding layer, the attributes merged into a node are independently orthogonal transformed, using the transform matrix of the region adaptive hierarchical transform. After the independent orthogonal transform, a DC coefficient and a series of AC coefficients can be obtained. The DC coefficient of the current node is taken as the attribute information of the current node.

[0340] After getting the RAHT coding layer, a rate-distortion optimization algorithm is introduced to adaptively select the coding mode of the current layer, and four coding modes are introduced: forward inter-prediction + transform, bi-directional inter-prediction + transform, backward inter-prediction + transform, and transform only.

[0341] The rate-distortion optimization algorithm is used at the encoding end to predict and encode the attribute information of the current layer node using the four coding modes. Finally, the rate-distortion optimization algorithm is used to obtain the best coding mode of the current layer, and the best coding mode is passed to the decoding end. The decoding end uses the decoded mode obtained by analysis to reconstruct and recover the attribute information of the nodes of the current layer to be decoded.

[0342] 3. Repeat step 1 to downsample the point geometry information of the current slice, and downsample the geometry information along the Z direction, Y direction and X direction once each time, to obtain a RAHT coding layer unit. Then, the attribute information is transformed and encoded using the RAHT coding layer as a unit.

[0343] 4. Until the root node is reached, i.e. the node is downsampled, and the attribute encoding is stopped.

[0344] The coding unit to be encoded is taken as an example of a coding group (CG).

[0345] The RAHT coding layer unit is divided to obtain different CGs. Similarly, a best prediction coding mode needs to be passed to reconstruct the attribute information of each node in the current CG. Here, the CG group can be adaptively divided, such as fixing the maximum number of nodes of each CG to N, dividing each N node of each RAHT coding layer into a CG group, and then selecting the best coding mode for each CG at the encoding end. The attribute encoding method of the coding group to be encoded can include:

[0346] 1. When encoding the attribute information of the current slice, the geometry information and attribute information of the current slice are obtained. The points of the current slice are downsampled using the same algorithm principle as the octree. Specifically, the geometry information is downsampled along the Z direction, Y direction and X direction once each time, to obtain a RAHT coding layer.

[0347] 2. After obtaining the RAHT coding layer, assuming that the number of nodes of the current RAHT coding layer is Q. Divide every N nodes of the current RAHT coding layer into a coding group CG, based on such a manner, obtain M coding groups.

[0348] 3. After dividing the current RAHT coding layer, obtain different coding groups, and perform mode adaptive selection on the properties of different coding groups. Perform independent orthogonal transformation on the properties merged into a node, and the specific transformation kernel adopts the transformation matrix of the background "region adaptive hierarchical transformation". After the independent orthogonal transformation, a DC coefficient and a series of AC coefficients can be obtained. The DC coefficient of the current node is taken as the property information of the current node.

[0349] After obtaining the coding group, introduce the rate-distortion optimization algorithm to obtain the best coding mode of the current coding group, and pass the best coding mode to the decoding end, and the decoding end uses the obtained decoding mode to reconstruct and recover the property information of the nodes of the current to-be-decoded group. In the rate-distortion optimization algorithm, first, the distortion D of the reconstructed property and the original property of each prediction mode is calculated, and then the code stream R required for encoding of each prediction mode is obtained, and the rate-distortion value J determines the best coding mode.

[0350] 4. Repeat step 1 to downsample the point geometry information of the current slice, and downsample the geometry information along the Z direction, the Y direction and the X direction once each time, that is, obtain a RAHT coding layer unit, and then divide the RAHT coding layer into coding groups, and transform and encode the property information in units of coding groups.

[0351] 5. Until the root node is reached, that is, the attribute coding is stopped.

[0352] Further, when the inter-frame RAHT prediction of the property is performed, if the current to-be-coded layer can be predicted, four coding modes will be introduced for the current to-be-coded layer first, and then the rate-distortion optimization algorithm is used to select the best prediction coding mode for prediction coding, so that the coding efficiency of the point cloud property can be improved. As shown in Table 6, it can be seen that for the sequence that can use the inter-frame prediction of the property, on the qnxadas-junction-approach sequence, the attribute coding Bd-rate is improved by about 7.6%, which significantly improves the coding efficiency of the point cloud property.

[0353] Table 6

[0354] In still another embodiment of the present application, based on the same inventive concept as the foregoing embodiments, referring to FIG. 27, a schematic diagram of a composition structure of an encoder provided in an embodiment of the present application is shown. As shown in FIG. 27, the encoder 110 can include a first prediction unit 111, a first determination unit 112, and an encoding unit 113; wherein,

[0355] The first prediction unit 111 is configured to perform attribute encoding on a node of a to-be-encoded unit based on a candidate encoding mode of the to-be-encoded unit, and determine an attribute reconstruction value of the node of the to-be-encoded unit.

[0356] The first determination unit 112 is configured to perform encoding decision on the candidate encoding mode based on the attribute reconstruction value of the node of the to-be-encoded unit, determine an encoding mode of the to-be-encoded unit, and determine a value of a first syntax element of the to-be-encoded unit based on the encoding mode of the to-be-encoded unit.

[0357] The encoding unit 113 is configured to perform encoding processing on the first syntax element of the to-be-encoded unit, and write the obtained encoding bits into a bitstream.

[0358] It can be understood that the functional units of the encoder also perform the encoding method of any one of the foregoing embodiments, which will not be described here.

[0359] It can be understood that in the embodiments of the present application, the "unit" can be a part of circuit, a part of processor, a part of program or software, etc., and of course can also be a module, and can also be non-modular. Moreover, the constituent parts in the embodiments can be integrated in one processing unit, or can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function module.

[0360] When the integrated unit is realized in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments can be embodied in the form of a software product, and the computer software product is stored in a storage medium, includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the embodiments. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0361] Therefore, the embodiment of the present application provides a computer readable storage medium applied to the encoder 110, the computer readable storage medium stores a computer program, and the computer program is executed by the first processor to implement the method in any one of the foregoing embodiments.

[0362] Based on the components of the encoder 110 and the computer readable storage medium, referring to FIG. 28, a specific hardware structure schematic diagram of the encoder 110 provided by the embodiment of the present application is shown. As shown in FIG. 28, the encoder 110 can include a first memory 115 and a first processor 116, a first communication interface 117 and a first bus system 118. The first memory 115, the first processor 116 and the first communication interface 117 are coupled together through the first bus system 118. It can be understood that the first bus system 118 is used to realize the connection communication between the components. In addition to the data bus, the first bus system 118 also includes a power bus, a control bus and a status signal bus. However, in order to clearly illustrate, all kinds of buses are marked as the first bus system 118 in FIG. 20. Among them,

[0363] The first communication interface 117 is used for receiving and sending signals in the process of transmitting information with other external network elements;

[0364] The first memory 115 is used for storing a computer program capable of running on the first processor;

[0365] Based on the candidate encoding mode of the to-be-encoded unit, the attribute of the node of the to-be-encoded unit is encoded, and the attribute reconstruction value of the node of the to-be-encoded unit is determined;

[0366] Based on the attribute reconstruction value of the node of the to-be-encoded unit, the candidate encoding mode is encoded and decided, and the encoding mode of the to-be-encoded unit is determined;

[0367] Based on the encoding mode of the to-be-encoded unit, the value of the first syntax element of the to-be-encoded unit is determined;

[0368] The first syntax element of the to-be-encoded unit is encoded and processed, and the obtained encoding bits are written into the code stream.

[0369] It is to be appreciated that the first memory 115 in the embodiments of the application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Where the nonvolatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically EPROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as external cache. By way of example, and not limitation, many forms of RAM are available, for example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The first memory 115 of the system and method described herein are intended to include, without being limited to, these and any other suitable types of memory.

[0370] The first processor 116 can be an integrated circuit chip with processing capability. In implementation, each step of the above method can be completed by integrated logic circuit of hardware or instruction in the form of software in the first processor 116. The first processor 116 described above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block diagram disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor to execute, or be executed by a combination of hardware and software modules in the code processor. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register, or other mature storage medium in the art. The storage medium is located in the first memory 115, and the first processor 116 reads the information in the first memory 115 and combines the hardware to complete the steps of the above method.

[0371] It can be understood that the embodiments described in the present application can be realized by hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for executing functions of the present application, or a combination thereof. For software implementation, the techniques of the present application can be implemented by modules (for example, procedures, functions, and so on) for performing functions of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0372] Alternatively, as another embodiment, the first processor 116 is further configured to execute the encoding method of any one of the preceding embodiments when running a computer program.

[0373] The embodiment provides an encoder, in which, in the encoder, the node attribute distribution characteristics of a to-be-encoded unit are fully considered, an optimal encoding mode is adaptively selected for the to-be-encoded unit, a first syntax element indicating the optimal encoding mode is written into a code stream, and the to-be-encoded unit is subjected to point cloud attribute reconstruction by using the optimal encoding mode, so that the point cloud attribute encoding efficiency can be improved.

[0374] The embodiment of the present application further provides a computer readable storage medium, which stores the code stream generated by the encoding method in any one of the foregoing embodiments.

[0375] The embodiment of the present application further provides a code stream, which is generated by bit encoding according to to-be-encoded information; wherein,

[0376] The code stream is generated by bit encoding according to to-be-encoded information; wherein, the to-be-encoded information comprises at least one of the following: the first syntax element, the second syntax element, the third syntax element, the fourth syntax element, the fifth syntax element and the sixth syntax element, and the coefficient residual;

[0377] The value of the first syntax element is used for indicating the decoding mode of the to-be-decoded unit;

[0378] The value of the second syntax element is used for indicating whether the to-be-decoded unit is allowed to enable the region adaptive hierarchical transform prediction;

[0379] The third syntax element is used for indicating whether the to-be-decoded unit is allowed to refer to the first inter-frame reference unit;

[0380] The fourth syntax element is used for indicating whether the to-be-decoded unit is allowed to refer to the second inter-frame reference unit;

[0381] The fifth syntax element is used for indicating the number of to-be-decoded units;

[0382] The sixth syntax element is used for indicating the node number N of the to-be-decoded group.

[0383] In still another embodiment of the present application, based on the same inventive concept of the foregoing embodiments, referring to FIG. 29, a composition structure schematic diagram of a decoder provided by the embodiment of the present application is shown. As shown in FIG. 29, the decoder 120 can comprise a decoding unit 121, a second determining unit 122 and a second prediction unit 123; wherein,

[0384] The decoding unit 121 is configured to parse the code stream and determine the value of the first syntax element of the to-be-decoded unit;

[0385] The second determining unit 122 is configured to determine the decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit.

[0386] The second prediction unit 123 is configured to perform attribute decoding on the node of the to-be-decoded unit based on the determined decoding mode, and determine an attribute reconstruction value of the node of the to-be-decoded unit.

[0387] It can be understood that each functional unit of the decoder also performs the decoding method of any one of the preceding embodiments, which will not be described here.

[0388] It can be understood that in the present embodiment, the "unit" can be a partial circuit, a partial processor, a partial program or software, etc., and of course can also be a module, and can also be non-modular. Moreover, each component in the present embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function module.

[0389] If the integrated unit is realized in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the present embodiment provides a computer readable storage medium applied to the decoder 120, and the computer readable storage medium stores a computer program, and the computer program is executed by the second processor to implement the method of any one of the preceding embodiments.

[0390] Based on the components of the decoder 120 and the computer readable storage medium, referring to FIG. 30, a specific hardware structure schematic diagram of the decoder 120 provided by the present embodiment is shown. As shown in FIG. 30, the decoder 120 can include a second memory 127 and a second processor 124, a second communication interface 125 and a second bus system 126. The second memory 127 and the second processor 124, the second communication interface 125 are coupled together through the second bus system 126. It can be understood that the second bus system 126 is used to realize the connection communication between these components. In addition to including a data bus, the second bus system 126 also includes a power bus, a control bus and a state signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the second bus system 126 in FIG. 30. Among them,

[0391] The second communication interface 125 is configured to receive and send signals in the process of transceiving information with other external network elements.

[0392] The second memory 127 is configured to store a computer program capable of running on the second processor.

[0393] In some embodiments, the second processor 124 is configured to, when running the computer program, perform:

[0394] The code stream is parsed to determine a value of a first syntax element of the to-be-decoded unit;

[0395] Based on the value of the first syntax element of the to-be-decoded unit, a decoding mode of the to-be-decoded unit is determined.

[0396] Based on the determined decoding mode, attribute decoding is performed on the node of the to-be-decoded unit to determine a property reconstruction value of the node of the to-be-decoded unit.

[0397] Optionally, as another embodiment, the second processor 124 is further configured to execute the method in any one of the preceding embodiments when running the computer program.

[0398] It can be understood that the second memory 127 has a similar hardware function to the first memory 115, and the second processor 124 has a similar hardware function to the first processor 116; and details are not described herein.

[0399] The embodiment provides a decoder, in which a first syntax element is determined from a code stream to determine a decoding mode of a to-be-decoded unit, the decoding mode is a best decoding mode adaptively selected by fully considering node property distribution characteristics of the to-be-decoded unit, and the to-be-decoded unit is subjected to point cloud property reconstruction by using the decoding mode, so that the point cloud property decoding efficiency can be improved.

[0400] In still another embodiment of the present application, referring to FIG. 31, a schematic diagram of a composition structure of a coding system provided in the embodiment of the present application is shown. As shown in FIG. 31, the coding system 130 can include an encoder 131 and a decoder 132.

[0401] In the embodiment of the present application, the encoder 131 can be any one of the encoders in the preceding embodiments, and the decoder 132 can be any one of the decoders in the preceding embodiments.

[0402] It should be noted that, in the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement “comprises a” does not exclude the presence of another identical element in the process, method, article or device that includes the element.

[0403] The serial numbers of the above embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0404] The methods disclosed in the several method embodiments of the present application can be combined arbitrarily without conflict to obtain new method embodiments.

[0405] The features disclosed in several product embodiments of the present application can be combined arbitrarily without conflict, to obtain new product embodiments.

[0406] The features disclosed in several method or device embodiments of the present application can be combined arbitrarily without conflict, to obtain new method embodiments or device embodiments.

[0407] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims. Industrial applicability

[0408] In a point cloud coding method, a code stream, an encoder, a decoder and a storage medium in an embodiment of the present application, at the encoding end, the node attribute distribution characteristics of the to-be-encoded unit are fully considered, the optimal coding mode is adaptively selected for the to-be-encoded unit, the first syntax element indicating the optimal coding mode is written into the code stream, at the decoding end, the first syntax element is determined from the code stream, thereby determining the decoding mode of the to-be-decoded unit, at the encoding and decoding end, the point cloud attribute reconstruction is performed by using the decoding unit level coding and decoding mode, which can improve the point cloud attribute decoding efficiency.

Claims

1. A point cloud decoding method applied to a decoder, the method comprising: parsing a bitstream to determine a value of a first syntax element of a to-be-decoded unit; determining a decoding manner of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit; performing attribute decoding on a node of the to-be-decoded unit based on the determined decoding manner to determine an attribute reconstruction value of the node of the to-be-decoded unit.

2. The method of claim 1, wherein, The method further comprises: determining the value of the first syntax element of the to-be-decoded unit in a case where it is determined that the to-be-decoded unit is allowed to enable attribute prediction.

3. The method of claim 2, wherein, The method further comprises: determining a value of a second syntax element; determining that the to-be-decoded unit is allowed to enable attribute prediction in a case where it is determined that the to-be-decoded unit is allowed to enable region adaptive hierarchical transform prediction decoding based on the value of the second syntax element.

4. The method according to any one of claims 1 to 3, wherein, The method further comprises: determining the value of the first syntax element of the to-be-decoded unit in a case where it is determined that the to-be-decoded unit is allowed to enable attribute inter prediction.

5. The method of claim 4, wherein, The method further comprises: determining a value of a third syntax element and / or a value of a fourth syntax element; determining that the to-be-decoded unit is allowed to enable attribute inter prediction in a case where it is determined that the to-be-decoded unit is allowed to refer to a first inter reference unit based on the value of the third syntax element and / or the to-be-decoded unit is allowed to refer to a second inter reference unit based on the value of the fourth syntax element.

6. The method according to any one of claims 1 to 5, wherein, When the to-be-decoded unit is a to-be-decoded layer or a to-be-decoded group, the parsing a bitstream to determine a value of a first syntax element of a to-be-decoded unit comprises: parsing a bitstream to determine a value of a fifth syntax element, the fifth syntax element being used to indicate a number of to-be-decoded units; determining the value of the first syntax element of the to-be-decoded unit based on the value of the fifth syntax element.

7. The method of claim 1, wherein, When the to-be-decoded unit is a to-be-decoded group, the method further comprises: parsing a bitstream to determine a value of a sixth syntax element of the to-be-decoded group; determining a node number N of the to-be-decoded group based on the value of the sixth syntax element, N being a positive integer; performing node division on a to-be-decoded layer based on the node number N of the to-be-decoded group to determine a to-be-decoded group.

8. The method of claim 7, wherein, The performing node division on a to-be-decoded layer based on the node number N of the to-be-decoded group to determine a to-be-decoded group comprises: dividing every N nodes into a to-be-decoded group based on the node number N of the to-be-decoded group and a node order of the to-be-decoded layer, starting from a first node; if there are less than N nodes left in the to-be-decoded layer, dividing the less than N nodes left into a to-be-decoded group.

9. The method according to any one of claims 1 to 8, wherein, The decoding manner comprises a decoding order of one or more decoding modes.

10. The method of claim 9, wherein, The performing attribute decoding on a node of the to-be-decoded unit based on the determined decoding manner to determine an attribute reconstruction value of the node of the to-be-decoded unit comprises: determining a current decoding mode based on the decoding order of the determined decoding manner; determining whether the current decoding mode is available for the node of the to-be-decoded unit; In a case that the current decoding mode is available for the node of the to-be-decoded unit, performing attribute decoding on the node of the to-be-decoded unit based on the current decoding mode, and determining an attribute reconstruction value of the node of the to-be-decoded unit; In a case that the current decoding mode is not available for the node of the to-be-decoded unit, determining a next decoding mode based on the determined decoding mode, and continuing to judge whether the next decoding mode is available for the node of the to-be-decoded unit.

11. The method of claim 9, wherein, The attribute decoding on the node of the to-be-decoded unit based on the determined decoding mode, and the determination of the attribute reconstruction value of the node of the to-be-decoded unit, include: Determining the decoding mode of the to-be-decoded unit based on the determined decoding mode; Performing attribute decoding on the node of the to-be-decoded unit based on the determined decoding mode, and determining the attribute reconstruction value of the node of the to-be-decoded unit.

12. The method according to any one of claims 9 to 11, wherein, The determination of the decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit includes: In a case that the value of the first syntax element of the to-be-decoded unit is a first numerical value, it is determined that the decoding order corresponding to the decoding mode of the to-be-decoded unit is a first inter-prediction transform mode, a second inter-prediction transform mode, a third inter-prediction transform mode, an intra-prediction transform mode and a transform mode in sequence.

13. The method according to any one of claims 9 to 12, wherein, The determination of the decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit includes: In a case that the value of the first syntax element of the to-be-decoded unit is a second numerical value, it is determined that the decoding order corresponding to the decoding mode of the to-be-decoded unit is a second inter-prediction transform mode, an intra-prediction transform mode and a transform mode in sequence.

14. The method according to any one of claims 9 to 13, wherein, The determination of the decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit includes: In a case that the value of the first syntax element of the to-be-decoded unit is a third numerical value, it is determined that the decoding order corresponding to the decoding mode of the to-be-decoded unit is a third inter-prediction transform mode, an intra-prediction transform mode and a transform mode in sequence.

15. The method according to any one of claims 9 to 14, wherein, The determination of the decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit includes: In a case that the value of the first syntax element of the to-be-decoded unit is a fourth numerical value, it is determined that the decoding order corresponding to the decoding mode of the to-be-decoded unit is an intra-prediction transform mode and a transform mode in sequence.

16. The method according to any one of claims 9 to 15, wherein, The determination of the decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit includes: In a case that the value of the first syntax element of the to-be-decoded unit is a fifth numerical value, it is determined that the decoding order corresponding to the decoding mode of the to-be-decoded unit is a transform mode.

17. The method of any one of claims 9 to 16, wherein, The determination of the decoding mode of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit includes: In a case where the first syntax element of the to-be-decoded unit has a sixth value, it is determined that the decoding order corresponding to the decoding mode of the to-be-decoded unit is in the order of inter-intra fusion prediction transform mode, first inter prediction transform mode, second inter prediction transform mode, third inter prediction transform mode, intra prediction transform mode, and transform mode.

18. The method of any one of claims 12, 13, 14, or 17, wherein, The first inter prediction transform mode is a bi-directional inter prediction transform mode, the second inter prediction transform mode is a forward inter prediction transform mode, and the third inter prediction transform mode is a backward inter prediction transform mode.

19. A point cloud encoding method applied to an encoder, the method comprising: performing attribute coding on a node of the to-be-encoded unit based on a candidate coding mode of the to-be-encoded unit, to determine an attribute reconstruction value of the node of the to-be-encoded unit; performing coding decision on the candidate coding mode based on the attribute reconstruction value of the node of the to-be-encoded unit, to determine a coding mode of the to-be-encoded unit; determining a value of a first syntax element of the to-be-encoded unit based on the coding mode of the to-be-encoded unit; performing coding processing on the first syntax element of the to-be-encoded unit, and writing obtained coding bits into a bitstream.

20. The method of claim 19, wherein, The method further comprises: in a case where it is determined that the to-be-encoded unit is allowed to enable attribute prediction, performing the attribute coding on the node of the to-be-encoded unit based on the candidate coding mode of the to-be-encoded unit, to determine the attribute reconstruction value of the node of the to-be-encoded unit.

21. The method of claim 20, wherein, The method further comprises: determining a value of a second syntax element; in a case where it is determined that the to-be-encoded unit is allowed to enable region adaptive hierarchical transform prediction coding based on the value of the second syntax element, determining that the to-be-encoded unit is allowed to enable attribute prediction; performing coding processing on the second syntax element, and writing obtained coding bits into a bitstream.

22. The method of any one of claims 19 to 21, wherein, The method further comprises: in a case where it is determined that the to-be-encoded unit is allowed to enable attribute inter prediction, performing the attribute coding on the node of the to-be-encoded unit based on the candidate coding mode of the to-be-encoded unit, to determine the attribute reconstruction value of the node of the to-be-encoded unit.

23. The method of claim 22, wherein, The method further comprises: determining a value of a third syntax element and / or a value of a fourth syntax element; in a case where it is determined that the to-be-encoded unit is allowed to refer to a first inter reference unit based on the value of the third syntax element, and / or the to-be-encoded unit is allowed to refer to a second inter reference unit based on the value of the fourth syntax element, determining that the to-be-encoded unit is allowed to enable attribute inter prediction; performing coding processing on the third syntax element and / or the fourth syntax element, and writing obtained coding bits into a bitstream.

24. The method of any one of claims 19 to 23, wherein, In a case where the to-be-encoded unit is a to-be-encoded layer or a to-be-encoded group, the method further comprises: determining a value of a fifth syntax element based on a number of to-be-encoded units; performing coding processing on the fifth syntax element, and writing obtained coding bits into a bitstream.

25. The method of claim 19, wherein, In a case where the to-be-encoded unit is a to-be-encoded group, the method further comprises: performing node division on a to-be-encoded layer based on a node number N of the to-be-encoded group, to determine a to-be-encoded group. determining a value of the sixth syntax element of the to-be-encoded group based on a number N of nodes of the to-be-encoded group; encoding the sixth syntax element, and writing a coding bit obtained by the encoding into a bitstream.

26. The method of claim 25, wherein, The node partitioning of the to-be-encoded layer based on the number N of nodes of the to-be-encoded group comprises: dividing every N nodes into a to-be-encoded group from a first node based on the number N of nodes of the to-be-encoded group and a node sequence of the to-be-encoded layer; if there are less than N nodes left in the to-be-encoded layer, the less than N nodes left are divided into a to-be-encoded group.

27. The method of any one of claims 19 to 26, wherein, The coding mode comprises a coding sequence of one or more coding modes.

28. The method of claim 27, wherein, The attribute coding of the node of the to-be-encoded unit based on the candidate coding mode of the to-be-encoded unit comprises: determining a current coding mode based on the coding sequence of the candidate coding mode; determining whether the current coding mode is available for the node of the to-be-encoded unit; in a case where the current coding mode is available for the node of the to-be-encoded unit, attribute coding of the node of the to-be-encoded unit based on the current coding mode is performed to determine an attribute reconstruction value of the node of the to-be-encoded unit; in a case where the current coding mode is not available for the node of the to-be-encoded unit, a next coding mode is determined based on the determined coding mode, and it is determined whether the next coding mode is available for the node of the to-be-encoded unit.

29. The method of claim 27, wherein, The attribute coding of the node of the to-be-encoded unit based on the candidate coding mode of the to-be-encoded unit comprises: determining a coding mode of a to-be-encoded unit based on the candidate coding mode; performing attribute coding of the node of the to-be-encoded unit based on the determined coding mode of the to-be-encoded unit to determine an attribute reconstruction value of the node of the to-be-encoded unit.

30. The method of any one of claims 27 to 29, wherein, The determining of the value of the first syntax element of the to-be-encoded unit based on the coding mode of the to-be-encoded unit comprises: if a coding sequence corresponding to the coding mode of the to-be-encoded unit is a first inter-prediction transform mode, a second inter-prediction transform mode, a third inter-prediction transform mode, an intra-prediction transform mode and a transform mode in sequence, the value of the first syntax element of the to-be-encoded unit is determined as a first value.

31. The method of any one of claims 27 to 30, wherein, The determining of the value of the first syntax element of the to-be-encoded unit based on the coding mode of the to-be-encoded unit comprises: if a coding sequence corresponding to the coding mode of the to-be-encoded unit is the second inter-prediction transform mode, the intra-prediction transform mode and the transform mode in sequence, the value of the first syntax element of the to-be-encoded unit is determined as a second value.

32. The method of any one of claims 27 to 31, wherein, The determining of the coding mode based on the value of the first syntax element of the to-be-encoded unit comprises: if a coding sequence corresponding to the coding mode of the to-be-encoded unit is the third inter-prediction transform mode, the intra-prediction transform mode and the transform mode in sequence, the value of the first syntax element of the to-be-encoded unit is determined as a third value.

33. The method of any one of claims 27 to 32, wherein, The determining of the coding mode based on the value of the first syntax element of the to-be-encoded unit comprises: The coding order corresponding to the coding mode of the to-be-encoded unit is in turn intra prediction transform mode and transform mode, and the value of the first syntax element of the to-be-encoded unit is determined as a fourth value.

34. The method of any one of claims 27 to 33, wherein, The determining the coding mode based on the value of the first syntax element of the to-be-encoded unit comprises: The coding order corresponding to the coding mode of the to-be-encoded unit is in turn transform mode, and the value of the first syntax element of the to-be-encoded unit is determined as a fifth value.

35. The method of any one of claims 27 to 34, wherein, The determining the coding mode based on the value of the first syntax element of the to-be-encoded unit comprises: The coding order corresponding to the coding mode of the to-be-encoded unit is in turn inter-intra fusion prediction transform mode, first inter prediction transform mode, second inter prediction transform mode, third inter prediction transform mode, intra prediction transform mode and transform mode, and the value of the first syntax element of the to-be-encoded unit is determined as a sixth value. The first inter prediction transform mode is bi-directional inter prediction transform mode, the second inter prediction transform mode is forward inter prediction transform mode, and the third inter prediction transform mode is backward inter prediction transform mode.

36. The method of any one of claims 30, 31, 32, or 35, wherein, The code stream is generated according to to-be-encoded information by bit encoding; wherein the to-be-encoded information comprises at least one of the following: first syntax element, second syntax element, third syntax element, fourth syntax element, fifth syntax element and sixth syntax element; 37. A bitstream, wherein, The value of the first syntax element is used to indicate the decoding mode of the to-be-decoded unit. The value of the second syntax element is used to indicate whether the to-be-decoded unit is allowed to enable region adaptive hierarchical transform prediction; The third syntax element is used to indicate whether the to-be-decoded unit is allowed to refer to a first inter reference unit; The fourth syntax element is used to indicate whether the to-be-decoded unit is allowed to refer to a second inter reference unit; The fifth syntax element is used to indicate the number of to-be-decoded units. The sixth syntax element is used to indicate the node number N of a to-be-decoded group.

38. An encoder comprising a first prediction unit, a first determination unit and an encoding unit; wherein: The first prediction unit is configured to perform attribute encoding on a node of a to-be-encoded unit based on a candidate coding mode of the to-be-encoded unit, and determine an attribute reconstruction value of the node of the to-be-encoded unit; The first determination unit is configured to perform coding decision on the candidate coding mode based on the attribute reconstruction value of the node of the to-be-encoded unit, and determine the coding mode of the to-be-encoded unit; The value of the first syntax element of the to-be-encoded unit is determined based on the coding mode of the to-be-encoded unit; The encoding unit is configured to perform encoding processing on the first syntax element of the to-be-encoded unit, and write the obtained encoding bits into a code stream.

39. An encoder comprising a first memory and a first processor; wherein: The first memory is configured to store a computer program capable of running on the first processor; The first processor is configured to execute the method according to any one of claims 19 to 36 when running the computer program. ​ 40. A decoder comprising a decoding unit, a second determining unit and a second predicting unit; wherein: the decoding unit is configured to parse a bitstream and determine a value of a first syntax element of a to-be-decoded unit; the second determining unit is configured to determine a decoding manner of the to-be-decoded unit based on the value of the first syntax element of the to-be-decoded unit; and the second predicting unit is configured to perform attribute decoding on a node of the to-be-decoded unit based on the determined decoding manner, and determine an attribute reconstruction value of the node of the to-be-decoded unit.

41. A decoder comprising a second memory and a second processor; wherein: the second memory is configured to store a computer program capable of running on the second processor; and the second processor is configured to execute a method according to any one of claims 1 to 18 when running the computer program. The computer readable storage medium stores a bitstream generated by the encoding method according to any one of claims 19 to 36. The computer readable storage medium stores a computer program which, when executed, implements the method according to any one of claims 1 to 18, or implements the method according to any one of claims 19 to 36. ​ ​ ​ 42. A computer readable storage medium, wherein, ​ 43. A computer readable storage medium, wherein, ​

Citation Information

Patent Citations

  • Data encoding method and device, data decoding method and device, and storage medium

    CN112449754A

  • Attribute residual encoding in G-PCC

    CN115699771A

  • Improvements in attribute layers and indications in point cloud coding

    CN116708799A

  • Image encoding / decoding method and apparatus using syntax combination, and method for transmitting bitstream

    WO2021025416A1

  • Inter prediction coding for geometry point cloud compression

    WO2022147100A1