Encoding method, decoding method, encoder, decoder and storage medium

By comprehensively considering the neighborhood geometry and attribute distribution characteristics of the node during RAHT encoding or decoding, and selecting the best codec mode, the problem of low RAHT attribute coding efficiency in the prior art is solved, and more efficient point cloud codec performance is achieved.

WO2025076668A9PCT designated stage expired Publication Date: 2025-06-12GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/123639
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-09
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The existing G-PCC attribute RAHT encoding is less efficient intra-frame encoding because it does not encode with the distribution characteristics of each node attribute itself.

Method used

When encoding or decoding RAHT, a new encoding and decoding mode is introduced, taking into account the neighborhood geometric distribution characteristics and attribute distribution characteristics of the current layer nodes, and choosing the best encoding and decoding mode to improve the RAHT attribute encoding and decoding efficiency.

Benefits of technology

By comprehensively considering the neighborhood geometry and attribute distribution characteristics of nodes, the efficiency of RAHT attribute encoding and decoding is improved and the encoding and decoding performance of point clouds is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023123639_12062025_PF_FP_ABST
    Figure CN2023123639_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are an encoding method, a decoding method, an encoder, a decoder and a storage medium. The decoding method comprises: in case of initiating mode selection when determining to perform region adaptive hierarchical transform decoding at a current layer, determining neighborhood geometric distribution information of a current layer node; when the neighborhood geometric distribution information satisfies a first prediction condition, determining neighborhood attribute distribution information of the current layer node; according to the neighborhood attribute distribution information, determining a target decoding mode of the current layer node from candidate decoding modes; and, according to the target decoding mode, performing attribute decoding on the current layer node, so as to determine an attribute reconstructed value of the current layer node. The present application performs mode selection by comprehensively taking into account neighborhood geometric distribution characteristics and neighborhood attribute distribution characteristics of current layer nodes, so as to provide optimal encoding and decoding modes for current nodes, thus improving the RAHT attribute encoding and decoding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Coding and decoding method, encoder, decoder and storage medium Technical Field

[0001] The embodiments of the present application relate to the field of point cloud encoding and decoding technology, and in particular to an encoding and decoding method, an encoder, a decoder, and a storage medium. Background Art

[0002] In the geometry-based Point Cloud Compression (G-PCC) codec framework provided by the Moving Picture Experts Group (MPEG), the geometric information and attribute information of the point cloud are encoded separately. Among them, the attribute coding of G-PCC can include: Predicting Transform (PT), Lifting Transform (LT) and Region Adaptive Hierarchical Transform (RAHT). The first two predict the point cloud based on the generation order of the level of detail (LOD), while RAHT adaptively transforms the attribute information from bottom to top based on the construction level of the octree.

[0003] In G-PCC attribute RAHT coding, the syntax elements in the Attribute Parameters Set (APS) can be used to determine whether the current sequence adopts a predictive coding scheme to perform intra-frame prediction on the attributes. However, this attribute coding scheme does not consider the distribution characteristics of each node attribute itself for attribute coding, resulting in low intra-frame coding efficiency of point cloud attributes.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a coding and decoding method, an encoder, a decoder, and a storage medium, which can improve the coding efficiency of point cloud attributes and thereby improve the coding and decoding performance of point clouds.

[0006] The technical solution of the embodiment of the present application can be implemented as follows:

[0007] In a first aspect, an embodiment of the present application provides a decoding method, applied to a decoder, the method comprising:

[0008] Decoding the code stream and determining a first syntax element identifier;

[0009] If it is determined according to the first syntax element identifier that the mode selection is started when the current layer performs region adaptive layered transform decoding, determining the neighborhood geometric distribution information of the current layer node;

[0010] In the case where the neighborhood geometric distribution information satisfies the first prediction condition, determining the neighborhood attribute distribution information of the current layer node;

[0011] Determining a target decoding mode for the current layer node from candidate decoding modes according to the neighborhood attribute distribution information; the candidate decoding modes include: an attribute prediction and transformation mode, and an attribute transformation mode;

[0012] Attribute decoding is performed on the current layer node according to the target decoding mode to determine an attribute reconstruction value of the current layer node.

[0013] In a second aspect, an embodiment of the present application provides an encoding method, applied to an encoder, the method comprising:

[0014] determining a first syntax element identifier;

[0015] If it is determined according to the first syntax element identifier that mode selection is enabled when region-adaptive layered transform coding is performed on the current layer, determining neighborhood geometric distribution information of nodes in the current layer;

[0016] In the case where the neighborhood geometric distribution information satisfies the first prediction condition, determining the neighborhood attribute distribution information of the current layer node;

[0017] Determining a target coding mode for the current layer node from candidate coding modes according to the neighborhood attribute distribution information; wherein the candidate decoding modes include: attribute prediction and transformation mode, and attribute transformation mode;

[0018] Performing attribute encoding on the current layer node according to the target coding mode to determine an attribute reconstruction value of the current layer node;

[0019] The first syntax element identifier is encoded, and the obtained encoded bits are written into a bitstream.

[0020] In a third aspect, an embodiment of the present application provides an encoder, the encoder comprising a first determining unit, a second determining unit, and an encoding unit; wherein,

[0021] The first determining unit is configured to determine a first syntax element identifier;

[0022] The first determining unit is further configured to determine neighborhood geometric distribution information of nodes in the current layer if it is determined according to the first syntax element identifier that mode selection is enabled when region-adaptive layered transform coding is performed on the current layer;

[0023] The first determining unit is further configured to determine the neighborhood attribute distribution information of the current layer node if the neighborhood geometric distribution information satisfies the first prediction condition;

[0024] The second determining unit is configured to determine the target coding mode of the current layer node from candidate coding modes according to the neighborhood attribute distribution information; wherein the candidate coding modes include: attribute prediction and transformation mode, attribute transformation mode;

[0025] The second determining unit is configured to perform attribute encoding on the current layer node according to the target coding mode to determine an attribute reconstruction value of the current layer node;

[0026] The encoding unit is configured to encode the first syntax element identifier and write the obtained encoded bits into a bitstream.

[0027] In a fourth aspect, an embodiment of the present application provides an encoder, comprising a first memory and a first processor; wherein,

[0028] a first memory for storing a computer program capable of running on the first processor;

[0029] The first processor is configured to execute the method according to the second aspect when running a computer program.

[0030] In a fifth aspect, an embodiment of the present application provides a decoder, comprising a decoding unit, a third determining unit, and a fourth determining unit; wherein,

[0031] The decoding unit is configured to decode the code stream and determine a first syntax element identifier;

[0032] The third determining unit is configured to determine the neighborhood geometric distribution information of the node in the current layer if it is determined according to the first syntax element identifier that the mode selection is started when the current layer performs region adaptive layered transform decoding;

[0033] The third determining unit is further configured to determine neighborhood attribute distribution information of the current layer node if the neighborhood geometric distribution information satisfies the first prediction condition;

[0034] The fourth determining unit is configured to determine a target decoding mode of the current layer node from candidate decoding modes according to the neighborhood attribute distribution information; wherein the candidate decoding modes include: attribute prediction and transformation mode, and attribute transformation mode;

[0035] The fourth determining unit is further configured to perform attribute decoding on the current layer node according to the target decoding mode to determine the attribute reconstruction value of the current layer node.

[0036] In a sixth aspect, an embodiment of the present application provides a decoder, the decoder comprising a second memory and a second processor; wherein,

[0037] a second memory for storing a computer program capable of running on the second processor;

[0038] The second processor is configured to execute the method according to the first aspect when running a computer program.

[0039] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a code stream generated by the encoding method as described.

[0040] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed, implements the method described in the first aspect or the method described in the second aspect.

[0041] The embodiments of the present application provide a coding and decoding method, an encoder, a decoder, and a storage medium. When performing adaptive layered transform (RAHT) encoding or decoding, one or more new coding and decoding modes are introduced, and the mode is selected by comprehensively considering the neighborhood geometric distribution characteristics and neighborhood attribute distribution characteristics of the current layer node to provide the optimal coding and decoding mode for the current node, thereby improving the RAHT attribute coding and decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] FIG1A is a schematic diagram of a three-dimensional point cloud image;

[0043] FIG1B is a partially enlarged view of a three-dimensional point cloud image;

[0044] FIG2A is a schematic diagram of six viewing angles of a point cloud image;

[0045] FIG2B is a schematic diagram of a data storage format corresponding to a point cloud image;

[0046] FIG3 is a schematic diagram of a network architecture for point cloud encoding and decoding;

[0047] FIG4A is a schematic diagram of a composition framework of a G-PCC encoder;

[0048] FIG4B is a schematic diagram of a composition framework of a G-PCC decoder;

[0049] FIG5A is a schematic diagram of a low plane position in the Z-axis direction;

[0050] FIG5B is a schematic diagram of a high plane position in the Z-axis direction;

[0051] FIG6 is a schematic diagram of a node coding sequence;

[0052] FIG7A is a schematic diagram of plane identification information;

[0053] FIG7B is a schematic diagram of another type of planar identification information;

[0054] FIG8 is a schematic diagram of sibling nodes of a current node;

[0055] Figure 9 is a schematic diagram of the intersection of a laser radar and a node;

[0056] FIG10 is a schematic diagram of neighborhood nodes at the same partition depth and the same coordinates;

[0057] FIG11 is a schematic diagram of a current node being located at a low plane position of a parent node;

[0058] FIG12 is a schematic diagram showing a current node being located at a high plane position of a parent node;

[0059] FIG13 is a schematic diagram of predictive coding of planar position information of a laser radar point cloud;

[0060] FIG14 is a schematic diagram of IDCM encoding;

[0061] FIG15 is a schematic diagram of coordinate transformation of a rotating laser radar to obtain a point cloud;

[0062] FIG16 is a schematic diagram of predictive coding in the X-axis or Y-axis direction;

[0063] FIG17A is a schematic diagram showing an angle of the Y plane predicted by the horizontal azimuth angle;

[0064] FIG17B is a schematic diagram showing an angle of the X-plane predicted by the horizontal azimuth angle;

[0065] FIG18 is another schematic diagram of predictive coding in the X-axis or Y-axis direction;

[0066] FIG19A is a schematic diagram of three intersection points included in a sub-block;

[0067] FIG19B is a schematic diagram of a triangular facet set fitted using three intersection points;

[0068] FIG19C is a schematic diagram of upsampling of a triangle face set;

[0069] FIG20 is a schematic diagram of a distance-based LOD construction process;

[0070] FIG21 is a schematic diagram of a visualization result of an LOD generation process;

[0071] FIG22 is a schematic diagram of an encoding process for attribute prediction;

[0072] FIG23 is a schematic diagram of the composition of a pyramid structure;

[0073] FIG24 is a schematic diagram showing the composition of another pyramid structure;

[0074] FIG25 is a schematic diagram of an LOD structure for inter-layer nearest neighbor search;

[0075] FIG26 is a schematic diagram of a nearest neighbor search structure based on spatial relationships;

[0076] FIG27A is a schematic diagram of a coplanar spatial relationship;

[0077] FIG27B is a schematic diagram of a coplanar and colinear spatial relationship;

[0078] FIG27C is a schematic diagram of a spatial relationship of coplanarity, colinearity, and copointness;

[0079] FIG28 is a schematic diagram of inter-layer prediction based on fast search;

[0080] FIG29 is a schematic diagram of an LOD structure for nearest neighbor search within an attribute layer;

[0081] FIG30 is a schematic diagram of intra-layer prediction based on fast search;

[0082] FIG31A is a schematic diagram of attribute inter-frame prediction based on fast search;

[0083] FIG31B is a schematic diagram of a block-based neighborhood search structure;

[0084] FIG32 is a schematic diagram of an encoding process of a lifting transform;

[0085] FIG33 is a schematic diagram of a RAHT transformation structure;

[0086] FIG34 is a schematic diagram of a RAHT transformation process along the x, y, and z directions;

[0087] FIG35A is a schematic diagram of a RAHT forward transformation process;

[0088] FIG35B is a schematic diagram of a RAHT inverse transformation process;

[0089] FIG36 is a schematic diagram of an attribute coding block;

[0090] FIG37 is a schematic diagram of the principle of attribute prediction transform coding based on RAHT;

[0091] FIG38 is a schematic diagram of a neighborhood prediction relationship for attribute prediction;

[0092] FIG39 is a schematic diagram of a flowchart of a decoding method provided in an embodiment of the present application;

[0093] FIG40 is a schematic diagram of a RAHT coding layer;

[0094] FIG41 is a schematic diagram of a flow chart of an encoding method provided in an embodiment of the present application;

[0095] FIG42 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application;

[0096] FIG43 is a schematic diagram of a specific hardware structure of an encoder provided in an embodiment of the present application;

[0097] FIG44 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application;

[0098] FIG45 is a schematic diagram of a specific hardware structure of a decoder provided in an embodiment of the present application;

[0099] Figure 46 is a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0100] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.

[0101] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0102] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0103] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0104] Point cloud is a three-dimensional representation of the surface of an object. Point cloud (data) of the surface of an object can be collected through acquisition equipment such as photoelectric radar, lidar, laser scanner, and multi-view camera.

[0105] A point cloud is a set of irregularly distributed discrete points in space that express the spatial structure and surface properties of a three-dimensional object or scene. Figure 1A shows a three-dimensional point cloud image and Figure 1B shows a partially enlarged view of the three-dimensional point cloud image. It can be seen that the point cloud surface is composed of densely distributed points.

[0106] In a two-dimensional image, each pixel contains information and is distributed regularly, so there's no need to record its location. However, the distribution of points in a point cloud in three-dimensional space is random and irregular, so recording the location of each point in space is necessary to fully represent the point cloud. Similar to a two-dimensional image, each location in the acquisition process has corresponding attribute information, typically an RGB color value, which reflects the object's color. For a point cloud, in addition to color information, each point's corresponding attribute information often includes reflectance values, which reflect the surface texture of the object. Therefore, point cloud data typically includes both point location information and point attribute information. Point location information can also be referred to as point geometric information. For example, point geometric information can be the point's three-dimensional coordinates (x, y, z). Point attribute information can include color information and / or reflectance. For example, reflectance can be one-dimensional reflectance information (r). Color information can be information in any color space, or it can be three-dimensional color information, such as RGB. Here, R represents red (R), G represents green (G), and B represents blue (B). For another example, the color information may be luminance and chrominance (YCbCr, YUV) information, where Y represents brightness (Luma), Cb (U) represents blue color difference, and Cr (V) represents red color difference.

[0107] For example, a point cloud generated using laser measurement principles can include both its 3D coordinate information and its reflectivity. For another example, a point cloud generated using photogrammetry principles can include both its 3D coordinate information and its 3D color information. For another example, a point cloud generated using a combination of laser measurement and photogrammetry principles can include both its 3D coordinate information, its reflectivity value, and its 3D color information.

[0108] Figures 2A and 2B show a point cloud image and its corresponding data storage format. Figure 2A provides six viewing angles of the point cloud image, while Figure 2B consists of a file header and data. The header includes the data format, data representation type, the total number of points in the point cloud, and the content represented by the point cloud. For example, the point cloud is in ".ply" format, represented by ASCII code, with a total of 207,242 points. Each point has 3D coordinate information (x, y, z) and 3D color information (r, g, b).

[0109] Point clouds can be divided into the following categories according to the acquisition method:

[0110] Static point cloud: the object is stationary and the device that obtains the point cloud is also stationary;

[0111] Dynamic point cloud: The object is moving, but the device that obtains the point cloud is stationary;

[0112] Dynamic point cloud acquisition: The device used to acquire the point cloud is in motion.

[0113] For example, point clouds can be divided into two categories according to their usage:

[0114] Category 1: Machine perception point cloud, which can be used in scenarios such as autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, and disaster relief robots;

[0115] Category 2: Human eye perception point cloud, which can be used in point cloud application scenarios such as digital cultural heritage, free viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.

[0116] Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes. Moreover, since point clouds are obtained by directly sampling real objects, they can provide a strong sense of reality while ensuring accuracy. Therefore, they are widely used, including virtual reality games, computer-aided design, geographic information systems, automatic navigation systems, digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive remote presentation, and three-dimensional reconstruction of biological tissues and organs.

[0117] Point clouds are primarily collected through computer generation, 3D laser scanning, and 3D photogrammetry. Computers can generate point clouds of virtual 3D objects and scenes; 3D laser scanning can obtain point clouds of static real-world 3D objects or scenes, generating millions of point clouds per second; and 3D photogrammetry can obtain point clouds of dynamic real-world 3D objects or scenes, generating tens of millions of point clouds per second. These technologies reduce the cost and time required to acquire point cloud data while improving data accuracy. While changes in point cloud data acquisition methods have made it possible to acquire large amounts of point cloud data, the processing of this massive amount of 3D point cloud data is facing bottlenecks due to storage space and transmission bandwidth constraints, as application demands grow.

[0118] For example, taking a point cloud video with a frame rate of 30 frames per second (fps), each frame contains 700,000 points, and each point has coordinate information (xyz, float) and color information (RGB, uchar). Therefore, the data volume of a 10-second point cloud video is approximately 0.7 million × (4 bytes × 3 + 1 byte × 3) × 30 fps × 10 seconds = 3.15 GB, where 1 byte is 10 bits. For a 1280 × 720 2D video with a YUV sampling format of 4:2:0 and a frame rate of 24 fps, the data volume for 10 seconds is approximately 1280 × 720 × 12 bits × 24 fps × 10 seconds ≈ 0.33 GB. The data volume of a 10-second two-view 3D video is approximately 0.33 × 2 = 0.66 GB. This shows that the data volume of a point cloud video far exceeds that of 2D and 3D videos of the same length. Therefore, in order to better realize data management, save server storage space, and reduce the transmission traffic and transmission time between the server and the client, point cloud compression has become a key issue in promoting the development of the point cloud industry.

[0119] That is to say, since the point cloud is a collection of massive points, storing the point cloud not only consumes a lot of memory, but is also not conducive to transmission. There is also not enough bandwidth to support direct transmission of the point cloud at the network layer without compression. Therefore, the point cloud needs to be compressed.

[0120] Currently, the point cloud coding framework that can compress point clouds can be the geometry-based Point Cloud Compression (G-PCC) codec framework or the video-based Point Cloud Compression (V-PCC) codec framework provided by the Moving Picture Experts Group (MPEG), or the AVS-PCC codec framework provided by AVS. The G-PCC codec framework can be used to compress the first type of static point clouds and the third type of dynamically acquired point clouds, which can be based on the Point Cloud Compression Test Platform (Test Model Compression 13, TMC13). The V-PCC codec framework can be used to compress the second type of dynamic point clouds, which can be based on the Point Cloud Compression Test Platform (Test Model Compression 2, TMC2). Therefore, the G-PCC codec framework is also called the point cloud codec TMC13, and the V-PCC codec framework is also called the point cloud codec TMC2.

[0121] An embodiment of the present application provides a network architecture of a point cloud encoding and decoding system including a decoding method and an encoding method. FIG3 is a schematic diagram of a network architecture of a point cloud encoding and decoding system provided by an embodiment of the present application. As shown in FIG3 , the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01. During the implementation process, the electronic device can be various types of devices with point cloud encoding and decoding functions. For example, the electronic device can include a mobile phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensor device, a server, etc., which is not limited by the embodiment of the present application. Among them, the decoder or encoder in the embodiment of the present application can be the above-mentioned electronic device.

[0122] Among them, the electronic device in the embodiment of the present application has a point cloud encoding and decoding function, generally including a point cloud encoder (ie, encoder) and a point cloud decoder (ie, decoder).

[0123] The following describes the related technologies using the G-PCC encoding and decoding framework as an example.

[0124] As you can understand, in the point cloud G-PCC codec framework, the point cloud data to be encoded is first divided into multiple slices through slice partitioning. In each slice, the geometric information of the point cloud and the attribute information corresponding to each point are encoded separately.

[0125] Figure 4A shows a schematic diagram of the G-PCC encoder's architecture. As shown in Figure 4A, during the geometry encoding process, the geometric information is transformed so that the entire point cloud is contained within a bounding box. Quantization is then performed. This quantization step primarily serves a scaling purpose. Due to quantization rounding, the geometric information of some point clouds becomes identical. Parameters are then used to determine whether to remove duplicate points. This process of quantization and removing duplicate points is also known as voxelization. The bounding box is then partitioned into an octree or constructed as a prediction tree. During this process, arithmetic coding is performed on the points within the leaf nodes of the partition to generate a binary geometry bitstream. Alternatively, arithmetic coding is performed on the intersection points (vertices) generated by the partition (surface fitting is performed based on the intersection points) to generate a binary geometry bitstream. During the attribute encoding process, after the geometry encoding is completed and the geometry information is reconstructed, color conversion is performed to convert the color information (i.e., attribute information) from the RGB color space to the YUV color space. The reconstructed geometry information is then used to recolor the point cloud so that the unencoded attribute information corresponds to the reconstructed geometry information. Attribute encoding is mainly performed on color information. In the color information encoding process, there are two main transformation methods. One is the distance-based lifting transformation that relies on the level of detail (LOD) division, and the other is the direct region adaptive hierarchical transformation (RAHT). Both methods convert color information from the spatial domain to the frequency domain, obtain high-frequency coefficients and low-frequency coefficients through transformation, and finally quantize the coefficients. Then, the quantized coefficients are arithmetically encoded to generate a binary attribute bit stream.

[0126] Figure 4B shows a schematic diagram of the composition framework of a G-PCC decoder. As shown in Figure 4B, for the acquired binary bit stream, the geometric bit stream and attribute bit stream in the binary bit stream are first decoded independently. When decoding the geometric bit stream, the geometric information of the point cloud is obtained through arithmetic decoding-reconstruction of the octree / reconstruction of the prediction tree-reconstruction of the geometry-coordinate inverse conversion; when decoding the attribute bit stream, the attribute information of the point cloud is obtained through arithmetic decoding-inverse quantization-LOD partitioning / RAHT-color inverse conversion, and the point cloud data to be encoded (i.e., the output point cloud) is restored based on the geometric information and attribute information.

[0127] It should be noted that, as shown in FIG4A or FIG4B , the current geometric coding and decoding of G-PCC can be divided into octree-based geometric coding and decoding (marked by a dotted box) and prediction tree-based geometric coding and decoding (marked by a dotted box).

[0128] For Octree geometry encoding (OctGeomEnc), octree-based geometry encoding includes: first, coordinate conversion of the geometric information so that all point clouds are contained in a bounding box. Then quantization is performed. This step of quantization mainly plays a role of scaling. Due to the quantization rounding, the geometric information of some points is the same. The parameters are used to decide whether to remove duplicate points. The process of quantization and removal of duplicate points is also called voxelization. Next, the bounding box is continuously divided into trees (such as octrees, quadtrees, binary trees, etc.) in the order of breadth-first traversal, and the placeholder code of each node is encoded. In related technologies, a company proposed an implicit geometric division method. First, the bounding box of the point cloud is calculated. Assume d x >d y >d z , the bounding box corresponds to a cuboid. When geometrically partitioning, the binary tree partitioning will be performed based on the x-axis to obtain two child nodes; until d is satisfied x =d y >d z When the conditions are met, the quadtree partitioning will be performed based on the x and y axes to obtain four child nodes; when d is finally satisfied x =d y =d z When the conditions are met, the octree partitioning will continue until the leaf node obtained by the partitioning is a 1×1×1 unit cube. The partitioning will stop and the points in the leaf node will be encoded to generate a binary code stream. In the process of binary tree / quadtree / octree partitioning, two parameters are introduced: K and M. Parameter K indicates the maximum number of binary tree / quadtree partitions before octree partitioning; parameter M is used to indicate that the minimum block side length corresponding to binary tree / quadtree partitioning is 2 M . At the same time, K and M must meet the following conditions: Assume d max =max(d x ,d y ,d z ), d min =min(d x ,d y ,d z ), parameter K satisfies: K ≥ d max -d min ; Parameter M satisfies: M≥d minThe reason why the parameters K and M meet the above conditions is that in the current G-PCC geometric implicit partitioning process, the priority of the partitioning method is binary tree, quadtree and octree. When the node block size does not meet the conditions of binary tree / quadtree, the node will be divided into octree until the minimum unit of leaf node is 1×1×1. The octree-based geometric information coding mode can effectively encode the geometric information of the point cloud by utilizing the correlation between adjacent points in space. However, for some relatively flat nodes or nodes with planar characteristics, the coding efficiency of the point cloud geometric information can be further improved by using plane coding.

[0129] Exemplarily, Figure 5A and Figure 5B provide a kind of plane position schematic diagram.Wherein, Figure 5A shows a kind of low plane position schematic diagram in the Z-axis direction, and Figure 5B shows a kind of high plane position schematic diagram in the Z-axis direction.As shown in Figure 5A, here (a), (a0), (a1), (a2), (a3) ​​all belong to the low plane position in the Z-axis direction. Taking (a) as an example, it can be seen that the four child nodes occupied in the current node are all located at the low plane position of the current node in the Z-axis direction, so it can be considered that the current node belongs to a Z plane and is a low plane in the Z-axis direction. Similarly, as shown in Figure 5B, here (b), (b0), (b1), (b2), (b3) all belong to the high plane position in the Z-axis direction. Taking (b) as an example, it can be seen that the four child nodes occupied in the current node are located at the high plane position of the current node in the Z-axis direction, so it can be considered that the current node belongs to a Z plane and is a high plane in the Z-axis direction.

[0130] Taking (a) in Figure 5A as an example, the efficiency of octree encoding and plane encoding is compared. Figure 6 provides a schematic diagram of the node encoding sequence, that is, node encoding is performed in the order of 0, 1, 2, 3, 4, 5, 6, and 7 as shown in Figure 6. Here, if octree encoding is used for (a) in Figure 5A, the placeholder information of the current node is represented as: 11001100. However, if plane encoding is used, first, an identifier must be encoded to indicate that the current node is a plane in the Z-axis direction. Secondly, if the current node is a plane in the Z-axis direction, the plane position of the current node must also be represented. Finally, only the placeholder information of the lower plane node in the Z-axis direction needs to be encoded (i.e., the placeholder information of the four child nodes 0, 2, 4, and 6). Therefore, encoding the current node based on plane encoding only requires 6 bits, which can reduce the representation by 2 bits compared to the related art octree encoding. Based on this analysis, plane encoding has significantly higher coding efficiency than octree encoding. Therefore, for an occupied node, if a plane encoding method is used in a certain dimension, it is first necessary to represent the plane identification (planarMode) and plane position (PlanePos) information of the current node in that dimension, and then encode the occupancy information of the current node based on the plane information of the current node. For example, FIG7A shows a schematic diagram of plane identification information. As shown in FIG7A, there is a low plane in the Z-axis direction; correspondingly, the value of the plane identification information is true (true) or 1, that is, planarMode_ Z = true; the plane position information is the low plane (low), that is, PlanePosition_ Z =low. FIG7B shows another schematic diagram of plane identification information. As shown in FIG7B, here it is not a plane in the Z-axis direction; correspondingly, the value of the plane identification information is false or 0, that is, planarMode_ Z =false.

[0131] It should be noted that for PlaneMode_ i :0 means the current node is not a plane in the i-axis direction, 1 means the current node is a plane in the i-axis direction. If the current node is a plane in the i-axis direction, then for PlanePosition_ i : 0 means the current node is a plane in the i-axis direction and the plane position is low, 1 means the current node is a high plane in the i-axis direction. Here, i represents the coordinate dimension, which can be the X-axis direction, Y-axis direction, or Z-axis direction, so i = 0, 1, 2.

[0132] In the G-PCC standard, when determining whether a node meets the conditions for planar coding and when the node meets the conditions for planar coding, predictive coding of the planar identifier and planar position information of the node is required.

[0133] In the embodiments of the present application, there are three judgment conditions in the current G-PCC standard for determining whether a node meets the conditions for planar coding. The following is a detailed description of each of them.

[0134] First, judge according to the planar probability of the node in each dimension.

[0135] (1) Determine the local area density (local_node_density) of the current node;

[0136] (2) Determine the probability Prob(i) of the current node in each dimension.

[0137] When the local area density of the node is less than the threshold Th (for example, Th = 3), compare the planar probability Prob(i) of the current node in the three coordinate dimensions with the thresholds Th0, Th1, and Th2, where Th0 < Th1 < Th2 (for example, Th0 = 0.6, Th1 = 0.77, Th2 = 0.88). Here, Eligible i (i = 0, 1, 2) represents whether planar coding is started in each dimension: Eligible i = Prob(i) >= threshold.

[0138] It should be noted that the threshold is adaptively changed. For example, when Prob(0) > Prob(1) > Prob(2), the settings of Eligible i are as follows: Eligible0 = Prob(0) >= Th0; Eligible1 = Prob(1) >= Th1; Eligible2 = Prob(2) >= Th2.

[0139] When Prob(1) > Prob(0) > Prob(2), the settings of Eligible i are as follows: Eligible0 = Prob(0) >= Th1; Eligible1 = Prob(1) >= Th0; Eligible2 = Prob(2) >= Th2.

[0140] Here, the update of Prob(i) is specifically as follows: Prob(i) new = (L × Prob(i) + δ(coded node)) / L + 1 (1)

[0141] Where L = 255; in addition, if the coded node is a plane, δ(coded node) is 1; otherwise, δ(coded node) is 0.

[0142] Here, the update of local_node_density is as follows: local_node_density new =local_node_density+4*numSiblings (2)

[0143] Where local_node_density is initialized to 4, and numSiblings is the number of sibling nodes of the node. For example, FIG8 shows a schematic diagram of the sibling nodes of the current node. As shown in FIG8 , the current node is a node filled with slashes, and the nodes filled with grids are sibling nodes. Then, the number of sibling nodes of the current node is 5 (including the current node itself).

[0144] Second, determine whether the current layer nodes meet the plane coding requirements based on the point cloud density of the current layer.

[0145] The density of the current layer points is used to determine whether to perform planar coding on the nodes of the current layer. Assuming that the number of points in the current point cloud to be coded is pointCount, the number of points reconstructed by the inferred direct coding model (IDCM) coding is numPointCountRecon, and because the octree is coded based on the order of breadth-first traversal, the number of nodes to be coded in the current layer can be obtained as nodeCount. Then, the assumption to determine whether to start planar coding in the current layer is planarEligibleKOctreeDepth, specifically: planarEligibleK OctreeDepth = (pointCount-numPointCountRecon) <nodeCount×1.3。

[0146] Among them, if (pointCount-numPointCountRecon) is less than nodeCount×1.3, then planarEligibleK OctreeDepth is true; if (pointCount-numPointCountRecon) is not less than nodeCount×1.3, then planarEligibleKOctreeDepth is false. In this way, when planarEligibleKOctreeDepth is true, all nodes in the current layer are planar coded; otherwise, all nodes in the current layer are not planar coded and only octree coding is used.

[0147] 3. Determine whether the current node meets the plane coding requirements based on the acquisition parameters of the lidar point cloud.

[0148] Figure 9 shows a schematic diagram of the intersection of a laser radar and a node. As shown in Figure 9, a node filled with a grid is simultaneously traversed by two laser beams, so the current node is not a plane in the direction perpendicular to the Z axis. A node filled with a diagonal line is small enough to be traversed by two laser beams simultaneously, so it is possible that the node filled with a diagonal line is a plane in the direction perpendicular to the Z axis.

[0149] Furthermore, for nodes that meet the plane coding conditions, predictive coding may be performed on the plane identification information and the plane position information.

[0150] First, predictive coding of plane identification information.

[0151] Here, only three context information are used for encoding, that is, the plane identification in each coordinate dimension is designed separately for context.

[0152] Secondly, predictive coding of plane position information.

[0153] It should be understood that for the encoding of non-lidar point cloud planar position information, the predictive encoding of the planar position information may include:

[0154] (a) Using the occupancy information of neighboring nodes, the plane position information of the current node is predicted to be three elements: predicted as low plane, predicted as high plane, and unpredictable;

[0155] (b) Spatial distance between nodes at the same partition depth and coordinates as the current node and the current node: “near” and “far”;

[0156] (c) If the node at the same partition depth and the same coordinates as the current node is a plane, determine the plane position of the node;

[0157] (d) Coordinate dimension (i=0, 1, 2).

[0158] It should be noted that in an embodiment of the present application, after determining the spatial distance between the node at the same division depth and the same coordinates as the current node and the current node, if the spatial distance is less than the preset distance threshold, then the spatial distance can be determined to be "near"; or, if the spatial distance is greater than the preset distance threshold, then the spatial distance can be determined to be "far".

[0159] For example, Figure 10 shows a schematic diagram of neighboring nodes at the same partition depth and coordinates. As shown in Figure 10, the bold large cube represents the parent node, the small grid-filled cube inside it represents the current node, and the vertex position of the current node is shown. The small white-filled cube represents neighboring nodes at the same partition depth and coordinates. The distance between the current node and the neighboring node is the spatial distance, which can be judged as "near" or "far." In addition, if the neighboring node is a plane, the planar position of the neighboring node is also required.

[0160] In this way, as shown in Figure 10, the current node is a small cube filled with a grid, and the neighboring node is a small cube filled with white at the same octree partition depth level and the same vertical coordinate, and the distance between the two nodes is judged as "near" and "far", and the plane position of the reference node is used.

[0161] Furthermore, in an embodiment of the present application, FIG11 shows a schematic diagram of a current node being located at a low plane position of a parent node. As shown in FIG11 , (a), (b), and (c) show three examples of the current node being located at a low plane position of a parent node. Specific descriptions are as follows:

[0162] ① If any of the child nodes 4 to 7 of the point fill node is occupied, and all the grid fill nodes are not occupied, it is very likely that there is a plane in the current node (filled with a slash), and the plane is located lower.

[0163] ② If the child nodes 4 to 7 of the point fill node are not occupied, and any grid fill node is occupied, it is very likely that there is a plane in the current node (filled with diagonal lines), and the plane is located higher.

[0164] ③ If the child nodes 4 to 7 of the point fill node are all empty nodes and the grid fill nodes are all empty nodes, the plane position cannot be inferred and is therefore marked as unknown.

[0165] ④ If any of the child nodes 4 to 7 of the point fill node is occupied and any of the grid fill nodes is occupied, the plane position cannot be inferred at this time, so it is marked as unknown.

[0166] In an embodiment of the present application, FIG12 shows a schematic diagram of a current node being located at a high plane position of a parent node. As shown in FIG12, (a), (b), and (c) show three examples of the current node being located at a high plane position of a parent node. The specific description is as follows:

[0167] ① If any of the child nodes 4 to 7 of the grid fill node is occupied, and the point fill node is not occupied, it is very likely that there is a plane in the current node (filled with diagonal lines), and the plane position is low.

[0168] ② If the child nodes 4 to 7 of the grid fill node are not occupied, and the point fill node is occupied, it is very likely that there is a plane in the current node (filled with a slash), and the plane position is higher.

[0169] ③If the child nodes 4 to 7 of the grid fill node are all unoccupied, and the point fill node is unoccupied, the plane position cannot be inferred at this time, so it is marked as unknown.

[0170] ④ If one of the child nodes 4 to 7 of the grid fill node is occupied and the point fill node is occupied, the plane position cannot be inferred at this time and is therefore marked as unknown.

[0171] It should also be understood that, with respect to the coding of the laser radar point cloud plane position information, FIG13 shows a schematic diagram of the predictive coding of the laser radar point cloud plane position information. As shown in FIG13, when the laser radar emission angle is θ bottom When , it can be mapped to the bottom virtual plane; when the laser radar emission angle is θ top At this time, it can be mapped to the high plane (Top virtual plane).

[0172] That is, by using the laser radar acquisition parameters to predict the plane position of the current node, and by using the position where the current node intersects with the laser ray to quantize the position into multiple intervals, the final result is the context information of the plane position of the current node. The specific calculation process is as follows: Assume that the coordinates of the laser radar are (x Lidar ,y Lidar ,z Lidar ), the geometric coordinates of the current node are (x, y, z), then first calculate the vertical tangent value tanθ of the current node relative to the lidar, the calculation formula is as follows:

[0173] Furthermore, because each laser has a certain offset angle relative to the laser radar, it is also necessary to calculate the relative tangent value tanθ of the current node relative to the laser corr,L , the specific calculation is as follows:

[0174] Finally, the relative tangent value tanθ of the current node will be used corr,L To predict the plane position of the current node, as follows, assuming that the tangent value of the lower boundary of the current node is tan(θ bottom ), the tangent value of the upper boundary is tan(θ top ), according to tanθ corr,L The plane position is quantized into four quantization intervals, that is, the context information of the plane position is determined.

[0175] However, the octree-based geometric information coding mode only has an efficient compression rate for points that are correlated in space. For points that are isolated in the geometric space, the use of the Direct Coding Model (DCM) can greatly reduce the complexity. For all nodes in the octree, the use of DCM is not indicated by flag information, but is inferred from the parent node and neighbor information of the current node. There are three ways to determine whether the current node is eligible for DCM coding, as follows:

[0176] (1) The current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.

[0177] (2) The parent node of the current node has only one child node, the current node, and the six neighbor nodes that share a face with the current node are all empty nodes.

[0178] (3) The number of sibling nodes of the current node is greater than 1.

[0179] Exemplarily, FIG14 provides a schematic diagram of IDCM coding. If the current node does not have the DCM coding qualification, it will be divided into octrees. If it has the DCM coding qualification, the number of points contained in the node will be further determined. When the number of points is less than a threshold (for example, 2), the node will be DCM-encoded, otherwise the octree division will continue. When the DCM coding mode is applied, it is first necessary to encode whether the current node is a true isolated point, that is, IDCM_flag. When IDCM_flag is true, the current node is DCM-encoded, otherwise octree coding is still used. When the current node meets the DCM coding requirements, the DCM coding mode of the current node needs to be encoded. There are currently two DCM modes: (a) there is only one point (or multiple points, but they are duplicate points); (b) there are two points. Finally, the geometric information of each point needs to be encoded. Assuming that the side length of the node is 2 d When encoding each component of the node's geometric coordinates, d bits are required, and these bits are directly encoded into the bitstream. It is important to note that when encoding LiDAR point clouds, the efficiency of geometric information coding can be further improved by predictively encoding the three-dimensional coordinate information using LiDAR acquisition parameters.

[0180] Furthermore, the IDCM encoding process is described in detail below.

[0181] When the current node meets the DCM encoding mode, the number of points of the current node, numPoints, is encoded first; the number of points of the current node is encoded according to different DirectModes:

[0182] (1) If the current node does not meet the requirements of the DCM node, exit directly (that is, the number of points is greater than 2 points and is not a duplicate point).

[0183] (2) If the number of points numPonts in the current node is less than or equal to 2, the encoding process is as follows:

[0184] i) First encode whether the numPonts of the current node is greater than 1;

[0185] ii) If the current node has only one point and the geometry coding environment is geometry lossless coding, it is necessary to encode that the second point of the current node is not a duplicate point.

[0186] (3) If the number of points numPonts in the current node is greater than 2, the encoding process is as follows:

[0187] i) First encode the numPonts of the current node to be less than or equal to 1;

[0188] ii) Secondly, it is encoded that the second point of the current node is a repeated point, and then it is encoded whether the number of repeated points of the current node is greater than 1. When the number of repeated points is greater than 1, it is necessary to perform exponential Golomb decoding on the remaining number of repeated points.

[0189] After encoding the number of points in the current node, the coordinate information of the points contained in the current node is encoded. The following will introduce the lidar point cloud and the human eye point cloud in detail.

[0190] (1) Point cloud facing the human eye.

[0191] (1) If the current node contains only one point, the geometric information of the point in three dimensions will be directly encoded (Bypass coding);

[0192] (2) If the current node contains two points, the first coordinate axis dirextAxis will be obtained by using the geometric coordinates of the points. It should be noted that the coordinate axes currently compared only include the x-axis and y-axis, not the z-axis. Assuming that the geometric coordinates of the current node are nodePos, the judgment method is as follows: dirextAxis = ! (nodePos[0] <nodePos[1]) (5)

[0193] That is, the axis with the smallest node coordinate geometry position will be used as the priority encoding axis dirextAxis, and then the geometry information of the priority encoding axis dirextAxis will be encoded as follows. Assume that the encoding geometry bit depth corresponding to the priority encoding axis is nodeSizeLog2, and assume that the coordinates of the two points are pointPos[0] and pointPos[1]. The specific encoding process is as follows:

[0194] After encoding the priority axis dirextAxis, the geometric coordinates of the current node are directly encoded. Assuming that the remaining encoding bit depth of each point is nodeSizeLog2, the specific encoding process is as follows:

[0195] for(int axisIdx=0; axisIdx<3; ++axisIdx)

[0196] for(int mask=(1<<nodeSizeLog2[axisIdx])> >1;mask;mask>>1)

[0197] encodePosBit(!!(pointPos[axisIdx]&mask)).

[0198] (2) LiDAR point cloud.

[0199] If the current node contains two points, the priority coded coordinate axis dirextAxis will be obtained first by using the geometric coordinates of the points. Assuming that the geometric coordinates of the current node are nodePos, the judgment method is as follows:

[0200] dirextAxis=!(nodePos[0] <nodePos[1])

[0201] That is, the axis with the smaller node coordinate geometry position will be used as the priority encoding axis dirextAxis. It should be noted that the currently compared coordinate axes only include the x-axis and y-axis, not the z-axis. Then, the geometric information of the priority encoded coordinate axis dirextAxis is first encoded as follows, assuming that the encoding geometry bit depth corresponding to the priority encoded axis is nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1]. The specific encoding process is as follows:

[0202] After encoding the priority-encoded coordinate axis dirextAxis, the geometric coordinates of the current node are encoded.

[0203] Since the LiDAR point cloud can obtain the acquisition parameters of the LiDAR point cloud, the geometric coordinate information of the current node can be predicted, thereby further improving the efficiency of the geometric information encoding of the point cloud. Similarly, the geometric information nodePos of the current node is first used to obtain a directly encoded main axis direction, and then the geometric information of the already encoded direction is used to predict the geometric information of another dimension. Assuming that the directly encoded axis direction is directAxis and the bit depth of the direct encoding is nodeSizeLog2, the encoding method is as follows:

[0204] for(int mask=(1<<nodeSizeLog2)> >1;mask;mask>>1);

[0205] encodePosBit(!!(pointPos[directAxis]&mask)).

[0206] It should be noted here that all geometric accuracy information in the directAxis direction will be encoded here.

[0207] For example, Figure 15 provides a schematic diagram of coordinate transformation for obtaining point clouds using a rotating laser radar. In the Cartesian coordinate system, the (x, y, z) coordinates of each node can be converted to Indicates. In addition, the laser scanner can perform laser scanning at a preset angle, and different θ(i) can be obtained under different values ​​of i. For example, when i is equal to 1, θ(1) can be obtained, and the corresponding scanning angle is -15°; when i is equal to 2, θ(2) can be obtained, and the corresponding scanning angle is -13°; when i is equal to 10, θ(10) can be obtained, and the corresponding scanning angle is +13°; when i is equal to 9, θ(19) can be obtained, and the corresponding scanning angle is +15°.

[0208] In this way, after encoding all the precision of the directAxis coordinate direction, the LaserIdx corresponding to the current point will be calculated first, that is, the pointLaserIdx number in Figure 15, and the LaserIdx of the current node, that is, nodeLaserIdx; secondly, the LaserIdx of the node, that is, nodeLaserIdx, will be used to predict the LaserIdx of the point, that is, pointLaserIdx. The calculation method of the LaserIdx of the node or point is as follows. Assuming that the geometric coordinates of the point are pointPos, the starting coordinates of the laser ray are LidarOrigin, and assuming that the number of Lasers is LaserNum, the tangent value of each Laser is tanθ i , the vertical offset position of each Laser is Z i ,but:

[0209] After calculating the current point's LaserIdx, the LaserIdx of the current node is first used to predictively encode the pointLaserIdx. After encoding the current point's LaserIdx, the three-dimensional geometric information of the current point is predictively encoded using the LiDAR acquisition parameters.

[0210] For example, FIG16 shows a schematic diagram of predictive coding in the X-axis or Y-axis direction. As shown in FIG16 , the box filled with a grid represents the current node, and the box filled with a slash represents the already coded node. Here, the LaserIdx corresponding to the current point is first used to obtain the corresponding predicted value of the horizontal azimuth, that is, Secondly, the node geometry information corresponding to the current point is used to obtain the horizontal azimuth angle corresponding to the node Assuming that the geometric coordinates of the node are nodePos, the horizontal azimuth angle The calculation method between the node geometry information is as follows:

[0211] By using the acquisition parameters of the laser radar, we can get the number of rotation points of each laser, numPoints, which represents the number of points obtained by each laser ray rotating one circle. Then, we can use the number of rotation points of each laser to calculate the rotation angular velocity deltaPhi of each laser. The calculation method is as follows:

[0212] Furthermore, using the horizontal azimuth angle of the node And the horizontal azimuth of the previous Laser code point corresponding to the current point Calculate the predicted horizontal azimuth angle corresponding to the current point That is, the predicted value of the horizontal azimuth angle as shown in Figure 17A and Figure 17B. Figure 17A shows a schematic diagram of predicting the angle of the Y plane through the horizontal azimuth angle, and Figure 17B shows a schematic diagram of predicting the angle of the X plane through the horizontal azimuth angle. Here, the predicted value of the horizontal azimuth angle corresponding to the current point is The calculation is as follows:

[0213] For example, FIG18 shows another schematic diagram of predictive coding in the X-axis or Y-axis direction. As shown in FIG18 , the portion filled with a grid (left side) represents a low plane, and the portion filled with dots (right side) represents a high plane. Indicates the low plane horizontal azimuth of the current node, Indicates the horizontal azimuth of the high plane of the current node, Indicates the predicted horizontal azimuth angle corresponding to the current node.

[0214] Thus, by using the predicted value of the horizontal azimuth and the low plane horizontal azimuth of the current node and the high plane horizontal azimuth To predict the geometric information of the current node. The details are as follows:

[0215] int context=(angLel≥0&&angLeR≥0)||(angLel<0&&angLeR<0)? 0:2;

[0216] int minAngle=std∷min(abs(angLel),abs(angLeR));

[0217] int maxAngle=std∷max(abs(angLel),abs(angLeR));

[0218] context+=maxAngle>minAngle? 0:1;

[0219] context+=maxAngle>minAngle? 0:4.

[0220] After encoding the LaserIdx of the point, the Z-axis direction of the current point will be predicted using the LaserIdx corresponding to the current point. That is, the depth information radius of the radar coordinate system is calculated by using the x and y information of the current point. Then, the tangent value of the current point and the vertical offset are obtained using the LaserIdx of the current point. The predicted value of the Z-axis direction of the current point, namely Z_pred, can be obtained. The details are as follows:

[0221] int tanTheta=tanθ laserIdx ;

[0222] int zOffset = Z laserIdx ;

[0223] Z_pred=radius×tanTheta-zOffset.

[0224] Furthermore, Z_pred is used to perform predictive coding on the geometric information of the current point in the Z-axis direction to obtain the prediction residual Z_res, and finally Z_res is encoded.

[0225] It is important to note that when partitioning nodes into leaf nodes, the number of duplicate points in the leaf nodes must be encoded in the case of lossless geometric coding. Ultimately, the placeholder information for all nodes is encoded to generate a binary bitstream. Furthermore, G-PCC currently introduces a plane coding mode. During the geometric partitioning process, it determines whether the child nodes of the current node are in the same plane. If the child nodes of the current node meet the condition of being in the same plane, the child nodes of the current node are represented by that plane.

[0226] For octree-based geometric decoding, the decoder follows a breadth-first traversal. Before decoding each node's occupancy information, it first uses the reconstructed geometric information to determine whether the current node is for plane decoding or IDCM decoding. If the current node meets the requirements for plane decoding, it first decodes the plane identifier and plane position information of the current node. Then, based on the plane information, it decodes the current node's occupancy information. If the current node meets the requirements for IDCM decoding, it first decodes whether the current node is a true IDCM node. If so, it continues to parse the DCM decoding mode of the current node, then obtains the number of points in the current DCM node, and finally decodes the geometric information of each point. For nodes that do not meet either plane decoding or DCM decoding requirements, the current node's occupancy information is decoded. By continuously parsing in this way, the placeholder code of each node is obtained, and the node is continuously partitioned until a 1×1×1 unit cube is obtained. The number of points contained in each leaf node is parsed, and the geometrically reconstructed point cloud information is finally recovered.

[0227] The following is a detailed introduction to the IDCM decoding process.

[0228] Similar to the processing at the encoding end, we first use prior information to determine whether the node should start IDCM. The starting conditions of IDCM are as follows:

[0229] (1) The current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.

[0230] (2) The parent node of the current node has only one child node, the current node, and the six neighbor nodes that share a face with the current node are all empty nodes.

[0231] (3) The number of sibling nodes of the current node is greater than 1.

[0232] Furthermore, when a node meets the conditions for DCM coding, it is first decoded to determine whether the current node is a true DCM node, that is, IDCM_flag; when IDCM_flag is true, the current node adopts DCM coding, otherwise it still adopts octree coding.

[0233] Next, decode the number of points numPoints of the current node. The specific decoding method is as follows:

[0234] i) First decode whether numPonts of the current node is greater than 1;

[0235] ii) If the numPonts of the current node is greater than 1, continue decoding to see if the second point is a duplicate point; if the second point is not a duplicate point, it can be implicitly inferred that the second type of DCM mode contains only two points;

[0236] iii) If the numPonts of the current node obtained by decoding is less than or equal to 1, continue decoding to see if the second point is a repeated point; if the second point is not a repeated point, it can be implicitly inferred that the second type of DCM pattern is satisfied, which contains only one point; if the second point obtained by decoding is a repeated point, it can be inferred that the third type of DCM pattern is satisfied, which contains multiple points, but they are all repeated points, then continue decoding to see if the number of repeated points is greater than 1 (entropy decoding), and if it is greater than 1, continue decoding the number of remaining repeated points (using exponential Columbus decoding).

[0237] If the current node does not meet the requirements of the DCM node, it will exit directly (that is, the number of points is greater than 2 points and it is not a duplicate point).

[0238] After decoding the number of points in the current node, the coordinate information of the points contained in the current node is decoded. The following will introduce the lidar point cloud and the human eye point cloud in detail.

[0239] (1) Point cloud facing the human eye.

[0240] (1) If the current node contains only one point, the geometric information of the point in three dimensions will be directly decoded (Bypass coding);

[0241] (2) If the current node contains two points, the first decoded coordinate axis dirextAxis will be obtained by using the geometric coordinates of the points. It should be noted that the coordinate axes currently compared only include the x and y axes, not the z axis. Assuming that the geometric coordinates of the current node are nodePos, the judgment method is as follows: dirextAxis = ! (nodePos[0] <nodePos[1]) (9)

[0242] That is, the axis with the smallest node coordinate geometry position will be used as the priority decoding axis dirextAxis, and then the geometry information of the priority decoding axis dirextAxis will be decoded first in the following way. Assume that the geometry bit depth to be decoded corresponding to the priority decoding axis is nodeSizeLog2, and assume that the coordinates of the two points are pointPos[0] and pointPos[1]. The specific encoding process is as follows:

[0243] After decoding the prioritized axis dirextAxis, the geometric coordinates of the current point are directly decoded. Assuming the remaining encoding bit depth of each point is nodeSizeLog2 and the coordinate information of the point is pointPos, the specific decoding process is as follows:

[0244] (2) LiDAR point cloud.

[0245] If the current node contains two points, the priority decoding coordinate axis dirextAxis will be obtained first by using the geometric coordinates of the points. Assuming that the geometric coordinates of the current node are nodePos, the judgment method is as follows: dirextAxis = ! (nodePos[0] <nodePos[1]) (10)

[0246] That is, the axis with the smaller node coordinate geometry position will be used as the priority decoding axis dirextAxis. It should be noted that the currently compared coordinate axes only include the x-axis and y-axis, not the z-axis. Secondly, the priority encoded coordinate axis dirextAxis geometry information is first decoded as follows, assuming that the encoding geometry bit depth corresponding to the priority decoded axis is nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1]. The specific encoding process is as follows:

[0247] It should be noted here that all geometric accuracy information in the directAxis direction will be decoded here.

[0248] After decoding all the precision of the directAxis coordinate direction, the LaserIdx of the current node, i.e., nodeLaserIdx, is calculated first. Then, the LaserIdx of the node, i.e., nodeLaserIdx, is used to predict and decode the LaserIdx of the point, i.e., pointLaserIdx. The calculation method of the LaserIdx of the node or point is the same as that of the encoder. Finally, the predicted residual information of the LaserIdx of the current point and the LaserIdx of the node is decoded to obtain ResLaserIdx. The decoding method is as follows: PointLaserIdx = nodeLaserIdx + ResLaserIdx (11)

[0249] After decoding the LaserIdx of the current point, the three-dimensional geometric information of the current point is predicted and decoded using the acquisition parameters of the laser radar. The specific algorithm is as follows:

[0250] As shown in Figure 11, the LaserIdx corresponding to the current point is first used to obtain the corresponding predicted value of the horizontal azimuth angle, that is, Secondly, the node geometry information corresponding to the current point is used to obtain the horizontal azimuth angle corresponding to the node Assuming that the geometric coordinates of the node are nodePos, the horizontal azimuth angle The calculation method between the node geometry information is as follows:

[0251] By using the acquisition parameters of the laser radar, we can get the number of rotation points of each laser, numPoints, which represents the number of points obtained by each laser ray rotating one circle. Then, we can use the number of rotation points of each laser to calculate the rotation angular velocity deltaPhi of each laser. The calculation method is as follows:

[0252] Furthermore, using the horizontal azimuth angle of the node And the horizontal azimuth of the previous Laser code point corresponding to the current point Calculate the predicted horizontal azimuth angle corresponding to the current point That is, the predicted value of the horizontal azimuth angle as shown in Figures 17A and 17B. The calculation method is as follows:

[0253] Thus, by using the predicted value of the horizontal azimuth and the low plane horizontal azimuth of the current node and the horizontal azimuth of the high plane To predict and decode the geometric information of the current node. The details are as follows:

[0254] int context=(angLel≥0&&angLeR≥0)||(angLel<0&&angLeR<0)? 0:2;

[0255] int absAngleL=abs(angLel);

[0256] int absAngleR=abs(angLeR);

[0257] context+=absAngleL>absAngleR? 0:1;

[0258] context+=maxAngle>minAngle<<1?4:0.

[0259] After decoding the LaserIdx of the completed point, the Z-axis direction of the current point will be predicted and decoded using the LaserIdx corresponding to the current point. That is, the depth information radius of the radar coordinate system is calculated by using the x and y information of the current point. Then, the tangent value of the current point and the vertical offset are obtained using the LaserIdx of the current point. The predicted value of the Z-axis direction of the current point, namely Z_pred, can be obtained. The details are as follows:

[0260] int tanTheta=tanθ laserIdx ;

[0261] int zOffset = Z laserIdx ;

[0262] Z_pred=radius×tanTheta-zOffset.

[0263] Furthermore, the decoded Z_res and Z_pred are used to reconstruct and restore the geometric information of the current point in the Z-axis direction.

[0264] For triangle soup (trisoup)-based geometric information coding, geometric partitioning must also be performed first in the trisoup-based geometric information coding framework. However, unlike geometric information coding based on binary trees, quadtrees, and octrees, this method does not require step-by-step partitioning of the point cloud into unit cubes with side lengths of 1×1×1. Instead, the partitioning stops when the sub-blocks (blocks) have a side length of W. Based on the surface formed by the distribution of the point cloud in each block, the surface and the twelve edges of the block are obtained. The vertex coordinates of each block are encoded in sequence to generate a binary code stream.

[0265] For trisoup-based point cloud geometry reconstruction, when performing point cloud geometry reconstruction at the decoding end, the vertex coordinates are first decoded to complete the triangle face reconstruction. This process is shown in Figures 19A, 19B, and 19C. Among them, there are three intersection points (v1, v2, v3) in the block shown in Figure 19A. The set of triangle faces formed by these three intersection points in a certain order is called triangle soup, or trisoup, as shown in Figure 19B. Afterwards, sampling is performed on the triangle face set, and the obtained sampling points are used as the reconstructed point cloud within the block, as shown in Figure 19C.

[0266] For predictive geometry coding (PredGeomTree), the following steps are involved: first, sort the input point cloud. Currently, the sorting methods used include unordered, Morton order, azimuth order, and radial distance order. At the encoding end, the prediction tree structure is established using two different methods: KD-Tree (high-latency slow mode) and low-latency fast mode (using lidar calibration information). When using lidar calibration information, each point is divided into different lasers, and a prediction tree structure is established according to the different lasers. Next, based on the structure of the prediction tree, each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter. Finally, through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameter are encoded to generate a binary code stream.

[0267] For geometric decoding based on the prediction tree, the decoding end reconstructs the prediction tree structure by continuously parsing the bit stream. Secondly, the geometric position prediction residual information and quantization parameters of each prediction node are obtained through parsing, and the prediction residual is dequantized to restore the reconstructed geometric position information of each node, finally completing the geometric reconstruction at the decoding end.

[0268] After the geometric encoding is completed, the geometric information needs to be reconstructed. At present, attribute encoding is mainly performed on color information. First, the color information is converted from the RGB color space to the YUV color space. Then, the point cloud is recolored using the reconstructed geometric information so that the unencoded attribute information corresponds to the reconstructed geometric information. In color information encoding, there are two main transformation methods. One is the distance-based lifting transformation that relies on LOD partitioning, and the other is to directly perform RAHT transformation. Both methods will convert the color information from the spatial domain to the frequency domain, and obtain high-frequency coefficients and low-frequency coefficients through transformation. Finally, the coefficients are quantized and encoded to generate a binary code stream, as shown in Figures 4A and 4B.

[0269] Furthermore, when using geometric information to predict attribute information, Morton codes can be used to perform nearest neighbor search. The Morton code corresponding to each point in the point cloud can be obtained from the geometric coordinates of the point. The specific method for calculating the Morton code is described below. For a three-dimensional coordinate represented by a d-bit binary number for each component, its three components can be expressed as:

[0270] Among them, x l ,y l ,z l∈{0,1} are the binary values ​​corresponding to the highest bit (l=1) to the lowest bit (l=d) of x, y, and z respectively. The Morton code M is to cross-arrange x, y, and z starting from the highest bit. l ,y l ,z l To the lowest bit, the calculation formula of M is as follows:

[0271] Among them, m l′ ∈{0,1} are the values ​​of the highest bit (l′=1) to the lowest bit (l′=3d) of M. After obtaining the Morton code M of each point in the point cloud, the points in the point cloud are arranged in ascending order of the Morton code, and the weight value w of each point is set to 1.

[0272] It can also be understood that for the G-PCC codec framework, the general test conditions are as follows:

[0273] (1) There are 4 test conditions:

[0274] Condition 1: The geometric position is limited and the attributes are lost;

[0275] Condition 2: Geometric position lossless, attribute lossy;

[0276] Condition 3: Geometric position lossless, attribute loss limited;

[0277] Condition 4: Geometric position and attributes are lossless.

[0278] (2) The general test sequence includes four categories: Cat1A, Cat1B, Cat3-fused, and Cat3-frame. Among them, Cat2-frame point cloud only contains reflectance attribute information, Cat1A and Cat1B point clouds only contain color attribute information, and Cat3-fused point cloud contains both color and reflectance attribute information.

[0279] (3) Technical routes: There are two types in total, which are distinguished by the algorithm used for geometric compression.

[0280] Technical route 1: Octree encoding branch.

[0281] At the encoding end, the bounding box is divided into sub-cubes in sequence, and the non-empty sub-cubes (containing points in the point cloud) are divided again until the leaf node obtained is a 1×1×1 unit cube. In the case of geometric lossless coding, the number of points contained in the leaf node needs to be encoded, and finally the geometric octree encoding is completed to generate a binary code stream.

[0282] At the decoding end, the decoding end obtains the placeholder code of each node by continuously parsing in the order of breadth-first traversal, and continuously divides the nodes in turn until a 1×1×1 unit cube is obtained. In the case of geometric lossless decoding, it is necessary to parse the number of points contained in each leaf node and finally recover the geometrically reconstructed point cloud information.

[0283] Technical route 2: prediction tree encoding branch.

[0284] On the encoding side, the prediction tree structure is established using two different methods: based on KD-Tree (high-latency slow mode) and using lidar calibration information (low-latency fast mode). Using lidar calibration information, each point can be divided into different lasers, and the prediction tree structure is established according to different lasers. Next, based on the structure of the prediction tree, each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter. Finally, through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream.

[0285] At the decoding end, the decoding end reconstructs the prediction tree structure by continuously parsing the bit stream. Secondly, the geometric position prediction residual information and quantization parameters of each prediction node are obtained through parsing, and the prediction residual is dequantized to restore the reconstructed geometric position information of each node, finally completing the geometric reconstruction at the decoding end.

[0286] It should also be noted that, as shown in Figure 4A or Figure 4B, the current G-PCC coding framework includes three attribute coding methods: Predicting Transform (PT), Lifting Transform (LT), and Region Adaptive Hierarchical Transform (RAHT). Among them, the first two predictively encode the point cloud based on the generation order of LOD, while RAHT adaptively transforms the attribute information from bottom to top based on the construction level of the octree. The following will describe these three point cloud attribute coding methods in detail.

[0287] (a) Predictive coding of point cloud attribute information.

[0288] Currently, the attribute prediction module of G-PCC adopts a nearest neighbor attribute prediction coding scheme based on a hierarchical (Level-of-details, LoDs) structure. The construction methods of LOD include the distance-based LOD construction scheme, the fixed sampling rate-based LOD construction scheme, and the octree-based LOD construction scheme, etc. In the distance threshold-based LOD construction scheme, before constructing LOD, the point cloud is first sorted by Morton to ensure strong attribute correlation between adjacent points. Figure 20 is a schematic diagram of a distance-based LOD construction process. As shown in Figure 20, according to L Manhattan (Manhattan) distances (dl) preset by the user in advance, l = 0, 1, … L-1; the point cloud is divided into L different point cloud detail levels (Rl), l = 0, 1, … L-1, where (dl)l = 0, 1, … L-1 satisfies dl < dl-1. The construction process of LOD is described as follows:

[0289] (1) First, mark all points in the point cloud as unvisited, and establish a set V to store the set of points that have been visited; (2) For each iteration l, by traversing the points in the point cloud, if the current point has been visited, ignore the point, otherwise calculate the minimum distance D from the current point to the point set V. If D < dl, ignore the point; otherwise, mark the current point as visited and add the current point to the refinement level Rl and the point set V; (3) The points in the level of detail LODl are composed of the points in the refinement levels R0, R1, R2…Rl; (4) Continuously repeat the above steps until all points are marked as visited.

[0290] Based on the LOD structure, the attribute value of each point is linearly weighted predicted by using the reconstructed attribute values of points in the same level or a higher level of LOD. Among them, the maximum number of reference prediction neighbors is determined by the high-level syntax elements of the encoder. For the attribute of each point, at the encoding end, the rate-distortion optimization algorithm is used to select to perform weighted prediction by using the attributes of the N nearest neighbor points searched or select the attribute of a single nearest neighbor point for prediction, and finally encode the selected prediction mode and the prediction residual.

[0291] Among them, N represents the number of prediction points in the nearest neighbor point set of point i, Pi represents the sum of the N nearest neighbor points of point i, Dm represents the spatial geometric distance from the nearest neighbor point m to the current point i, Attrm represents the attribute value after reconstruction of the nearest neighbor point m, and Attr i ′ represents the predicted attribute value of the current point i, and the number of points N is a preset value in advance.

[0292] To balance attribute coding efficiency and parallel processing between different LOD layers, a switch is introduced in the encoder's high-level syntax elements to control whether to use intra-LOD prediction. If turned on, intra-LOD prediction is enabled, allowing predictions to be made using points within the same LOD layer. Note that when the number of LOD layers is 1, intra-LOD prediction is always used.

[0293] Figure 21 shows a visualization of the LOD generation process. This provides a subjective example of the distance-based LOD generation process. Specifically (from left to right): points in the first layer represent the outer contours of the point cloud; as the number of detail layers increases, the point cloud details become increasingly clear.

[0294] Figure 22 is a schematic diagram of the attribute prediction encoding process. As shown in Figure 22, for the specific process of G-PCC attribute prediction, for the original point cloud, the three nearest neighbors of the Kth point are first searched, and then attribute prediction is performed. The difference between the attribute prediction value of the Kth point and the original attribute value of the Kth point is calculated to obtain the prediction residual of the Kth point. Quantization and arithmetic coding are then performed to finally generate the attribute bit rate.

[0295] (i) Optimal prediction value selection:

[0296] After the LOD is constructed, according to the generation order of LOD, the three nearest neighboring points of the current point to be encoded are first found from the encoded data points. The attribute reconstruction values ​​of these three nearest neighboring points are used as candidate prediction values ​​of the current point to be encoded; then, the optimal prediction value is selected from them according to the rate-distortion optimization (RDO). For example, when encoding the attribute value of point P2 in Figure 20, the prediction variable index of the attribute value of the nearest neighbor point P4 is set to 1; the attribute prediction variable indexes of the second nearest neighbor point P5 and the third nearest neighbor point P0 are set to 2 and 3 respectively; the prediction variable index of the weighted average of points P0, P5 and P4 is set to 0, as shown in Table 1; finally, RDO is used to select the best prediction variable. The formula for weighted average is as follows:

[0297] in, Represents the spatial geometric weight of the neighboring point j to the current point i:

[0298] Represents the attribute prediction value of the current point i, j represents the index of the three neighboring points, Represents the attribute value after reconstruction of the neighboring points, x i ,y i ,z i is the geometric position coordinate of the current point i, x ij ,yij ,z ij is the geometric coordinate of the neighboring point j.

[0299] For example, Table 1 provides an example of candidate prediction item samples for an attribute code.

[0300] Table 1

[0301] (ii) Attribute prediction residuals and quantification:

[0302] The attribute prediction value of the current point i is obtained through the above prediction (k is the total number of points in the point cloud). Let (a i ) i∈0…k-1 is the original value of the attribute of the current point, then the attribute residual (r i ) i∈0…k-1 Denoted as:

[0303] Further quantify the prediction residuals:

[0304] Among them, Q i It represents the quantized attribute residual of the current point i, Qs is the quantization step (Qs), which can be calculated by the quantization parameter QP (QP) specified by CTC.

[0305] (iii) The encoding end reconstructs the attribute value:

[0306] The purpose of reconstruction at the encoding end is to predict the subsequent points. Before reconstructing the attribute value, the residual must be dequantized. is the residual after inverse quantization:

[0307] and predicted value Add up to get the reconstruction value of point i

[0308] There are currently two main types of algorithms for attribute nearest neighbor search based on LOD partitioning: intra-frame nearest neighbor search and inter-frame nearest neighbor search. The inter-frame nearest neighbor search algorithm is detailed below, while the intra-frame nearest neighbor search can be divided into inter-layer nearest neighbor search and intra-layer nearest neighbor search.

[0309] (i) Intra-frame nearest neighbor search:

[0310] Intra-frame nearest neighbor search is divided into two algorithms: inter-layer nearest neighbor search and intra-layer nearest neighbor search. After LOD division, it resembles a pyramid structure, as shown in Figure 23.

[0311] In a specific implementation, for inter-layer nearest neighbor search, the pyramid structure is shown in FIG24. FIG25 is a pyramid structure for inter-layer nearest neighbor search.

[0312] Schematic diagram of the LOD construction process of neighbor search. As shown in Figure 25, different LOD layers are obtained based on geometric information division.

[0313] LOD0, LOD1 and LOD2 use the points in LOD0 to predict the attributes of the points in the next layer of LOD in the nearest neighbor search between layers

[0314] In the process.

[0315] The entire process of searching for the nearest neighbor within a frame is described in detail below.

[0316] During the entire LOD partitioning process, there are three sets: O(k), L(k), and I(k). Among them, k is the index of the LOD layer during LOD partitioning, and I(k) is the input point set during the current LOD layer partitioning. After LOD partitioning, the O(k) set and L(k) set are obtained. The O(k) set stores the sampling point set, and L(k) is the point set in the current LOD layer. The entire LOD partitioning process is as follows:

[0317] (1) Initialization.

[0318] if k=0,L(k)←{}; otherwise,L(k)←L(k-1);

[0319] O(k)←{};

[0320] (2) Using the LOD partitioning algorithm, the sampling points are stored in O(k), and the remaining points are divided into L(k);

[0321] (3) When the next iteration is performed, I←O(k).

[0322] It should be noted here that since the entire LOD division process is based on the Morton code, O(k), L(k) and I(k) store the Morton code index corresponding to the point.

[0323] When performing inter-layer nearest neighbor search, that is, the points in the L(k) set perform nearest neighbor search in the O(k) set. The specific search algorithm is as follows:

[0324] Taking the nearest neighbor search based on spatial relationships as an example, when predicting the current point P, the neighbor search is performed by using the parent block (Block B) corresponding to point P. As shown in Figure 26, points in the neighboring blocks that are coplanar or colinear with the current parent block are searched for attributes.

[0325] Figure 27A shows a schematic diagram of a coplanar spatial relationship, where there are 6 spatial blocks that have a relationship with the current parent block. Figure 27B shows a schematic diagram of a coplanar and colinear spatial relationship, where there are 18 spatial blocks that have a relationship with the current parent block. Figure 27C shows a schematic diagram of a coplanar, colinear, and co-point spatial relationship, where there are 26 spatial blocks that have a relationship with the current parent block.

[0326] First, the coordinates of the current point are used to obtain the corresponding spatial block. Second, a nearest neighbor search is performed in the previously encoded LOD layer to find the spatial blocks that are coplanar, colinear, and co-point with the current block to obtain the N nearest neighbors of the current point.

[0327] After performing coplanar, colinear, and co-point nearest neighbor searches, if the N nearest neighbors of the current point are still not found, the N nearest neighbors of the current point will be found based on a fast search algorithm. The specific algorithm is as follows:

[0328] As shown in Figure 28, when performing inter-attribute layer prediction, the geometric coordinates of the current point to be encoded are first used to obtain the Morton code corresponding to the current point. Secondly, based on the Morton code of the current point, the first reference point (j) with a value greater than the Morton code of the current point is found in the reference frame. Then, the nearest neighbor search is performed within the range of [j-searchRange, j+searchRange].

[0329] The rest of the specific algorithms for updating the nearest neighbor are the same as the inter-frame nearest neighbor search algorithm and will not be described in detail here. The specific algorithms will be mentioned in the inter-frame nearest neighbor search algorithm.

[0330] In another specific implementation, for the nearest neighbor search within a layer, Figure 29 shows a schematic diagram of the LOD structure of the nearest neighbor search within an attribute layer. As shown in Figure 29, if the intra-layer prediction algorithm is turned on, that is, the syntax element EnableRefferingSameLoD=1, then the nearest neighbor search within the layer can be allowed. For example, for the LOD1 layer, the nearest neighbor point of the current point P6 can be P1, which is not allowed in other layers; if the syntax element EnableRefferingSameLoD=0, then inter-layer search is allowed in other layers. For example, for the LOD1 layer, the nearest neighbor point of the current point P6 can be P4. That is to say, when the intra-layer prediction algorithm is turned on, the nearest neighbor search will be performed in the same layer LOD and the set of encoded points in the same layer to obtain the N nearest neighbors of the current point (the inter-layer nearest neighbor search is also performed).

[0331] When performing prediction within the attribute layer, a nearest neighbor search is performed based on a fast search algorithm. The specific algorithm is shown in Figure 30. The current point is represented by a grid. Assuming the Morton code index of the current point is i, the nearest neighbor search is performed in [i+1, i+searchRange]. The specific nearest neighbor search algorithm is consistent with the inter-frame block-based fast search algorithm and is not described in detail here.

[0332] (ii) Nearest neighbor search between frames:

[0333] Figure 31A is a schematic diagram of attribute inter-frame prediction based on fast search. As shown in Figure 31A, when performing attribute inter-frame prediction, the geometric coordinates of the current point to be encoded are first used to obtain the Morton code corresponding to the current point, and then based on the Morton code of the current point, the first reference point (j) that is larger than the Morton code of the current point is found in the reference frame, and then the nearest neighbor search is performed within the range of [j-searchRange, j+searchRange].

[0334] The current nearest neighbor search within and between frames is based on block-based neighborhood search, as shown in Figure 31B. As shown in Figure 31B, when performing neighborhood search for the current point (Morton code index is i), the points in the reference frame are first divided into N (N=3) layers according to the Morton code. The specific division algorithm is as follows:

[0335] First layer: Assume that the points of the reference frame are numPoints, first divide the points in the reference frame into M (M=2 5 =32) points are divided into one block;

[0336] Second layer: Based on the first layer, the blocks of the first layer are also processed in the order of Morton code every M (M=2 5 =32) blocks are divided into one block;

[0337] The third layer: Based on the second layer, the blocks of the second layer are also processed in the order of the Morton code every M (M=2 5 =32) blocks are divided into one block;

[0338] Finally, the predicted structure shown in FIG31B is obtained.

[0339] When performing attribute prediction based on the prediction structure shown in FIG31B , assuming that the Morton code index of the current point to be encoded is i, first obtain the first point in the reference frame whose Morton code is greater than or equal to the current point, with index j. Then, the block index of the reference point is calculated based on j. The specific calculation method is as follows:

[0340] First layer: BucketSize_0 = 2 5 =32;

[0341] Second layer: BucketSize_1 = 2 5 =32×BucketSize_0=1024;

[0342] Third layer: BucketSize_2 = 2 5 =32×BucketSize_1=32768.

[0343] Assume that the reference range in the prediction frame of the current point is [j-searchRange, j+searchRange], use j-searchRange to calculate the starting index of the third layer, and use j+searchRange to calculate the ending index of the third layer; secondly, first determine whether some blocks in the second layer need to be searched for the nearest neighbor in the blocks of the third layer, and then go to the second layer, and determine whether a search is needed for each block in the first layer. If some blocks in the first layer need to be searched for the nearest neighbor, then the midpoints of some blocks in the first layer will be judged point by point to update the nearest neighbor.

[0344] The following is an introduction to the algorithm based on index calculation block. Assuming that the Morton code index corresponding to the current point is index, then the index of the corresponding third-level block is: idx_2=index / BucketSize_2 (24)

[0345] After obtaining the block index idx_2 of the third layer, the start index and end index of the block corresponding to the current block in the second layer can be obtained using idx_2: startIdx1=idx_2×BucketSize_1 (25) endIdx=idx_2×BucketSize_1+BucketSize_1-1 (26)

[0346] Similarly, the index of the first layer block is obtained based on the index of the second layer block based on the same algorithm.

[0347] When performing a block-based nearest neighbor search, we first determine whether the current block needs to be searched for the nearest neighbor. This is called filtering the nearest neighbor search for the block. Each spatial block can be obtained through two variables: minPos and maxPos. MinPos represents the minimum value of the block, and maxPos represents the maximum value of the block.

[0348] Assume that the distance to the farthest point among the N nearest neighbors of the current point is Dist, the coordinates of the point to be encoded are (x, y, z), and the current block is represented by (minPos, maxPos), where minPos is the minimum value of the bounding box in three dimensions and maxPos is the maximum value of the bounding box in three dimensions. The distance D between the current point and the bounding box is calculated as follows:

[0349] int dx=int(std::max(std::max(minPos[0]-point[0],0),point[0]-maxPos[0]));

[0350] int dy=int(std::max(std::max(minPos[1]-point[1],0),point[1]-maxPos[1]));

[0351] int dz=int(std::max(std::max(minPos[2]-point[2],0),point[2]-maxPos[2]));

[0352] D = dx + dy + dz;

[0353] When D is less than or equal to Dist, the points in the current block will be traversed.

[0354] (b) Lifting transform encoding of point cloud attribute information.

[0355] Figure 32 is a schematic diagram of the encoding process of a lifting transform. The lifting transform also predicts the attributes of the point cloud based on LOD. The difference from the predictive transform is that the lifting transform first divides the LOD into high and low layers, predicts in the reverse order of the LOD generation layer, and introduces an update operator in the prediction process to update the quantization weights of the low-level LOD midpoints to improve the accuracy of the prediction. This is because the attribute values ​​of the low-level LOD midpoints are frequently used to predict the attribute values ​​of the high-level LOD midpoints, and the points in the low-level LOD should have greater influence.

[0356] Step 1: Segmentation process.

[0357] The segmentation process is to divide the complete LOD layer into a low LOD layer L(N) and a high LOD layer H(N). If a point cloud has three LOD layers, namely (LOD l ) l=0,1,2 , after segmentation, LOD2 is the high LOD layer, denoted as H(N), (LOD l ) l=0,1 It is the low LOD layer, denoted as L(N).

[0358] Step 2: Prediction process.

[0359] The point in the high-level LOD selects the attribute information of the nearest neighbor point from the low-level LOD as the attribute prediction value P(N) of the current point to be coded, and the prediction residual D(N) is recorded as: D(N) = H(N) - P(N) (27)

[0360] Step 3: Update process.

[0361] Update the attribute prediction residual D(N) in the high-level LOD to obtain U(N), and use U(N) to improve the attribute value of the midpoint of the low-level LOD, as shown in the following formula: L′(N)=L(N)+U(N) (28)

[0362] The above process will iterate continuously until the lowest LOD according to the order of LOD from high to low.

[0363] Because the LOD-based prediction scheme makes points in the lower LOD layers more influential, the transformation scheme based on the lifting wavelet transform introduces quantization weights and updates the prediction residual based on the prediction residual D(N) and the distance between the prediction point and the adjacent points. Finally, the quantization weights used in the transformation process are used to adaptively quantize the prediction residual. It is important to note that the quantization weight value of each point can be determined by geometric reconstruction at the decoding end, so the quantization weights should not be encoded.

[0364] (c) Region-adaptive hierarchical transformation.

[0365] The Regional Adaptive Hierarchical Transform (RAHT) is a Haar wavelet transform that transforms point cloud attribute information from the spatial domain to the frequency domain, further reducing the correlation between point cloud attributes. Its main concept is to transform the nodes in each layer in the X, Y, and Z dimensions in a bottom-up manner according to the octree structure (as shown in Figure 34), and iterate until the root node of the octree. As shown in Figure 33, the basic concept is to perform a wavelet transform based on the hierarchical structure of the octree, associate attribute information with the octree nodes, and recursively transform the attributes of occupied nodes under the same parent node in a bottom-up manner, transforming the nodes in each layer in the X, Y, and Z dimensions until the root node of the octree is reached. During the hierarchical transformation process, the low-pass / low-frequency (DC) coefficients obtained after the transformation of the nodes in the same layer are passed to the nodes in the next layer for further transformation, while all high-pass / high-frequency (AC) coefficients can be encoded using an arithmetic coder.

[0366] During the transformation process, the DC coefficients (direct current components) of the transformed nodes at the same layer are passed to the previous layer for further transformation, while the AC coefficients (alternating current components) of each layer are quantized and encoded. The main transformation processes are described below.

[0367] FIG35A is a schematic diagram of a RAHT forward transformation process, and FIG35B is a schematic diagram of a RAHT inverse transformation process. For the transformation and inverse transformation process corresponding to RAHT, assuming that g′ L,2x,y,z and g′L,2x+1,y,z are the DC coefficients of two neighboring points in the L layer. After linear transformation, the information of the L-1 layer is the AC coefficient f′ L-1,x,y,z and DC coefficient g′ L-1,x,y,z ; Then, f′ L-1,x,y,z No more transformation will be performed, and quantization coding will be performed directly, g′ L-1,x,y,z The nearest neighbor will continue to be searched for transformation. If no neighbor is found, it will be directly passed to the L-2 layer. That is, the RAHT transformation is only effective for nodes with neighbor points. Nodes without neighbor points will be directly passed to the previous layer. In the above transformation process, g′ L,2x,y,z The weights (the number of non-empty child nodes in the node) corresponding to g′L, 2x+2, y, and z are w′ respectively. L,2x,y,z and w′L,2x+1,y,z (abbreviated as w′0 and w′1), g′ L-1,x,y,z The weight is w′ L-1,x,y,z , then the general transformation formula is:

[0368] Among them, T w0,w1 is the transformation matrix:

[0369] The transformation matrix will be updated as the weights corresponding to each point change adaptively. The above process will be iterated and updated continuously according to the partitioning structure of the octree until the root node of the octree is reached.

[0370] Region Adaptive Hierarchical Intra Prediction Transform Coding

[0371] Regional adaptive hierarchical predictive transform coding is based on RAHT transform coding for prediction. As shown in Figure 33, RAHT attribute transform is based on the order of the octree hierarchy, and the transformation is continuously performed from the voxel level until the root node is obtained, thereby completing the hierarchical transform coding of the entire attribute. In predictive transform coding, attribute predictive transform coding is also performed based on the hierarchical order of the octree, but the transformation is continuously performed from the root node to the voxel level. In each RAHT attribute transformation process, attribute predictive transform coding is performed based on a 2x2x2 block. Figure 36 is a schematic diagram of an attribute coding block. As shown in Figure 36, it can be seen that the dark-filled block is the current block to be encoded, and the light-filled block is some neighboring blocks that are coplanar and colinear with the current block to be encoded.

[0372] FIG37 is a schematic diagram of a RAHT-based attribute prediction transform coding principle. First, the attribute A of the current block can be obtained by the attribute of the point contained in the current block. node , as shown in Figure 37(a), specifically as follows: A node =∑ p∈node attribute(p) (31)

[0373] Secondly, the attributes of the current block and the number of points in the current block are normalized to obtain the mean value a of the current block attributes. node , as shown in Figure 37(b), the specific normalization process is as follows: node =∑ p∈node 1=#{p∈node} (32) a node =A node / w node (33)

[0374] Attribute transform coding is performed using the mean of the current block attributes.

[0375] Furthermore, by using the spatial geometric distance between the neighboring blocks of the current block and each sub-block of the current block, a linear weighted prediction is performed on the attributes of each sub-block to obtain the predicted attributes of the sub-block of the current block, also known as the predicted attribute block, as shown in Figure 37(c). The predicted attributes of the current block are denormalized to obtain the final predicted attributes, as shown in Figure 37(e). Figure 37(d) shows the original attributes of the current block, also known as the attribute original block. Finally, the predicted attributes and original attributes of the sub-blocks are transformed to obtain the original AC coefficients and predicted AC coefficients. The AC coefficient residual is obtained based on the original AC coefficients and the predicted AC coefficients. Figure 37(f) shows the AC coefficient residual of the current block, and the AC coefficient parameters are encoded.

[0376] FIG38 is a schematic diagram of a neighborhood prediction relationship for attribute prediction. As shown in FIG38 , first, the neighborhood blocks of the current block are determined (a maximum of 19 neighborhood blocks). Second, the neighborhood blocks are used to predict the relationship between the neighborhood blocks and each sub-block a of the current block. up The spatial geometric distance between the two blocks is used to perform linear weighted prediction on the attributes of each sub-block, and finally the predicted block attributes are transformed. The specific attribute transformation method is as follows:

[0377] Region Adaptive Hierarchical Inter-frame Predictive Transform Coding

[0378] In the existing G-PCC attribute inter-frame prediction coding, if inter-frame prediction coding is started, the RAHT attribute transform coding structure will be constructed first based on the geometric information of the current node to be coded, that is, the nodes will be continuously merged at the voxel level until the root node of the entire RAHT transform tree is obtained, thereby completing the transform coding hierarchical structure of the entire attribute. Secondly, according to the RAHT transform structure, the root node is divided to obtain N child nodes of each node (N is less than or equal to 8). In the second inter-frame prediction coding scheme, the attributes of the N child nodes will first be independently orthogonally transformed using the RAHT transform to obtain DC and AC coefficients. Then, the AC coefficients of the N child nodes will be predicted for attribute inter-frame according to the following method:

[0379] The inter-frame prediction node of the current node is valid: that is, if the same-position node exists, the attribute of the prediction node is directly used as the attribute prediction value of the current node to be encoded.

[0380] The current node can find a node with exactly the same position as the current node in the cache of the reference frame: that is, if the same-position node exists, the AC coefficients of the M child nodes contained in the same-position node will be directly used as the AC coefficient attribute prediction values ​​of the N child nodes of the current node.

[0381] 1. If the AC coefficient of the predicted node is not zero: the AC coefficient of the predicted node is directly used as the predicted value;

[0382] 2. If the AC coefficient of the prediction node is zero, the AC coefficient of the corresponding child node of the intra-frame prediction will be used as the prediction value

[0383] The inter-frame prediction node of the current node is invalid: that is, the co-located node does not exist, so the attribute prediction value of the adjacent node in the frame is used as the attribute prediction value of the node to be encoded

[0384] And on this basis, the existing RAHT inter-frame coding will select the best RAHT coding mode for each layer: intra-frame prediction coding or inter-frame prediction coding. When the cost of the intra-frame prediction coding mode is less than the cost of the inter-frame prediction coding mode, RAHT intra-frame prediction will be performed on the current layer, otherwise RAHT inter-frame prediction will be performed.

[0385] In the existing G-PCC attribute RAHT intra-frame coding, the attribute is intra-frame predicted by deciding whether to adopt a prediction coding scheme in the high-level aps syntax element, where the prediction coding scheme includes: parent node prediction and child node prediction. Specifically, in the current RAHT intra-frame coding scheme, if the prediction coding scheme is turned on, the number of neighboring nodes N of the current node will be used first to determine whether it is greater than a certain threshold. Only when the number of neighboring nodes is greater than a certain threshold will the AC coefficient of the current node be predicted (intra-frame prediction). If the number of neighboring nodes N of the current node is less than a certain threshold, it is considered that the current node does not meet the conditions for prediction coding, and only the attribute transformation of the current node will be performed. The biggest advantage of this coding scheme is that it effectively considers the neighborhood geometric space correlation of the current node by utilizing the spatial correlation of the current node, especially the neighborhood distribution characteristics of the current node, thereby effectively improving the coding efficiency of point cloud attribute information. However, this coding scheme does not take into account the inherent distribution characteristics of the AC coefficient attribute information of each node, but directly determines the attribute prediction coding scheme of the current sequence in the sequence set. Secondly, when starting the prediction coding scheme, the coding method of the current node is determined only based on the neighborhood geometric spatial correlation of the current node, and does not effectively combine the attribute distribution characteristics of the current node, resulting in low intra-frame coding efficiency of the attribute information.

[0386] In existing G-PCC attribute RAHT inter-frame coding, attribute information is predictively coded by determining whether to use inter-frame prediction or intra-frame prediction in the high-level aps syntax element. The treeDepth syntax element determines the number of layers to enable inter-frame prediction. In lower RAHT layers, only RAHT intra-frame prediction is used. This attribute coding scheme has two major issues: 1. It does not analyze the distribution of AC coefficients across different RAHT layers in different slices, but instead directly determines the attribute inter-frame coding scheme for the current sequence within the sequence set. 2. By determining the number of inter-frame prediction layers in the aps, inter-frame coding is often only enabled in upper RAHT layers, as the intra-frame correlation of AC coefficients in lower RAHT layers is stronger than that between AC coefficients. However, this coding scheme does not fully and effectively utilize the distribution of AC coefficients across different RAHT attribute layers, resulting in low coding efficiency for attribute information.

[0387] Based on this, an embodiment of the present application provides a coding and decoding method, which introduces one or more new coding and decoding modes when performing adaptive layered transform RAHT encoding or decoding, and comprehensively considers the neighborhood geometric distribution characteristics and neighborhood attribute distribution characteristics of the current layer node to select the mode, providing the optimal coding and decoding mode for the current node, thereby improving the RAHT attribute coding and decoding efficiency.

[0388] To facilitate understanding of the technical solutions of the embodiments of the present application, the technical solutions of the present application are described in detail below through specific embodiments. The above related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application. The embodiments of the present application include at least part of the following contents. The present application provides a coding and decoding method, and more specifically provides a point cloud coding and decoding technology.

[0389] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0390] In one embodiment of the present application, referring to FIG39 , a flowchart of a decoding method provided by an embodiment of the present application is shown. As shown in FIG39 , the method may include:

[0391] S3901: Decode the code stream and determine the first syntax element identifier;

[0392] It should be noted that when encoding point cloud attribute information, a first syntax element identifier is used to indicate whether the current sequence, current slice, or other decoding unit activates the mode selection function during RAHT decoding. In some embodiments, the first syntax element identifier is a high-level syntax element and is set in the attribute header (ABH) parameter set. In other words, the bitstream is decoded, the attribute header parameter set is determined, and the first syntax element identifier is determined from the attribute header parameter set.

[0393] Exemplarily, when the first syntax element identifier is a first value, it is determined that a mode selection function is enabled when performing RAHT decoding, that is, when performing RAHT decoding, a target decoding mode is selected from candidate decoding modes. When the first syntax element identifier is a second value, it is determined that a mode selection function is not enabled when performing RAHT decoding, and attribute decoding is performed using an existing RAHT decoding scheme.

[0394] In some embodiments, a first syntax element identifier is used to indicate whether the current sequence, current slice, or other decoding unit starts the intra-frame prediction mode selection function during RAHT decoding. In other embodiments, the first syntax element identifier can also indicate whether the current sequence, current slice, or other decoding unit starts the inter-frame prediction mode selection function during RAHT decoding. In other embodiments, the first syntax element identifier can also indicate whether the current sequence, current slice, or other decoding unit starts the inter-frame intra-frame prediction mode selection function during RAHT decoding. Here, the inter-frame intra-frame prediction mode can be understood as any one of the intra-frame prediction mode, inter-frame prediction mode, or inter-frame intra-frame fusion prediction mode that can be used when performing attribute prediction.

[0395] The inter-frame and intra-frame fusion prediction mode can merge the inter-frame prediction value and intra-frame prediction value of the RAHT decoding layer attributes, and finally obtain the best prediction value according to different weights, thereby further improving the RAHT coding efficiency of point cloud attributes. Specifically, assuming that the RAHT intra-frame prediction value of the current node is predIntraVal and the inter-frame prediction value is predInterVal, the final prediction value predVal is: predVal = w1*predIntraVal + w2*predIntraVal.

[0396] It should be noted that RAHT decoding is based on the octree hierarchy to divide the reconstructed geometric information of the current decoding unit into layers, and attribute decoding is performed on each RAHT decoding layer. The current attribute RAHT decoding order is to divide the layers from the root node to the voxel level (1x1x1), thereby completing the encoding and attribute reconstruction of the entire point cloud attributes. As shown in Figure 40, we define the layer obtained by downsampling along the Z direction, Y direction, and X direction as a RAHT decoding layer, that is, layer.

[0397] S3902: If it is determined according to the first syntax element identifier that the mode selection is enabled when performing region-adaptive layered transform decoding on the current layer, determine the neighborhood geometric distribution information of the node on the current layer;

[0398] It should be noted that the neighborhood geometric distribution information is specifically used to characterize the geometric distribution characteristics of the reference node of the current layer node, and to determine whether the neighborhood geometric distribution information meets the first prediction condition. Specifically, it can be used to determine whether the neighborhood geometric distribution characteristics of the current layer node match the preset neighborhood geometric distribution characteristics. If so, the target decoding mode is further determined based on the neighborhood attribute distribution information of the current layer node.

[0399] In some embodiments, the method may further include: decoding the code stream to determine a second syntax element identifier; determining a target prediction mode for the current layer startup mode selection based on the second syntax element identifier; and determining neighborhood geometric distribution information based on the target prediction mode. Specifically, the second syntax element identifier can be used to indicate the target prediction mode for the startup mode selection of any decoding unit to which the current layer belongs, and the neighborhood geometric distribution information corresponding to different prediction modes may be different. The decoding unit in the embodiment of the present application can also start mode selection in a specific prediction mode.

[0400] In some embodiments, the target prediction mode may include one of the following: intra-frame prediction mode, inter-frame prediction mode and inter-frame intra-frame prediction mode. Exemplarily, according to the target prediction mode, determining the neighborhood geometric distribution information includes: when the target prediction mode is intra-frame prediction mode, determining the neighborhood geometric distribution information includes the number of neighboring nodes of the parent node of the current layer node and the number of neighboring nodes of the grandparent node; when the target prediction mode is inter-frame prediction mode, determining the neighborhood geometric distribution information includes the occupancy information of the same-position parent node of the current layer node and the occupancy information of the parent node; when the target prediction mode is inter-frame intra-frame prediction mode, determining the neighborhood geometric distribution information includes the occupancy information of the same-position parent node of the current layer node and the occupancy information of the parent node, as well as the number of neighboring nodes of the parent node of the current layer node and the number of neighboring nodes of the grandparent node.

[0401] Exemplarily, the target prediction mode is an intra-frame prediction mode, and the first prediction condition includes: the number of neighboring nodes of the parent node is greater than a first threshold, and the number of neighboring nodes of the grandparent node is greater than a second threshold; the target prediction mode is an inter-frame prediction mode, and the first prediction condition includes: the co-located parent node exists, and the placeholder information of the co-located parent node is the same as the placeholder information of the parent node; the target prediction mode is an inter-frame or intra-frame prediction mode, and the first prediction condition includes: the number of neighboring nodes of the parent node is greater than a first threshold, the number of neighboring nodes of the grandparent node is greater than a second threshold, the co-located parent node exists, and the placeholder information of the co-located parent node is the same as the placeholder information of the parent node.

[0402] It should be noted that, for the intra prediction mode, different thresholds may be set according to different decoding layers and hierarchical depths.

[0403] S3903: When the neighborhood geometric distribution information satisfies the first prediction condition, determine the neighborhood attribute distribution information of the current layer node;

[0404] It should be noted that the current layer node is any node in the current layer. In the embodiment of the present application, for the node whose neighborhood geometric distribution information meets the first prediction condition, the optimal decoding mode is further determined by mode selection based on the neighborhood attribute distribution information.

[0405] The neighborhood attribute distribution information is specifically used to characterize the attribute distribution characteristics of the reference node of the current layer node. In some embodiments, the attribute distribution characteristics can be attribute change characteristics or attribute difference characteristics of the reference node. When the attribute change is small, it indicates that the current layer node can use attribute prediction. When the attribute change is large, it indicates that the current layer node cannot use attribute prediction.

[0406] S3904: Determine a target decoding mode for the current layer node from candidate decoding modes based on neighborhood attribute distribution information; wherein the candidate decoding modes include: attribute prediction and transformation mode, and attribute prediction mode;

[0407] In some embodiments, the code stream is decoded to determine the neighborhood attribute distribution information of the current layer node. The neighborhood attribute distribution information can be determined by the encoding end and transmitted to the decoding end. Exemplarily, the encoding end makes a decision on the candidate decoding mode based on the neighborhood attribute distribution information of the current layer node to determine the target encoding mode (the corresponding decoding end is the target decoding mode), and encodes the target encoding mode index. The decoding end determines the target decoding mode based on the decoding mode index. It can be understood that since the target decoding mode is determined based on the neighborhood attribute distribution information of the current layer node, the target decoding mode index can also be used as information that characterizes the neighborhood attribute distribution characteristics.

[0408] The neighborhood attribute distribution information may include a third syntax element identifier for indicating the target decoding mode. The third syntax element identifier may be a syntax element identifier corresponding to each node of the current layer, or a syntax element identifier corresponding to some or all nodes of the current layer, or a syntax element identifier corresponding to multiple RAHT decoding layers.

[0409] In some embodiments, the third syntax element identifier is used to indicate the target decoding mode of the current layer, that is, all nodes in the current layer select the target decoding mode for decoding when performing mode selection.

[0410] In some other embodiments, the third syntax element identifier is used to indicate the target decoding mode of the coefficient group of the current layer, that is, all nodes in the coefficient group select the target decoding mode for decoding when performing mode selection. In this scheme, the coefficient groups of the nodes of the current to-be-coded layer are first split to obtain different coefficient groups. The number of nodes in each coefficient group is N, and the encoder selects the optimal coding mode for each coefficient group. For each coefficient group, the corresponding coding mode also needs to be transmitted to the decoder, so that the decoder reconstructs the attribute information using the decoding mode corresponding to each coefficient group.

[0411] It should be noted that the length of the third syntax element identifier can be determined based on the number of candidate decoding modes. The third syntax element identifier can be a high-level syntax element and is set in an attribute header parameter set. In other words, the bitstream is decoded, the attribute header parameter set is determined, and the third syntax element identifier is determined from the attribute header parameter set.

[0412] In other embodiments, determining the neighborhood attribute distribution information of the current layer node includes: obtaining an attribute reconstruction value of a first reference node of the current layer node; and determining the neighborhood attribute distribution information of the current layer node based on the attribute reconstruction value of the first reference node. In other words, the neighborhood attribute distribution information can be determined by the decoder based on the attribute reconstruction value of the reconstructed first reference node of the current layer node, and the optimal decoding mode is implicitly derived based on the attribute reconstruction value of the reconstructed node, without the need for the encoder to transmit a mode index for determination.

[0413] Determine the neighborhood attribute distribution information of the current layer node, including: determining the attribute reconstruction value of the first reference node of the current layer node according to the target prediction mode; determine the neighborhood attribute distribution information of the current layer node according to the attribute reconstruction value of the first reference node.

[0414] The neighborhood attribute distribution information is related to the target prediction mode. Exemplarily, when the target prediction mode is an intra-frame prediction mode, the first reference node includes at least one of the following: the neighborhood node of the parent node, and the neighborhood node of the same layer that has been reconstructed; when the target prediction mode is an inter-frame prediction mode, the first reference node includes at least one of the following: the parent node and the same-position parent node in the reference frame, the neighborhood node of the parent node and the neighborhood node of the same-position parent node in the reference frame, the neighborhood node of the same layer that has been reconstructed, and the same-position neighborhood node of the same layer that has been reconstructed in the reference frame; when the target prediction mode is an inter-frame or intra-frame prediction mode, the first reference node includes at least one of the following: the parent node and the same-position parent node in the reference frame, the neighborhood node of the parent node and the neighborhood node of the same-position parent node in the reference frame, the neighborhood node of the same layer that has been reconstructed, and the same-position neighborhood node of the same layer that has been reconstructed in the reference frame.

[0415] Exemplarily, for intra prediction mode, the neighborhood attribute distribution information includes the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the parent node. For inter prediction mode, the neighborhood attribute distribution information includes the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the co-located parent node. For inter and intra prediction modes, the neighborhood attribute distribution information includes the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the co-located parent node, as well as the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the parent node.

[0416] The calculation function used for the difference calculation can be the absolute error sum SAD, the absolute transformation difference sum SATD, the mean square error MSE, the error square sum SSD, the mean absolute difference MAD, the mean square error sum MSD, the absolute value DCT of the transformation coefficient, the Hadamard transform, etc., which are not specifically limited here. For example, if the SAD of the attribute reconstruction value of the parent node of the current node and the attribute reconstruction value of the neighboring node of the parent node of the current node are within a certain range, it is considered that the neighborhood attribute distribution characteristics of the current node are relatively flat. Based on such distribution characteristics, it can be implicitly deduced that the current node adopts the predictive coding mode; otherwise, it is considered that the attribute distribution of the neighborhood range of the current node is relatively jittery, and the attribute transformation mode will be adopted.

[0417] In other embodiments, when the neighborhood geometric distribution information satisfies the first prediction condition, one method may be determined from the explicit indexing method and the implicit derivation method, and then the optimal decoding mode may be determined.

[0418] It should also be noted that in the embodiment of the present application, when the neighborhood geometric distribution information does not meet the first prediction condition, the target decoding mode of the current layer node is determined to be a preset decoding mode, where the preset decoding mode can be one of the candidate decoding modes provided in the embodiment of the present application, or another decoding mode. Exemplarily, the preset decoding mode is an attribute transformation mode.

[0419] In some embodiments, the method further includes: decoding the bitstream to determine a second syntax element identifier; determining a target prediction mode for current layer activation mode selection based on the second syntax element identifier; and determining a candidate decoding mode based on the target prediction mode. Different candidate decoding modes may correspond to different candidate decoding modes.

[0420] In some embodiments, determining a candidate decoding mode based on a target prediction mode includes: if the target prediction mode is an intra prediction mode, determining the attribute prediction and transform mode includes the attribute intra prediction and transform mode; if the target prediction mode is an inter prediction mode, determining the attribute prediction and transform mode includes the attribute inter prediction and transform mode; if the target prediction mode is an inter-intra prediction mode, determining the attribute prediction and transform mode includes the attribute inter prediction and transform mode, the attribute intra prediction and transform mode, and the attribute inter-intra prediction and transform mode. In other embodiments, if the target prediction mode is an inter-intra prediction mode, determining the attribute prediction and transform mode includes the attribute inter prediction and transform mode and the attribute intra prediction and transform mode. In other embodiments, if the target prediction mode is an inter-intra prediction mode, determining the attribute prediction and transform mode includes the attribute inter prediction and transform mode and the attribute intra prediction and transform mode. In other embodiments, if the target prediction mode is an inter-intra prediction mode, determining the attribute prediction and transform mode includes the attribute inter prediction and transform mode and the attribute inter-intra prediction and transform mode.

[0421] In some embodiments, the target decoding mode of the current layer node is determined from the candidate decoding modes according to the third syntax element identifier.

[0422] In other embodiments, if the neighborhood attribute distribution information satisfies a second prediction condition, the target decoding mode is determined to be the attribute prediction and transformation mode; if the neighborhood attribute distribution information does not satisfy the second prediction condition, the target decoding mode is determined to be the attribute transformation mode. The second prediction condition serves as a second judgment condition for whether to perform attribute prediction, specifically for determining whether the neighborhood attribute distribution characteristics of the current layer node meet preset neighborhood attribute distribution characteristics for attribute prediction. When both the neighborhood geometric distribution characteristics and the neighborhood attribute distribution characteristics meet the prediction condition, the target decoding mode is determined.

[0423] It should be noted that when the attribute prediction and transformation mode include two or more, a corresponding number of second prediction conditions may also be set. Exemplarily, when the target prediction mode is the intra-frame prediction mode, the second prediction condition includes the prediction condition corresponding to the attribute intra-frame prediction and the transformation mode; when the target prediction mode is the inter-frame prediction mode, the second prediction condition includes the prediction condition corresponding to the attribute inter-frame prediction and the transformation mode; when the target prediction mode is the inter-frame intra-frame prediction mode, the second prediction condition includes the prediction condition corresponding to the attribute intra-frame prediction and the transformation mode, the prediction condition corresponding to the attribute inter-frame prediction and the transformation mode, and the prediction condition corresponding to the attribute inter-frame intra-frame prediction and the transformation mode.

[0424] For the intra-frame prediction mode, the second prediction condition may include that the difference between the reconstruction attribute value of the neighboring node of the parent node and the reconstruction attribute value of the parent node is less than a first difference threshold. For the inter-frame prediction mode, the second prediction condition may include that the difference between the reconstruction attribute value of the neighboring node of the co-located parent node and the reconstruction attribute value of the co-located parent node is less than a second difference threshold. For the inter-frame and intra-frame fusion prediction mode, the second prediction condition may include that the difference between the reconstruction attribute value of the neighboring node of the parent node and the reconstruction attribute value of the parent node is less than a first difference threshold, and the difference between the reconstruction attribute value of the neighboring node of the co-located parent node and the reconstruction attribute value of the co-located parent node is less than a second difference threshold.

[0425] S3905: Decode the attributes of the current layer node according to the target decoding mode to determine the attribute reconstruction value of the current layer node.

[0426] In some embodiments, when the target decoding mode is the attribute prediction and transformation mode, the specific decoding process may include: determining the attribute prediction value of the current layer node; performing attribute transformation based on the attribute prediction value of the current layer node to obtain the AC coefficient prediction value of the current layer node; decoding the code stream to determine the AC coefficient residual value of the current layer node; determining the AC coefficient reconstruction value of the current layer node based on the AC coefficient prediction value and the AC coefficient residual value of the current layer node; performing inverse transformation based on the AC coefficient reconstruction value of the current layer node and the DC coefficient reconstruction value of the parent node to determine the attribute reconstruction value of the current layer node.

[0427] Exemplarily, determining the attribute prediction value of the current layer node includes: when the attribute prediction and transformation mode is attribute intra-frame prediction, determining a first intra-frame prediction mode; determining a first attribute prediction value of the current node based on the first intra-frame prediction mode; or, when the attribute prediction and transformation mode is attribute inter-frame prediction, determining a first inter-frame prediction mode; determining a second attribute prediction value of the current node based on the first inter-frame prediction mode; or, when the attribute prediction and transformation mode is attribute inter-frame intra-frame prediction, determining a first inter-frame intra-frame prediction mode; determining a third attribute prediction value of the current node based on the first inter-frame intra-frame prediction mode. That is, under different prediction modes, different attribute prediction methods can be used to determine the attribute prediction value of the current node.

[0428] Exemplarily, the first intra-frame prediction mode includes at least one of the following: intra-frame parent node prediction mode, intra-frame same-layer node prediction mode, and intra-frame parent node and same-layer node prediction mode; the first inter-frame prediction mode includes at least one of the following: inter-frame same-position parent node prediction mode, inter-frame same-position node prediction mode, and inter-frame same-position parent node and same-position node prediction mode; the first inter-frame intra-frame prediction mode includes at least one of the following: intra-frame prediction mode, inter-frame prediction mode and inter-frame intra-frame fusion prediction mode.

[0429] For example, if the inter-frame prediction node of the current node is valid, that is, the co-located node exists, the attribute of the prediction node is directly used as the attribute prediction value of the current node to be encoded;

[0430] The current node can find a node with exactly the same position as the current node in the cache of the reference frame: that is, if the same-position node exists, the AC coefficients of the M child nodes contained in the same-position node will be directly used as the AC coefficient attribute prediction values ​​of the N child nodes of the current node.

[0431] 1. If the AC coefficient of the predicted node is not zero: the AC coefficient of the predicted node is directly used as the predicted value;

[0432] 2. If the AC coefficient of the prediction node is zero, the AC coefficient of the corresponding child node of the intra-frame prediction will be used as the prediction value

[0433] The inter-frame prediction node of the current node is invalid: that is, the co-located node does not exist, then the attribute prediction value of the adjacent node in the frame is used as the attribute prediction value of the node to be encoded.

[0434] In another embodiment of the present application, see Figure 41, which shows a schematic flow chart of an encoding method provided by an embodiment of the present application. As shown in Figure 41, the method may include:

[0435] S4101: Determine a first syntax element identifier;

[0436] It should be noted that when encoding point cloud attribute information, a first syntax element identifier is used to indicate whether the mode selection function is enabled for the current sequence, current slice, or other coding mode during RAHT encoding. In some embodiments, the first syntax element identifier is a high-level syntax element and is set in an attribute header (ABH) parameter set. In other words, the first syntax element identifier is encoded in the attribute header parameter set.

[0437] Exemplarily, when the first syntax element identifier is a first value, it is determined that a mode selection function is enabled when performing RAHT encoding, that is, when performing RAHT encoding, a target coding mode is selected from candidate coding modes. When the first syntax element identifier is a second value, it is determined that a mode selection function is not enabled when performing RAHT encoding, and an existing RAHT coding scheme is used for attribute decoding.

[0438] In some embodiments, a first syntax element identifier is used to indicate whether the intra-frame prediction mode selection function is enabled for the current sequence, current slice, or other coding mode during RAHT encoding. In other embodiments, a first syntax element identifier may also be used to indicate whether the inter-frame prediction mode selection function is enabled for the current sequence, current slice, or other coding mode during RAHT encoding. In other embodiments, a first syntax element identifier may also be used to indicate whether the inter-frame or intra-frame prediction mode selection function is enabled for the current sequence, current slice, or other coding mode during RAHT encoding. Here, the inter-frame or intra-frame prediction mode can be understood as any one of the intra-frame prediction mode, the inter-frame prediction mode, or the inter-frame or intra-frame fusion prediction mode that can be used when performing attribute prediction.

[0439] It should be noted that RAHT encoding divides the reconstructed geometric information of the current encoding mode into layers based on the octree hierarchy, and decodes the attributes of each RAHT encoding layer. The current attribute RAHT encoding order is to divide the layers from the root node to the voxel level (1x1x1), thereby completing the encoding and attribute reconstruction of the entire point cloud attributes. As shown in Figure 40, we define the layer obtained by downsampling along the Z direction, Y direction, and X direction as a RAHT encoding layer, or layer.

[0440] S4102: If it is determined according to the first syntax element identifier that the mode selection is enabled when performing region-adaptive layered transform decoding on the current layer, determine the neighborhood geometric distribution information of the node on the current layer;

[0441] It should be noted that the neighborhood geometric distribution information is specifically used to characterize the geometric distribution characteristics of the reference node of the current layer node, and to determine whether the neighborhood geometric distribution information meets the first prediction condition. Specifically, it can be used to determine whether the neighborhood geometric distribution characteristics of the current layer node match the preset neighborhood geometric distribution characteristics. If so, the target coding mode is further determined based on the neighborhood attribute distribution information of the current layer node.

[0442] In some embodiments, the method may further include: according to the second syntax element identifier, starting the target prediction mode selected according to the current layer mode; determining a candidate coding mode according to the target prediction mode; encoding the second syntax element identifier, and writing the resulting coded bits into the bitstream. Specifically, the second syntax element identifier can be used to indicate the target prediction mode for starting mode selection for any coding mode to which the current layer belongs, and the corresponding neighborhood geometric distribution information may be different under different prediction modes. The coding mode of the embodiment of the present application can also start mode selection under a specific prediction mode.

[0443] In some embodiments, the target prediction mode may include one of the following: intra-frame prediction mode, inter-frame prediction mode and inter-frame intra-frame prediction mode. Exemplarily, according to the target prediction mode, determining the neighborhood geometric distribution information includes: when the target prediction mode is intra-frame prediction mode, determining the neighborhood geometric distribution information includes the number of neighboring nodes of the parent node of the current layer node and the number of neighboring nodes of the grandparent node; when the target prediction mode is inter-frame prediction mode, determining the neighborhood geometric distribution information includes the occupancy information of the same-position parent node of the current layer node and the occupancy information of the parent node; when the target prediction mode is inter-frame intra-frame prediction mode, determining the neighborhood geometric distribution information includes the occupancy information of the same-position parent node of the current layer node and the occupancy information of the parent node, as well as the number of neighboring nodes of the parent node of the current layer node and the number of neighboring nodes of the grandparent node.

[0444] Exemplarily, the target prediction mode is an intra-frame prediction mode, and the first prediction condition includes: the number of neighboring nodes of the parent node is greater than a first threshold, and the number of neighboring nodes of the grandparent node is greater than a second threshold; the target prediction mode is an inter-frame prediction mode, and the first prediction condition includes: the co-located parent node exists, and the placeholder information of the co-located parent node is the same as the placeholder information of the parent node; the target prediction mode is an inter-frame or intra-frame prediction mode, and the first prediction condition includes: the number of neighboring nodes of the parent node is greater than a first threshold, the number of neighboring nodes of the grandparent node is greater than a second threshold, the co-located parent node exists, and the placeholder information of the co-located parent node is the same as the placeholder information of the parent node.

[0445] It should be noted that, for the intra prediction mode, different thresholds may be set according to different decoding layers and hierarchical depths.

[0446] S4103: When the neighborhood geometric distribution information satisfies the first prediction condition, determine the neighborhood attribute distribution information of the current layer node;

[0447] It should be noted that the current layer node is any node in the current layer. In the embodiment of the present application, for the node whose neighborhood geometric distribution information meets the first prediction condition, the optimal coding mode is further determined by mode selection based on the neighborhood attribute distribution information.

[0448] The neighborhood attribute distribution information is specifically used to characterize the attribute distribution characteristics of the reference node of the current layer node. In some embodiments, the attribute distribution characteristics can be attribute change characteristics or attribute difference characteristics of the reference node. When the attribute change is small, it indicates that the current layer node can use attribute prediction. When the attribute change is large, it indicates that the current layer node cannot use attribute prediction.

[0449] S4104: Determine the target coding mode of the current layer node from the candidate coding modes according to the neighborhood attribute distribution information; wherein the candidate coding modes include: attribute prediction and transformation mode, attribute prediction mode; exemplary,

[0450] In some embodiments, the method further includes: encoding neighborhood attribute distribution information of the current layer node, and the neighborhood attribute distribution information can be determined by the encoding end and transmitted to the decoding end. Exemplarily, the encoding end makes a decision on the candidate encoding mode based on the neighborhood attribute distribution information of the current layer node to determine the target encoding mode, and encodes the target encoding mode index, and the decoding end determines the target decoding mode by decoding the mode index. It can be understood that since the target decoding mode is determined based on the neighborhood attribute distribution information of the current layer node, the target decoding mode index can also be used as information that characterizes the neighborhood attribute distribution characteristics.

[0451] In some embodiments, the neighborhood attribute distribution information includes the original value of the attribute of the current layer node; based on the neighborhood attribute distribution information, the target coding mode of the current layer node is determined from the candidate coding modes, including: encoding the attribute of the current layer node according to each candidate coding mode, and determining the attribute reconstruction value of the current layer node; performing cost calculation based on the attribute reconstruction value and the original value of the attribute of the current layer node, and determining the cost value corresponding to each candidate coding mode; based on the cost value corresponding to each candidate coding mode, determining the candidate coding mode corresponding to the minimum cost value as the target coding mode of the current layer node; based on the target coding mode, determining a third syntax element identifier; wherein the third syntax element identifier is used to indicate the target coding mode; encoding the third syntax element identifier, and writing the obtained coded bits into the bitstream.

[0452] In some embodiments, the cost function used may be a rate-distortion optimization cost RDO, and a rate-distortion optimization algorithm is used to select an optimal coding mode for predictive coding, thereby improving the coding efficiency of point cloud attributes.

[0453] Exemplarily, a cost calculation is performed based on the attribute reconstruction values ​​and original attribute values ​​of the current layer nodes to determine the cost value corresponding to each candidate coding mode, including: a cost calculation is performed based on the attribute reconstruction values ​​and original attribute values ​​of all nodes in the current layer or the coefficient group of the current layer to determine the cost value corresponding to the current layer or the coefficient group of the current layer in each candidate coding mode.

[0454] The rate-distortion optimization algorithm uses two prediction modes to predict and encode the attribute information of the current layer node at the encoding end. Finally, the rate-distortion optimization algorithm obtains the optimal encoding mode for the current layer and passes the optimal encoding mode to the decoding end. The decoding end uses the decoded mode obtained by parsing to reconstruct and restore the attribute information of the current layer point to be decoded. Among them, in the rate-distortion optimization algorithm, the distortion D between the reconstructed attribute and the original attribute of each prediction mode is first calculated. Then, the code stream R required for encoding in each decoding mode is obtained. The rate-distortion cost is calculated as follows: J = D + λxR (37)

[0455] Among them, λ can be calculated through attribute quantization parameters. The current λ calculation method is as follows:

[0456] The parameter N is currently set to different values ​​depending on reflectivity and color.

[0457] The coding mode of each layer is finally added to the ABH (Attribute Brick Header) parameter set.

[0458] In other embodiments, the cost function used may also be the sum of absolute error SAD, the sum of absolute transformation difference SATD, the mean square error MSE, the sum of squared errors SSD, the mean absolute difference MAD, the mean sum of squared errors MSD, the absolute value of the transformation coefficient DCT, the Hadamard transform, etc., which are not specifically limited here.

[0459] The neighborhood attribute distribution information may include a third syntax element identifier for indicating the target coding mode. The third syntax element identifier may be a syntax element identifier corresponding to each node of the current layer, or a syntax element identifier corresponding to some or all nodes of the current layer, or a syntax element identifier corresponding to multiple RAHT coding layers.

[0460] In some embodiments, the third syntax element identifier is used to indicate the target coding mode of the current layer, that is, all nodes in the current layer select the target coding mode for encoding when performing mode selection.

[0461] In some other embodiments, the third syntax element identifier is used to indicate the target coding mode of the coefficient group of the current layer, that is, all nodes in the coefficient group select the target coding mode for encoding when performing mode selection. In this scheme, the coefficient groups of the nodes of the current to-be-encoded layer are first split to obtain different coefficient groups. The number of nodes in each coefficient group is N, and the encoder selects the optimal coding mode for each coefficient group. For each coefficient group, the corresponding coding mode also needs to be transmitted to the decoder, so that the decoder reconstructs the attribute information using the coding mode corresponding to each coefficient group.

[0462] It should be noted that the length of the third syntax element identifier can be determined according to the number of candidate coding modes. The third syntax element identifier can be a high-level syntax element and is set in the attribute header information parameter set.

[0463] In other embodiments, determining the neighborhood attribute distribution information of the current layer node includes: obtaining an attribute reconstruction value of a first reference node of the current layer node; and determining the neighborhood attribute distribution information of the current layer node based on the attribute reconstruction value of the first reference node. In other words, the neighborhood attribute distribution information can be determined by the decoder based on the attribute reconstruction value of the reconstructed first reference node of the current layer node, and the optimal coding mode is implicitly derived based on the attribute reconstruction value of the reconstructed node, without the need for the encoder to transmit a mode index for determination.

[0464] Determine the neighborhood attribute distribution information of the current layer node, including: determining the attribute reconstruction value of the first reference node of the current layer node according to the target prediction mode; determine the neighborhood attribute distribution information of the current layer node according to the attribute reconstruction value of the first reference node.

[0465] The neighborhood attribute distribution information is related to the target prediction mode. Exemplarily, when the target prediction mode is an intra-frame prediction mode, the first reference node includes at least one of the following: the neighborhood node of the parent node, and the neighborhood node of the same layer that has been reconstructed; when the target prediction mode is an inter-frame prediction mode, the first reference node includes at least one of the following: the parent node and the same-position parent node in the reference frame, the neighborhood node of the parent node and the neighborhood node of the same-position parent node in the reference frame, the neighborhood node of the same layer that has been reconstructed, and the same-position neighborhood node of the same layer that has been reconstructed in the reference frame; when the target prediction mode is an inter-frame or intra-frame prediction mode, the first reference node includes at least one of the following: the parent node and the same-position parent node in the reference frame, the neighborhood node of the parent node and the neighborhood node of the same-position parent node in the reference frame, the neighborhood node of the same layer that has been reconstructed, and the same-position neighborhood node of the same layer that has been reconstructed in the reference frame.

[0466] Exemplarily, for intra prediction mode, the neighborhood attribute distribution information includes the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the parent node. For inter prediction mode, the neighborhood attribute distribution information includes the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the co-located parent node. For inter and intra prediction modes, the neighborhood attribute distribution information includes the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the co-located parent node, as well as the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the parent node.

[0467] The calculation functions used for difference calculation can be absolute error sum SAD, absolute transformation difference sum SATD, mean square error MSE, error square sum SSD, mean absolute difference MAD, mean error square sum MSD, absolute value of transformation coefficient DCT, Hadamard transform, etc., which are not specifically limited here.

[0468] In other embodiments, when the neighborhood geometric distribution information satisfies the first prediction condition, one method may be determined from the explicit indexing method and the implicit derivation method, and then the optimal encoding mode may be determined.

[0469] It should also be noted that in the embodiment of the present application, when the neighborhood geometric distribution information does not meet the first prediction condition, the target coding mode of the current layer node is determined to be a preset coding mode, where the preset coding mode can be one of the candidate coding modes provided in the embodiment of the present application, or other coding modes. Exemplarily, the preset coding mode is an attribute transform mode.

[0470] In some embodiments, the method further comprises: determining a candidate coding mode according to the target prediction mode. Different prediction modes may correspond to different candidate coding modes.

[0471] In some embodiments, candidate coding modes are determined based on the target prediction mode, including: if the target prediction mode is an intra-frame prediction mode, determining the attribute prediction and transformation mode includes the attribute intra-frame prediction and transformation mode; if the target prediction mode is an inter-frame prediction mode, determining the attribute prediction and transformation mode includes the attribute inter-frame prediction and transformation mode; if the target prediction mode is an inter-frame intra-frame prediction mode, determining the attribute prediction and transformation mode includes the attribute inter-frame prediction and transformation mode, the attribute intra-frame prediction and transformation mode, and the attribute inter-frame intra-frame prediction and transformation mode.

[0472] In some embodiments, the target coding mode of the current layer node is determined from the candidate coding modes according to the third syntax element identifier.

[0473] In other embodiments, if the neighborhood attribute distribution information satisfies the second prediction condition, the target coding mode is determined to be the attribute prediction and transformation mode; if the neighborhood attribute distribution information does not satisfy the second prediction condition, the target coding mode is determined to be the attribute transformation mode. The second prediction condition serves as a second judgment condition for whether to perform attribute prediction, specifically for determining whether the neighborhood attribute distribution characteristics of the current layer node meet the preset neighborhood attribute distribution characteristics for attribute prediction. When both the neighborhood geometric distribution characteristics and the neighborhood attribute distribution characteristics meet the prediction condition, the target coding mode is determined.

[0474] It should be noted that when the attribute prediction and transformation mode include two or more, a corresponding number of second prediction conditions may also be set. Exemplarily, when the target prediction mode is the intra-frame prediction mode, the second prediction condition includes the prediction condition corresponding to the attribute intra-frame prediction and the transformation mode; when the target prediction mode is the inter-frame prediction mode, the second prediction condition includes the prediction condition corresponding to the attribute inter-frame prediction and the transformation mode; when the target prediction mode is the inter-frame intra-frame prediction mode, the second prediction condition includes the prediction condition corresponding to the attribute intra-frame prediction and the transformation mode, the prediction condition corresponding to the attribute inter-frame prediction and the transformation mode, and the prediction condition corresponding to the attribute inter-frame intra-frame prediction and the transformation mode.

[0475] For the intra-frame prediction mode, the second prediction condition may include that the difference between the reconstruction attribute value of the neighboring node of the parent node and the reconstruction attribute value of the parent node is less than a first difference threshold. For the inter-frame prediction mode, the second prediction condition may include that the difference between the reconstruction attribute value of the neighboring node of the co-located parent node and the reconstruction attribute value of the co-located parent node is less than a second difference threshold. For the inter-frame and intra-frame fusion prediction mode, the second prediction condition may include that the difference between the reconstruction attribute value of the neighboring node of the parent node and the reconstruction attribute value of the parent node is less than a first difference threshold, and the difference between the reconstruction attribute value of the neighboring node of the co-located parent node and the reconstruction attribute value of the co-located parent node is less than a second difference threshold.

[0476] S4105: Perform attribute encoding on the current layer node according to the target coding mode to determine the attribute reconstruction value of the current layer node.

[0477] S4106: Encode the first syntax element identifier, and write the obtained coded bits into the bitstream.

[0478] In some embodiments, when the target coding mode is the attribute prediction and transformation mode, the specific coding process may include: determining the attribute prediction value of the current layer node; performing attribute transformation based on the attribute prediction value of the current layer node to obtain the AC coefficient prediction value of the current layer node; determining the AC coefficient residual value of the current layer node; determining the AC coefficient reconstruction value of the current layer node based on the AC coefficient prediction value and the AC coefficient residual value of the current layer node; performing inverse transformation based on the AC coefficient reconstruction value of the current layer node and the DC coefficient reconstruction value of the parent node to determine the attribute reconstruction value of the current layer node; encoding the AC coefficient residual value, and writing the obtained coding bits into the bitstream.

[0479] Exemplarily, determining the attribute prediction value of the current layer node includes: when the attribute prediction and transformation mode is attribute intra-frame prediction, determining a first intra-frame prediction mode; determining a first attribute prediction value of the current node based on the first intra-frame prediction mode; or, when the attribute prediction and transformation mode is attribute inter-frame prediction, determining a first inter-frame prediction mode; determining a second attribute prediction value of the current node based on the first inter-frame prediction mode; or, when the attribute prediction and transformation mode is attribute inter-frame intra-frame prediction, determining a first inter-frame intra-frame prediction mode; determining a third attribute prediction value of the current node based on the first inter-frame intra-frame prediction mode. That is, under different prediction modes, different attribute prediction methods can be used to determine the attribute prediction value of the current node.

[0480] Exemplarily, the first intra-frame prediction mode includes at least one of the following: intra-frame parent node prediction mode, intra-frame same-layer node prediction mode, and intra-frame parent node and same-layer node prediction mode; the first inter-frame prediction mode includes at least one of the following: inter-frame same-position parent node prediction mode, inter-frame same-position node prediction mode, and inter-frame same-position parent node and same-position node prediction mode; the first inter-frame intra-frame prediction mode includes at least one of the following: intra-frame prediction mode, inter-frame prediction mode and inter-frame intra-frame fusion prediction mode.

[0481] For example, if the inter-frame prediction node of the current node is valid, that is, the co-located node exists, the attribute of the prediction node is directly used as the attribute prediction value of the current node to be encoded;

[0482] The current node can find a node with exactly the same position as the current node in the cache of the reference frame: that is, if the same-position node exists, the AC coefficients of the M child nodes contained in the same-position node will be directly used as the AC coefficient attribute prediction values ​​of the N child nodes of the current node.

[0483] 1. If the AC coefficient of the predicted node is not zero: the AC coefficient of the predicted node is directly used as the predicted value;

[0484] 2. If the AC coefficient of the prediction node is zero, the AC coefficient of the corresponding child node of the intra-frame prediction will be used as the prediction value

[0485] The inter-frame prediction node of the current node is invalid: that is, the co-located node does not exist, then the attribute prediction value of the adjacent node in the frame is used as the attribute prediction value of the node to be encoded.

[0486] In summary, the embodiments of the present application have made improvements to the G-PCC attribute RAHT intra-frame coding by introducing one or more new intra-frame coding and decoding modes. First, by combining the attribute intra-frame prediction coding and the attribute transform coding scheme: attribute prediction (parent node intra-frame prediction and same-layer node intra-frame prediction) coding and attribute transform coding scheme, and before encoding the AC coefficients of different RAHT coding layers, the encoding end uses the rate-distortion optimization algorithm to obtain the optimal coding mode of the current RAHT coding layer, namely: prediction coding, transform coding, and finally passes the optimal coding mode of the current RAHT coding layer to the decoding end. The decoding end uses the coding mode of the current layer RAHT to adaptively restore the AC coefficient of the current layer, thereby completing the entire attribute RAHT coding, and ultimately improving the RAHT attribute coding efficiency.

[0487] The specific algorithm on the encoding side is as follows:

[0488] Step 1: Adaptively determine whether the nodes in the current layer can use attribute intra prediction based on the number of neighboring nodes in the current layer and the number of neighboring nodes of the parent node;

[0489] Step 2: If the nodes in the current layer can use attribute intra-frame prediction, then introduce the rate-distortion optimization algorithm for the current layer. By encoding each node in the current layer, the cost corresponding to each coding mode is calculated to obtain the optimal coding mode; or the optimal coding mode is determined by deducing the attribute distribution information of the reconstructed reference node.

[0490] Step 3: Finally, the optimal coding mode is used to predictively encode the attributes of the current layer node to obtain the attribute reconstruction value. It should be noted that if the rate-distortion optimization algorithm is introduced for mode selection in step 2, and if the reconstruction value of the optimal coding mode is cached, the cached attribute reconstruction value can be directly obtained here.

[0491] Step 4: Write the indication information of the optimal coding mode into the bitstream. It should be noted that this step can be omitted if step 2 adopts an implicit derivation method.

[0492] The specific algorithm of the decoding end is as follows:

[0493] Step 1: Adaptively determine whether the nodes in the current layer can use attribute intra prediction based on the number of neighboring nodes in the current layer and the number of neighboring nodes of the parent node;

[0494] Step 2: If the nodes in the current layer can use attribute intra prediction, the decoded bitstream obtains the best decoding mode for the current layer.

[0495] Step 3: Finally, use the best decoding mode to decode the attributes of the current layer node.

[0496] In the embodiment of the present application, when performing RAHT intra-frame encoding on attributes, a coding mode is introduced at each RAHT coding layer to adaptively select the attribute prediction and transformation mode or attribute transformation mode, and the coding mode is ultimately transmitted to the decoder, which uses the coding mode to reconstruct the attributes of the point cloud. In this solution, the focus is on introducing a coding mode at each RAHT coding layer, obtaining the optimal coding mode by utilizing a rate-distortion optimization algorithm at the encoder, and then using the decoding mode at the decoder to reconstruct the attributes of the point cloud. Currently, the coding mode of each layer is stored in the ABH, and the decoding mode of the RAHT coding layer is obtained at the decoder through the ABH. There is no restriction on the form in which this parameter is encoded. In addition, there is no restriction on the prediction coding method of each node (intra-frame parent node prediction coding method and intra-frame same-layer node prediction coding method). Instead, the intra-frame coding mode of the current node attribute is determined by considering the distribution characteristics of the AC coefficients obtained after the transformation.

[0497] The embodiment of the present application improves the G-PCC attribute RAHT inter-frame coding by introducing one or more new inter-frame coding and decoding modes. First, by combining three attribute prediction coding schemes: inter-frame prediction, inter-frame prediction + intra-frame prediction, and transform coding scheme, and before encoding the AC coefficients of different RAHT coding layers, the encoding end uses a rate-distortion optimization algorithm to obtain the optimal coding mode of the current RAHT coding layer, namely: inter-frame prediction coding scheme 2 + transform coding, intra-frame prediction coding + transform coding, and only transform coding. Finally, the optimal coding mode of the current RAHT coding layer is passed to the decoding end. The decoding end uses the coding mode of the current layer RAHT to adaptively restore the AC coefficients of the current layer, thereby completing the entire attribute RAHT coding, and ultimately improving the RAHT attribute coding efficiency.

[0498] The specific algorithm on the encoding side is as follows:

[0499] Step 1: Adaptively determine whether the nodes in the current layer can use attribute inter-frame prediction based on the number of neighboring nodes in the current layer and the number of neighboring nodes of the parent node;

[0500] Step 2: If the nodes in the current layer can use attribute prediction and attribute inter-frame prediction, the rate-distortion optimization algorithm is introduced for the current layer. By encoding each node in the current layer, the cost corresponding to each prediction coding mode is calculated to obtain the optimal coding mode.

[0501] Step 3: Finally, use the best encoding mode to encode the attributes of the current layer node.

[0502] Step 4: Write the indication information of the optimal coding mode into the bitstream.

[0503] The specific algorithm of the decoding end is as follows:

[0504] Step 1: Adaptively determine whether the nodes in the current layer can use attribute inter-frame prediction based on the number of neighboring nodes in the current layer and the number of neighboring nodes of the parent node;

[0505] Step 2: If the node in the current layer can use attribute prediction and can perform attribute inter-frame prediction, the node obtains the best prediction decoding mode of the current layer.

[0506] Step 3: Finally, the best prediction decoding mode is used to predict and decode the attributes of the current layer node.

[0507] In the embodiment of the present application, when performing RAHT inter-frame coding on attributes, a coding mode is introduced in each RAHT coding layer to adaptively select inter-frame intra-frame prediction and transform mode, intra-frame prediction and transform mode, and transform coding mode, and finally the coding mode is transmitted to the decoding end, which uses the coding mode to reconstruct the attributes of the point cloud. In this solution, the focus is on introducing a coding mode in each RAHT coding layer, obtaining the optimal coding mode by utilizing the rate-distortion optimization selection algorithm at the encoding end, and then using the decoding mode at the decoding end to reconstruct the attributes of the point cloud. Currently, the coding mode of each layer is stored in the ABH, and the decoding end uses the ABH to obtain the decoding mode of the RAHT coding layer. There is no restriction on the form in which this parameter is encoded.

[0508] Table 1 shows the test results of the attribute coding efficiency using the embodiment of the present application. It can be seen that after the introduction of the rate-distortion optimization algorithm, for sequences that can use inter-frame attribute prediction, the attribute coding pixel depth (bit per pixel, BPP) is reduced by about 3.9%, significantly improving the coding efficiency of point cloud attributes.

[0509] Table 1 shows the test results of the encoding efficiency of the attributes using the embodiment of the present application.

[0510] In the embodiment of the present application, the description of the syntax elements (Attribute data unit header syntax) in the attribute header information is shown in Table 2.

[0511] Table 2

[0512] Among them, attr_coding_type is used to indicate the attribute coding type. Attr_coding_type == 0 can be understood as indicating that the current slice uses RAHT coding. disableAttrInterPred is used to indicate that the inter-frame prediction mode is enabled. When it is false, it indicates that the inter-frame prediction mode is enabled. When it is true, it indicates that the inter-frame prediction mode is disabled. raht_prediction_enabled is used to indicate that the intra-frame prediction mode is enabled. When it is true, it indicates that the intra-frame prediction mode is enabled. When it is false, it indicates that the intra-frame prediction mode is disabled. attr_code_mode[i] is used to indicate the target coding mode corresponding to the i-th layer of the current slice.

[0513] By adopting the above technical solution, when performing adaptive layered transform RAHT encoding or decoding, one or more new encoding and decoding modes are introduced, and the mode selection is performed by comprehensively considering the neighborhood geometric distribution characteristics and neighborhood attribute distribution characteristics of the current layer node, providing the optimal encoding and decoding mode for the current node, thereby improving the RAHT attribute encoding and decoding efficiency.

[0514] In another embodiment of the present application, based on the same inventive concept as the above embodiment, see Figure 42, which shows a schematic diagram of the composition structure of an encoder provided by an embodiment of the present application. As shown in Figure 42, the encoder 110 may include a first determination unit 111, a second determination unit 112 and an encoding unit 113; wherein,

[0515] A first determining unit 111 is configured to determine a first syntax element identifier;

[0516] The first determining unit 111 is further configured to determine the neighborhood geometric distribution information of the node in the current layer if it is determined according to the first syntax element identifier that mode selection is enabled when region adaptive layered transform coding is performed on the current layer;

[0517] The first determining unit 111 is further configured to determine neighborhood attribute distribution information of the current layer node if the neighborhood geometric distribution information satisfies the first prediction condition;

[0518] The second determining unit 112 is configured to determine a target coding mode for the current layer node from candidate coding modes based on the neighborhood attribute distribution information; wherein the candidate coding modes include: attribute prediction and transformation mode, and attribute transformation mode;

[0519] The second determining unit 112 is configured to perform attribute encoding on the current layer node according to the target coding mode and determine the attribute reconstruction value of the current layer node;

[0520] The encoding unit 113 is configured to encode the first syntax element identifier and write the obtained coded bits into the bitstream.

[0521] It can be understood that each functional unit of the encoder also executes the encoding method of any one of the aforementioned embodiments, which will not be repeated here.

[0522] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.

[0523] If the integrated unit is implemented as a software functional module and not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0524] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 110. The computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, it implements the method of any one of the aforementioned embodiments.

[0525] Based on the composition of the encoder 110 and the computer-readable storage medium, refer to Figure 43, which shows a specific hardware structure diagram of the encoder 110 provided in an embodiment of the present application. As shown in Figure 43, the encoder 110 may include: a first memory 115 and a first processor 116, a first communication interface 117 and a first bus system 118. The first memory 115, the first processor 116, and the first communication interface 117 are coupled together through the first bus system 118. It can be understood that the first bus system 118 is used to achieve connection and communication between these components. In addition to the data bus, the first bus system 118 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as the first bus system 118 in Figure 20. Among them,

[0526] The first communication interface 117 is used to receive and send signals when sending and receiving information with other external network elements;

[0527] A first memory 115, configured to store computer programs that can be run on the first processor;

[0528] determining a first syntax element identifier;

[0529] If it is determined according to the first syntax element identifier that mode selection is enabled when region-adaptive layered transform coding is performed on the current layer, determining neighborhood geometric distribution information of nodes in the current layer;

[0530] When the neighborhood geometric distribution information satisfies the first prediction condition, determining the neighborhood attribute distribution information of the current layer node;

[0531] Determine the target coding mode of the current layer node from the candidate coding modes based on the neighborhood attribute distribution information; wherein the candidate decoding modes include: attribute prediction and transformation mode, attribute transformation mode;

[0532] Perform attribute encoding on the current layer node according to the target encoding mode to determine the attribute reconstruction value of the current layer node;

[0533] The first syntax element identifier is encoded, and the obtained encoded bits are written into the bitstream.

[0534] It is understood that the first memory 115 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The first memory 115 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0535] The first processor 116 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the first processor 116. The above-mentioned first processor 116 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the first memory 115 , and the first processor 116 reads the information in the first memory 115 and completes the steps of the above method in combination with its hardware.

[0536] It is understood that the embodiments described herein can be implemented with hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSP devices, DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions of the present application or a combination thereof. For software implementation, the technology of the present application can be implemented by a module (such as a process, a function, etc.) that performs the functions of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0537] Optionally, as another embodiment, the first processor 116 is further configured to execute the encoding method of any one of the aforementioned embodiments when running the computer program.

[0538] This embodiment provides an encoder in which a mode is selected by comprehensively considering the neighborhood geometric distribution characteristics and neighborhood attribute distribution characteristics of the current layer node, thereby providing an optimal coding mode for the current node, thereby improving the RAHT attribute coding efficiency.

[0539] An embodiment of the present application further provides a computer-readable storage medium, which stores a code stream generated by the encoding method of any one of the aforementioned embodiments.

[0540] An embodiment of the present application further provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following: a first syntax element identifier, a second syntax element identifier, a third syntax element identifier, and coefficient residual information.

[0541] In another embodiment of the present application, based on the same inventive concept as the above embodiment, see Figure 44, which shows a schematic diagram of the composition structure of a decoder provided by the embodiment of the present application. As shown in Figure 44, the decoder 120 may include: a decoding unit 121, a third determining unit 122 and a fourth determining unit 123; wherein,

[0542] The decoding unit 121 is configured to decode the code stream and determine a first syntax element identifier;

[0543] The third determining unit 122 is configured to determine the neighborhood geometric distribution information of the node in the current layer if it is determined according to the first syntax element identifier that the mode selection is started when the current layer performs region adaptive layered transform decoding;

[0544] The third determining unit 122 is further configured to determine the neighborhood attribute distribution information of the current layer node if the neighborhood geometric distribution information satisfies the first prediction condition;

[0545] The fourth determining unit 123 is configured to determine a target decoding mode for the current layer node from candidate decoding modes based on the neighborhood attribute distribution information; wherein the candidate decoding modes include: attribute prediction and transformation mode, attribute transformation mode;

[0546] The fourth determining unit 123 is further configured to perform attribute decoding on the current layer node according to the target decoding mode, and determine the attribute reconstruction value of the current layer node.

[0547] It can be understood that each functional unit of the decoder also executes the decoding method of any one of the aforementioned embodiments, which will not be repeated here.

[0548] It is understood that in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system. Furthermore, the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The aforementioned integrated units can be implemented in the form of hardware or software functional modules.

[0549] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, this embodiment provides a computer-readable storage medium for use in decoder 120. The computer-readable storage medium stores a computer program that, when executed by the second processor, implements any of the methods in the aforementioned embodiments.

[0550] Based on the composition of the above-mentioned decoder 120 and the computer-readable storage medium, refer to Figure 45, which shows a specific hardware structure diagram of the decoder 120 provided in an embodiment of the present application. As shown in Figure 45, the decoder 120 may include: a second memory 127 and a second processor 124, a second communication interface 125 and a second bus system 126. The second memory 127 and the second processor 124, and the second communication interface 125 are coupled together through the second bus system 126. It can be understood that the second bus system 126 is used to realize the connection and communication between these components. In addition to the data bus, the second bus system 126 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as the second bus system 126 in Figure 22. Among them,

[0551] The second communication interface 125 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;

[0552] A second memory 127, for storing a computer program that can be run on the second processor;

[0553] In some embodiments, the second processor 124 is configured to, when running the computer program, execute:

[0554] Decoding the code stream and determining a first syntax element identifier;

[0555] If it is determined according to the first syntax element identifier that the mode selection is started when the current layer performs region adaptive layered transform decoding, determining the neighborhood geometric distribution information of the node in the current layer;

[0556] When the neighborhood geometric distribution information satisfies the first prediction condition, determining the neighborhood attribute distribution information of the current layer node;

[0557] According to the neighborhood attribute distribution information, the target decoding mode of the current layer node is determined from the candidate decoding modes; the candidate decoding modes include: attribute prediction and transformation mode, attribute transformation mode;

[0558] The attributes of the current layer nodes are decoded according to the target decoding mode to determine the attribute reconstruction value of the current layer nodes.

[0559] Optionally, as another embodiment, the second processor 124 is further configured to execute any one of the methods in the foregoing embodiments when running a computer program.

[0560] It can be understood that the hardware functions of the second memory 127 and the first memory 115 are similar, and the hardware functions of the second processor 124 and the first processor 116 are similar; they are not described in detail here.

[0561] This embodiment provides a decoder in which a mode is selected by comprehensively considering the neighborhood geometric distribution characteristics and neighborhood attribute distribution characteristics of the current layer node, thereby providing an optimal decoding mode for the current node, thereby improving the RAHT attribute decoding efficiency.

[0562] In yet another embodiment of the present application, see Figure 46 , which shows a schematic diagram of the structure of a coding and decoding system provided by an embodiment of the present application. As shown in Figure 46 , the coding and decoding system 130 may include an encoder 131 and a decoder 132 .

[0563] In the embodiment of the present application, the encoder 131 may be the encoder described in any one of the aforementioned embodiments, and the decoder 132 may be the decoder described in any one of the aforementioned embodiments.

[0564] It should be noted that, in this application, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0565] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0566] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0567] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0568] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0569] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims. Industrial Applicability

[0570] In an embodiment of the present application, a coding and decoding method, an encoder, a decoder, and a storage medium are provided. The method includes: determining the neighborhood geometric distribution information of the current layer node when the mode selection is started when the current layer performs regional adaptive layered transform decoding; determining the neighborhood attribute distribution information of the current layer node when the neighborhood geometric distribution information meets the first prediction condition; determining the target decoding mode of the current layer node from the candidate decoding modes based on the neighborhood attribute distribution information; performing attribute decoding on the current layer node according to the target decoding mode to determine the attribute reconstruction value of the current layer node. By comprehensively considering the neighborhood geometric distribution characteristics and neighborhood attribute distribution characteristics of the current layer node for mode selection, the optimal coding and decoding mode is provided for the current node, thereby improving the RAHT attribute coding and decoding efficiency.

Claims

1. A decoding method, applied to a decoder, the method comprises: decoding a bitstream to determine a first syntax element identifier; when starting mode selection for region adaptive hierarchical transform decoding in the current layer is determined according to the first syntax element identifier, determining neighborhood geometric distribution information of a current layer node; when the neighborhood geometric distribution information satisfies a first prediction condition, determining neighborhood attribute distribution information of the current layer node; determining a target decoding mode of the current layer node from candidate decoding modes according to the neighborhood attribute distribution information; wherein the candidate decoding modes include: attribute prediction and transform mode, attribute transform mode; performing attribute decoding on the current layer node according to the target decoding mode to determine an attribute reconstruction value of the current layer node.

2. The method according to claim 1, wherein, the method further comprises: decoding a bitstream to determine a second syntax element identifier; determining a target prediction mode for starting mode selection in the current layer according to the second syntax element identifier; determining the candidate decoding modes according to the target prediction mode.

3. The method according to claim 2, wherein, the determining the candidate decoding modes according to the target prediction mode includes: when the target prediction mode is an intra prediction mode, determining that the attribute prediction and transform mode includes an attribute intra prediction and transform mode; when the target prediction mode is an inter prediction mode, determining that the attribute prediction and transform mode includes an attribute inter prediction and transform mode; when the target prediction mode is an inter-intra prediction mode, determining that the attribute prediction and transform mode includes an attribute inter prediction and transform mode, an attribute intra prediction and transform mode, and an attribute inter-intra prediction and transform mode.

4. The method according to claim 2 or 3, wherein, the determining the neighborhood attribute distribution information of the current layer node includes: determining an attribute reconstruction value of a first reference node of the current layer node according to the target prediction mode; determining the neighborhood attribute distribution information of the current layer node according to the attribute reconstruction value of the first reference node.

5. The method according to claim 4, wherein, when the target prediction mode is an intra prediction mode, the first reference node includes at least one of the following: a neighborhood node of a parent node, a neighborhood node that has been reconstructed in the same layer; when the target prediction mode is an inter prediction mode, the first reference node includes at least one of the following: a parent node and a co-located parent node in a reference frame, a neighborhood node of a parent node and a neighborhood node of a co-located parent node in a reference frame, a neighborhood node that has been reconstructed in the same layer and a co-located neighborhood node that has been reconstructed in the same layer in a reference frame; when the target prediction mode is an inter-intra prediction mode, the first reference node includes at least one of the following: a parent node and a co-located parent node in a reference frame, a neighborhood node of a parent node and a neighborhood node of a co-located parent node in a reference frame, a neighborhood node that has been reconstructed in the same layer and a co-located neighborhood node that has been reconstructed in the same layer in a reference frame.

6. The method according to claim 5, wherein, The neighborhood attribute distribution information includes the difference between the reconstruction attribute value of the first reference node and the reconstruction attribute value of the parent node, and the difference between the reconstruction attribute of the first reference node and the reconstruction attribute value of the co-parent node.

7. The method according to any one of claims 4 to 6, wherein, determining the target decoding mode of the current layer node from the candidate decoding modes according to the neighborhood attribute distribution information includes: if the neighborhood attribute distribution information satisfies the second prediction condition, determining the target decoding mode as the attribute prediction and transformation mode; if the neighborhood attribute distribution information does not satisfy the second prediction condition, determining the target decoding mode as the attribute transformation mode.

8. The method according to claim 7, wherein, the target prediction mode is the intra prediction mode, and the second prediction condition includes the prediction conditions corresponding to the intra prediction and transformation mode of attributes; the target prediction mode is the inter prediction mode, and the second prediction condition includes the prediction conditions corresponding to the inter prediction and transformation mode of attributes; the target prediction mode is the inter-intra prediction mode, and the second prediction condition includes the prediction conditions corresponding to the intra prediction and transformation mode of attributes, the prediction conditions corresponding to the inter prediction and transformation mode of attributes, and the prediction conditions corresponding to the inter-intra prediction and transformation mode of attributes.

9. The method according to any one of claims 1 to 3, wherein, determining the neighborhood attribute distribution information of the current layer node includes: decoding the bitstream to determine the neighborhood attribute distribution information.

10. The method according to claim 9, wherein, the neighborhood attribute distribution information includes a third syntax element identifier for indicating the target decoding mode.

11. The method according to claim 10, wherein, the third syntax element identifier is used to indicate the target decoding mode of the current layer; or the third syntax element identifier is used to indicate the target decoding mode of the coefficient group of the current layer.

12. The method according to any one of claims 1 to 11, wherein, the first syntax element identifier, the second syntax element identifier, and the third syntax element identifier are set in the attribute header information parameter set.

13. The method according to claim 2, wherein, determining the neighborhood geometric distribution information of the current layer node includes: determining the neighborhood geometric distribution information according to the target prediction mode.

14. The method according to claim 13, wherein, determining the neighborhood geometric distribution information according to the target prediction mode includes: when the target prediction mode is the intra prediction mode, determining the neighborhood geometric distribution information includes the number of neighborhood nodes of the parent node of the current layer node and the number of neighborhood nodes of the grandparent node; when the target prediction mode is the inter prediction mode, determining the neighborhood geometric distribution information includes the occupancy information of the co-parent node of the current layer node and the occupancy information of the parent node; when the target prediction mode is the inter-intra prediction mode, determining the neighborhood geometric distribution information includes the occupancy information of the co-parent node of the current layer node and the occupancy information of the parent node, and the number of neighborhood nodes of the parent node of the current layer node and the number of neighborhood nodes of the grandparent node.

15. The method according to claim 13, wherein, when the target prediction mode is an intra prediction mode, the first prediction condition includes: the number of neighboring nodes of the parent node is greater than a first threshold, and the number of neighboring nodes of the grandparent node is greater than a second threshold; when the target prediction mode is an inter prediction mode, the first prediction condition includes: the co-located parent node exists, and the occupancy information of the co-located parent node is the same as the occupancy information of the parent node; when the target prediction mode is an inter-intra prediction mode, the first prediction condition includes: the number of neighboring nodes of the parent node is greater than a first threshold, the number of neighboring nodes of the grandparent node is greater than a second threshold, the co-located parent node exists, and the occupancy information of the co-located parent node is the same as the occupancy information of the parent node.

16. The method according to claim 1, wherein, when the target decoding mode is the attribute prediction and transform mode, the attribute decoding of the current layer node according to the target decoding mode to determine the attribute reconstruction value of the current layer node includes: determining an attribute prediction value of the current layer node; performing attribute transformation according to the attribute prediction value of the current layer node to obtain an AC coefficient prediction value of the current layer node; decoding the bitstream to determine an AC coefficient residual value of the current layer node; determining an AC coefficient reconstruction value of the current layer node according to the AC coefficient prediction value and the AC coefficient residual value of the current layer node; performing inverse transformation according to the AC coefficient reconstruction value of the current layer node and the DC coefficient reconstruction value of the parent node to determine the attribute reconstruction value of the current layer node.

17. The method according to claim 16, wherein, the determining of the attribute prediction value of the current layer node includes: when the attribute prediction and transform mode is intra attribute prediction, determining a first intra prediction mode; determining a first attribute prediction value of the current node according to the first intra prediction mode; or, when the attribute prediction and transform mode is inter attribute prediction, determining a first inter prediction mode; determining a second attribute prediction value of the current node according to the first inter prediction mode; or, when the attribute prediction and transform mode is inter-intra attribute prediction, determining a first inter-intra prediction mode; determining a third attribute prediction value of the current node according to the first inter-intra prediction mode.

18. The method according to claim 16, wherein, the first intra prediction mode includes at least one of the following: intra parent node prediction mode, intra same-layer node prediction mode, and intra parent node and same-layer node prediction mode; the first inter prediction mode includes at least one of the following: inter co-located parent node prediction mode, inter co-located node prediction mode, and inter co-located parent node and co-located node prediction mode; the first inter-intra prediction mode includes at least one of the following: intra prediction mode, inter prediction mode, and inter-intra fusion prediction mode.

19. The method according to claim 1, wherein, When the target decoding mode is the attribute transformation mode, performing attribute decoding on the current layer node according to the target decoding mode to determine the attribute reconstruction value of the current layer node includes: Decoding the bitstream to determine the reconstructed value of the AC coefficient of the current layer node; Performing an inverse transform based on the reconstructed value of the AC coefficient of the current layer node and the reconstructed value of the DC coefficient of the parent node to determine the attribute reconstruction value of the current layer node.

20. The method according to any one of claims 1 to 19, wherein, the method further includes: When the neighborhood geometric distribution information does not satisfy the first prediction condition, determining that the target decoding mode of the current layer node is the attribute transformation mode; Performing attribute decoding on the current layer node according to the attribute transformation mode to determine the attribute reconstruction value of the current layer node.

21. An encoding method applied to an encoder, the method includes: Determining a first syntax element identifier; If it is determined that mode selection is started when performing region adaptive hierarchical transform coding on the current layer according to the first syntax element identifier, determining the neighborhood geometric distribution information of the current layer node; When the neighborhood geometric distribution information satisfies the first prediction condition, determining the neighborhood attribute distribution information of the current layer node; According to the neighborhood attribute distribution information, determining the target coding mode of the current layer node from the candidate coding modes; wherein, the candidate coding modes include: attribute prediction and transformation mode, attribute transformation mode; Performing attribute coding on the current layer node according to the target coding mode to determine the attribute reconstruction value of the current layer node; Encoding the first syntax element identifier and writing the obtained encoded bits into the bitstream.

22. The method according to claim 21, wherein, the method further includes: According to the second syntax element identifier, determining the target prediction mode for starting mode selection on the current layer; According to the target prediction mode, determining the candidate coding modes; Encoding the second syntax element identifier and writing the obtained encoded bits into the bitstream.

23. The method according to claim 22, wherein, the determining the candidate coding modes according to the target prediction mode includes: When the target prediction mode is the intra prediction mode, determining that the attribute prediction and transformation mode includes the attribute intra prediction and transformation mode; When the target prediction mode is the inter prediction mode, determining that the attribute prediction and transformation mode includes the attribute inter prediction and transformation mode; When the target prediction mode is the inter-intra prediction mode, determining that the attribute prediction and transformation mode includes the attribute inter prediction and transformation mode, the attribute intra prediction and transformation mode, and the attribute inter-intra prediction and transformation mode.

24. The method according to claim 22 or 23, wherein, the determining the neighborhood attribute distribution information of the current layer node includes: According to the target prediction mode, determining the attribute reconstruction value of the first reference node of the current layer node; According to the attribute reconstruction value of the first reference node, determining the neighborhood attribute distribution information of the current layer node.

25. The method according to claim 24, wherein, The target prediction mode is an intra prediction mode, and the first reference node includes at least one of the following: neighborhood nodes of the parent node, neighborhood nodes that have been reconstructed in the same layer; When the target prediction mode is an inter prediction mode, the first reference node includes at least one of the following: the parent node and the co-located parent node in the reference frame, the neighborhood nodes of the parent node and the neighborhood nodes of the co-located parent node in the reference frame, the neighborhood nodes that have been reconstructed in the same layer and the co-located neighborhood nodes that have been reconstructed in the same layer in the reference frame; The target prediction mode is an inter-intra prediction mode, and the first reference node includes at least one of the following: the parent node and the co-located parent node in the reference frame, the neighborhood nodes of the parent node and the neighborhood nodes of the co-located parent node in the reference frame, the neighborhood nodes that have been reconstructed in the same layer and the co-located neighborhood nodes that have been reconstructed in the same layer in the reference frame.

26. The method according to claim 25, wherein, The neighborhood attribute distribution information includes the difference between the reconstruction attribute value of the first reference node and the reconstruction attribute value of the parent node, and the difference between the reconstruction attribute of the first reference node and the reconstruction attribute value of the co-located parent node.

27. The method according to any one of claims 24 to 26, wherein, Determining the target coding mode of the current layer node from the candidate coding modes according to the neighborhood attribute distribution information includes: If the neighborhood attribute distribution information meets the second prediction condition, determining that the target coding mode is an attribute prediction and transform mode; If the neighborhood attribute distribution information does not meet the second prediction condition, determining that the target coding mode is an attribute transform mode.

28. The method according to claim 27, wherein, When the target prediction mode is an intra prediction mode, the second prediction condition includes the prediction conditions corresponding to the attribute intra prediction and transform mode; When the target prediction mode is an inter prediction mode, the second prediction condition includes the prediction conditions corresponding to the attribute inter prediction and transform mode; When the target prediction mode is an inter-intra prediction mode, the second prediction condition includes the prediction conditions corresponding to the attribute intra prediction and transform mode, the prediction conditions corresponding to the attribute inter prediction and transform mode, and the prediction conditions corresponding to the attribute inter-intra prediction and transform mode.

29. The method according to any one of claims 21 to 23, wherein, The neighborhood attribute distribution information includes the original attribute value of the current layer node; Determining the target coding mode of the current layer node from the candidate coding modes according to the neighborhood attribute distribution information includes: Performing attribute coding on the current layer node according to each candidate coding mode to determine the attribute reconstruction value of the current layer node; Calculating the cost according to the attribute reconstruction value and the original attribute value of the current layer node to determine the cost value corresponding to each candidate coding mode; Determining the candidate coding mode corresponding to the minimum cost value as the target coding mode of the current layer node according to the cost value corresponding to each candidate coding mode; Determining a third syntax element identifier according to the target coding mode; wherein, the third syntax element identifier is used to indicate the target coding mode; Encoding the third syntax element identifier and writing the obtained encoded bits into the bitstream.

30. The method according to claim 29, wherein, the third syntax element identifier is used to indicate the target coding mode of the current layer; or the third syntax element identifier is used to indicate the target coding mode of the coefficient group of the current layer.

31. The method according to claim 29, wherein, the calculating the cost value corresponding to each candidate coding mode according to the reconstructed value of the attribute and the original value of the attribute of the current layer node includes: calculating the cost value according to the reconstructed value of the attribute and the original value of the attribute of all nodes in the current layer or the coefficient group of the current layer, and determining the cost value corresponding to each candidate coding mode of the current layer or the coefficient group of the current layer.

32. The method according to any one of claims 21 to 31, wherein, the first syntax element identifier, the second syntax element identifier, and the third syntax element identifier are set in the attribute header information parameter set.

33. The method according to claim 32, wherein, the determining the neighborhood geometric distribution information of the current layer node includes: determining the neighborhood geometric distribution information according to the target prediction mode.

34. The method according to claim 33, wherein, the determining the neighborhood geometric distribution information according to the target prediction mode includes: when the target prediction mode is an intra prediction mode, determining the neighborhood geometric distribution information includes the number of neighborhood nodes of the parent node of the current layer node and the number of neighborhood nodes of the grandparent node; when the target prediction mode is an inter prediction mode, determining the neighborhood geometric distribution information includes the occupancy information of the co-located parent node of the current layer node and the occupancy information of the parent node; when the target prediction mode is an inter-intra prediction mode, determining the neighborhood geometric distribution information includes the occupancy information of the co-located parent node of the current layer node and the occupancy information of the parent node, and the number of neighborhood nodes of the parent node of the current layer node and the number of neighborhood nodes of the grandparent node.

35. The method according to claim 33, wherein, when the target prediction mode is an intra prediction mode, the first prediction condition includes: the number of neighborhood nodes of the parent node is greater than a first threshold, and the number of neighborhood nodes of the grandparent node is greater than a second threshold; when the target prediction mode is an inter prediction mode, the first prediction condition includes: the co-located parent node exists, and the occupancy information of the co-located parent node is the same as the occupancy information of the parent node; when the target prediction mode is an inter-intra prediction mode, the first prediction condition includes: the number of neighborhood nodes of the parent node is greater than a first threshold, the number of neighborhood nodes of the grandparent node is greater than a second threshold, the co-located parent node exists, and the occupancy information of the co-located parent node is the same as the occupancy information of the parent node.

36. The method according to claim 21, wherein, when the target coding mode is the attribute prediction and transformation mode, the encoding the attribute of the current layer node according to the target coding mode and determining the reconstructed value of the attribute of the current layer node includes: determining the attribute prediction value of the current layer node; Perform attribute transformation based on the predicted value of the attribute of the current layer node to obtain the predicted value of the AC coefficient of the current layer node; Determine the residual value of the AC coefficient of the current layer node; Determine the reconstructed value of the AC coefficient of the current layer node according to the predicted value of the AC coefficient and the residual value of the AC coefficient of the current layer node; Perform inverse transformation according to the reconstructed value of the AC coefficient of the current layer node and the reconstructed value of the DC coefficient of the parent node to determine the reconstructed value of the attribute of the current layer node; Encode the residual value of the AC coefficient and write the obtained encoded bits into the bitstream.

37. The method according to claim 36, wherein, the determining the predicted value of the attribute of the current layer node includes: when the attribute prediction and transformation mode is intra-attribute prediction, determining the first intra-frame prediction mode; determining the first predicted value of the attribute of the current node according to the first intra-frame prediction mode; or, when the attribute prediction and transformation mode is inter-attribute prediction, determining the first inter-frame prediction mode; determining the second predicted value of the attribute of the current node according to the first inter-frame prediction mode; or, when the attribute prediction and transformation mode is intra-inter-attribute prediction, determining the first intra-inter-frame prediction mode; determining the third predicted value of the attribute of the current node according to the first intra-inter-frame prediction mode.

38. The method according to claim 36, wherein, the first intra-frame prediction mode includes at least one of the following: intra-frame parent node prediction mode, intra-frame same-layer node prediction mode, and intra-frame parent node and same-layer node prediction mode; the first inter-frame prediction mode includes at least one of the following: inter-frame co-located parent node prediction mode, inter-frame co-located node prediction mode, and inter-frame co-located parent node and co-located node prediction mode; the first intra-inter-frame prediction mode includes at least one of the following: intra-frame prediction mode, inter-frame prediction mode, and intra-inter-frame fusion prediction mode.

39. The method according to claim 21, wherein, when the target coding mode is the attribute transformation mode, the performing attribute coding on the current layer node according to the target coding mode to determine the reconstructed value of the attribute of the current layer node includes: determining the reconstructed value of the AC coefficient of the current layer node; performing inverse transformation according to the reconstructed value of the AC coefficient of the current layer node and the reconstructed value of the DC coefficient of the parent node to determine the reconstructed value of the attribute of the current layer node.

40. The method according to any one of claims 21 to 39, wherein, the method further includes: when the neighborhood geometric distribution information does not satisfy the first prediction condition, determining that the target coding mode of the current layer node is the attribute transformation mode; performing attribute coding on the current layer node according to the attribute transformation mode to determine the reconstructed value of the attribute of the current layer node.

41. An encoder, the encoder includes a first determination unit, a second determination unit, and a coding unit; wherein, the first determination unit is configured to determine a first syntax element identifier; the first determination unit is further configured to determine the neighborhood geometric distribution information of the current layer node if mode selection is started when it is determined according to the first syntax element identifier that region adaptive hierarchical transform coding is performed on the current layer; The first determining unit is further configured to determine the neighborhood attribute distribution information of the current layer node when the neighborhood geometric distribution information meets the first prediction condition; The second determining unit is configured to determine the target coding mode of the current layer node from the candidate coding modes according to the neighborhood attribute distribution information; wherein the candidate coding modes include: attribute prediction and transformation mode, and attribute transformation mode; The second determining unit is configured to perform attribute coding on the current layer node according to the target coding mode to determine the attribute reconstruction value of the current layer node; The coding unit is configured to code the first syntax element identifier and write the obtained coded bits into the code stream.

42. An encoder, the encoder comprising a first memory and a first processor; Wherein, The first memory is used to store a computer program that can run on the first processor; The first processor is configured to execute the method according to any one of claims 21 to 40 when running the computer program.

43. A decoder, the decoder comprising a decoding unit, a third determining unit and a fourth determining unit; Wherein, The decoding unit is configured to decode the code stream to determine the first syntax element identifier; The third determining unit is configured to determine the neighborhood geometric distribution information of the current layer node when starting mode selection according to the first syntax element identifier to determine that the current layer performs region adaptive hierarchical transform decoding; The third determining unit is further configured to determine the neighborhood attribute distribution information of the current layer node when the neighborhood geometric distribution information meets the first prediction condition; The fourth determining unit is configured to determine the target decoding mode of the current layer node from the candidate decoding modes according to the neighborhood attribute distribution information; wherein the candidate decoding modes include: attribute prediction and transformation mode, and attribute transformation mode; The fourth determining unit is further configured to perform attribute decoding on the current layer node according to the target decoding mode to determine the attribute reconstruction value of the current layer node.

44. A decoder, the decoder comprising a second memory and a second processor; Wherein, The second memory is used to store a computer program that can run on the second processor; The second processor is configured to execute the method according to any one of claims 1 to 20 when running the computer program.

45. A computer-readable storage medium, Wherein, The computer-readable storage medium stores the code stream generated by the coding method according to any one of claims 21 to 40.

46. A computer-readable storage medium, Wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the method according to any one of claims 1 to 20, or implements the method according to any one of claims 21 to 40.