Encoding method, decoding method, decoder, encoder, and computer readable storage medium
By adaptively dividing point cloud nodes in point cloud compression and using sparse convolution networks, the problem of low encoding and decoding efficiency of sparse scene-based point clouds is solved, and more efficient encoding and decoding performance is achieved.
Patent Information
- Application Number
- PCT/CN2024/071847
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-17
AI Technical Summary
The existing point cloud geometric compression coding scheme of sparse convolutional neural network model has low encoding and decoding efficiency on sparse scene-like point clouds, resulting in a degradation in the encoding and decoding performance.
By determining the effective accuracy or density of the current scale point cloud in three spatial dimensions, adaptively divide the current scale node into child nodes, and using the sparse convolution network and the appropriate convolution kernel K, K nearest neighbor search algorithm is performed to extract geometric feature values, avoid increasing the size of the convolution kernel K, and achieve more efficient encoding and decoding.
It improves the geometric encoding and decoding efficiency of point clouds and improves the encoding and decoding performance, especially on sparse scene point clouds.
Smart Images

Figure CN2024071847_17072025_PF_FP_ABST
Abstract
Description
Coding and decoding method, decoder, encoder and computer-readable storage medium Technical Field
[0001] The present application relates to point cloud compression coding and decoding technology, and in particular to a coding and decoding method, a decoder, an encoder and a computer-readable storage medium. Background Art
[0002] A point cloud is a collection of points that stores the geometric position and associated attribute information of each point, thereby accurately describing objects in space. Point cloud data is massive; a single point cloud frame can contain millions of points. This poses significant challenges to its efficient storage and transmission. Therefore, compression technology is used to reduce redundant information in point cloud storage, facilitating subsequent processing.
[0003] In the current sparse convolution-based neural network model point cloud geometry compression (Sparse PCGC) encoding scheme, the encoding and decoding efficiency of sparse scene point clouds, such as lidar point clouds, is low, thereby reducing the encoding and decoding performance.
[0004] Summary of the Invention
[0005] The embodiments of the present application provide a coding and decoding method, a decoder, an encoder, and a computer-readable storage medium, which can improve coding and decoding efficiency, thereby improving coding and decoding performance.
[0006] The technical solution of this application is achieved as follows:
[0007] The present invention provides a decoding method, which is applied to a decoder and includes:
[0008] Parsing the bitstream to determine the effective accuracy of the current-scale point cloud in three spatial dimensions, or a target dimension identifier corresponding to the current-scale point cloud; the target dimension identifier is determined based on the density of the current-scale point cloud in the three spatial dimensions;
[0009] Determining a child node corresponding to a current-scale node in the current-scale point cloud according to the effective accuracies in the three spatial dimensions or according to the target dimension identifier;
[0010] Decode the child node corresponding to the current scale node to determine the placeholder information corresponding to the child node.
[0011] This embodiment of the present application provides an encoding method, which is applied to an encoder, including:
[0012] Determine the effective accuracy or density of the point cloud at the current scale in three spatial dimensions;
[0013] Determining a child node corresponding to a current-scale node in the current-scale point cloud according to the effective precision or density in the three spatial dimensions;
[0014] The placeholder information of the child nodes corresponding to the current scale node is encoded to determine the encoding information of the current scale point cloud.
[0015] An embodiment of the present application provides a decoder, including:
[0016] a parsing portion configured to parse the code stream to determine the effective accuracy of the current-scale point cloud in three spatial dimensions, or a target dimension identifier corresponding to the current-scale point cloud; the target dimension identifier is determined based on the density of the current-scale point cloud in the three spatial dimensions;
[0017] A first determining part is configured to determine a child node corresponding to a current scale node in the current scale point cloud according to the effective precision in the three spatial dimensions or according to the target dimension identifier;
[0018] The decoding part is configured to decode the child node corresponding to the current scale node and determine the placeholder information corresponding to the child node.
[0019] An embodiment of the present application provides an encoder, including:
[0020] The second determining part is configured to determine the effective accuracy or density of the current scale point cloud in three spatial dimensions; and determine the child node corresponding to the current scale node in the current scale point cloud according to the effective accuracy or density in the three spatial dimensions;
[0021] The encoding part is configured to encode the placeholder information of the child nodes corresponding to the current scale node to determine the encoding information of the current scale point cloud.
[0022] The embodiment of the present application provides a code stream, which is generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following:
[0023] Encoded information of the current scale point cloud;
[0024] The encoding information of the scale point cloud is determined by determining the child nodes corresponding to the current scale node in the current scale point cloud based on the effective accuracy or density of the current scale point cloud in three spatial dimensions, and encoding the occupancy information of the child nodes corresponding to the current scale node.
[0025] An embodiment of the present application provides a decoder, including:
[0026] a first memory configured to store executable instructions;
[0027] The first processor is configured to implement any one of the above decoding methods when executing the executable instructions stored in the first memory.
[0028] An embodiment of the present application provides an encoder, including:
[0029] a second memory configured to store executable instructions;
[0030] The second processor is configured to implement any one of the encoding methods described above when executing the executable instructions stored in the second memory.
[0031] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a first processor to execute the above-mentioned decoding method, or for causing a second processor to execute the above-mentioned encoding method.
[0032] An embodiment of the present application provides a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a first processor, the decoding method provided by the embodiment of the present application is implemented; or, when the computer program or instructions are executed by a second processor, the encoding method provided by the embodiment of the present application is implemented.
[0033] The embodiment of the present application provides a coding and decoding method, a decoder, an encoder and a computer-readable storage medium. The effective accuracy or density of the current scale point cloud in the three spatial dimensions characterizes the distribution of the nodes in the current scale point cloud in the three spatial dimensions. The embodiment of the present application divides the current scale node in the current scale point cloud according to the distribution of the nodes in the current scale point cloud in the three spatial dimensions, determines the child nodes corresponding to the current scale node, and ensures that the distribution of the divided child nodes is relatively uniform. When performing the K nearest neighbor search algorithm (KNN) based on the above-mentioned child node division method to extract neighborhood feature values, it is not necessary to increase the size of the convolution kernel K to enhance the feature value. It is only necessary to continuously reduce the accuracy of the longest dimension through the child node division method, and finally achieve the ideal state of cube or octree division, so that the geometric feature values can be effectively extracted with the help of a sparse convolutional network and an appropriate convolution kernel K, thereby improving the geometric coding and decoding efficiency of the point cloud, and further improving the coding and decoding performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] FIG1A is a schematic diagram of a node at a low plane position in the Z-axis direction;
[0035] FIG1B is a second schematic diagram of a node at a low plane position in the Z-axis direction;
[0036] FIG1C is a third schematic diagram of a node at a low plane position in the Z-axis direction;
[0037] FIG1D is a fourth schematic diagram of a node at a low plane position in the Z-axis direction;
[0038] FIG1E is a fifth schematic diagram of a node at a low plane position in the Z-axis direction;
[0039] FIG2A is a schematic diagram of a node at a low plane position in the Z-axis direction;
[0040] FIG2B is a second schematic diagram of a node at a low plane position in the Z-axis direction;
[0041] FIG2C is a third schematic diagram of a node at a low plane position in the Z-axis direction;
[0042] FIG2D is a fourth schematic diagram of a node at a low plane position in the Z-axis direction;
[0043] FIG2E is a fifth schematic diagram of a node at a low plane position in the Z-axis direction;
[0044] FIG3 is a schematic diagram of the sequence numbers of child nodes at different positions;
[0045] FIG4A is a schematic diagram of the plane identification and plane position of the current node in the dimension corresponding to the Z axis;
[0046] FIG4B is a schematic diagram of a plane identification of the current node in the dimension corresponding to the Z axis;
[0047] Figure 5 is a schematic diagram of IDCM encoding;
[0048] FIG6 is a schematic diagram of geometric information reconstruction in a sub-block;
[0049] FIG7A is a schematic diagram of the SparsePCGC encoding framework;
[0050] FIG7B is a schematic diagram of the IRN module structure in the SparsePCGC coding framework;
[0051] FIG8 is a schematic diagram of the SparsePCGC encoding process of the node occupancy information and the eigenvalue encoding;
[0052] FIG9A is a schematic diagram of octree partitioning of a Multi-stage encoding scheme;
[0053] FIG9B is a schematic diagram of node classification of the Multi-stage encoding scheme 1;
[0054] FIG9C is a second schematic diagram of node classification of the Multi-stage encoding scheme;
[0055] FIG9D is a third schematic diagram of node classification of the Multi-stage encoding scheme;
[0056] Figure 10 is a schematic diagram of the encoding framework of SparsePCGC combined with Multi-stage;
[0057] FIG11 is a flow chart of G-PCC coding;
[0058] FIG12 is a flow chart of G-PCC decoding;
[0059] FIG13 is a schematic diagram of an optional flow chart of an encoding method provided in an embodiment of the present application;
[0060] FIG14 is a schematic diagram of an optional process of multi-tree partitioning provided in an embodiment of the present application;
[0061] FIG15A is a schematic diagram of an optional binary tree partitioning method provided in an embodiment of the present application;
[0062] FIG15B is a schematic diagram of an optional division of two sub-nodes provided in an embodiment of the present application;
[0063] FIG15C is a schematic diagram of an optional division of two sub-nodes provided in an embodiment of the present application;
[0064] FIG16A is a schematic diagram of an optional quadtree partitioning method provided in an embodiment of the present application;
[0065] FIG16B is a schematic diagram of an optional division of four sub-nodes provided in an embodiment of the present application;
[0066] FIG16C is a schematic diagram of an optional division of four sub-nodes provided in an embodiment of the present application;
[0067] FIG16D is a schematic diagram of an optional division of four sub-nodes provided in an embodiment of the present application;
[0068] FIG16E is a schematic diagram of an optional division of four sub-nodes provided in an embodiment of the present application;
[0069] FIG17 is a schematic diagram of an optional process of octree partitioning provided in an embodiment of the present application;
[0070] FIG18 is a schematic diagram of an optional flow chart of applying the encoding method provided in an embodiment of the present application to an actual scenario;
[0071] FIG19 is a schematic diagram of an optional flow chart of a decoding method provided in an embodiment of the present application;
[0072] FIG20 is a schematic diagram of an optional structure of a decoder provided in an embodiment of the present application;
[0073] FIG21 is a schematic diagram of an optional structure of an encoder provided in an embodiment of the present application
[0074] FIG22 is a schematic diagram of an optional structure of a decoder provided in an embodiment of the present application
[0075] Figure 23 is a schematic diagram of an optional structure of the encoder provided in an embodiment of the present application. DETAILED DESCRIPTION
[0076] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0077] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0078] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0079] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0080] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0081] 1) Voxel: Voxel, short for volume element, is the smallest unit of digital data used in three-dimensional space. Voxels allow for the gridding of 3D space and the assignment of characteristics to each grid. For example, a voxel is a fixed-size cube in three-dimensional space. Voxels are widely used in fields such as 3D imaging, scientific data, and medical imaging.
[0082] Point cloud compression algorithms include two schemes developed by the Moving Picture Experts Group (MPEG): Video-based Point Cloud Compression (V-PCC) and Geometry-based Point Cloud Compression (G-PCC). G-PCC primarily implements geometry compression using an octree model and / or a triangular surface model. V-PCC primarily uses 3D-to-2D projection and video compression.
[0083] In the point cloud G-PCC encoder framework, the geometric information of the point cloud and the attribute information corresponding to the points in the point cloud are encoded separately. Currently, G-PCC's geometric encoding and decoding can be divided into octree-based geometric encoding and decoding, triangle soup-based geometric encoding and decoding, and prediction tree-based geometric encoding and decoding. The following are introduced respectively:
[0084] 1. Octree-based geometric encoding and decoding:
[0085] 1.1. Octree-based geometric coding:
[0086] First, the coordinates of the geometric information are transformed so that all the point clouds are contained in a bounding box. Then, quantization is performed. This step of quantization mainly plays a role in scaling. Due to the quantization rounding, the geometric information of some points is the same. The parameters are used to decide whether to remove duplicate points. The process of quantization and removal of duplicate points is also called voxelization. Next, the bounding box is continuously divided into multiple trees (octree / quadtree / binary tree) in the order of breadth-first traversal, and the placeholder code of each node is encoded. Among them, an implicit geometric division method includes: calculating the bounding box of the point cloud (2 dx ,2 dy ,2 dz ), assuming that the bounding box corresponds to a cuboid d x >d y >d z In geometric partitioning, first, binary tree partitioning is performed based on the x-axis to obtain two child nodes; until d is satisfied x =d y >d z When the conditions are met, the quadtree partitioning will be performed based on the X and Y axes to obtain four child nodes; when d is finally satisfied x =d y =d zWhen the conditions are met, the octree partitioning will continue until the leaf node obtained by the partitioning is a 1x1x1 unit cube. The partitioning will stop and the points in the leaf node will be encoded to generate a binary code stream. In the process of binary tree / quadtree / octree partitioning, two parameters are introduced: K and M. Parameter K indicates the maximum number of binary tree / quadtree partitions before octree partitioning; parameter M is used to indicate that the minimum block side length corresponding to binary tree / quadtree partitioning is 2 M . At the same time, K and M must meet the following conditions: Assume d max =max(dx,dy,dz),d min =min(dx,dy,dz), parameter K satisfies: K>=d max -d min ; Parameter M satisfies: M>=d min The reason why the parameters K and M meet the above conditions is that in the current G-PCC geometric implicit partitioning process, the priority of the partitioning method is binary tree, quadtree and octree. When the node block size does not meet the conditions of binary tree / quadtree, the node will be divided into octree until the minimum unit of leaf node 1x1x1 is reached. The octree-based geometric information coding mode can use the correlation between adjacent points in space to effectively encode the geometric information of the point cloud. However, for some relatively flat nodes or nodes with planar characteristics, the coding efficiency of the point cloud geometric information can be further improved through plane coding.
[0087] As shown in Figures 1A to 1E, Figures 1A to 1E show nodes at low plane positions in the Z-axis direction. As shown in Figures 2A to 2E, Figures 2A to 2E show nodes at high plane positions in the Z-axis direction. Taking Figure 1A as an example, it can be seen that the four occupied child nodes in the current node are all located at low plane positions of the current node in the Z-axis direction, so it can be considered that the current node belongs to a Z plane and is a low plane in the Z-axis direction. Similarly, Figure 2A shows that the occupied child nodes in the current node are located at high plane positions of the current node in the Z-axis direction. Taking Figure 1A as an example, the efficiency of octree coding and plane coding is compared.
[0088] If the octree encoding method is used for Figure 1A, the placeholder information of the current node is represented as: 11001100. If the plane encoding method is used, an identifier needs to be encoded to indicate that the current node is a plane in the Z-axis direction. If the current node is a plane in the Z-axis direction, the plane position of the current node needs to be represented. Secondly, only the placeholder information of the low-plane node in the Z-axis direction needs to be encoded (i.e., the placeholder information of the four sub-nodes 0, 2, 4, and 6 in Figure 3). Encoding the current node based on the plane encoding method requires encoding 6 bits, which can reduce the representation of 2 bits compared to the original octree encoding. Based on this analysis, plane encoding has a more obvious encoding advantage over octree encoding. For an occupied node, if the plane encoding method is used for encoding in a certain dimension, first, the plane identifier (planeMode) and plane position (planePositon) information of the current node in that dimension are represented. Secondly, the placeholder information of the current node is encoded based on the plane information of the current node. Taking the dimension corresponding to the Z axis as an example, the plane identifier and plane position of the current node in the dimension corresponding to the Z axis may be as shown in FIG. 4A and FIG. 4B .
[0089] Note that: for the plane identifier planeMode i (i=0,1,2): 0 means the current node is not a plane in the i-axis direction. When the node is a plane in the i-axis direction, the plane position planePosition i : 0 means that the current node is a plane in the i-axis direction and the plane position is a low plane, 1 means that the current node is a high plane in the i-axis direction.
[0090] Octree-based geometric information coding only achieves efficient compression rates for highly correlated points in space. For isolated points in geometric space, using Direct Coding Mode (DCM) can significantly reduce complexity. For all nodes in the octree, DCM usage is not indicated by a flag bit but is inferred from the current node's parent and neighbor information. There are three ways to determine whether a node is eligible for DCM encoding, as shown in Figure 5.
[0091] The first way: the current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.
[0092] The second method: The parent node of the current node has only one child node, the current node, and the six neighbor nodes that share a face with the current node are all empty nodes.
[0093] The third method: The number of sibling nodes of the current node is greater than 1.
[0094] If the current node is not eligible for DCM encoding, it will be partitioned using an octree. If it is eligible, the number of points in the node will be further determined. If the number of points is less than a threshold of 2, the node will be DCM-encoded; otherwise, the octree partitioning will continue. When applying the DCM encoding mode, it is necessary to encode whether the current node is a true isolated point, namely the IDCM_flag. If the IDCM_flag is true, the current node will be DCM-encoded; otherwise, octree encoding will continue. If the current node meets the DCM encoding requirements, the DCM encoding mode for the current node will be encoded. Currently, there are two DCM modes: 1: only one point exists (or multiple points, but they are duplicates); 2: two points exist. Finally, the geometric information of each point needs to be encoded. Assuming that the edge length of the node is 2^d, encoding each component of the node's geometric coordinates requires d bits, and these bits are directly encoded into the bitstream. It is important to note that when encoding lidar point clouds, the three-dimensional coordinate information is predictively encoded using lidar acquisition parameters, further improving the encoding efficiency of the geometric information.
[0095] It should be noted that when nodes are divided into leaf nodes, in the case of geometric lossless coding, the number of repeated points in the leaf nodes needs to be encoded. Finally, the placeholder information of all nodes is encoded to generate a binary code stream.
[0096] 1.2. Octree-based geometric decoding:
[0097] The decoder follows a breadth-first traversal order. Before decoding each node's placeholder information, it uses the reconstructed geometric information to determine whether the current node should be decoded using plane decoding or inferred direct coding mode (IDCM). If the current node meets the requirements for plane decoding, the plane identifier and plane position information are decoded, and the placeholder information of the current node is decoded based on the plane information. If the current node meets the requirements for IDCM decoding, it is necessary to determine whether the current node is a true IDCM node. If so, the DCM decoding mode of the current node is parsed to obtain the number of points in the current DCM node, and finally the geometric information of each point is decoded. For nodes that do not meet either plane decoding or DCM decoding requirements, the placeholder information of the current node is decoded. This process continues to parse the placeholder code of each node, and the node is continuously partitioned until a 1x1x1 unit cube is obtained. The parsing process then determines the number of points contained in each leaf node, ultimately recovering the geometrically reconstructed point cloud information.
[0098] 2. Geometric information encoding and decoding based on triangle face sets:
[0099] 2.1. Geometric information encoding based on triangle patch sets:
[0100] In the geometric information encoding framework based on triangle face sets, geometric partitioning is also required first. However, unlike geometric information encoding based on binary trees, quad trees, or octrees, this method does not require step-by-step partitioning of the point cloud into unit cubes with sides of length 1x1x1. Instead, the partitioning stops when the block (sub-block) has a side length of W. Based on the surface formed by the distribution of the point cloud in each block, up to twelve vertices (intersection points) generated by this surface and the twelve edges of the block are obtained. The vertex coordinates of each block are encoded in sequence to generate a binary code stream.
[0101] 2.2 Geometric Information Decoding Based on Triangle Face Sets
[0102] When reconstructing point cloud geometry on the decoder, vertex coordinates are first decoded to complete triangle reconstruction, as shown in Figure 6. The block shown in part (a) of Figure 6 contains three vertices (v1, v2, and v3). The set of triangles constructed using these three vertices in a certain order is called triangle soup, or trisoup, as shown in part (b) of Figure 6. Next, sampling is performed on this triangular set, and the resulting sampled points serve as the reconstructed point cloud within the block, as shown in part (c) of Figure 6.
[0103] 3. Geometric information encoding and decoding based on prediction tree:
[0104] 2.1. Geometric information encoding based on prediction tree:
[0105] The input point cloud is sorted. The currently used sorting methods include unordered, Morton order, azimuth order, and radial distance order. At the encoding end, the prediction tree structure is established in two different ways, including: KD-Tree (high-latency slow mode) and using the lidar calibration information to divide each point into different Lasers and establish a prediction structure according to different Lasers (low-latency fast mode). Next, based on the structure of the prediction tree, each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter. Through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream.
[0106] 2.2. Geometric information decoding based on prediction tree:
[0107] The decoding end continuously parses the bit stream to reconstruct the prediction tree structure. Then, it obtains the geometric position prediction residual information and quantization parameters of each prediction node through parsing, and dequantizes the prediction residual to restore the reconstructed geometric position information of each node, finally completing the geometric reconstruction at the decoding end.
[0108] After the geometric encoding is completed, the geometric information is reconstructed. At present, attribute encoding is mainly performed on color information. First, the color information is converted from the RGB color space to the YUV color space. Secondly, the point cloud is recolored using the reconstructed geometric information so that the unencoded attribute information corresponds to the reconstructed geometric information. In color information encoding, there are two main transformation methods. One is the distance-based lifting transformation that relies on the level of detail (LOD) division, and the other is to directly perform the region adaptive hierarchical transform (RAHT) transformation. Both methods will convert the color information from the spatial domain to the frequency domain, and obtain high-frequency coefficients and low-frequency coefficients through transformation. Finally, the coefficients are quantized and encoded to generate a binary code stream.
[0109] When using geometric information to predict attribute information, Morton codes can be used to perform nearest neighbor searches. The Morton code corresponding to each point in the point cloud can be obtained from the geometric coordinates of the point. The specific method for calculating the Morton code is described below. For a three-dimensional coordinate represented by a d-bit binary number for each component, its three components can be expressed as:
[0110] Among them, x l ,y l ,z l ∈{0,1} are the binary values corresponding to the highest bit (l=1) to the lowest bit (l=d) of x, y, and z respectively. The Morton code M is to cross-arrange x, y, and z starting from the highest bit. l ,y l ,z l To the lowest bit, the calculation formula of M is as follows:
[0111] Among them, m l′ ∈{0,1} are the values of the highest bit (l′=1) to the lowest bit (l′=3d) of M. After obtaining the Morton code M of each point in the point cloud, the points in the point cloud are arranged in ascending order of the Morton code, and the weight w of each point is set to 1.
[0112] The G-PCC framework currently includes three attribute encoding methods for attribute information: Predicting Transform (PT), Lifting Transform (LT), and Region Adaptive Hierarchical Transform (RAHT). The first two predictively encode point clouds based on the order in which LODs are generated, while RAHT adaptively transforms attribute information from the bottom up based on the octree hierarchy.
[0113] The general test conditions of GPCC may include the following four types:
[0114] Condition 1: The geometric position is limited and the attributes are lost;
[0115] Condition 2: Geometric position lossless, attribute lossy;
[0116] Condition 3: Geometric position lossless, attribute loss limited;
[0117] Condition 4: Geometric position and attributes are lossless.
[0118] GPCC's general test sequences include four categories: Cat1A, Cat1B, Cat3-fused, and Cat3-frame. Among them, Cat2-frame point clouds only contain reflectance attribute information, Cat1A and Cat1B point clouds only contain color attribute information, and Cat3-fused point clouds contain both color and reflectance attribute information.
[0119] GPCC's 3. Technical routes: There are 2 types in total, distinguished by the algorithm used for geometric compression.
[0120] Technical Route 1: Octree Encoding Branch:
[0121] At the encoding end, the bounding box is divided into sub-cubes in sequence, and the non-empty sub-cubes (containing points in the point cloud) are continued to be divided until the leaf node obtained is a 1x1x1 unit cube. In the case of geometric lossless coding, the number of points contained in the leaf node needs to be encoded, and finally the geometric octree encoding is completed to generate a binary code stream.
[0122] At the decoding end, the decoding end obtains the placeholder code of each node by continuously parsing in the order of breadth-first traversal, and continuously divides the nodes in turn until a 1x1x1 unit cube is obtained. In the case of geometric lossless decoding, it is necessary to parse the number of points contained in each leaf node and finally recover the geometric reconstructed point cloud information.
[0123] Technical Route 2: Prediction Tree Encoding Branch:
[0124] At the encoding end, the prediction tree structure is established in two different ways, including: KD-Tree (high-latency slow mode) and using lidar calibration information to divide each point into different lasers and establish a prediction structure according to different lasers (low-latency fast mode). Based on the structure of the prediction tree, each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter. Finally, through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream;
[0125] At the decoding end, the decoding end reconstructs the prediction tree structure by continuously parsing the bit stream. Secondly, the geometric position prediction residual information and quantization parameters of each prediction node are obtained through parsing, and the prediction residual is dequantized to restore the reconstructed geometric position information of each node, finally completing the geometric reconstruction at the decoding end.
[0126] With the development of artificial intelligence (AI) technology, neural networks have been applied to geometry-based point cloud compression technology. Neural network-based point cloud geometry compression technology can be broadly categorized into lossy and lossless compression. Lossless compression algorithms primarily focus on the design of prediction models for voxel occupancy probabilities. Voxel data representations typically utilize octree models, volumetric models, sparse tensor representations, and other methods. For lossless geometry compression, the encoder typically uses surrounding context, such as parent nodes and neighboring nodes, as input. After processing through neural network layers (e.g., convolutional and fully connected), the encoder outputs the occupancy probability of each voxel in the point cloud's geometric data. An entropy encoder then converts the voxel occupancy symbol corresponding to each voxel's occupancy probability into a bitstream. Correspondingly, on the decoder side, the occupancy probability of each voxel is predicted using the same process. Based on the predicted occupancy probability, an entropy decoder is used to decode the voxel occupancy symbol from the bitstream to reconstruct the point cloud's geometric data.
[0127] For example, there are volumetric model compression techniques based on 3D convolutional neural networks (3D CNNs), neural network compression techniques that directly apply multi-layer perceptron (MLP) algorithms to point cloud geometric coordinate sets, compression techniques that use MLPs or 3D CNNs for probability estimation and entropy coding of octree node occupancy information, and compression techniques based on 3D sparse convolutional neural networks. Point clouds can be divided into sparse point clouds and dense point clouds based on point density. Sparse point clouds have a large representation range in 3D space and are sparsely distributed, making them suitable for representing a scene. Dense point clouds, on the other hand, have a small representation range and are densely distributed, making them suitable for representing an object. The compression performance of these compression techniques on these two types of point clouds often differs significantly, with superior performance on dense point clouds and poorer performance on sparse point clouds.
[0128] In the neural network-based point cloud geometry compression technology, the coding framework of SparsePCGC is shown in Figure 7A. In Figure 7A, AE represents Arithmetic Encoder and AD represents Arithmetic Decoder.
[0129] As shown in Figure 7A, the encoder-decoder transformation of SparsePCGC consists of a multi-layer sparse convolutional neural network. An Inception-Residual Network (IRN) is used to enhance the network's feature analysis capabilities. The IRN structure is shown in Figure 7B. After each upsampling and downsampling, there is a feature extraction module consisting of three IRN units. Downsampling is achieved through convolution with a stride of 2 and a kernel size of 2, aggregating the features of voxels within each 2×2×2 spatial unit into a single voxel. After each downsampling, the length, width, and height of the point cloud are halved, resulting in a total of three downsampling operations. Upsampling in the decoder is achieved through transposed convolution with a stride of 2 and a kernel size of 2, dividing a single voxel into 2×2×2 voxels, doubling the length, width, and height of the point cloud. After each upsampling, binary classification is used to retain the voxels predicted to be occupied from the generated voxels and remove the voxels predicted to be empty and their attributes to reconstruct geometric details. Through layered and progressive reconstruction, the rough point cloud gradually recovers the detailed structure. In Figure 7A and Figure 7B, ReLU represents the activation function.
[0130] As shown in Figure 7A, at the encoding end, the geometric information of the point cloud is continuously downsampled, starting from the initial voxel-level point cloud, and eventually the entire point cloud is merged into a node, namely the root node. For each layer of nodes, the geometric data of the first-scale point cloud is voxel downsampled twice. The first-scale point cloud can be regarded as the original-scale point cloud to be encoded. After performing a voxel downsampling on the geometric data of the first-scale point cloud, the geometric data of the second-scale point cloud can be obtained; after performing a voxel downsampling on the geometric data of the second-scale point cloud again, the geometric data of the third-scale point cloud can be obtained. Assuming that the current layer node is the geometric data of the first-scale point cloud, after another downsampling, the occupancy information of the second-scale node needs to be entropy encoded. Before encoding the node occupancy information, first, the geometric data of the second-scale node is used to extract feature data, and secondly, the extracted feature data is used to entropy encode the second-scale node occupancy information.
[0131] Specifically, at the encoding end, the encoder network performs at least two voxel downsampling and feature extraction on the geometric data of the second-scale point cloud to obtain feature data used to enhance the geometric data of the third-scale point cloud. As shown in Figure 8, two encoders each perform voxel downsampling (with a step size of 2×2×2) and feature extraction once to extract feature data that is truly helpful for reconstruction and reduce the amount of data to be transmitted. The feature data extracted by the neural network is called latent feature data. The encoder network quantizes the output feature data, entropy encodes it, and writes it into the bitstream. It can also be directly entropy-encoded and written into the bitstream, thereby realizing the use of the occupancy information of the encoded nodes in the same layer to encode the placeholder information.
[0132] At the decoding end, as shown in Figure 8, entropy decoding is performed to obtain lossless geometric data of the third-scale point cloud and feature data used to enhance the point cloud geometric data. This lossless geometric data is the geometric data to be enhanced for the third-scale point cloud. After passing through the decoder network, the feature data is subjected to voxel upsampling and feature inference. The output feature data is then concatenated with the geometric data to be enhanced for the third-scale point cloud to obtain the enhanced geometric data for the third-scale point cloud, which is the third-scale point cloud geometric data + feature data shown in Figure 8. After obtaining the enhanced geometric data for the third-scale point cloud, another decoder in the decoder network performs voxel upsampling and feature inference on the enhanced geometric data for the third-scale point cloud. The data output by the decoder is then subjected to probability prediction and entropy decoding (or point cloud cropping) to obtain the reconstructed geometric data for the second-scale point cloud. After feature enhancement, the feature values derived from the inference of the geometric data for the third-scale point cloud can be used to effectively entropy decode the node occupancy information, thereby reducing the bitrate of the node occupancy information and significantly improving decoding performance. The above encoder network and decoder network belong to the same autoencoder model, and the network parameters of the two are obtained through joint training.
[0133] The above scheme is the core encoding module of SparsePCGC, which can derive the encoding of the current layer to any layer, and each layer shares a network parameter. Assuming that the i+1th scale point cloud is the smallest scale point cloud, the geometric data of the point cloud is losslessly compressed through entropy coding. After obtaining the geometric data of the i+1th scale point cloud, the network entropy decoding is used to obtain the eigenvalues, and the decoding network is needed to upsample the eigenvalues to enhance the geometric data of the i+1th scale point cloud. Finally, the decoder network is used to restore the i+1th enhanced point cloud to obtain the geometric data of the i-th scale point cloud. The point cloud data is restored layer by layer in sequence until it is restored to the voxel level, thus completing the encoding and decoding of the geometric information of the entire point cloud.
[0134] Based on Figures 7A and 8, it can be seen that the SparsePCGC encoding network has two core modules: feature extraction and voxel downsampling. Specifically, feature extraction is performed on the data using at least one of a first residual network based on sparse convolution and a first self-attention network; the output data of the first residual network and the first attention network are subjected to a voxel downsampling using a sparse convolution layer with a stride of 2×2×2; and the features output by the sparse convolution layer are extracted using at least one of a second residual network based on sparse convolution and a second self-attention network. More specifically, the first residual network and the second residual network include one or more residual layers based on sparse convolution. Each residual layer, as shown in Figure 7A, includes three or more branches. Branch one directly outputs the output data, while the other branches perform feature inference on the input data through different numbers of sparse convolution layers. Finally, the outputs of the other branches are concatenated and added to the output of branch one to obtain the data of the residual layer. As shown in Figure 7A, three branches are shown, one of which includes two sparse convolutional layers, and the other includes three sparse convolutional layers, with activation functions set between adjacent sparse convolutional layers.
[0135] Therefore, in the SparsePCGC encoding framework shown in Figure 7A, the encoder sequentially includes: a first sparse convolutional network, a first self-attention network, a first residual network, a sparse convolutional layer with a stride of 2×2×2, a second residual network, a second attention network, and a second sparse convolutional network. Activation functions are provided between the first sparse convolutional network and the first self-attention network, and between the first residual network and the sparse convolutional layer. The first sparse convolutional network and the second sparse convolutional network include one or more sparse convolutional layers. The processing for each self-attention layer includes: for each point in the point cloud, searching for neighboring points based on the point's coordinate data, performing a linear transformation on the distance information from the point to the neighboring points to obtain a position feature, and finally adding the position feature to the neighboring point features to obtain a position-encoded aggregate feature. The purpose of the self-attention layer is to perform a first linear transformation on the input feature data to obtain a first vector, perform a second linear transformation on the first vector and the aggregated feature to obtain a second vector for matrix multiplication, and then, after passing the activation function, obtain the attention weight of each point relative to its neighboring points.
[0136] Assume that the point cloud data of the input point cloud neighborhood self-attention layer includes feature data and coordinate C in ∈R n×3 (n represents the number of points, d in Represents the dimension of the input data feature information, 3 is the dimension of the data geometric coordinates), coordinate C in ∈R n×3 Used to find neighboring points.
[0137] For each point P in the point cloud i , use KNN to find the K nearest neighboring points {P i1 ,P i2 ,P i3 ….P ik}, and gather the coordinates C of K neighborhood points knn ∈R n×k×3 and features
[0138] For each point P i As the center point, find P i The K neighboring points {P i1 ,P i2 ,P i3 ….P ik} and the center point P i The relative distance {dist i1 ,dist i2 ,dist i3 ….dist ik}, get relative distance information dist knn ∈R n×k×1 . Through the linear layer W dist dist knn The dimension is mapped from 1 to d in Dimension, get the relative position feature and feature F knn Add (i.e. add to feature F knn Above), realize position encoding: F′ knn =dist knn W dist +F knn
[0139] in, F′ knn is the aggregated feature after position encoding, Through position encoding, the features are given the perceptual information of the relative positions between corresponding points, and the features of each neighborhood point have spatial position information.
[0140] QKV vector generation:
[0141] The input feature data F in Through the linear layer W Q Transform to obtain the Q vector, and transform the position-encoded aggregate feature F′ knn Pass through the linear layer W respectively k and linear layer W v , get K vector and V vector, that is, Q = F in W Q (K,V)=F′knn ·(W K ,W V )
[0142] in, and Represents three different linear transformations. The Q vector represents the query vector, the K vector represents the vector of the correlation between the query information and other information, and the V vector represents the vector of the queried information. Among them, the dimension parameter d a and d out Can be equal to d in , for example, both are set to 32. a and d out It may not be equal to d in , that is, dimension transformation.
[0143] After obtaining the Q vector, K vector, and V vector, perform matrix multiplication on the Q vector and the K vector. The result is activated by the Softmax activation function. When each point is used as the center point, the attention weight A relative to its neighboring points is output. Finally, the attention weight A is multiplied by the V vector to obtain the feature data F of the output point cloud. out ,Right now:
[0144] Based on this, the geometric position information to be encoded / decoded in the current layer and the geometric position information of the neighborhood can be used to derive the eigenvalue of the node in the current layer, and finally the eigenvalue is used to perform entropy coding on the occupancy information of the node.
[0145] The above is the basic coding framework of SparsePCGC. The position information of the parent node layer nodes is used to extract the eigenvalues of the parent node layer geometry, and the eigenvalues are used to entropy encode the placeholder information. Since SparsePCGC encodes based on the breadth-first criterion of the octree, the current layer nodes encode or decode the placeholder information of each node according to certain criteria at the codec end. The correlation between nodes in the same layer is often more similar than the correlation between the current layer nodes and the parent node layer. Therefore, SparsePCGC proposes an enhanced coding scheme, namely the Multi-stage coding scheme, which is specifically shown in Figures 9A-9D.
[0146] As shown in Figure 9A, there are three nodes. Each node will have eight child nodes if divided according to the octree. When decoding the placeholder information of each node's child nodes, the nodes in the current layer can first be divided into N categories according to certain criteria, such as 1 category (1 stage: 8 child nodes are decoded at once) as shown in Figure 9B; 4 categories (4 stages: 1 and 2 are the first category, 3 and 4 are the second category, 5 and 6 are the third category, and 7 and 8 are the eighth category) as shown in Figure 9C; or even 8 categories (8 stages) as shown in Figure 9D. After being divided into different categories, encoding or decoding is performed in order from the first category to the Nth category. This encoding scheme allows parallel processing of the placeholder information decoding of each child node within the same category, while serial encoding or decoding is performed between different categories.
[0147] Secondly, based on SparsePCGC, we use two types of node geometry to enhance the geometry of nodes in the current layer. The first type is the geometry of the parent node layer; the second type is the geometry of nodes in the same layer that have been encoded or decoded in the current layer. The specific encoding framework is shown in Figure 10.
[0148] Taking category 8 as an example, when encoding or decoding the placeholder information of the first category of child nodes, only the position information of the parent node of the current layer node can be used to extract the eigenvalue. After completing the placeholder information of the first category of child nodes, the geometric position information of the child nodes actually occupied in the first category of child nodes can be obtained. When decoding the position information of the second category of child nodes, two types of information can be used: the parent node layer position information and the first category of child node position information that has been encoded or decoded in the same layer. Therefore, the two types of information can be combined to further enhance the eigenvalue of the geometric position information of the second category of child nodes, so that the eigenvalue can be more effectively used to entropy encode the node placeholder information, effectively improving the geometric coding efficiency. After completing the encoding of the first and second categories of node placeholder information, when decoding the third category of child node placeholder information, the parent node layer node position information and the geometric position information of the node that has been encoded or decoded in the same layer can also be combined to obtain the eigenvalue of the geometric position of the third category of child nodes. Then, the enhanced eigenvalue is used to effectively encode the placeholder information of the third category of child nodes. According to this encoding or decoding method, until the placeholder information of the last type of child node, that is, the eighth type of child node, is encoded, the geometric position information of all child nodes that have been encoded in the parent node layer and the current layer can be used to obtain the eigenvalue of the current node. Finally, the eigenvalue is used to encode or decode the placeholder information of the last type of child node, thereby completing the encoding / decoding of the geometric position information of the current layer node.
[0149] Based on the above analysis, we can see that in the current SparsePCGC encoding scheme, the feature value of the current node to be encoded is extracted using the geometric information of the parent node layer and the already encoded nodes at the same layer. The feature value is used to obtain the occupancy probability of different child nodes in the node to be encoded, and the occupancy information of each child node in the current node is entropy encoded. The current SparsePCGC encoding scheme limits the node upsampling and downsampling partitioning methods to octave, namely, only using octree partitioning. For dense point clouds or point clouds with regular distributions, octree partitioning is an effective encoding partitioning method. However, for sparse scene point clouds, such as LiDAR point clouds, these data tend to have a flat distribution. If the simple octree partitioning is still used for such scene point clouds, it is difficult to effectively obtain the occupancy probability of the current node using the KNN algorithm. This is because the distribution is extremely sparse, resulting in large distances between neighbors. Therefore, for the compression of such point cloud data, the KNN neighborhood search range is generally expanded, that is, the size of the convolution kernel K in the deep learning network is increased. This can effectively solve the problem of being unable to search for valid neighborhoods due to sparsity to a certain extent, but it also brings another more serious problem: as the convolution kernel K increases, the space occupied by the neural network model parameters grows exponentially, and the model complexity also grows exponentially. This reduces encoding and decoding efficiency and, in turn, performance.
[0150] The embodiments of the present application provide a coding and decoding method, a decoder, an encoder and a computer-readable storage medium, which can improve coding and decoding efficiency, and thus improve coding and decoding performance. In order to facilitate the understanding of the technical solution provided by the embodiments of the present application, a flow chart of G-PCC encoding and a flow chart of G-PCC decoding are first provided. It should be noted that the flow chart of G-PCC encoding and the flow chart of G-PCC decoding described in the embodiments of the present application are only for more clearly illustrating the technical solution of the embodiments of the present application, and do not constitute a limitation on the technical solution provided by the embodiments of the present application. It is known to those skilled in the art that with the evolution of point cloud compression technology and the emergence of new business scenarios, the technical solution provided by the embodiments of the present application is also applicable to point cloud coding and decoding architectures similar to G-PCC. The point cloud compressed by the embodiments of the present application can be a point cloud in a video, but is not limited to this.
[0151] In the point cloud G-PCC encoder framework, the point cloud of the input 3D image model is sliced and each slice is encoded independently.
[0152] The G-PCC encoding process flow diagram shown in Figure 11 is applied to the encoder. For the point cloud data to be encoded, it is first divided into multiple slices using striping. Within each slice, the geometric information and attribute information of the point cloud are encoded separately. During the geometric encoding process, the geometric information is transformed so that the entire point cloud is contained within a bounding box. Quantization is then performed. Quantization primarily serves a scaling purpose. Due to quantization rounding, the geometric information of some point clouds remains the same. The decision to remove duplicate points can be made based on parameters. This process of quantization and removing duplicate points is also known as voxelization. The bounding box is then partitioned into eight equal parts using an octree. In the octree-based geometric information encoding process, the bounding box is divided into eight equal parts. The sub-cubes that are not empty (contain points in the point cloud) are further divided into eight equal parts until the resulting leaf nodes are 1x1x1 unit cubes. The points in the leaf nodes are then arithmetic-coded to generate a binary geometric bitstream, also known as the geometry codestream. In the process of geometric information encoding based on triangle soup (trisoup), octree partitioning must also be performed first. However, unlike the geometric information encoding based on octree, the trisoup does not need to divide the point cloud into unit cubes with a side length of 1x1x1 step by step. Instead, the division stops when the sub-block (block) has a side length of W. Based on the surface formed by the distribution of the point cloud in each block, the surface and the twelve edges of the block are obtained. The vertices are arithmetically encoded (surface fitting is performed based on the intersections) to generate a binary geometric bit stream, i.e., a geometric code stream. Vertex is also used to implement the process of geometric reconstruction, and the reconstructed geometric information is used when encoding the attributes of the point cloud.
[0153] During the attribute encoding process, color conversion is performed to convert the color information (i.e., attribute information) from the RGB color space to the YUV color space. The reconstructed geometric information is then used to recolor the point cloud so that the unencoded attribute information corresponds to the reconstructed geometric information. There are two main transformation methods for color information encoding: a distance-based lifting transform that relies on level of detail (LOD) partitioning, and a direct region-adaptive hierarchical transform (RAHT) transformation. Both methods convert color information from the spatial domain to the frequency domain, obtaining high-frequency and low-frequency coefficients. These coefficients are then quantized (i.e., quantized coefficients). Finally, the geometrically encoded data (octree partitioning and surface fitting) is sliced and synthesized with the attribute encoded data (quantized coefficients). The vertex coordinates of each block are then encoded (i.e., arithmetic encoding) to generate a binary attribute bitstream, i.e., the attribute codestream.
[0154] The flowchart of G-PCC decoding shown in Figure 12 is applied to the decoder. The decoder obtains the binary code stream and independently decodes the geometric bit stream (i.e., geometric code stream) and attribute bit stream in the binary code stream. When decoding the geometric bit stream, the geometric information of the point cloud is obtained through arithmetic decoding-octree synthesis-surface fitting-reconstruction geometry-inverse coordinate transformation; when decoding the attribute bit stream, the attribute information of the point cloud is obtained through arithmetic decoding-inverse quantization-LOD-based inverse lifting or RAHT-based inverse transformation-inverse color conversion, and the three-dimensional image model of the point cloud data to be encoded is restored based on the geometric information and attribute information.
[0155] The encoding method of the embodiment of the present application can be applied to the geometric information encoding process of the G-PCC as shown in Figure 11. After the voxelization process is completed, when encoding the current-scale point cloud, the effective accuracy or density of the current-scale point cloud in three spatial dimensions is determined; based on the effective accuracy or density in the three spatial dimensions, the child nodes corresponding to the current-scale node in the current-scale point cloud are determined, thereby replacing the process of octree partitioning of the bounding box in Figure 11. The encoding process of the embodiment of the present application can adopt the arithmetic coding method shown in Figure 11, such as entropy coding. The placeholder information of the child nodes corresponding to the current-scale node is encoded through the arithmetic coding process, and the encoding information of the current-scale point cloud is determined and written into the geometry bitstream (codestream).
[0156] The decoding method of the embodiment of the present application can be applied to the geometric information decoding process of the G-PCC shown in Figure 12. By parsing the bitstream, the effective accuracy of the current-scale point cloud in the three spatial dimensions, or the target dimension identifier corresponding to the current-scale point cloud, is determined; the target dimension identifier is determined based on the density of the current-scale point cloud in the three spatial dimensions; based on the effective accuracy in the three spatial dimensions or the target dimension identifier, the child nodes corresponding to the current-scale node in the current-scale point cloud are determined, thereby replacing the octree synthesis in Figure 12. The decoding process of the embodiment of the present application can adopt the arithmetic decoding method shown in Figure 12, such as entropy decoding. Through the arithmetic decoding process, the child nodes corresponding to the current-scale node are decoded to determine the corresponding placeholder information of the child nodes. Then, based on the placeholder information of the child nodes, the geometric information corresponding to the next-scale point cloud of the current scale can be recovered.
[0157] It should be noted that the encoding method and decoding method of the embodiment of the present application can also be used in other point cloud encoding and decoding processes other than G-PCC. The encoding method and decoding method of the embodiment of the present application can also be used in encoding and decoding schemes for placeholder information. When dividing the nodes, based on this scheme, the nodes to be encoded can be adaptively divided into spatial geometries with regular distribution according to the effective accuracy or density of the current scale point cloud in each spatial dimension. Based on such a spatial distribution, the use of a deep learning network model can often obtain more effective eigenvalues, thereby further leveraging the efficiency of the deep learning network coding model.
[0158] The following describes an encoding method applied to an encoder provided in an embodiment of the present application.
[0159] See Figure 13, which is an optional flow chart of the encoding method provided in the embodiment of the present application, which will be described in conjunction with the steps shown in Figure 13. The current scale point cloud in the following embodiments specifically refers to the geometric information of the current scale point cloud.
[0160] S101. Determine the effective accuracy or density of the current scale point cloud in three spatial dimensions.
[0161] In S101, the encoder downsamples the original voxel-level point cloud to obtain a point cloud at a specific scale. The encoder encodes the point cloud at that scale and obtains the encoding information for that scale. The encoder can downsample the point cloud at that scale to obtain a point cloud at the next scale, encode the point cloud at the next scale, and determine the encoding information for the point cloud at the next scale. In this way, through multiple downsampling and encoding processes, the encoder compresses the original voxel-level point cloud layer by layer until it reaches the root node, completing the encoding of the entire point cloud.
[0162] Here, the point cloud to be encoded at a scale determined by each downsampling is considered the current scale point cloud. The current scale point cloud can also be considered the current layer point cloud, and the current scale nodes in the current scale point cloud are the same-layer nodes at the same scale. The encoder downsamples the current layer point cloud, and the next scale point cloud determined can be considered the next layer point cloud. The current layer point cloud is the parent node layer of the next layer point cloud, and the current layer nodes in the current layer point cloud are the parent nodes of the nodes in the next layer point cloud.
[0163] In S101 , when encoding the current-scale point cloud (ie, the current-layer point cloud), the encoder determines the effective precision or density of the current-scale point cloud in three spatial dimensions.
[0164] In some embodiments, the encoder may identify a bounding box corresponding to the current-scale point cloud, and determine effective precision in the three spatial dimensions based on the sizes of the bounding box in the three spatial dimensions (eg, bit information corresponding to the side lengths).
[0165] In some embodiments, the encoder may perform principal component analysis (PCA) on the current-scale point cloud in three spatial dimensions to determine the density in the three spatial dimensions, where the density represents the density of nodes in the current-scale point cloud in the corresponding dimensions.
[0166] S102 : Determine a child node corresponding to a current-scale node in the current-scale point cloud according to effective precision or density in three spatial dimensions.
[0167] In S102, the encoder determines a subnode division method corresponding to the current-scale point cloud based on the effective accuracy of the current-scale point cloud in three spatial dimensions. Each current-scale node in the current-scale point cloud is divided into subnodes according to the determined subnode division method, and the subnodes corresponding to each current-scale node are determined. This means that the subnodes to be encoded for each current-scale node are determined. Thus, by encoding the placeholder information of the subnodes corresponding to each current-scale node, the encoding information for each current-scale node is determined, thereby determining the encoding information for the current-scale point cloud.
[0168] Alternatively, the encoder determines a subnode partitioning scheme for the current-scale point cloud based on the density of the current-scale point cloud in the three spatial dimensions. Each current-scale node in the current-scale point cloud is partitioned into subnodes according to the determined subnode partitioning scheme, and the subnodes corresponding to each current-scale node are determined. This also determines the subnodes to be encoded for each current-scale node. Thus, by encoding the placeholder information for the subnodes corresponding to each current-scale node, the encoding information for each current-scale node can be determined, thereby determining the encoding information for the current-scale point cloud.
[0169] In some embodiments, the above sub-node division method can represent the number of sub-nodes to be divided (ie, the current scale node is divided into two sub-nodes, or four sub-nodes, or eight sub-nodes) and the dimension of division (ie, along which spatial dimension the division is performed).
[0170] In S102, regarding the sub-node division method based on effective precision, since both the encoder and decoder can locally determine the effective precision of the current-scale point cloud in the three spatial dimensions, in some embodiments, the encoder and decoder can adaptively perform sub-node division based on the effective precision of the current-scale point cloud, without having to communicate the determined sub-node division method via syntax identification information. In other words, the encoder and decoder can use the effective precision of the current-scale point cloud in the three spatial dimensions to adaptively and implicitly determine the sub-node division method corresponding to the current-scale point cloud.
[0171] In some embodiments, the encoder may determine a maximum effective precision, a median effective precision, and a minimum effective precision based on the effective precision in the three spatial dimensions.
[0172] For example, the encoder can determine the maximum effective precision according to formula (1); determine the median effective precision according to formula (1); and determine the minimum effective precision according to formula (3). max =max(dim x ,dim y ,dim z ) (1) dim mid =mid(dim x ,dim y ,dim z ) (2) dim min = min (dim x ,dim y ,dim z ) (3)
[0173] In formula (1) to formula (3), dim x is the effective accuracy of the current scale point cloud in the spatial dimension x; dim y is the effective accuracy of the current scale point cloud in the spatial dimension y; dim z dim is the effective accuracy of the current scale point cloud in the spatial dimension z; max The maximum effective precision; dim mid is the median of effective precision; dim min The minimum effective precision.
[0174] In this way, the encoder can divide the current scale node according to the maximum effective precision value, the median effective precision value, and the minimum effective precision value, and determine the child nodes corresponding to the current scale node.
[0175] For example, as shown in Figure 14, the ratio of the effective accuracy of the voxel-level point cloud in the three spatial dimensions is (8:2:2). It can be seen that the maximum effective accuracy corresponds to the X dimension, and the effective accuracy of the Y and Z dimensions is the same. In this case, a binary tree merge can be performed along the X dimension, that is, the sampling rate of the sparse neural network is (2,1,1). After the merge, after encoding the occupancy information of the two child nodes of each node, according to the ratio of the effective accuracy of the nodes in the next layer in the three spatial dimensions is (4:2:2), the binary tree is still used for merging, and the sampling rate of the sparse convolutional neural network is (2,1,1); after continuing to merge, the ratio of the effective accuracy of the three spatial dimensions is (2:2:2). At this time, an octree is used for merging, that is, it is necessary to encode the occupancy information of the 8 child nodes of each node in the current layer.
[0176] In S102, regarding the method of dividing sub-nodes based on density, since the density needs to be determined based on the geometric information of the current-scale point cloud, the encoder determines the density in three spatial dimensions using PCA, and determines the target dimension based on the density in the three spatial dimensions; based on the target dimension, the current-scale node is divided to determine the sub-nodes corresponding to the current-scale node. For example, if the density in the three spatial dimensions indicates that the nodes in the current-scale point cloud are evenly distributed in the x and y dimensions and scattered in the z dimension, then the z dimension can be determined as the target dimension, that is, the division is performed along the z dimension to ensure that the sub-nodes after the division can search for a valid neighborhood prediction point set within a certain range. The encoder determines the target dimension identifier corresponding to the target dimension, encodes the target dimension identifier, and writes the obtained encoding bits and the encoding information of the current-scale point cloud into the bitstream, so that the decoder can determine the target dimension based on the target dimension identifier and perform the same sub-node division on the current-scale point cloud in the decoder.
[0177] That is, by encoding the target dimension identifier and sending it to the decoding end, the encoder can enable the decoding end to explicitly determine the division method of the sub-nodes according to the target dimension identifier.
[0178] S103 : Encode the placeholder information of the child nodes corresponding to the current scale node to determine the encoding information of the current scale point cloud.
[0179] In S103 , the encoder may use the geometric information of the current-scale point cloud to predict the occupancy information of the child nodes corresponding to the current-scale node, such as the occupancy probability, and encode the occupancy information of the child nodes corresponding to the current-scale node to determine the encoding information of the current-scale point cloud.
[0180] In some embodiments, the geometric eigenvalues of the child nodes can be extracted based on the SparsePCGC encoding scheme described above, combined with the geometric position information of the current scale node and the geometric position information of the already encoded nodes in the current scale point cloud. The eigenvalues are then used to entropy encode the placeholder information of the child nodes. Based on the spatial distribution of the child nodes after division in the embodiments of the present application, a deep learning network model can often obtain more effective eigenvalues, thereby further enhancing the efficiency of the deep learning network coding model.
[0181] In some embodiments, the encoder writes the encoding information of the current-scale point cloud into the bitstream and continues to encode the point cloud of the next scale: the encoder can downsample the current-scale point cloud to determine the next-scale point cloud (i.e., the next layer of point cloud); similarly, according to the above method, based on the effective accuracy or density of the next-scale point cloud in the three spatial dimensions, the next-scale point cloud is divided into sub-nodes and the occupancy information is encoded to determine the encoding information of the next-scale point cloud, until the maximum effective accuracy in the three spatial dimensions reaches a preset accuracy threshold, or until the minimum density in the three spatial dimensions reaches a preset density threshold, and the point cloud encoding is completed.
[0182] It can be understood that the effective accuracy or density of the current scale point cloud in the three spatial dimensions characterizes the distribution of nodes in the current scale point cloud in the three spatial dimensions. The embodiment of the present application divides the current scale node in the current scale point cloud according to the distribution of nodes in the current scale point cloud in the three spatial dimensions, determines the sub-nodes corresponding to the current scale node, and ensures that the distribution of the divided sub-nodes is relatively uniform. When performing KNN neighborhood feature value extraction based on the above-mentioned sub-node division method, there is no need to increase the size of the convolution kernel K to enhance the feature value. It is only necessary to continuously reduce the accuracy of the longest dimension through the sub-node division method, and finally achieve the ideal state of cube or octree division, so that the geometric feature values can be effectively extracted with the help of sparse convolutional networks and appropriate convolution kernels K, thereby improving the geometric coding efficiency of the point cloud and thus improving the coding performance.
[0183] In some embodiments, regarding the method of dividing sub-nodes according to effective precision, the process of dividing the current scale node according to the maximum effective precision, the median effective precision, and the minimum effective precision in S102 and determining the sub-nodes corresponding to the current scale node may include the following situations:
[0184] In the first case, when the maximum effective precision is greater than the median effective precision, the current scale node is divided to determine the two child nodes corresponding to the current scale node.
[0185] In the embodiment of the present application, when the maximum effective precision is greater than the median effective precision, the current scale node is divided in a binary manner to determine two child nodes corresponding to the current scale node.
[0186] For determining the two child nodes corresponding to the current scale node, the regular division method and the oblique division method can be used as follows:
[0187] For a regular partitioning method, the encoder can perform binary tree partitioning on the current scale node along the dimension corresponding to the maximum effective precision to determine two child nodes corresponding to the current scale node. FIG15A shows a binary tree partitioning method.
[0188] The oblique angle partitioning method can generally be applied to point clouds with oblique angle distribution. For the oblique angle partitioning method, the current scale node corresponds to eight candidate subnodes in three spatial dimensions, as shown in Figure 3. In Figure 3, the subnode numbered 0 represents the first candidate subnode; the subnode numbered 0 represents the first candidate subnode; the subnode numbered 1 represents the second candidate subnode; the subnode numbered 2 represents the third candidate subnode; the subnode numbered 3 represents the fourth candidate subnode; the subnode numbered 4 represents the fifth candidate subnode; the subnode numbered 5 represents the sixth candidate subnode; the subnode numbered 6 represents the seventh candidate subnode; and the subnode numbered 7 represents the eighth candidate subnode.
[0189] As shown in FIG15B , the encoder may divide the first, fourth, fifth, and eighth candidate subnodes among the eight candidate subnodes into one subnode, and divide the second, third, sixth, and seventh candidate subnodes into one subnode, thereby determining the two subnodes corresponding to the current scale node. Alternatively, as shown in FIG15C , the encoder may divide the first, third, sixth, and eighth candidate subnodes among the eight candidate subnodes into one subnode, and divide the second, fourth, fifth, and seventh candidate subnodes into one subnode, thereby determining the two subnodes corresponding to the current scale node.
[0190] In the second case, when the maximum effective precision is equal to the median effective precision, and the median effective precision is greater than the minimum effective precision, the current scale node is divided to determine the four child nodes corresponding to the current scale node.
[0191] In the embodiment of the present application, when the maximum effective precision is equal to the median effective precision and the median effective precision is greater than the minimum effective precision, the current scale node is divided into four parts to determine the four child nodes corresponding to the current scale node.
[0192] For determining the four child nodes corresponding to the current scale node, the regular division method and the oblique division method can be used as follows:
[0193] For the regular partitioning method, the encoder can perform quadtree partitioning on the current scale node along the dimension corresponding to the maximum effective precision and the dimension corresponding to the median effective precision, respectively, to determine the four child nodes corresponding to the current scale node. FIG16A shows a quadtree partitioning method.
[0194] For the oblique angle division method, the following situations can be included:
[0195] As shown in Figure 16B , the first and eighth candidate subnodes of the eight candidate subnodes are divided into one subnode, the second and seventh candidate subnodes are divided into one subnode, the third and sixth candidate subnodes are divided into one subnode, and the fourth and fifth candidate subnodes are divided into one subnode, thereby determining the four subnodes corresponding to the current scale node. This oblique division method is equivalent to the diagonal division of the verification cube.
[0196] Alternatively, as shown in FIG16C , the first candidate subnode and the sixth candidate subnode of the eight candidate subnodes are divided into one subnode, the second candidate subnode and the fifth candidate subnode are divided into one subnode, the third candidate subnode and the eighth candidate subnode are divided into one subnode, and the fourth candidate subnode and the seventh candidate subnode are divided into one subnode, thereby determining the four nodes corresponding to the current scale node. It can be seen that this oblique division method is equivalent to dividing along the plane diagonals of the X and Z dimensions, respectively. In some embodiments, this oblique division method can be used when the effective precision maximum value is equal to the effective precision median value, the effective precision median value is greater than the effective precision minimum value, and the effective precision maximum value and the effective precision median value correspond to the X and Z dimensions.
[0197] Alternatively, as shown in FIG16D , the first and seventh candidate subnodes of the eight candidate subnodes are divided into one subnode, the second and eighth candidate subnodes are divided into one subnode, the third and fifth candidate subnodes are divided into one subnode, and the fourth and sixth candidate subnodes are divided into one subnode, thereby determining the four subnodes corresponding to the current scale node. It can be seen that this oblique division method is equivalent to dividing along the plane diagonals of the X and Y dimensions, respectively. In some embodiments, this oblique division method can be used when the effective precision maximum value is equal to the effective precision median value, the effective precision median value is greater than the effective precision minimum value, and the effective precision maximum value and the effective precision median value correspond to the X and Y dimensions.
[0198] Alternatively, as shown in FIG16E , the first and fourth candidate subnodes of the eight candidate subnodes are divided into one subnode, the second and third candidate subnodes are divided into one subnode, the fifth and eighth candidate subnodes are divided into one subnode, and the sixth and seventh candidate subnodes are divided into one subnode, thereby determining the four subnodes corresponding to the current scale node. It can be seen that this oblique division method is equivalent to dividing along the plane diagonals of the Y and Z dimensions, respectively. In some embodiments, this oblique division method can be used when the effective precision maximum value is equal to the effective precision median value, the effective precision median value is greater than the effective precision minimum value, and the effective precision maximum value and the effective precision median value correspond to the Y and Z dimensions.
[0199] In the third case, when the maximum effective precision is equal to the median effective precision, and the median effective precision is equal to the minimum effective precision, the current scale node is partitioned into octrees to determine the eight child nodes corresponding to the current scale node.
[0200] In the embodiment of the present application, when the maximum effective precision is equal to the median effective precision, and the median effective precision is equal to the minimum effective precision, the current scale node is partitioned using an octree partitioning method to determine the eight child nodes corresponding to the current scale node. Figure 17 shows an octree partitioning method.
[0201] It is understandable that in the embodiment of the present application, the current scale node in the current scale point cloud is divided according to the effective accuracy of the nodes in the three spatial dimensions in the current scale point cloud, and the child nodes corresponding to the current scale node are determined. When encoding the occupancy information for sparse point clouds or irregularly distributed point clouds, even if the initial distribution of the point cloud is relatively flat or irregular, after several divisions along the longest dimension, it can quickly become a cube. In this way, when performing KNN neighborhood feature value extraction based on this encoding scheme, there is no need to increase the size of the convolution kernel K to enhance the feature value. It is only necessary to use a partitioning method such as a binary tree or a quadtree to continuously reduce the accuracy of the longest dimension, and finally achieve the ideal state of cube or octree partitioning, so that the sparse convolution network and the appropriate convolution kernel K can be used to effectively extract the geometric feature values, improve the geometric coding efficiency of the point cloud, and thus improve the coding performance.
[0202] For example, the above encoding method may be as shown in FIG18 , as follows:
[0203] S201, determine dim max 、dim mid , and dim min The value of .
[0204] In S201, the encoder determines the effective precision dim of the current scale point cloud in three spatial dimensions. x ,dim y ,dim z ; According to dim x ,dim y ,dim z Determine the maximum effective precision dim max , effective precision median dim mid and the minimum effective precision dim min .
[0205] S202, determine dim max Whether it is 0.
[0206] In S202, if dim max If dim is 0, the encoding ends. max If it is not 0, execute S203.
[0207] S203, determine dim max Is it equal to dim? mid .
[0208] In S203, if dim max equal to dim mid , then execute S205, if not, then execute S204.
[0209] S204: Determine two child nodes through binary tree partitioning.
[0210] In S204, in dim max >dim mid &&dim mid >dim min , or, dim max >dim mid &&dim mid =dim min In the case of , a binary tree is used for partitioning, and the partitioning is performed along the dimension of the maximum effective precision to determine the two child nodes corresponding to the current scale node. That is, each current scale node has two child nodes whose placeholder information needs to be encoded.
[0211] S205, determine dim mid Is it equal to dim? min .
[0212] In S205, if dim mid equal to dim min , then execute S207, if not, then execute S206.
[0213] S206: Determine two child nodes through quadtree partitioning.
[0214] In S204, in dim max = = dim mid &&dim mid >dim min In the case of , a quadtree is used for partitioning, and the partitioning is performed along the dimension of the maximum effective precision and the dimension of the median effective precision to determine the four child nodes corresponding to the current scale node. That is, each current scale node has four child nodes whose placeholder information needs to be encoded.
[0215] S207: Determine eight child nodes through octree partitioning.
[0216] In S207, in dim max = = dim mid &&dim mid = = dim min In the case of , octree is used for partitioning to determine the eight child nodes corresponding to the current scale node, that is, each current scale node has eight child nodes whose placeholder information needs to be encoded.
[0217] S208, SparsePCGC encoding.
[0218] In S208, encoding is performed based on the placeholder information of the divided sub-nodes to determine the encoding information corresponding to the current scale node, and thus the encoding information of the current scale point cloud. The process of S201-S208 is then repeated to encode the next scale point cloud until the root node is reached, indicating that the current point cloud has been merged into a root node.
[0219] It is understandable that the geometric coding scheme based on the multi-tree structure division method proposed in the embodiment of the present application, when encoding the current scale point cloud, adaptively determines the division method of the current layer nodes according to the geometric distribution of the nodes in the current scale point cloud, including a variety of combined division schemes such as binary tree, quadtree and octree. This can effectively solve the limitations of the traditional octree division method on sparse point clouds (scene point clouds). Due to the irregular distribution of sparse point clouds, it is usually difficult to search for a valid neighborhood point set based on the KNN neighborhood search algorithm. In most traditional coding algorithms and deep learning coding algorithms, neighborhood points are generally searched at the cost of expanding the neighborhood search range or increasing the convolution kernel size, and this solution is generally achieved at the cost of exponential growth in coding complexity and network parameters. In an embodiment of the present application, the spatial geometric distribution of the nodes in the layer to be encoded is used to adaptively determine the division method of the nodes in the current layer. The binary tree division method is used to continuously divide the longest precision dimension, and the quadtree division method is used to continuously divide the longest and medium dimensions until the accuracy of the three dimensions is consistent. That is, when the point cloud is regularly distributed in space, the traditional octree is used for division. This scheme can ensure that after the division, the nodes to be encoded are within a certain range in space, and an effective neighborhood prediction point set can be obtained based on the KNN neighborhood search, thereby avoiding the need to expand the neighborhood search range by increasing the convolution kernel size K. In this way, the point cloud coding efficiency is improved, and the point cloud encoding and decoding performance is thereby improved.
[0220] See Figure 19, which is an optional flow chart of the decoding method provided in an embodiment of the present application, and will be explained in conjunction with the steps shown in Figure 19.
[0221] S301 , parsing the code stream to determine the effective accuracy of the current scale point cloud in three spatial dimensions, or the target dimension identifier corresponding to the current scale point cloud.
[0222] In S301, the decoder determines the initial scale of the point cloud to be decoded (such as the root node point cloud) by parsing the bitstream, performs upsampling and placeholder information decoding on the point cloud to be decoded scale by scale, and restores the geometric information of the point cloud scale by scale.
[0223] In the decoder, entropy decoding is performed on the initial-scale point cloud to determine the geometric data of the initial-scale point cloud. Based on the geometric data of the initial-scale point cloud, the decoder performs scale-by-scale upsampling and occupancy information decoding to restore the geometric information of the point cloud scale by scale.
[0224] Here, the point cloud to be decoded at each scale determined by upsampling is considered the current scale point cloud. The current scale point cloud can also be considered the current layer point cloud, and the current scale nodes in the current scale point cloud are nodes at the same layer and scale. The decoder upsamples the current layer point cloud, and the next scale point cloud determined can be considered the next layer point cloud. The current layer point cloud is the parent node layer of the next layer point cloud, and the current layer nodes in the current layer point cloud are the parent nodes of the nodes in the next layer point cloud.
[0225] In this embodiment of the present application, for the current-scale point cloud (i.e., the current layer point cloud), the decoder determines the effective accuracy of the current-scale point cloud in three spatial dimensions based on the bounding box size corresponding to the current-scale point cloud. Alternatively, the decoder determines the target dimension identifier by parsing the point cloud target dimension identifier in the bitstream. The target dimension identifier is determined by the encoder based on the density of the current-scale point cloud in three spatial dimensions and encoded and sent to the decoder.
[0226] S302 : Determine a child node corresponding to a current-scale node in the current-scale point cloud according to effective accuracies in three spatial dimensions or according to a target dimension identifier.
[0227] In S302, the decoder determines the child nodes corresponding to the current scale node in the current scale point cloud based on the effective precision in the three spatial dimensions. That is, the decoder determines the division method of the current scale point cloud through implicit deduction, and uses the division method to divide each current scale node in the current scale point cloud to determine the child nodes corresponding to each current scale node.
[0228] Alternatively, the decoder divides the current-scale node in the current-scale point cloud along the spatial dimension corresponding to the target dimension identifier and determines the child nodes corresponding to the current-scale node. In other words, the division method of the current-scale point cloud is explicitly determined, and each current-scale node in the current-scale point cloud is divided using this division method to determine the child nodes corresponding to each current-scale node.
[0229] S303: Decode the child node corresponding to the current scale node to determine the placeholder information corresponding to the child node.
[0230] In S303, the decoder decodes the child node corresponding to the current scale node and determines the occupancy information corresponding to the child node. The occupancy information corresponding to the determined child node can then be used to restore the geometric information of the next scale point cloud (i.e., the next layer of point cloud). The decoder performs the next decoding and reconstruction based on the geometric information corresponding to the next scale point cloud until the target scale point cloud is restored. Here, the decoding order of the decoder is from the root node to the leaf node. The target scale can be set according to actual needs. It can be the original scale point cloud corresponding to the encoder or other scales. The specific selection is based on the actual situation and is not limited in the embodiment of the present application.
[0231] It is understandable that the effective accuracy or density of the current-scale point cloud in the three spatial dimensions characterizes the distribution of nodes in the current-scale point cloud in the three spatial dimensions. The embodiment of the present application divides the current-scale node in the current-scale point cloud according to the distribution of nodes in the current-scale point cloud in the three spatial dimensions, determines the subnodes corresponding to the current-scale node, and ensures that the subnodes obtained by the division are distributed more evenly. When performing KNN neighborhood feature value extraction based on the above-mentioned subnode division method, there is no need to increase the size of the convolution kernel K to enhance the feature value. It is only necessary to continuously reduce the accuracy of the longest dimension through the subnode division method, and finally achieve the ideal state of cube or octree division, so that the geometric feature values can be effectively extracted with the help of sparse convolutional networks and appropriate convolution kernels K, thereby improving the geometric decoding efficiency of the point cloud and thus improving the decoding performance.
[0232] In some embodiments, determining, based on the effective precision in the three spatial dimensions, a child node corresponding to a current-scale node in the current-scale point cloud includes:
[0233] Determine the maximum effective precision, the median effective precision, and the minimum effective precision according to the effective precision in the three spatial dimensions;
[0234] The current scale node in the current scale point cloud is divided according to the maximum effective precision value, the median effective precision value, and the minimum effective precision value, and a child node corresponding to the current scale node is determined.
[0235] In some embodiments, dividing the current scale node in the current scale point cloud according to the maximum effective precision value, the median effective precision value, and the minimum effective precision value, and determining the child node corresponding to the current scale node includes:
[0236] When the maximum effective precision value is greater than the median effective precision value, the current scale node is divided to determine two child nodes corresponding to the current scale node.
[0237] In some embodiments, dividing the current scale node in the current scale point cloud according to the maximum effective precision value, the median effective precision value, and the minimum effective precision value, and determining the child node corresponding to the current scale node includes:
[0238] When the maximum effective precision value is equal to the median effective precision value, and the median effective precision value is greater than the minimum effective precision value, the current scale node is divided to determine four child nodes corresponding to the current scale node.
[0239] In some embodiments, dividing the current scale node in the current scale point cloud according to the maximum effective precision value, the median effective precision value, and the minimum effective precision value, and determining the child node corresponding to the current scale node includes:
[0240] When the maximum effective precision value is equal to the median effective precision value, and the median effective precision value is equal to the minimum effective precision value, octree partitioning is performed on the current scale node to determine eight child nodes corresponding to the current scale node.
[0241] In some embodiments, dividing the current scale node to determine two child nodes corresponding to the current scale node includes:
[0242] Perform binary tree partitioning on the current scale node along the dimension corresponding to the maximum effective precision to determine two child nodes corresponding to the current scale node.
[0243] In some embodiments, the current scale node corresponds to eight candidate child nodes in the three spatial dimensions; and dividing the current scale node to determine the two child nodes corresponding to the current scale node includes:
[0244] Splitting the first, fourth, fifth, and eighth candidate subnodes among the eight candidate subnodes into one subnode, and splitting the second, third, sixth, and seventh candidate subnodes into one subnode, thereby determining two subnodes corresponding to the current scale node;
[0245] or,
[0246] The first candidate subnode, the third candidate subnode, the sixth candidate subnode, and the eighth candidate subnode among the eight subnodes are divided into one subnode, and the second candidate subnode, the fourth candidate subnode, the fifth candidate subnode, and the seventh candidate subnode are divided into one subnode, thereby determining two subnodes corresponding to the current scale node.
[0247] In some embodiments, dividing the current scale node to determine four child nodes corresponding to the current scale node includes:
[0248] Perform quadtree partitioning on the current scale node along the dimension corresponding to the maximum effective precision and the dimension corresponding to the median of the effective precision, respectively, to determine four child nodes corresponding to the current scale node.
[0249] In some embodiments, the current scale node corresponds to eight sub-nodes in the three spatial dimensions; dividing the current scale node to determine the four sub-nodes corresponding to the current scale node includes:
[0250] Splitting the first candidate subnode and the eighth candidate subnode among the eight candidate subnodes into one subnode, splitting the second candidate subnode and the seventh candidate subnode into one subnode, splitting the third candidate subnode and the sixth candidate subnode into one subnode, and splitting the fourth candidate subnode and the fifth candidate subnode into one subnode, thereby determining four subnodes corresponding to the current scale node;
[0251] or,
[0252] Dividing the first candidate subnode and the sixth candidate subnode of the eight candidate subnodes into one subnode, dividing the second candidate subnode and the fifth candidate subnode into one subnode, dividing the third candidate subnode and the eighth candidate subnode into one subnode, and dividing the fourth candidate subnode and the seventh candidate subnode into one subnode, thereby determining four nodes corresponding to the current scale node;
[0253] or,
[0254] Splitting the first candidate subnode and the seventh candidate subnode among the eight candidate subnodes into one subnode, splitting the second candidate subnode and the eighth candidate subnode into one subnode, splitting the third candidate subnode and the fifth candidate subnode into one subnode, and splitting the fourth candidate subnode and the sixth candidate subnode into one subnode, thereby determining four subnodes corresponding to the current scale node;
[0255] or,
[0256] The first candidate subnode and the fourth candidate subnode among the eight candidate subnodes are divided into one subnode, the second candidate subnode and the third candidate subnode are divided into one subnode, the fifth candidate subnode and the eighth candidate subnode are divided into one subnode, and the sixth candidate subnode and the seventh candidate subnode are divided into one subnode, thereby determining four subnodes corresponding to the current scale node.
[0257] It should be noted that the sub-node division method and basis on the decoder side are consistent with those on the encoder side. The implementation details of the above process can be referred to the description of the corresponding part in the encoder embodiment, which will not be repeated here.
[0258] The embodiment of the present application provides a decoder 1, as shown in FIG20 , including:
[0259] The parsing section 11 is configured to parse the code stream to determine the effective accuracy of the current-scale point cloud in three spatial dimensions, or a target dimension identifier corresponding to the current-scale point cloud; the target dimension identifier is determined based on the density of the current-scale point cloud in the three spatial dimensions;
[0260] A first determining part 12 is configured to determine a child node corresponding to a current scale node in the current scale point cloud according to the effective precision in the three spatial dimensions or according to the target dimension identifier;
[0261] The decoding part 13 is configured to decode the child node corresponding to the current scale node and determine the placeholder information corresponding to the child node.
[0262] In some embodiments, the first determining part 12 is further configured to determine the maximum effective precision, the median effective precision and the minimum effective precision according to the effective precision in the three spatial dimensions;
[0263] The current scale node in the current scale point cloud is divided according to the maximum effective precision value, the median effective precision value, and the minimum effective precision value, and a child node corresponding to the current scale node is determined.
[0264] In some embodiments, the first determining part 12 is further configured to divide the current scale node and determine two child nodes corresponding to the current scale node when the maximum effective precision value is greater than the median effective precision value.
[0265] In some embodiments, the first determining part 12 is further configured to divide the current scale node and determine four child nodes corresponding to the current scale node when the maximum effective precision value is equal to the median effective precision value and the median effective precision value is greater than the minimum effective precision value.
[0266] In some embodiments, the first determining part 12 is further configured to perform octree partitioning on the current scale node to determine eight child nodes corresponding to the current scale node when the maximum effective precision value is equal to the median effective precision value and the median effective precision value is equal to the minimum effective precision value.
[0267] In some embodiments, the first determining part 12 is further configured to perform binary tree partitioning on the current scale node along the dimension corresponding to the maximum effective precision to determine two child nodes corresponding to the current scale node.
[0268] In some embodiments, the current scale node corresponds to eight subnodes in the three spatial dimensions; the first determining part 12 is further configured to divide the first candidate subnode, the fourth candidate subnode, the fifth candidate subnode, and the eighth candidate subnode among the eight candidate subnodes into one subnode, and divide the second candidate subnode, the third candidate subnode, the sixth candidate subnode, and the seventh candidate subnode into one subnode, thereby determining the two subnodes corresponding to the current scale node;
[0269] or,
[0270] The first candidate subnode, the third candidate subnode, the sixth candidate subnode, and the eighth candidate subnode among the eight subnodes are divided into one subnode, and the second candidate subnode, the fourth candidate subnode, the fifth candidate subnode, and the seventh candidate subnode are divided into one subnode, thereby determining two subnodes corresponding to the current scale node.
[0271] In some embodiments, the first determining part 12 is further configured to perform quadtree partitioning on the current scale node along the dimension corresponding to the maximum effective precision and the dimension corresponding to the median effective precision, respectively, to determine four child nodes corresponding to the current scale node.
[0272] In some embodiments, the current scale node corresponds to eight subnodes in the three spatial dimensions; the first determining portion 12 is further configured to divide the first candidate subnode and the eighth candidate subnode of the eight candidate subnodes into one subnode, divide the second candidate subnode and the seventh candidate subnode into one subnode, divide the third candidate subnode and the sixth candidate subnode into one subnode, and divide the fourth candidate subnode and the fifth candidate subnode into one subnode, thereby determining the four subnodes corresponding to the current scale node;
[0273] or,
[0274] Dividing the first candidate subnode and the sixth candidate subnode of the eight candidate subnodes into one subnode, dividing the second candidate subnode and the fifth candidate subnode into one subnode, dividing the third candidate subnode and the eighth candidate subnode into one subnode, and dividing the fourth candidate subnode and the seventh candidate subnode into one subnode, thereby determining four nodes corresponding to the current scale node;
[0275] or,
[0276] Splitting the first candidate subnode and the seventh candidate subnode among the eight candidate subnodes into one subnode, splitting the second candidate subnode and the eighth candidate subnode into one subnode, splitting the third candidate subnode and the fifth candidate subnode into one subnode, and splitting the fourth candidate subnode and the sixth candidate subnode into one subnode, thereby determining four subnodes corresponding to the current scale node;
[0277] or,
[0278] The first candidate subnode and the fourth candidate subnode among the eight candidate subnodes are divided into one subnode, the second candidate subnode and the third candidate subnode are divided into one subnode, the fifth candidate subnode and the eighth candidate subnode are divided into one subnode, and the sixth candidate subnode and the seventh candidate subnode are divided into one subnode, thereby determining four subnodes corresponding to the current scale node.
[0279] In some embodiments, the first determining part 12 is further configured to divide the current scale node in the current scale point cloud along the spatial dimension corresponding to the target dimension identifier, and determine the child nodes corresponding to the current scale node.
[0280] In some embodiments, the decoding part 3 is further configured to restore geometric information corresponding to the next scale point cloud based on the occupancy information corresponding to the sub-node; and perform the next decoding and reconstruction based on the geometric information corresponding to the next scale point cloud until the target scale point cloud is restored.
[0281] The embodiment of the present application provides an encoder 2, as shown in FIG21 , including:
[0282] The second determining part 21 is configured to determine the effective accuracy or density of the current scale point cloud in three spatial dimensions; and determine the child node corresponding to the current scale node in the current scale point cloud according to the effective accuracy or density in the three spatial dimensions;
[0283] The encoding part 22 is configured to encode the placeholder information of the child nodes corresponding to the current-scale node to determine the encoding information of the current-scale point cloud.
[0284] In some embodiments, the second determining part 21 is further configured to determine the maximum effective precision, the median effective precision and the minimum effective precision according to the effective precision in the three spatial dimensions;
[0285] The current scale node is divided according to the maximum effective precision value, the median effective precision value, and the minimum effective precision value, and a child node corresponding to the current scale node is determined.
[0286] In some embodiments, the second determining part 21 is further configured to divide the current scale node and determine two child nodes corresponding to the current scale node when the maximum effective precision value is greater than the median effective precision value.
[0287] In some embodiments, the second determining part 21 is further configured to divide the current scale node and determine four child nodes corresponding to the current scale node when the maximum effective precision value is equal to the median effective precision value and the median effective precision value is greater than the minimum effective precision value.
[0288] In some embodiments, the second determining part 21 is further configured to perform octree partitioning on the current scale node to determine eight child nodes corresponding to the current scale node when the maximum effective precision value is equal to the median effective precision value and the median effective precision value is equal to the minimum effective precision value.
[0289] In some embodiments, the second determining part 21 is further configured to perform binary tree partitioning on the current scale node along the dimension corresponding to the maximum effective precision to determine two child nodes corresponding to the current scale node.
[0290] In some embodiments, the current scale node corresponds to eight candidate subnodes in the three spatial dimensions; the second determining portion 21 is further configured to divide the first candidate subnode, the fourth candidate subnode, the fifth candidate subnode, and the eighth candidate subnode among the eight candidate subnodes into one subnode, and divide the second candidate subnode, the third candidate subnode, the sixth candidate subnode, and the seventh candidate subnode into one subnode, thereby determining two subnodes corresponding to the current scale node;
[0291] or,
[0292] The first, third, sixth, and eighth candidate subnodes among the eight candidate subnodes are divided into one subnode, and the second, fourth, fifth, and seventh candidate subnodes are divided into one subnode, thereby determining two subnodes corresponding to the current scale node.
[0293] In some embodiments, the second determining part 21 is further configured to perform quadtree partitioning on the current scale node along the dimension corresponding to the maximum effective precision and the dimension corresponding to the median of the effective precision, respectively, to determine four child nodes corresponding to the current scale node.
[0294] In some embodiments, the current scale node corresponds to eight candidate subnodes in the three spatial dimensions; the second determining portion 21 is further configured to divide the first candidate subnode and the eighth candidate subnode of the eight candidate subnodes into one subnode, divide the second candidate subnode and the seventh candidate subnode into one subnode, divide the third candidate subnode and the sixth candidate subnode into one subnode, and divide the fourth candidate subnode and the fifth candidate subnode into one subnode, thereby determining the four subnodes corresponding to the current scale node;
[0295] or,
[0296] Dividing the first candidate subnode and the sixth candidate subnode of the eight candidate subnodes into one subnode, dividing the second candidate subnode and the fifth candidate subnode into one subnode, dividing the third candidate subnode and the eighth candidate subnode into one subnode, and dividing the fourth candidate subnode and the seventh candidate subnode into one subnode, thereby determining four nodes corresponding to the current scale node;
[0297] or,
[0298] Splitting the first candidate subnode and the seventh candidate subnode among the eight candidate subnodes into one subnode, splitting the second candidate subnode and the eighth candidate subnode into one subnode, splitting the third candidate subnode and the fifth candidate subnode into one subnode, and splitting the fourth candidate subnode and the sixth candidate subnode into one subnode, thereby determining four subnodes corresponding to the current scale node;
[0299] or,
[0300] The first candidate subnode and the fourth candidate subnode among the eight candidate subnodes are divided into one subnode, the second candidate subnode and the third candidate subnode are divided into one subnode, the fifth candidate subnode and the eighth candidate subnode are divided into one subnode, and the sixth candidate subnode and the seventh candidate subnode are divided into one subnode, thereby determining four subnodes corresponding to the current scale node.
[0301] In some embodiments, the encoding part 22 is further configured to write the encoding information of the current-scale point cloud into a bitstream.
[0302] In some embodiments, the second determining part 21 is further configured to perform principal component analysis on the current-scale point cloud in the three spatial dimensions to determine the density in the three spatial dimensions.
[0303] In some embodiments, the second determining portion 21 is further configured to determine the target dimension based on the densities in the three spatial dimensions;
[0304] The current scale node is divided according to the target dimension, and child nodes corresponding to the current scale node are determined.
[0305] In some embodiments, the second determining part 21 is further configured to determine a target dimension identifier corresponding to the target dimension;
[0306] The encoding part 22 is further configured to encode the target dimension identifier and write the obtained encoding bits and the encoding information of the current-scale point cloud into a bitstream.
[0307] In some embodiments, the encoding part 22 is further configured to downsample the current scale point cloud to determine a next scale point cloud;
[0308] Based on the effective accuracy or density of the next-scale point cloud in the three spatial dimensions, the next-scale point cloud is divided into sub-nodes and the occupancy information is encoded, and the encoding information of the next-scale point cloud is determined until the maximum effective accuracy in the three spatial dimensions reaches a preset accuracy threshold, or until the minimum density in the three spatial dimensions reaches a preset density threshold, the point cloud encoding is completed.
[0309] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0310] In some embodiments, the present application also provides a decoder. Figure 22 is a schematic diagram of an optional structure of a decoder 3 provided in the present application. As shown in Figure 22, the decoder 3 includes a first memory 32 and a first processor 33. The first memory 32 and the first processor 33 are connected via a first communication bus 34. The first memory 32 is used to store executable instructions. The first processor 33 is used to implement the decoding method provided in the present application when executing the executable instructions stored in the first memory 32.
[0311] In some embodiments, the present application also provides an encoder. Figure 23 is a schematic diagram of an optional structure of an encoder 4 provided in the present application. As shown in Figure 23, encoder 4 includes a second memory 42 and a second processor 43. The second memory 42 and the second processor 43 are connected via a second communication bus 44. The second memory 42 is used to store executable instructions. The second processor 43 is used to implement the encoding method provided in the present application when executing the executable instructions stored in the second memory 42.
[0312] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a first processor, the first processor will be caused to execute any one of the decoding methods provided in the embodiments of the present application; or, when the executable instructions are executed by a second processor, the second processor will be caused to execute any one of the encoding methods provided in the embodiments of the present application.
[0313] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0314] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0315] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0316] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0317] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0318] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.
[0319] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0320] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0321] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application. Industrial Applicability
[0322] The embodiment of the present application provides a coding and decoding method, a decoder, an encoder and a computer-readable storage medium. The effective accuracy or density of the current scale point cloud in the three spatial dimensions characterizes the distribution of the nodes in the current scale point cloud in the three spatial dimensions. The embodiment of the present application divides the current scale node in the current scale point cloud according to the distribution of the nodes in the current scale point cloud in the three spatial dimensions, determines the sub-nodes corresponding to the current scale node, and ensures that the distribution of the divided sub-nodes is relatively uniform. When performing KN N neighborhood feature value extraction based on the above sub-node division method, there is no need to increase the size of the convolution kernel K to enhance the feature value. It is only necessary to continuously reduce the accuracy of the longest dimension through the sub-node division method, and finally achieve the ideal state of cube or octree division, so that the geometric feature values can be effectively extracted with the help of sparse convolutional networks and appropriate convolution kernel K, thereby improving the geometric coding and decoding efficiency of the point cloud, and further improving the coding and decoding performance.
Claims
1. A decoding method, applied to a decoder, comprising: Analyzing a bitstream to determine the effective precision of the current scale point cloud in three spatial dimensions, or the target dimension identifier corresponding to the current scale point cloud; The target dimension identifier is determined according to the density of the current scale point cloud in the three spatial dimensions; Determining the child nodes corresponding to the current scale nodes in the current scale point cloud according to the effective precision in the three spatial dimensions or according to the target dimension identifier; Decoding the child nodes corresponding to the current scale nodes to determine the occupancy information corresponding to the child nodes.
2. The method according to claim 1, wherein, The determining the child nodes corresponding to the current scale nodes in the current scale point cloud according to the effective precision in the three spatial dimensions includes: Determining the maximum effective precision, the median effective precision, and the minimum effective precision according to the effective precision in the three spatial dimensions; Dividing the current scale nodes in the current scale point cloud according to the maximum effective precision, the median effective precision, and the minimum effective precision to determine the child nodes corresponding to the current scale nodes.
3. The method according to claim 2, wherein The dividing the current scale nodes in the current scale point cloud according to the maximum effective precision, the median effective precision, and the minimum effective precision to determine the child nodes corresponding to the current scale nodes includes: When the maximum effective precision is greater than the median effective precision, dividing the current scale nodes to determine two child nodes corresponding to the current scale nodes.
4. The method according to claim 2, wherein The dividing the current scale nodes in the current scale point cloud according to the maximum effective precision, the median effective precision, and the minimum effective precision to determine the child nodes corresponding to the current scale nodes includes: When the maximum effective precision is equal to the median effective precision and the median effective precision is greater than the minimum effective precision, dividing the current scale nodes to determine four child nodes corresponding to the current scale nodes.
5. The method according to claim 2, wherein The dividing the current scale nodes in the current scale point cloud according to the maximum effective precision, the median effective precision, and the minimum effective precision to determine the child nodes corresponding to the current scale nodes includes: When the maximum effective precision is equal to the median effective precision and the median effective precision is equal to the minimum effective precision, performing an octree division on the current scale nodes to determine eight child nodes corresponding to the current scale nodes.
6. The method according to claim 3, wherein The dividing the current scale nodes to determine two child nodes corresponding to the current scale nodes includes: Performing a binary tree division on the current scale nodes along the dimension corresponding to the maximum effective precision to determine two child nodes corresponding to the current scale nodes.
7. The method according to claim 3, wherein The current scale nodes correspond to eight candidate child nodes in the three spatial dimensions; The dividing the current scale nodes to determine two child nodes corresponding to the current scale nodes includes: Divide the first candidate child node, the fourth candidate child node, the fifth candidate child node, and the eighth candidate child node among the eight candidate child nodes into one child node, and divide the second candidate child node, the third candidate child node, the sixth candidate child node, and the seventh candidate child node into one child node, so as to determine the two child nodes corresponding to the current scale node; Or, Divide the first candidate child node, the third candidate child node, the sixth candidate child node, and the eighth candidate child node among the eight child nodes into one child node, and divide the second candidate child node, the fourth candidate child node, the fifth candidate child node, and the seventh candidate child node into one child node, so as to determine the two child nodes corresponding to the current scale node.
8. The method according to claim 4, wherein The dividing the current scale node to determine the four child nodes corresponding to the current scale node includes: Perform a quadtree division on the current scale node along the dimension corresponding to the maximum effective precision and the dimension corresponding to the median of the effective precision respectively, so as to determine the four child nodes corresponding to the current scale node.
9. The method according to claim 4, wherein, The current scale node corresponds to eight child nodes on the three spatial dimensions; The dividing the current scale node to determine the four child nodes corresponding to the current scale node includes: Divide the first candidate child node and the eighth candidate child node among the eight candidate child nodes into one child node, divide the second candidate child node and the seventh candidate child node into one child node, divide the third candidate child node and the sixth candidate child node into one child node, and divide the fourth candidate child node and the fifth candidate child node into one child node, so as to determine the four child nodes corresponding to the current scale node; Or, Divide the first candidate child node and the sixth candidate child node among the eight candidate child nodes into one child node, divide the second candidate child node and the fifth candidate child node into one child node, divide the third candidate child node and the eighth candidate child node into one child node, and Divide the fourth candidate child node and the seventh candidate child node into one child node, so as to determine the four nodes corresponding to the current scale node; Or, Divide the first candidate child node and the seventh candidate child node among the eight candidate child nodes into one child node, divide the second candidate child node and the eighth candidate child node into one child node, divide the third candidate child node and the fifth candidate child node into one child node, and divide the fourth candidate child node and the sixth candidate child node into one child node, so as to determine the four child nodes corresponding to the current scale node; Or, Divide the first candidate child node and the fourth candidate child node among the eight candidate child nodes into one child node, divide the second candidate child node and the third candidate child node into one child node, divide the fifth candidate child node and the eighth candidate child node into one child node, and divide the sixth candidate child node and the seventh candidate child node into one child node, so as to determine the four child nodes corresponding to the current scale node.
10. The method according to claim 1, wherein, Determining the child nodes corresponding to the current scale node in the current scale point cloud according to the target dimension identifier includes: Identify the corresponding spatial dimension along the target dimension, divide the current scale nodes in the current scale point cloud, and determine the child nodes corresponding to the current scale nodes.
11. The method according to claim 2 or 10, wherein The method further includes: Restore the geometric information corresponding to the next scale point cloud according to the occupancy information corresponding to the child nodes; Perform the next decoding and reconstruction based on the geometric information corresponding to the next scale point cloud until the target scale point cloud is restored.
12. An encoding method applied to an encoder, including: Determine the effective precision or density of the current scale point cloud in three spatial dimensions; Determine the child nodes corresponding to the current scale nodes in the current scale point cloud according to the effective precision or density in the three spatial dimensions; Encode the occupancy information of the child nodes corresponding to the current scale nodes to determine the encoded information of the current scale point cloud.
13. The method according to claim 12, wherein, Determining the child nodes corresponding to the current scale nodes in the current scale point cloud according to the effective precision in the three spatial dimensions includes: Determine the maximum effective precision, the median effective precision, and the minimum effective precision according to the effective precision in the three spatial dimensions; Divide the current scale node according to the maximum effective precision, the median effective precision, and the minimum effective precision to determine the child nodes corresponding to the current scale node.
14. The method according to claim 13, wherein, The dividing the current scale node according to the maximum effective precision, the median effective precision, and the minimum effective precision to determine the child nodes corresponding to the current scale node includes: When the maximum effective precision is greater than the median effective precision, divide the current scale node to determine two child nodes corresponding to the current scale node.
15. The method according to claim 13, wherein, The dividing the current scale node according to the maximum effective precision, the median effective precision, and the minimum effective precision to determine the child nodes corresponding to the current scale node includes: When the maximum effective precision is equal to the median effective precision and the median effective precision is greater than the minimum effective precision, divide the current scale node to determine four child nodes corresponding to the current scale node.
16. The method according to claim 13, wherein, The dividing the current scale node according to the maximum effective precision, the median effective precision, and the minimum effective precision to determine the child nodes corresponding to the current scale node includes: When the maximum effective precision is equal to the median effective precision and the median effective precision is equal to the minimum effective precision, perform an octree division on the current scale node to determine eight child nodes corresponding to the current scale node.
17. The method according to claim 14, wherein The dividing the current scale node to determine two child nodes corresponding to the current scale node includes: Perform a binary tree division on the current scale node along the dimension corresponding to the maximum effective precision to determine two child nodes corresponding to the current scale node.
18. The method according to claim 14, wherein The current scale node corresponds to eight candidate child nodes in the three spatial dimensions; The dividing the current scale node to determine two child nodes corresponding to the current scale node includes: Divide the first candidate child node, the fourth candidate child node, the fifth candidate child node, and the eighth candidate child node among the eight candidate child nodes into one child node, and divide the second candidate child node, the third candidate child node, the sixth candidate child node, and the seventh candidate child node into one child node, so as to determine two child nodes corresponding to the current scale node; Or, Divide the first candidate child node, the third candidate child node, the sixth candidate child node, and the eighth candidate child node among the eight candidate child nodes into one child node, and divide the second candidate child node, the fourth candidate child node, the fifth candidate child node, and the seventh candidate child node into one child node, so as to determine two child nodes corresponding to the current scale node.
19. The method according to claim 15, wherein The dividing the current scale node to determine four child nodes corresponding to the current scale node includes: Perform a quadtree division on the current scale node respectively along the dimension corresponding to the maximum effective precision and the dimension corresponding to the median of the effective precision to determine four child nodes corresponding to the current scale node.
20. The method according to claim 15, wherein, The current scale node corresponds to eight candidate child nodes in the three spatial dimensions; The dividing the current scale node to determine four child nodes corresponding to the current scale node includes: Divide the first candidate child node and the eighth candidate child node among the eight candidate child nodes into one child node, divide the second candidate child node and the seventh candidate child node into one child node, divide the third candidate child node and the sixth candidate child node into one child node, and divide the fourth candidate child node and the fifth candidate child node into one child node, so as to determine four child nodes corresponding to the current scale node; Or, Divide the first candidate child node and the sixth candidate child node among the eight candidate child nodes into one child node, divide the second candidate child node and the fifth candidate child node into one child node, divide the third candidate child node and the eighth candidate child node into one child node, and divide the fourth candidate child node and the seventh candidate child node into one child node, so as to determine four nodes corresponding to the current scale node; Or, Divide the first candidate child node and the seventh candidate child node among the eight candidate child nodes into one child node, divide the second candidate child node and the eighth candidate child node into one child node, divide the third candidate child node and the fifth candidate child node into one child node, and divide the fourth candidate child node and the sixth candidate child node into one child node, so as to determine four child nodes corresponding to the current scale node; Or, Divide the first candidate child node and the fourth candidate child node among the eight candidate child nodes into one child node, divide the second candidate child node and the third candidate child node into one child node, divide the fifth candidate child node and the eighth candidate child node into one child node, and divide the sixth candidate child node and the seventh candidate child node into one child node, so as to determine four child nodes corresponding to the current scale node.
21. The method according to any one of claims 12-20, wherein, The method further includes: Write the encoding information of the current scale point cloud into the bitstream.
22. The method according to claim 12, wherein, Determine the density of the current scale point cloud in the three spatial dimensions, including: Perform principal component analysis on the current-scale point cloud in the three spatial dimensions to determine the density in the three spatial dimensions.
23. The method according to claim 22, wherein Determine the child nodes corresponding to the current-scale nodes in the current-scale point cloud according to the density in the three spatial dimensions, including: Determine the target dimension according to the density in the three spatial dimensions; Divide the current-scale nodes according to the target dimension to determine the child nodes corresponding to the current-scale nodes.
24. The method according to claim 23, wherein, The method further includes: Determine the target dimension identifier corresponding to the target dimension; Encode the target dimension identifier, and write the obtained encoded bits and the encoded information of the current-scale point cloud into the code stream.
25. The method according to any one of claims 12 - 20 or any one of claims 22 - 24, wherein, The method further includes: Downsample the current-scale point cloud to determine the next-scale point cloud; According to the effective precision or density of the next-scale point cloud in the three spatial dimensions, perform child node division and occupancy information encoding on the next-scale point cloud to determine the encoded information of the next-scale point cloud until the maximum effective precision in the three spatial dimensions reaches a preset precision threshold, or until the minimum density in the three spatial dimensions reaches a preset density threshold, and complete point cloud encoding.
26. A decoder, comprising: A parsing part configured to parse the code stream to determine the effective precision of the current-scale point cloud in three spatial dimensions, or the target dimension identifier corresponding to the current-scale point cloud; The target dimension identifier is determined according to the density of the current-scale point cloud in the three spatial dimensions; A first determination part configured to determine the child nodes corresponding to the current-scale nodes in the current-scale point cloud according to the effective precision in the three spatial dimensions or according to the target dimension identifier; A decoding part configured to decode the child nodes corresponding to the current-scale nodes to determine the occupancy information corresponding to the child nodes.
27. An encoder, comprising: A second determination part configured to determine the effective precision or density of the current-scale point cloud in three spatial dimensions; And determine the child nodes corresponding to the current-scale nodes in the current-scale point cloud according to the effective precision or density in the three spatial dimensions; An encoding part configured to encode the occupancy information of the child nodes corresponding to the current-scale nodes to determine the encoded information of the current-scale point cloud.
28. A decoder, the decoder includes a first memory and a first processor; wherein, The first memory is used to store a computer program that can run on the first processor; The first processor is configured to execute the method according to any one of claims 1 to 11 when running the computer program.
29. An encoder, the encoder includes a second memory and a second processor; wherein, The second memory is used to store a computer program that can run on the second processor; The second processor is configured to execute the method according to any one of claims 12 to 25 when running the computer program.
30. A bitstream, which is generated by performing bit encoding on information to be encoded; wherein, The information to be encoded at least includes at least one of the following: The encoded information of the current-scale point cloud; Among them, the encoding information of the scale point cloud is determined by determining the child nodes corresponding to the current scale nodes in the current scale point cloud according to the effective precision or density of the current scale point cloud in three spatial dimensions, and encoding the occupancy information of the child nodes corresponding to the current scale nodes.
31. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the method described in any one of claims 1 to 11, or implements the method described in any one of claims 12 to 25.
Citation Information
Patent Citations
Point cloud geometrical information encoding and decoding method
CN112565795A
Point cloud geometric lossless compression method based on sparse convolutional neural network
CN113613010A
Point cloud coding and decoding method and device and storage medium
CN115733990A
Point cloud compression method, encoder, decoder, and storage medium
US20230075442A1
Point cloud geometric information compression method and apparatus, point cloud geometric information decompression method and apparatus, point cloud video encoding method and apparatus, and point cloud video decoding method and apparatus
WO2023205969A1