Point cloud encoding method, point cloud decoding method, encoders, decoders and storage mediums
By introducing the condition of the difference in attribute values between the target node and the reference node in the RAHT inter-frame prediction, the problem of poor encoding and decoding performance caused by only considering the same geometric information in the prior art is solved, and the efficiency and effect of point cloud encoding and decoding are improved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2025-01-17
- Publication Date
- 2026-07-23
AI Technical Summary
In existing point cloud encoding and decoding technologies, RAHT inter-frame prediction only considers whether the geometric information of the target nodes is the same, without taking into account the differences in attribute values, resulting in poor encoding and decoding performance.
When determining whether a target node should perform RAHT inter-frame prediction, the difference in attribute values between the target node and the reference node is introduced as a condition to improve the encoding and decoding performance of attribute information.
By taking into account the differences in attribute values, the performance of point cloud encoding and decoding is improved, especially the point cloud encoding and decoding effect under the G-PCC and GES-TM encoding and decoding frameworks.
Smart Images

Figure CN2025073086_23072026_PF_FP_ABST
Abstract
Description
Point cloud encoding / decoding methods, codecs, and storage media Technical Field
[0001] This application relates to the field of point cloud encoding and decoding technology, and in particular to a point cloud encoding and decoding method, an encoder and decoder, and a storage medium. Background Technology
[0002] Within the point cloud encoding / decoding framework, during the encoding and decoding of point cloud attribute information, an encoding / decoding scheme based on region adaptive hierarchical transform (RAHT) can be initiated. How to improve the encoding / decoding performance of attribute information based on RAHT inter-frame prediction is a problem that needs to be solved. Summary of the Invention
[0003] This application provides a point cloud encoding / decoding method, an encoder / decoder, and a storage medium. The various aspects involved in this application are described below.
[0004] In a first aspect, a point cloud decoding method is provided, applied to a decoder, comprising: parsing the bitstream to determine the first reference node of the current node in the current frame; if the first reference node satisfies a first condition, performing RAHT inter-frame prediction on the attribute value of the current node, wherein the first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the second reference node in the reference frame.
[0005] Secondly, a point cloud encoding method is provided, applied to an encoder, comprising: determining a first reference node of the current node in the current frame; if the first reference node satisfies a first condition, performing RAHT inter-frame prediction on the attribute value of the current node, wherein the first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of a second reference node in the reference frame.
[0006] Thirdly, a decoder is provided, comprising: a determination unit configured to parse the bitstream and determine a first reference node of the current node in the current frame; and a prediction unit configured to perform RAHT inter-frame prediction on the attribute value of the current node if the first reference node satisfies a first condition, wherein the first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of a second reference node in the reference frame.
[0007] Fourthly, a decoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method of the first aspect when running the computer program.
[0008] Fifthly, an encoder is provided, comprising: a determining unit configured to determine a first reference node of a current node in a current frame; and a predicting unit configured to perform RAHT inter-frame prediction on the attribute values of the current node if the first reference node satisfies a first condition, wherein the first condition is related to the difference between the reconstructed attribute values of the first reference node and the reconstructed attribute values of a second reference node in the reference frame.
[0009] In a sixth aspect, an encoder is provided, the encoder comprising: a memory for storing a computer program; and a processor for executing the method of the second aspect when running the computer program.
[0010] In a seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program that, when executed, implements the method as described in the first or second aspect.
[0011] Eighthly, a non-volatile computer-readable storage medium is provided for storing a bit stream, the bit stream being generated by an encoding method using an encoder, or the bit stream being decoded by a decoding method using a decoder, wherein the decoding method is as described in the first aspect and the encoding method is as described in the second aspect.
[0012] Ninth aspect, a computer-readable storage medium is provided, which stores a bitstream generated according to the method of the second aspect.
[0013] During RAHT inter-frame prediction, the more similar the attribute values of the current node are to those of a reference node in the reference frame, the better the prediction performance. In related technologies, when determining whether the attribute values of a target node's child nodes are suitable for RAHT inter-frame prediction, only the presence of a reference target node with the same geometric information as the target node in the reference frame is considered, without taking into account the differences between the target node's attribute values and those of the reference target node. This application's embodiment incorporates the difference between the target node's attribute values and those of the reference target node into the determination condition (i.e., the first condition) when determining whether the attribute values of a child node are suitable for RAHT inter-frame prediction, thereby helping to improve the encoding and decoding performance of attribute information. Attached Figure Description
[0014] Figure 1 is a schematic diagram of a network architecture for point cloud encoding and decoding.
[0015] Figure 2A is a schematic diagram of the component framework of a G-PCC encoder.
[0016] Figure 2B is a schematic diagram of the component framework of a G-PCC decoder.
[0017] Figure 3A is a schematic diagram of an attribute information encoding process based on RAHT transformation.
[0018] Figure 3B is a schematic diagram of a decoding process for attribute information based on RAHT transformation.
[0019] Figure 4 is a schematic diagram of a two-point transformation process.
[0020] Figure 5 is a schematic diagram of another two-point transformation process.
[0021] Figure 6 is a schematic diagram of a RAHT transform that includes upsampling prediction.
[0022] Figure 7 is a schematic diagram of a RAHT inverse transform that includes upsampling prediction.
[0023] Figure 8 is a schematic diagram of a coding process based on RAHT inter-frame prediction.
[0024] Figure 9 is a schematic diagram of the coding process for RAHT inter-frame prediction (including judgment conditions).
[0025] Figure 10 is a schematic diagram of the coding process for RAHT inter-frame prediction (including judgment conditions that take into account occupancy code information).
[0026] Figure 11 is a flowchart illustrating the decoding method provided in an embodiment of this application.
[0027] Figure 12 is a flowchart illustrating the judgment conditions for RAHT inter-frame prediction provided in an embodiment of this application.
[0028] Figure 13 is a flowchart illustrating the encoding method provided in an embodiment of this application.
[0029] Figure 14 is a schematic diagram of the encoding process of RAHT inter-frame prediction (including the first condition) provided in the embodiment of this application.
[0030] Figure 15 is a schematic diagram of the structure of a decoder provided in an embodiment of this application.
[0031] Figure 16 is a schematic diagram of the structure of a decoder provided in another embodiment of this application.
[0032] Figure 17 is a schematic diagram of the encoder provided in an embodiment of this application.
[0033] Figure 18 is a schematic diagram of the encoder provided in another embodiment of this application. Detailed Implementation
[0034] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0036] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0037] It should also be noted that the terms "first, second, and third" used in the embodiments of this application are only used to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0038] A point cloud is a set of randomly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or scene. These points contain geometric information representing spatial location and attribute information representing the appearance and texture of the point cloud.
[0039] Two-dimensional images contain information at each pixel, and their distribution is regular, so there's no need to record their positional information separately. However, the distribution of points in a point cloud in three-dimensional space is random and irregular, so the position of each point in space needs to be recorded to fully represent a point cloud. Similar to two-dimensional images, each location during acquisition has corresponding attribute information, usually RGB color values, reflecting the color of an object. For point clouds, in addition to color information, the most common attribute information for each point is reflectance, which reflects the surface material of the object. Therefore, point cloud data typically includes point position information and point attribute information. Point position information can also be called point geometric information. For example, point geometric information can be the three-dimensional coordinates (x, y, z). Point attribute information can include color information and / or reflectance, etc. For example, reflectance can be one-dimensional reflectance information (r); color information can be information in any color space, or it can be three-dimensional color information, such as RGB information. Here, R represents red (red, R), G represents green (green, G), and B represents blue (blue, B). For example, color information can be luminance and chromaticity (YCbCr, YUV) information. Here, Y represents luminance (luma), Cb (U) represents blue color difference, and Cr (V) represents red color difference.
[0040] Point clouds obtained based on laser measurement principles can include the three-dimensional coordinates and reflectance values of each point. Similarly, point clouds obtained based on photogrammetry principles can include the three-dimensional coordinates and three-dimensional color information of each point. Furthermore, point clouds obtained by combining laser measurement and photogrammetry principles can include the three-dimensional coordinates, reflectance values, and three-dimensional color information of each point.
[0041] Currently, point cloud encoding frameworks capable of compressing point clouds can include the G-PCC codec framework provided by the Moving Picture Experts Group (MPEG) or the video-based point cloud compression (V-PCC) codec framework, as well as the AVS-PCC codec framework provided by AVS or the geometry-based solid content test model (GES-TM). The G-PCC codec framework can be used to compress both static point clouds (Type 1) and dynamically acquired point clouds (Type 3), and it can be based on a point cloud compression test platform (test model compression 13, TMC13). The V-PCC codec framework can be used to compress dynamic point clouds (Type 2), and it can be based on a point cloud compression test platform (test model compression 2, TMC2). Therefore, the G-PCC codec framework is also called the point cloud codec TMC13, and the V-PCC codec framework is also called the point cloud codec TMC2. GES-TM is an encoding and decoding framework proposed for dense point clouds (such as point clouds acquired in augmented reality (AR) or virtual reality (VR) scenes).
[0042] This application provides a network architecture for a point cloud encoding / decoding system that includes decoding and encoding methods. Figure 1 is a schematic diagram of such a network architecture. As shown in Figure 1, the network architecture includes one or more electronic devices 13 to 1N and a communication network 01. The electronic devices 13 to 1N can perform video interaction through the communication network 01. The electronic devices can be various types of devices with point cloud encoding / decoding capabilities, such as mobile phones, tablets, personal computers, personal digital assistants, navigators, digital phones, video phones, televisions, sensing devices, servers, etc. This application does not impose any limitations. The decoder or encoder in this application can be one of the aforementioned electronic devices.
[0043] The electronic device in this application embodiment has point cloud encoding and decoding functions, and generally includes a point cloud encoder (i.e., encoder) and a point cloud decoder (i.e. decoder).
[0044] The following section uses the G-PCC and AVS codec frameworks as examples to explain the relevant technologies.
[0045] As can be understood, in the G-PCC encoding and decoding framework for point clouds, the point cloud data to be encoded is first divided into multiple slices. Within each slice, the geometric information of the point cloud and the attribute information corresponding to each point are encoded separately.
[0046] Figure 2A illustrates a schematic diagram of the component framework of a G-PCC encoder. As shown in Figure 2A, during the geometric encoding process, coordinate transformation is performed on the geometric information to ensure that the entire point cloud is contained within a bounding box. Then, quantization is performed; this step primarily serves a scaling function. Due to quantization rounding, some point clouds have identical geometric information, so parameters are used to determine whether to remove duplicate points. This process of quantization and removing duplicate points is also known as voxelization. Next, the bounding box is partitioned into an octree or a prediction tree is constructed. During this process, arithmetic encoding is performed on the points in the leaf nodes of the partition to generate a binary geometric bitstream; or, arithmetic encoding is performed on the vertices generated by the partition (surface fitting based on the vertices) to generate a binary geometric bitstream. During the attribute encoding process, after geometric encoding is completed and the geometric information is reconstructed, color conversion is performed to convert the color information (i.e., attribute information) from the RGB color space to the YUV color space. Then, the reconstructed geometric information is used to recolor the point cloud, so that the unencoded attribute information corresponds to the reconstructed geometric information. Attribute encoding is mainly performed on color information. In the process of color information encoding, there are three main transformation methods. The first two methods rely on the level of detail (LOD) partitioning, namely distance-based lifting transformation and prediction transformation. The third method is to directly perform RAHT. All three methods will transform the color information from the spatial domain to the frequency domain, obtain high-frequency coefficients and low-frequency coefficients through transformation, and finally quantize the coefficients and then perform arithmetic encoding on the quantized coefficients to generate a binary attribute bit stream.
[0047] Figure 2B illustrates a schematic diagram of the G-PCC decoder's structural framework. As shown in Figure 2B, for the acquired binary bitstream, the geometric bitstream and attribute bitstream within the binary bitstream are first decoded independently. During the decoding of the geometric bitstream, arithmetic decoding—reconstructing the octree / reconstructing the prediction tree—reconstructing geometry—inverse coordinate transformation is used to obtain the geometric information of the point cloud. During the decoding of the attribute bitstream, arithmetic decoding—inverse quantization—LOD partitioning / RAHT—inverse color transformation is used to obtain the attribute information of the point cloud. Based on the geometric and attribute information, the point cloud data to be encoded (i.e., the output point cloud) is reconstructed.
[0048] It should be noted that, as shown in Figure 2A or Figure 2B, the current geometric coding and decoding of G-PCC can be divided into octree-based geometric coding and decoding (marked with dashed boxes) and prediction tree-based geometric coding and decoding (marked with dotted-dash boxes).
[0049] For octree-based geometry encoding (OctGeomEnc), the process involves: first, performing coordinate transformation on the geometric information to ensure that all points in the point cloud are contained within a single bounding box; then, quantization is performed, and due to rounding, some points may have identical geometric information. Whether to remove duplicate points is determined based on parameters; this process of quantization and removal of duplicate points is also known as voxelization. Next, the bounding boxes are continuously partitioned into tree types (e.g., octree, quadtree, binary tree) using a breadth-first search, and the placeholder code for each node is encoded. In related technologies, an implicit geometric partitioning method has been proposed, which first calculates the bounding box of the point cloud. Assume d x >d y >d z The bounding box corresponds to a cuboid. During geometric partitioning, a binary tree partition is first performed based on the x-axis, resulting in two child nodes; this continues until d is satisfied. x =d y >d z Only when the condition is met will the quadtree be continuously partitioned based on the x and y axes, resulting in four child nodes; when d is finally satisfied... x =d y =d z Under certain conditions, the octree partitioning will continue until the resulting leaf nodes form a 1×1×1 unit cube. The partitioning then stops, and the nodes in the leaf nodes are encoded to generate a binary code stream. In the binary / quadtree / octree partitioning process, two parameters are introduced: K and M. Parameter K indicates the maximum number of binary / quadtree partitions performed before octree partitioning; parameter M indicates that the minimum block side length during binary / quadtree partitioning is 2. MAt the same time, K and M must satisfy the following condition: Assume d max =max(d x ,d y ,d z ), d min =min(d x ,d y ,d z The parameter K satisfies: K ≥ d max ―d min The parameter M satisfies: M ≥ d min The reason why parameters K and M satisfy the above conditions is that in the current G-PCC implicit partitioning process, the priority of partitioning methods is binary tree, quadtree, and octree. Only when the node block size does not meet the binary / quadtree condition will the node be continuously partitioned into an octree until the smallest leaf node unit of 1×1×1 is reached. The octree-based geometric information encoding mode can effectively encode the geometric information of the point cloud by utilizing the correlation between nearest neighbors in space.
[0050] The following section uses the RAHT transform as an example to provide a detailed explanation of the encoding and decoding process of point cloud attribute information.
[0051] The principle of RAHT transform is Haar wavelet transform. Its core is to recursively transform attribute information from the root node to the child nodes in a top-down order for a hierarchical tree structure. The resulting direct current (DC) coefficients are passed to the next layer, while the alternating current (AC) coefficients are quantized and encoded. Figure 3A shows a schematic diagram of the attribute information encoding process based on RAHT transform. As shown in Figure 3, at the encoding end, firstly, based on the octree structure of the point cloud, it is determined whether to use upsampling prediction (or intra-frame prediction) or inter-frame prediction; secondly, recursively transforming from the root node to the child nodes from top to bottom yields layer-by-layer AC coefficients. If prediction occurs during the transformation process, the difference between the predicted AC coefficients and the original AC coefficients is used to obtain the residual value of the AC coefficients; finally, the residual value of the AC coefficients is quantized and encoded to generate the attribute information bitstream.
[0052] Figure 3B illustrates a schematic diagram of the attribute information decoding process based on RAHT transform. As shown in Figure 3B, at the decoding end, the attribute information decoding process based on RAHT transform is still performed in a top-down order from the root node to the child nodes. First, the decoder reads the attribute bitstream and performs entropy decoding, followed by inverse quantization to obtain the residual values of the AC coefficients. If there is intra-frame prediction or inter-frame prediction during the transformation process, the residual values of the AC coefficients are added to the predicted values of the AC coefficients to obtain the reconstructed values of the AC coefficients. Then, it is combined with the DC coefficients and subjected to inverse RAHT transform to obtain the reconstructed attribute values of this layer. The reconstructed attribute values of this layer can then be used to calculate the DC coefficients of the next layer, until all layers have been traversed to obtain the reconstructed attribute values. To aid further understanding, the concepts involved in the attribute information encoding process based on RAHT transform described above will be explained in detail below.
[0053] RAHT Transformation
[0054] The RAHT transform relies on a pre-partitioned octree structure, where each non-empty voxel block (hereinafter referred to as a transform block) contains 2×2×2 sub-blocks (hereinafter referred to as sub-blocks). Within each transform block, RAHT applies the Haar wavelet transform in the X, Y, and Z directions, respectively. The specific Haar wavelet transform formulas for two adjacent sub-blocks are shown below:
[0055] Among them, w i,j Let g be the weight of the j-th sub-block (or node to be transformed) in the i-th layer. i,j h is the attribute value of the leaf node. i,j Let AC be the coefficient of the alternation, and g′ be the coefficient of the alternation. i,j The DC coefficient is represented by .
[0056] The Haar wavelet transform of two adjacent sub-blocks can also be called a two-point transform. Figures 4 and 5 show a schematic diagram of a two-point transform process. As shown in Figure 4, after completing the two-point transform in one transform direction, AC coefficients and DC coefficients are obtained. As shown in Figure 5, the DC coefficients obtained in the first transform direction continue to be propagated along the remaining two directions (the first transform direction and the second transform direction) to continue the two-point transform until the entire transform block containing 2×2×2 sub-blocks has been traversed. Therefore, for a transform block containing 8 sub-blocks, 1 DC coefficient and 7 AC coefficients can be obtained.
[0057] For example, the two-point transformations in multiple directions described above can be simplified to the following formula:
[0058] Among them, A N ω represents the sum of attribute information within a sub-block. N T(ω1,ω2,…,ω) represents the number of points within the transformed sub-block.N ) represents the RAHT transformation matrix of the transform block.
[0059] It should be understood that the DC coefficients and AC coefficients represent the transform coefficients obtained after the transformation. During the encoding process, the N-1 AC coefficients obtained from each transformation are encoded, while only the DC coefficients of the root node are encoded; the DC coefficients of the remaining transform blocks are ignored. This is because these ignored DC coefficients can be calculated from the reconstruction attributes at the decoding end. The relationship between the DC coefficients and the reconstruction attribute values is shown in formula (1-5):
[0060] Among them, A N ω represents the sum of attribute information within a sub-block. N This indicates the number of points within the transformed sub-block.
[0061] In the current G-PCC encoding / decoding framework, the transformation process starts from the root node and is performed in a top-down order. Except for the bottom leaf nodes, each node in a layer can be considered a transform block, containing 2×2×2 sub-blocks. Each sub-block in the current layer corresponds to a transform block in the next layer, and the DC coefficients are passed down and decomposed layer by layer from the root node. Transform blocks within the same layer are transformed in ascending order of their coordinates in the Morton code, until all sub-blocks have been traversed.
[0062] Intra-frame upsampling prediction
[0063] To further remove redundancy and improve the compression efficiency of attribute information, an upsampling prediction method is introduced. The essence of upsampling prediction is intra-frame prediction, which uses the attribute information of already encoded points to predict the attribute information of the current point, thereby obtaining the predicted values of the AC coefficients. Then, the predicted values of the AC coefficients are subtracted from the original values to obtain the residual values of the AC coefficients. Next, the residual values of the AC coefficients are encoded to achieve the purpose of redundancy removal. The following sections, in conjunction with Figures 6 and 7, provide a more detailed description of the RAHT transform and inverse RAHT transform processes involving upsampling prediction.
[0064] Figure 6 illustrates a flowchart of a RAHT transform incorporating upsampling prediction. As shown in Figure 6, the RAHT transform in the tree structure proceeds from top to bottom, and the transform is performed within a 2×2×2 block. The reconstructed attribute values of the parent node and its coplanar and collinear neighboring nodes are used to predict the predicted attribute values of these child nodes. Then, the original and predicted attribute values of these child nodes are subjected to RAHT transform to obtain the corresponding DC and AC coefficients. Next, the AC coefficients obtained based on the original attribute values are differiated from those obtained based on the predicted attribute values to obtain the residual values of the AC coefficients, which are then quantized and entropy-encoded. Figure 7 illustrates a flowchart of an inverse RAHT transform incorporating upsampling prediction. As shown in Figure 7, during the inverse RAHT transform, the DC coefficients inherit the reconstructed attribute values of the parent node at the previous level. The residual values of the AC coefficients obtained at the decoding end are accumulated with the AC coefficients obtained by transforming the predicted attribute values from the upsampling prediction, and then combined with the inherited DC coefficients to undergo an inverse RAHT transform to obtain the reconstructed attribute values.
[0065] Inter-frame prediction
[0066] Inter-frame prediction utilizes the attribute information of already encoded point cloud frames to predict and encode the attribute information of the current frame, thereby reducing redundancy. Figure 8 illustrates a schematic diagram of an encoding process based on RAHT inter-frame prediction. As shown in Figure 8, the reference frame is first divided into an octree structure. After the reference frame undergoes RAHT transformation, the AC coefficients of each reference node are recorded for inter-frame prediction of the point cloud attribute information in the next frame. In the encoding of the target node's attribute information, the residual value of the AC coefficients is calculated using the AC coefficients of the reference nodes and the target node, and this residual value is then encoded. The decoding process is similar to the encoding process and will not be described in detail here.
[0067] It should be noted that inter-frame prediction will only be applied to the child nodes of the current frame and the reference frame if they have the same octree partitioning structure and the nodes are at the same position.
[0068] Coding coefficient residuals and residuals
[0069] After obtaining the predicted attribute values of the current 2×2×2 transform block through upsampling prediction or inter-frame prediction as described above, RAHT transformation is performed on the original attribute values and predicted attribute values of the sub-blocks within the transform block to obtain the corresponding DC coefficients and AC coefficients. For the obtained i-th predicted AC coefficient... (k is 8). Let (AC) i ) i∈1…k―1 For the i-th original AC coefficient, the residual value (r) of the AC coefficient is... i ) i∈0…k―1It can be represented as:
[0070] Then, the residual values of the AC coefficients are evaluated based on rate-distortion optimization (RDO): if this r... i If the rate-distortion cost RDcost set to 0 is less than the original rate-distortion cost RDcost, then r is set to 0. i Set to 0.
[0071] Further quantification of the prediction residuals based on RDO judgment:
[0072] Where, r i Q is the residual value of the AC coefficient. i This represents the quantized attribute residual value at the current point i. Qs is the quantization step size, which can be calculated from the quantization parameter (QP) specified by CTC.
[0073] Reconstructing attribute values at the encoding end
[0074] The purpose of encoding-side reconstruction is for predicting subsequent points. Before reconstructing attribute values, the residual values of the quantized AC coefficients need to be dequantized.
[0075] in, Q represents the residual value of the AC coefficients after dequantization. i Qs represents the quantized attribute residual value at the current point i, and Qs is the quantization step size.
[0076] With the predicted AC coefficient The sums are used to obtain the i-th reconstruction AC coefficient within the current block. Right now:
[0077] Finally, the reconstructed AC coefficients were analyzed. The reconstructed attribute values can be obtained by performing an inverse RAHT transformation together with the DC coefficients inherited from the parent node at the previous level.
[0078] As mentioned earlier, in the RAHT inter-frame prediction process, for a given target node, the condition for determining whether to perform RAHT inter-frame prediction on the coefficients of its child nodes is that the position of the target node in the octree structure of the current frame is the same as the position of the reference target node in the octree structure of the reference frame. Only when this condition is met is the RAHT inter-frame prediction coding method used to encode its child nodes. Figure 9 shows a schematic diagram of the RAHT inter-frame prediction coding process (including the judgment condition). As shown in Figure 9, if the target node can find a reference target node with the same position in the reference frame, then the target node's child nodes can undergo RAHT inter-frame prediction; otherwise, if the target node cannot find a reference target node with the same position in the reference frame, then the target node's child nodes will not undergo RAHT inter-frame prediction.
[0079] Figure 10 illustrates a schematic diagram of the encoding process for RAHT inter-frame prediction (including a judgment condition that considers occupancy code information). As shown in Figure 10, in a related technique, an additional judgment condition is proposed to be added to the existing judgment conditions for RAHT inter-frame prediction: the difference between the occupancy code of the target node's child nodes in the current frame and the occupancy code of the reference target node's child nodes in the reference frame is less than or equal to k, where k is 1. RAHT inter-frame prediction is only used when this condition and the existing judgment conditions are both met.
[0080] After adding this condition, the conditions for using RAHT inter-frame prediction on the child nodes of the target node become stricter, and this improvement improves the encoding and decoding performance. This proves that under the G-PCC encoding and decoding framework, it is not perfect to judge whether to perform RAHT inter-frame prediction solely based on whether the geometric position information of the nodes corresponds, which affects the encoding and decoding performance of attribute information.
[0081] To address the aforementioned issues, this application provides a point cloud encoding method, comprising: determining a first reference node of the current node in the current frame; and if the first reference node satisfies a first condition, performing RAHT inter-frame prediction on the attribute values of the current node, wherein the first condition is related to the difference between the reconstructed attribute values of the first reference node and the reconstructed attribute values of a second reference node in the reference frame.
[0082] This application embodiment also provides a point cloud decoding method, including: parsing the bitstream to determine a first reference node of the current node in the current frame; if the first reference node satisfies a first condition, then performing RAHT inter-frame prediction on the attribute value of the current node, wherein the first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of a second reference node in the reference frame.
[0083] During RAHT inter-frame prediction, the more similar the attribute values of the current node are to those of a reference node in the reference frame, the better the prediction performance. In related technologies, when determining whether the attribute values of a target node's child nodes are suitable for RAHT inter-frame prediction, only the presence of a reference target node with the same geometric information as the target node in the reference frame is considered, without taking into account the differences between the target node's attribute values and those of its reference target node. In this application, the difference between the target node's attribute values and those of the reference target node is incorporated into the determination condition (i.e., the first condition) when determining whether the attribute values of a child node (i.e., the current node in this application) are suitable for RAHT inter-frame prediction. This helps improve the encoding and decoding performance of attribute values.
[0084] The point cloud decoding method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0085] Figure 11 is a flowchart illustrating the point cloud decoding method provided in an embodiment of this application. The decoding method in Figure 11 can be applied to a decoder. The decoding method in Figure 11 can be used to decode the attribute information of a point cloud. In some implementations, this decoding method can be applied to G-PCC. Alternatively, in other implementations, this decoding method can be applied to a geometry-based solid content test model (GES-TM). GES-TM is an encoding and decoding framework proposed for dense point clouds (such as point clouds acquired in augmented reality (AR) or virtual reality (VR) scenes).
[0086] Referring to Figure 11, in step S1110, the bitstream is parsed to determine the first reference node of the current node in the current frame.
[0087] It is understandable that the nodes mentioned in step S1110 can be understood as nodes in an octree structure. During the RAHT transformation of attribute values, a node can also be called a transformation block, and its child nodes can be called child blocks. Each child block corresponds to a transformation block in the next layer. The attribute values corresponding to the transformation blocks can correspond to the attribute values of points in the point cloud.
[0088] The first reference node mentioned above can be of several types. For example, the first reference node can be the parent node of the current node, or the current node can be a child node of the first reference node. Another example is that the first reference node can be a neighbor node of the current node.
[0089] In step S1120, if the first reference node satisfies the first condition, then RAHT inter-frame prediction is performed on the attribute values of the current node. Here, the first condition is related to the difference between the reconstructed attribute values of the first reference node and the reconstructed attribute values of the second reference node in the reference frame.
[0090] In some implementations, the decoding method shown in Figure 11 may also include: if the first reference node does not meet the first condition, then no RAHT inter-frame prediction is performed on the attribute value of the current node.
[0091] The attribute values mentioned above can include various types. For example, the attribute value can be color information, such as luminance components or color components. Another example is reflectance information.
[0092] The geometric information of the first reference node mentioned above is the same as that of the second reference node. This geometric information can be coordinate information and / or hierarchical information. For example, the hierarchical position of the first reference node in the octree structure of the current frame, as well as its 3D coordinates, are the same as the hierarchical position and 3D coordinates of the second reference node in the octree structure of the reference frame. It should be noted that if no reference node with the same geometric information as the first reference node can be found in the reference frame, then the attribute values of the current node will not be used for RAHT inter-frame prediction.
[0093] During RAHT inter-frame prediction, the more similar the attribute values of the current node are to those of a reference node in the reference frame, the better the prediction performance. In related technologies, when determining whether the attribute values of a target node's child nodes are suitable for RAHT inter-frame prediction, only the presence of a reference target node with the same geometric information as the target node in the reference frame is considered, without taking into account the differences between the target node's attribute values and those of the reference target node. This application's embodiment incorporates the difference between the target node's attribute values and those of the reference target node into the determination condition (i.e., the first condition) when determining whether the attribute values of a child node are suitable for RAHT inter-frame prediction, thereby helping to improve the encoding and decoding performance of attribute values.
[0094] For example, in a RAHT-based encoding and decoding scheme, when decoding the attribute value of the current node, the attribute value of the parent node of the current node has already been reconstructed. Therefore, the difference between the attribute value of the current node and the attribute value of the reference node can be derived by the difference between the reconstructed attribute value of the parent node and the reconstructed attribute value of the reference node.
[0095] For example, the first reference node in the first condition above can be the parent node of the current node, and the second reference node can be the reference node of the parent node in the reference frame.
[0096] For example, in a RAHT-based encoding / decoding scheme, when decoding the attribute value of the current node, the attribute values of the current node's neighboring nodes have already been reconstructed, and the attribute values of the current node and its neighboring nodes in the same layer are likely similar. Therefore, the difference between the attribute values of the current node and the attribute values of its reference node can be derived by the difference between the reconstructed attribute values of the neighboring nodes and the reconstructed attribute values of their reference nodes.
[0097] For example, the first reference node in the first condition above can be a neighboring node of the current node, and the second reference node can be a reference node of that neighboring node in the reference frame. It should be understood that the neighboring node here can be a single node or multiple nodes.
[0098] For example, in a RAHT-based encoding / decoding scheme, if the current node includes multiple attribute values (such as a first attribute value and a second attribute value), then some attribute values will be decoded and reconstructed first, while others will be decoded and reconstructed later. Taking the first attribute value as the first to be decoded and reconstructed first, and the second attribute value as the second, the difference between the second attribute value of the current node and the second attribute value of the reference node can be derived from the difference between the first reconstructed attribute value of the current node and the first reconstructed attribute value of the reference node.
[0099] Alternatively, the parent node of the current node (such as the first reference node) includes at least one child node, and the difference between the second attribute value of the current node and the second attribute value of its reference node can be derived by the difference between the first reconstruction attribute value of the at least one child node and the first reconstruction attribute value of the parent node's reference node (such as the second parameter node).
[0100] For example, the first attribute value can be color information, and the second attribute value can be reflectance information.
[0101] Furthermore, the decoding method shown in Figure 11 can also set a preset threshold (i.e., the first threshold) in the first condition to measure the difference between the reconstructed attribute values of the first reference node and the reconstructed attribute values of the second reference node.
[0102] The value of the first threshold can be related to the reconstructed attribute value of the first reference node. That is, in a RAHT-based encoding / decoding scheme, the value of the first threshold can also change as the reconstructed attribute value of the first reference node changes. Since the attribute values of nodes in the upper layers of the octree are larger, and the differences in attribute values between different nodes are also larger, while the attribute values of nodes in the lower layers of the octree are smaller, and the differences in attribute values between different nodes are also smaller, setting a fixed threshold for the attribute value difference is not universally applicable to all layers. Therefore, the value of the first threshold is related to the reconstructed attribute value of the first reference node. For a first reference node with a larger reconstructed attribute value, the first threshold can be set larger; for a first reference node with a smaller reconstructed attribute value, the first threshold can be set smaller. This allows for a more accurate measurement of the difference between the reconstructed attribute values of the first and second reference nodes. Alternatively, different thresholds can be set for each layer of the octree in the current frame for more refined judgment.
[0103] For example, the first threshold can be determined based on the quotient of the reconstructed attribute value of the first reference node and a first parameter. Here, the first parameter can be a target preset value determined from multiple preset values.
[0104] For example, the first condition can be expressed as: Luma cur ≥λ×|Luma ref ―Luma cur | (1-10)
[0105] Alternatively, the first condition can also be expressed as: |Luma ref ―Luma cur |≤Luma cur / λ (1-11)
[0106] Among them, Luma cur Luma represents the reconstructed attribute value (such as the luminance component) of the first reference node. ref Luma represents the reconstruction attribute value of the second reference node. cur / λ represents the first threshold. For example, λ (i.e., the first parameter) can be 20.
[0107] The preceding text detailed how to determine whether to perform RAHT inter-frame prediction on the attribute values of the current node using the first condition. Further, embodiments of this application can further limit the scope of RAHT inter-frame prediction by setting other conditions (such as a second condition). For example, for multi-frame dynamic point clouds, in the upper layers of the RAHT transform closer to the root node, nodes with larger attribute values show less change between adjacent frames. Therefore, the inter-frame prediction results are relatively accurate in the upper layers closer to the root node; conversely, the inter-frame prediction results are less accurate in the lower layers farther from the root node. Therefore, in the process of determining whether to perform inter-frame prediction on the attribute values of the current node, limiting inter-frame prediction to the lower layers farther from the root node can improve the overall accuracy of inter-frame prediction for the point cloud. It should be understood that the last layer farther from the root node can be called layer 0, and the layer number increases from bottom to top along the octree.
[0108] In some implementations, the decoding method shown in Figure 11 can also set a second condition, which is related to the level information of the first reference node in the octree. For example, the second condition could be that the level of the first reference node in the octree is less than or equal to a second threshold. Exemplarily, the second threshold could be 1.
[0109] If the first and second conditions are combined, the decoding method shown in Figure 11 can further include: if the first reference node satisfies the second condition, then determine whether the first reference node satisfies the first condition; then, if the first reference node satisfies the first condition, perform RAHT inter-frame prediction on the attribute value of the current node. In other words, when the first reference node satisfies the second condition, i.e., when the first reference node is at a lower level in the octree, as mentioned earlier, the inter-frame prediction result is not accurate enough, and only then will the determination of whether the first reference node satisfies the first condition continue.
[0110] When the first reference node does not meet the second condition, that is, when the first reference node is at the upper level in the octree, as mentioned earlier, the inter-frame prediction result is relatively accurate. Therefore, RAHT inter-frame prediction can be performed directly on the attribute value of the current node.
[0111] For ease of understanding, the judgment conditions for RAHT inter-frame prediction provided in this application embodiment will be described in detail below with reference to Figure 12. The method shown in Figure 12 includes steps S1210 to S1280.
[0112] In step S1210, the RAHT inter-frame prediction judgment begins.
[0113] In step S1220, traverse each target node of each layer in the current frame.
[0114] In step S1230, it is determined whether the target node can find a reference target node on the reference frame. If the target node can find a reference target node on the reference frame, proceed to step S1240; if the target node cannot find a reference target node on the reference frame, proceed to step S1270.
[0115] In step S1240, it is determined whether the current level satisfies level≤level_constraint. This step corresponds to the second condition described above, which can be found in the previous text for details. If the current level satisfies level≤level_constraint, proceed to step S1250; if the current level does not satisfy level≤level_constraint, proceed to step S1260.
[0116] In step S1250, it is determined whether the target node satisfies Luma. cur ≥λ×|Luma ref ―Luma cur This step corresponds to the first condition described earlier; please refer to the previous text for details. If the target node satisfies Luma... cur ≥λ×|Luma ref ―Luma cur If the target node does not satisfy Luma, then proceed to step S1260; cur ≥λ×|Luma ref ―Luma cur If so, proceed to step S1270.
[0117] In step S1260, RAHT inter-frame prediction is performed on the attribute values of the child nodes of this target node.
[0118] In step S1270, the attribute values of the child nodes of this target node are not subject to RAHT inter-frame prediction.
[0119] In step S1280, the RAHT inter-frame prediction judgment ends.
[0120] For example, the specific implementation of the program can be as follows:
[0121] In some implementations, the method of performing RAHT inter-frame prediction on the attribute values of the current node may include: determining the reconstructed value of the first AC coefficient of the first child node of the second reference node; and then, based on the reconstructed value of the first AC coefficient, determining the predicted value of the second AC coefficient of the current node.
[0122] For example, determining the predicted value of the second AC coefficient may include using the reconstructed value of the first AC coefficient as the predicted value of the second AC coefficient.
[0123] In some implementations, the decoding method shown in Figure 11 may further include: parsing the bitstream to determine the residual value of the second AC coefficient; and then, determining the reconstructed value of the second AC coefficient based on the residual value of the second AC coefficient and the predicted value of the second AC coefficient.
[0124] For example, determining the reconstructed value of the second AC coefficient may include summing the residual value of the second AC coefficient and the predicted value of the second AC coefficient as the reconstructed value of the second AC coefficient.
[0125] In some implementations, the decoding method shown in Figure 11 may further include: performing an inverse RAHT transform based on the reconstructed values of the second AC coefficients and the DC coefficients to determine the reconstructed attribute values of the current node. Here, the DC coefficients can be obtained from the parent node of the current node (such as the first reference node).
[0126] It should be understood that the AC coefficient can also be called the attribute transformation coefficient, alternating current coefficient, high-frequency coefficient, AC high-frequency coefficient, or high-pass coefficient. The DC coefficient can also be called the direct current coefficient, low-frequency coefficient, DC low-frequency coefficient, or low-pass coefficient.
[0127] The test results obtained from testing the encoding and decoding method provided in the embodiments of this application will be introduced below to verify the performance improvement brought about by the embodiments of this application.
[0128] Table 1: Test results obtained from the encoding / decoding method based on the embodiments of this application under C1 conditions.
[0129] Table 2: Test results obtained by the encoding / decoding method based on the embodiments of this application under C2 conditions
[0130] Tables 1 and 2 present the test results obtained by testing the encoding / decoding method provided in the embodiments of this application on the GPCC reference software TMC13-v27-rc2. In this test, the multi-frame dynamic point cloud sequence required by MPEG was tested under C1 and C2 test conditions (Cat2). During the test, λ (i.e., the first parameter mentioned above) was set to 20, and level_constraint (i.e., the second threshold mentioned above) was set to 1.
[0131] Test condition C1 uses lossless geometry and lossy attribute encoding, while test condition C2 uses lossy geometry and lossy attribute encoding. In the table, End-to-End BD-AttrRate represents the end-to-end BD-Rate of attribute values relative to the attribute bitstream, and End-to-End BD-TotalRate represents the end-to-end BD-Rate of attribute values relative to the total bitstream. BD-Rate reflects the reduction in bitrate at average PSNR compared to the original method. A decrease in BD-Rate indicates a reduction in bitrate and improved performance while maintaining the same PSNR; conversely, a increase indicates a decrease in performance. In other words, a greater decrease in BD-Rate indicates better compression performance. Cat2-Aaverage, Cat2-B average, and Cat2-C average represent the average test results of point cloud sequences for each of the three datasets in Cat2. Finally, the Overall average is the average test result for all sequences.
[0132] In this embodiment of the application, tests were performed on the Luma, Chroma Cb, and Chroma Cr components of the Cat2 dataset under C1 test conditions, and average gains of -0.3%, -0.4%, and -0.4% were obtained, respectively. Tests were also performed on the Luma, Chroma Cb, and Chroma Cr components of the Cat2 dataset under CTC-C2 test conditions, and average gains of -0.5%, -0.5%, and -0.6% were obtained, respectively.
[0133] The point cloud decoding method provided by the embodiments of this application has been described in detail above with reference to Figure 11. The point cloud encoding method provided by the embodiments of this application will be described in detail below with reference to Figure 13.
[0134] Figure 13 is a flowchart illustrating the point cloud encoding method provided in an embodiment of this application. The encoding method in Figure 13 can be applied to an encoder. The encoding method in Figure 13 can be used to encode the attribute information of a point cloud. In some implementations, this encoding method can be applied to G-PCC. Alternatively, in other implementations, this encoding method can be applied to a geometry-based solid content test model (GES-TM). GES-TM is an encoding and decoding framework proposed for dense point clouds (such as point clouds acquired in augmented reality (AR) or virtual reality (VR) scenes).
[0135] Referring to Figure 13, in step S1310, the first reference node of the current node in the current frame is determined.
[0136] It is understandable that the nodes mentioned in step S1310 can be understood as nodes in an octree structure. During the RAHT transformation of attribute values, a node can also be called a transformation block, and its child nodes can be called child blocks. Each child block corresponds to a transformation block in the next layer. The attribute values corresponding to the transformation blocks can correspond to the attribute values of points in the point cloud.
[0137] The first reference node mentioned above can be of several types. For example, the first reference node can be the parent node of the current node, or the current node can be a child node of the first reference node. Another example is that the first reference node can be a neighbor node of the current node.
[0138] In step S1320, if the first reference node satisfies the first condition, then RAHT inter-frame prediction is performed on the attribute values of the current node. Here, the first condition is related to the difference between the reconstructed attribute values of the first reference node and the reconstructed attribute values of the second reference node in the reference frame.
[0139] In some implementations, the encoding method shown in Figure 13 may also include: if the first reference node does not meet the first condition, then no RAHT inter-frame prediction is performed on the attribute value of the current node.
[0140] The attribute values mentioned above can include various types. For example, the attribute value can be color information, such as luminance components or color components. Another example is reflectance information.
[0141] The geometric information of the first reference node mentioned above is the same as that of the second reference node. This geometric information can be coordinate information and / or hierarchical information. For example, the hierarchical position of the first reference node in the octree structure of the current frame, as well as its 3D coordinates, are the same as the hierarchical position and 3D coordinates of the second reference node in the octree structure of the reference frame. It should be noted that if no reference node with the same geometric information as the first reference node can be found in the reference frame, then the attribute values of the current node will not be used for RAHT inter-frame prediction.
[0142] During RAHT inter-frame prediction, the more similar the attribute values of the current node are to those of a reference node in the reference frame, the better the prediction performance. In related technologies, when determining whether the attribute values of a target node's child nodes are suitable for RAHT inter-frame prediction, only the presence of a reference target node with the same geometric information as the target node in the reference frame is considered, without taking into account the differences between the target node's attribute values and those of its reference target node. In this application, the difference between the target node's attribute values and those of the reference target node is incorporated into the determination condition (i.e., the first condition) when determining whether the attribute values of a child node (i.e., the current node in this application) are suitable for RAHT inter-frame prediction. This helps improve the encoding and decoding performance of attribute values.
[0143] For example, in a RAHT-based encoding and decoding scheme, when encoding the attribute value of the current node, the attribute value of the parent node of the current node has already been reconstructed. Therefore, the difference between the attribute value of the current node and the attribute value of the reference node can be derived by the difference between the reconstructed attribute value of the parent node and the reconstructed attribute value of the reference node.
[0144] For example, the first reference node in the first condition above can be the parent node of the current node, and the second reference node can be the reference node of the parent node in the reference frame.
[0145] For example, in a RAHT-based encoding / decoding scheme, when encoding the attribute value of the current node, the attribute values of the current node's neighboring nodes have already been reconstructed, and the attribute values of the current node and its neighboring nodes in the same layer are likely to be similar. Therefore, the difference between the attribute values of the current node and the attribute values of its reference node can be derived by the difference between the reconstructed attribute values of the neighboring nodes and the reconstructed attribute values of their reference nodes.
[0146] For example, the first reference node in the first condition above can be a neighboring node of the current node, and the second reference node can be a reference node of that neighboring node in the reference frame. It should be understood that the neighboring node here can be a single node or multiple nodes.
[0147] For example, in a RAHT-based encoding / decoding scheme, if the current node includes multiple attribute values (such as a first attribute value and a second attribute value), some attribute values will be encoded and reconstructed first, while others will be encoded and reconstructed later. Taking the first attribute value being encoded and reconstructed first, and the second attribute value being encoded and reconstructed later as an example, the difference between the second attribute value of the current node and the second attribute value of the reference node can be derived by the difference between the first reconstructed attribute value of the current node and the first reconstructed attribute value of the reference node.
[0148] Alternatively, the parent node of the current node (such as the first reference node) includes at least one child node, and the difference between the second attribute value of the current node and the second attribute value of its reference node can be derived by the difference between the first reconstruction attribute value of the at least one child node and the first reconstruction attribute value of the parent node's reference node (such as the second parameter node).
[0149] For example, the first attribute value can be color information, and the second attribute value can be reflectance information.
[0150] Furthermore, the encoding method shown in Figure 13 can also set a preset threshold (i.e., the first threshold) in the first condition to measure the difference between the reconstruction attribute values of the first reference node and the reconstruction attribute values of the second reference node.
[0151] The value of the first threshold can be related to the reconstructed attribute value of the first reference node. That is, in a RAHT-based encoding / decoding scheme, the value of the first threshold can also change as the reconstructed attribute value of the first reference node changes. Since the attribute values of nodes in the upper layers of the octree are larger, and the differences in attribute values between different nodes are also larger, while the attribute values of nodes in the lower layers of the octree are smaller, and the differences in attribute values between different nodes are also smaller, simply setting a fixed threshold for the attribute value difference is not universally applicable to all layers. Therefore, the value of the first threshold is related to the reconstructed attribute value of the first reference node. For a first reference node with a larger reconstructed attribute value, the first threshold can be set larger; for a first reference node with a smaller reconstructed attribute value, the first threshold can be set smaller. This allows for a more accurate measurement of the difference between the reconstructed attribute values of the first and second reference nodes. Alternatively, different thresholds can be set for each layer of the octree in the current frame for more refined judgment.
[0152] For example, the first threshold can be determined based on the quotient of the reconstructed attribute value of the first reference node and a first parameter. Here, the first parameter can be a target preset value determined from multiple preset values.
[0153] For example, the first condition can be expressed as: Luma cur ≥λ×|Luma ref ―Luma cur | (1-12)
[0154] Alternatively, the first condition can also be expressed as: |Luma ref ―Luma cur |≤Luma cur / λ (1-13)
[0155] Among them, Luma curLuma represents the reconstructed attribute value (such as the luminance component) of the first reference node. ref Luma represents the reconstruction attribute value of the second reference node. cur / λ represents the first threshold. For example, λ (i.e., the first parameter) can be 20.
[0156] For example, Figure 14 illustrates a schematic diagram of the encoding process for RAHT inter-frame prediction (including a first condition) provided in an embodiment of this application. As shown in Figure 14, based on determining whether the attribute values of the target node's child nodes use the initial condition for RAHT inter-frame prediction (i.e., the position of the target node in the octree structure of the current frame is the same as the position of the reference target node in the octree structure of the reference frame), a first condition is added: if the difference between the reconstructed attribute value of the target node and the reconstructed attribute value of the reference target node at its corresponding position in the reference frame is less than or equal to a first threshold, then RAHT inter-frame prediction is performed on the attribute values of the target node's child nodes. Since the AC coefficients are obtained by recursively changing from the root node to the child nodes from top to bottom during the RAHT encoding process, the reconstructed attribute values of the target node can be indexed at both the encoding and decoding ends.
[0157] The preceding text detailed how to determine whether to perform RAHT inter-frame prediction on the attribute values of the current node using the first condition. Further, embodiments of this application can further limit the scope of RAHT inter-frame prediction by setting other conditions (such as a second condition). For example, for multi-frame dynamic point clouds, in the upper layers of the RAHT transform closer to the root node, nodes with larger attribute values show less change between adjacent frames. Therefore, the inter-frame prediction results are relatively accurate in the upper layers closer to the root node; conversely, the inter-frame prediction results are less accurate in the lower layers farther from the root node. Therefore, in the process of determining whether to perform inter-frame prediction on the attribute values of the current node, limiting inter-frame prediction to the lower layers farther from the root node can improve the overall accuracy of inter-frame prediction for the point cloud. It should be understood that the last layer farther from the root node can be called layer 0, and the layer number increases from bottom to top along the octree.
[0158] Therefore, in some implementations, the encoding method shown in Figure 13 can also set a second condition, which is related to the level information of the first reference node in the octree. For example, the second condition could be that the level of the first reference node in the octree is less than or equal to a second threshold. Exemplarily, the second threshold could be 1.
[0159] If the first and second conditions are combined, the encoding method shown in Figure 13 can further include: if the first reference node satisfies the second condition, then determine whether the first reference node satisfies the first condition; then, if the first reference node satisfies the first condition, perform RAHT inter-frame prediction on the attribute value of the current node. In other words, when the first reference node satisfies the second condition, i.e., when the first reference node is at a lower level in the octree, as mentioned earlier, the inter-frame prediction result is not accurate enough, and only then will the determination of whether the first reference node satisfies the first condition continue.
[0160] When the first reference node does not meet the second condition, that is, when the first reference node is at the upper level in the octree, as mentioned earlier, the inter-frame prediction result is relatively accurate. Therefore, RAHT inter-frame prediction can be performed on the attribute value of the current node.
[0161] For ease of understanding, the judgment conditions for RAHT inter-frame prediction provided in this application embodiment will be described in detail below with reference to Figure 12. The method shown in Figure 12 includes steps S1210 to S1280.
[0162] In step S1210, the RAHT inter-frame prediction judgment begins.
[0163] In step S1220, traverse each target node of each layer in the current frame.
[0164] In step S1230, it is determined whether the target node can find a reference target node on the reference frame. If the target node can find a reference target node on the reference frame, proceed to step S1240; if the target node cannot find a reference target node on the reference frame, proceed to step S1270.
[0165] In step S1240, it is determined whether the current level satisfies level≤level_constraint. This step corresponds to the second condition described above, which can be found in the previous text for details. If the current level satisfies level≤level_constraint, proceed to step S1250; if the current level does not satisfy level≤level_constraint, proceed to step S1260.
[0166] In step S1250, it is determined whether the target node satisfies Luma. cur ≥λ×(Luma ref ―Luma cur This step corresponds to the first condition described earlier; please refer to the previous text for details. If the target node satisfies Luma... cur ≥λ×(Luma ref ―Luma curIf the target node does not satisfy Luma, then proceed to step S1260; cur ≥λ×(Luma ref ―Luma cur If ), then proceed to step S1270.
[0167] In step S1260, RAHT inter-frame prediction is performed on the attribute values of the child nodes of this target node.
[0168] In step S1270, the attribute values of the child nodes of this target node are not subject to RAHT inter-frame prediction.
[0169] In step S1280, the RAHT inter-frame prediction judgment ends.
[0170] For example, the specific implementation of the program is as follows:
[0171] In some implementations, the method of performing RAHT inter-frame prediction on the attribute values of the current node may include: performing RAHT transformation on the attribute values of the first child node of the second reference node to determine the first AC coefficient; then, performing RAHT transformation on the attribute values of the current node to determine the second AC coefficient; next, determining the predicted value of the second AC coefficient based on the reconstructed value of the first AC coefficient; and finally, determining the residual value of the second AC coefficient based on the original value of the second AC coefficient and the predicted value of the second AC coefficient.
[0172] For example, determining the predicted value of the second AC coefficient may include using the reconstructed value of the first AC coefficient as the predicted value of the second AC coefficient.
[0173] For example, determining the residual value of the second AC coefficient may include using the difference between the original value of the second AC coefficient and the predicted value of the second AC coefficient as the residual value of the second AC coefficient.
[0174] In some implementations, the encoding method shown in Figure 13 may also include writing the residual value of the second AC coefficient into the bitstream.
[0175] In some implementations, the encoding method shown in Figure 13 may further include: determining the reconstructed value of the second AC coefficient based on the predicted value of the second AC coefficient and the residual value of the second AC coefficient; then, performing an inverse RAHT transform based on the reconstructed value of the second AC coefficient and the DC coefficient to determine the reconstructed attribute value of the current node. Here, the DC coefficient can be obtained from the parent node of the current node (such as the first reference node).
[0176] For example, determining the reconstructed value of the second AC coefficient may include summing the predicted value of the second AC coefficient and the residual value of the second AC coefficient as the reconstructed value of the second AC coefficient.
[0177] It should be understood that the AC coefficient can also be called the attribute transformation coefficient, alternating current coefficient, high-frequency coefficient, AC high-frequency coefficient, or high-pass coefficient. The DC coefficient can also be called the direct current coefficient, low-frequency coefficient, DC low-frequency coefficient, or low-pass coefficient.
[0178] The method embodiments of this application have been described in detail above with reference to Figures 1 to 14. The apparatus embodiments of this application will be described in detail below with reference to Figures 15 to 18. It should be understood that the descriptions of the method embodiments correspond to the descriptions of the apparatus embodiments; therefore, any parts not described in detail can be referred to the preceding method embodiments.
[0179] Figure 15 is a schematic diagram of the structure of a decoder provided in an embodiment of this application. As shown in Figure 15, the decoder 1500 may include a determination unit 1510 and a prediction unit 1520.
[0180] Unit 1510 is configured to parse the bitstream and determine the first reference node of the current node in the current frame.
[0181] The prediction unit 1520 is configured to perform RAHT inter-frame prediction on the attribute value of the current node if the first reference node satisfies a first condition, wherein the first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the second reference node in the reference frame.
[0182] In some possible implementations, the relationship between the current node and the first reference node is one of the following: the current node is a child node of the first reference node; or, the first reference node is a neighbor node of the current node.
[0183] In some possible implementations, the attribute value of the current node includes a first attribute value and a second attribute value. The prediction unit 1520 is configured to perform RAHT inter-frame prediction on the second attribute value of the current node if the first reference node satisfies the first condition. The first condition is related to the difference between the first reconstructed attribute value of the first reference node and the first reconstructed attribute value of the second reference node.
[0184] In some possible implementations, the first condition includes: the difference between the reconstruction attribute value of the first reference node and the reconstruction attribute value of the second reference node is less than or equal to a first threshold.
[0185] In some possible implementations, the first reference node includes at least one child node, and the first condition includes: the difference between the first reconstruction attribute value of the at least one child node and the first reconstruction attribute value of the second reference node is less than or equal to the first threshold; wherein, the at least one child node includes the current node.
[0186] In some possible implementations, the value of the first threshold is related to the reconstruction attribute value of the first reference node.
[0187] In some possible implementations, the first threshold is determined based on the quotient of the reconstruction attribute value of the first reference node and the first parameter.
[0188] In some possible implementations, the prediction unit 1520 is configured to determine whether the first reference node satisfies the first condition if the first reference node satisfies the second condition, wherein the second condition is related to the hierarchical information of the first reference node in the octree; and to perform RAHT inter-frame prediction on the attribute value of the current node if the first reference node satisfies the first condition.
[0189] In some possible implementations, the second condition includes: the level of the first reference node in the octree is less than or equal to the second threshold.
[0190] In some possible implementations, the decoder 1500 is further configured to not perform RAHT inter-frame prediction on the attribute value of the current node if the first reference node does not meet the first condition.
[0191] In some possible implementations, the geometric information of the first reference node is the same as that of the second reference node.
[0192] In some possible implementations, the geometric information includes coordinate information and / or hierarchical information.
[0193] In some possible implementations, the prediction unit 1520 is configured to determine the reconstructed value of the first AC coefficient of the first child node of the second reference node; and, based on the reconstructed value of the first AC coefficient, determine the predicted value of the second AC coefficient of the current node.
[0194] In some possible implementations, the decoder 1500 is further configured to parse the bitstream, determine the residual value of the second AC coefficient, and determine the reconstructed value of the second AC coefficient based on the residual value of the second AC coefficient and the predicted value of the second AC coefficient.
[0195] In some possible implementations, the decoder 1500 is further configured to perform an inverse RAHT transform based on the reconstructed values of the second AC coefficients and the DC coefficients to determine the reconstructed attribute values of the current node.
[0196] Understandably, in the embodiments of this application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular one. Furthermore, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional module.
[0197] If the integrated unit is implemented as a software functional module and not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0198] Therefore, this application provides a computer-readable storage medium for use in a decoder 1500. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the decoding method described in any of the foregoing embodiments.
[0199] Based on the composition of the decoder 1500 and the computer-readable storage medium described above, refer to Figure 16, which shows a schematic diagram of the specific hardware structure of the decoder 1500 provided in this embodiment of the application. As shown in Figure 16, the decoder 1600 may include: a communication interface 1610, a memory 1620, and a processor 1630; the various components are coupled together through a bus system 1640. It is understood that the bus system 1640 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 1640 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 1640 in Figure 16.
[0200] The communication interface 1610 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;
[0201] Memory 1620 is used to store computer programs;
[0202] Processor 1630, when running the computer program, performs the following:
[0203] Parse the bitstream to determine the first reference node of the current node in the current frame;
[0204] If the first reference node satisfies the first condition, then RAHT inter-frame prediction is performed on the attribute value of the current node. The first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the second reference node in the reference frame.
[0205] It is understood that the memory 1620 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DRRAM). The memory 1620 of the systems and methods described in this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0206] The processor 1630 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 1630 or by instructions in software form. The processor 1630 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1620. Processor 1630 reads the information in memory 1620 and completes the steps of the above method in conjunction with its hardware.
[0207] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), DSP devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof. For software implementation, the technology described in this application can be implemented through modules (e.g., procedures, functions, etc.) that perform the functions described in this application. Software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or externally.
[0208] Alternatively, as another embodiment, the processor 1630 is also configured to execute the decoding method described in any of the foregoing embodiments when running the computer program.
[0209] Figure 17 is a schematic diagram of the structure of an encoder provided in an embodiment of this application. As shown in Figure 17, the encoder 1700 includes a determination unit 1710 and a prediction unit 1720.
[0210] The determination unit 1710 is configured to determine the first reference node of the current node in the current frame.
[0211] The prediction unit 1720 is configured to perform RAHT inter-frame prediction on the attribute value of the current node if the first reference node satisfies a first condition, wherein the first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the second reference node in the reference frame.
[0212] In some possible implementations, the relationship between the current node and the first reference node is one of the following: the current node is a child node of the first reference node; or, the first reference node is a neighbor node of the current node.
[0213] In some possible implementations, the attribute value of the current node includes a first attribute value and a second attribute value. The prediction unit 1720 is configured to perform RAHT inter-frame prediction on the second attribute value of the current node if the first reference node satisfies the first condition. The first condition is related to the difference between the first reconstructed attribute value of the first reference node and the first reconstructed attribute value of the second reference node.
[0214] In some possible implementations, the first condition includes: the difference between the reconstruction attribute value of the first reference node and the reconstruction attribute value of the second reference node is less than or equal to a first threshold.
[0215] In some possible implementations, the first reference node includes at least one child node, and the first condition includes: the difference between the first reconstruction attribute value of the at least one child node and the first reconstruction attribute value of the second reference node is less than or equal to the first threshold; wherein, the at least one child node includes the current node.
[0216] In some possible implementations, the value of the first threshold is related to the reconstruction attribute value of the first reference node.
[0217] In some possible implementations, the first threshold is determined based on the quotient of the reconstruction attribute value of the first reference node and the first parameter.
[0218] In some possible implementations, the prediction unit 1720 is configured to determine whether the first reference node satisfies the first condition if the first reference node satisfies the second condition, wherein the second condition is related to the hierarchical information of the first reference node in the octree; and if the first reference node satisfies the first condition, perform RAHT inter-frame prediction on the attribute value of the current node.
[0219] In some possible implementations, the second condition includes: the level of the first reference node in the octree is less than or equal to the second threshold.
[0220] In some possible implementations, the encoder 1700 is further configured to not perform RAHT inter-frame prediction on the attribute value of the current node if the first reference node does not meet the first condition.
[0221] In some possible implementations, the geometric information of the first reference node is the same as that of the second reference node.
[0222] In some possible implementations, the geometric information includes coordinate information and / or hierarchical information.
[0223] In some possible implementations, the prediction unit 1720 is configured to perform RAHT transformation on the attribute values of the first child node of the second reference node to determine the first AC coefficient; perform RAHT transformation on the attribute values of the current node to determine the second AC coefficient; determine the predicted value of the second AC coefficient based on the reconstructed value of the first AC coefficient; and determine the residual value of the second AC coefficient based on the original value and the predicted value of the second AC coefficient.
[0224] In some possible implementations, the encoder 1700 is also configured to write the residual value of the second AC coefficient into the bitstream.
[0225] In some possible implementations, the encoder 1700 is further configured to determine the reconstructed value of the second AC coefficient based on the predicted value of the second AC coefficient and the residual value of the second AC coefficient; and to determine the reconstructed attribute value of the current node by performing an inverse RAHT transform based on the reconstructed value of the second AC coefficient and the DC coefficient.
[0226] Understandably, in the embodiments of this application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular one. Furthermore, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional module.
[0227] If the integrated unit is implemented as a software functional module and not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, external hard drives, ROM, RAM, magnetic disks, or optical disks.
[0228] Therefore, this application provides a computer-readable storage medium for use in an encoder 1700. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the encoding method described in any of the foregoing embodiments.
[0229] Based on the composition of the encoder 1700 described above and the computer-readable storage medium, refer to Figure 18, which shows a schematic diagram of the specific hardware structure of the encoder 1700 provided in this embodiment of the application. As shown in Figure 18, the encoder 1700 may include: a communication interface 1810, a memory 1818, and a processor 1830; the various components are coupled together through a bus system 1840. It is understood that the bus system 1840 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 1840 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 1840 in Figure 18.
[0230] The communication interface 1810 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;
[0231] Memory 1818 is used to store computer programs;
[0232] Processor 1830, when running the computer program, performs the following:
[0233] Determine the first reference node of the current node in the current frame;
[0234] If the first reference node satisfies the first condition, then RAHT inter-frame prediction is performed on the attribute value of the current node. The first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the second reference node in the reference frame.
[0235] It is understood that the memory 1818 in the embodiments of this application may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be ROM, PROM, EPROM, EEPROM, or flash memory. Volatile memory may be RAM, which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as SRAM, DRAM, SDRAM, DDRSDRAM, ESDRAM, SLDRAM, and DRRAM. The memory 1818 of the systems and methods described in this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0236] The processor 1830 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the processor 1830 or by software instructions. The processor 1830 may be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 1818, and the processor 1830 reads the information in memory 1818 and, in conjunction with its hardware, completes the steps of the above method.
[0237] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more ASICs, DSPs, DSPDs, PLDs, FPGAs, general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof. For software implementation, the technology described in this application can be implemented through modules (e.g., procedures, functions, etc.) that perform the functions described in this application. Software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0238] Alternatively, as another embodiment, the processor 1830 is also configured to execute the encoding method described in any of the foregoing embodiments when running the computer program.
[0239] This application also provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium for storing bit streams. The bit streams can be generated by using an encoding method of an encoder, or the bit streams can be decoded by using a decoding method of a decoder. The decoding method can be the decoding method described in any of the preceding embodiments, and the encoding method can be the encoding method described in any of the preceding embodiments.
[0240] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0241] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0242] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0243] The features disclosed in the several product embodiments provided in this application are, without conflict, related to the following:
[0244] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0245] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A point cloud decoding method, applied to a decoder, comprising: Parse the bitstream to determine the first reference node of the current node in the current frame; If the first reference node satisfies the first condition, then the attribute value of the current node is subjected to Region Adaptive Hierarchical Transformation (RAHT) inter-frame prediction. The first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the second reference node in the reference frame.
2. The method according to claim 1, wherein, The relationship between the current node and the first reference node belongs to one of the following: The current node is a child node of the first reference node; or, The first reference node is a neighbor node of the current node.
3. The method according to claim 1 or 2, wherein, The first condition includes: The difference between the reconstruction attribute value of the first reference node and the reconstruction attribute value of the second reference node is less than or equal to a first threshold.
4. The method according to any one of claims 1 to 3, wherein, The attribute values of the current node include a first attribute value and a second attribute value. If the first reference node satisfies a first condition, then RAHT inter-frame prediction is performed on the attribute values of the current node, including: If the first reference node satisfies the first condition, then RAHT inter-frame prediction is performed on the second attribute value of the current node. The first condition is related to the difference between the first reconstruction attribute value of the first reference node and the first reconstruction attribute value of the second reference node.
5. The method according to claim 4, wherein, The first reference node includes at least one child node, and the first condition includes: The difference between the first reconstruction attribute value of the at least one child node and the first reconstruction attribute value of the second reference node is less than or equal to a first threshold. Wherein, the at least one child node includes the current node.
6. The method according to claim 3 or 5, wherein, The value of the first threshold is related to the reconstruction attribute value of the first reference node.
7. The method according to claim 6, wherein, The first threshold is determined based on the quotient of the reconstruction attribute value of the first reference node and the first parameter.
8. The method according to any one of claims 1 to 7, wherein, If the first reference node satisfies the first condition, then performing RAHT inter-frame prediction on the attribute value of the current node includes: If the first reference node satisfies the second condition, then it is determined whether the first reference node satisfies the first condition. The second condition is related to the hierarchical information of the first reference node in the octree. If the first reference node satisfies the first condition, then RAHT inter-frame prediction is performed on the attribute value of the current node.
9. The method according to claim 8, wherein, The second condition includes: The level of the first reference node in the octree is less than or equal to the second threshold.
10. The method according to claim 8 or 9, wherein, The method further includes: If the first reference node does not meet the second condition, then RAHT inter-frame prediction is performed on the attribute value of the current node.
11. The method according to any one of claims 1 to 10, wherein, The method further includes: If the first reference node does not meet the first condition, then no RAHT inter-frame prediction is performed on the attribute value of the current node.
12. The method according to any one of claims 1 to 11, wherein, The geometric information of the first reference node is the same as that of the second reference node.
13. The method according to claim 12, wherein, The geometric information includes coordinate information and / or hierarchical information.
14. The method according to any one of claims 1 to 13, wherein, The step of performing RAHT inter-frame prediction on the attribute values of the current node includes: Determine the reconstructed value of the first AC coefficient of the first child node of the second reference node; Based on the reconstructed value of the first AC coefficient, the predicted value of the second AC coefficient of the current node is determined.
15. The method according to claim 14, wherein, The method further includes: Analyze the bitstream to determine the residual value of the second AC coefficient; The reconstructed value of the second AC coefficient is determined based on the residual value of the second AC coefficient and the predicted value of the second AC coefficient.
16. The method according to claim 15, wherein, The method further includes: The reconstruction attribute value of the current node is determined by performing an inverse RAHT transformation based on the reconstructed value of the second AC coefficient and the DC coefficient.
17. A point cloud encoding method, applied to an encoder, comprising: Determine the first reference node of the current node in the current frame; If the first reference node satisfies the first condition, then the attribute value of the current node is subjected to Region Adaptive Hierarchical Transformation (RAHT) inter-frame prediction. The first condition is related to the difference between the reconstructed attribute value of the first reference node and the reconstructed attribute value of the second reference node in the reference frame.
18. The method according to claim 17, wherein, The relationship between the current node and the first reference node belongs to one of the following: The current node is a child node of the first reference node; or, The first reference node is a neighbor node of the current node.
19. The method according to claim 17 or 18, wherein, The first condition includes: The difference between the reconstruction attribute value of the first reference node and the reconstruction attribute value of the second reference node is less than or equal to a first threshold.
20. The method according to any one of claims 17 to 19, wherein, The attribute values of the current node include a first attribute value and a second attribute value. If the first reference node satisfies a first condition, then RAHT inter-frame prediction is performed on the attribute values of the current node, including: If the first reference node satisfies the first condition, then RAHT inter-frame prediction is performed on the second attribute value of the current node. The first condition is related to the difference between the first reconstruction attribute value of the first reference node and the first reconstruction attribute value of the second reference node.
21. The method according to claim 20, wherein, The first reference node includes at least one child node, and the first condition includes: The difference between the first reconstruction attribute value of the at least one child node and the first reconstruction attribute value of the second reference node is less than or equal to a first threshold. Wherein, the at least one child node includes the current node.
22. The method according to claim 19 or 21, wherein, The value of the first threshold is related to the reconstruction attribute value of the first reference node.
23. The method according to claim 22, wherein, The first threshold is determined based on the quotient of the reconstruction attribute value of the first reference node and the first parameter.
24. The method according to any one of claims 17 to 23, wherein, If the first reference node satisfies the first condition, then performing RAHT inter-frame prediction on the attribute value of the current node includes: If the first reference node satisfies the second condition, then it is determined whether the first reference node satisfies the first condition. The second condition is related to the hierarchical information of the first reference node in the octree. If the first reference node satisfies the first condition, then RAHT inter-frame prediction is performed on the attribute value of the current node.
25. The method according to claim 24, wherein, The second condition includes: The level of the first reference node in the octree is less than or equal to the second threshold.
26. The method according to claim 24 or 25, wherein, The method further includes: If the first reference node does not meet the second condition, then RAHT inter-frame prediction is performed on the attribute value of the current node.
27. The method according to any one of claims 17 to 26, wherein, The method further includes: If the first reference node does not meet the first condition, then no RAHT inter-frame prediction is performed on the attribute value of the current node.
28. The method according to any one of claims 17 to 27, wherein, The geometric information of the first reference node is the same as that of the second reference node.
29. The method according to claim 28, wherein, The geometric information includes coordinate information and / or hierarchical information.
30. The method according to any one of claims 17 to 29, wherein, The step of performing RAHT inter-frame prediction on the attribute values of the current node includes: Perform RAHT transformation on the attribute values of the first child node of the second reference node to determine the first AC coefficient; Perform RAHT transformation on the attribute values of the current node to determine the second AC coefficient; Based on the reconstructed value of the first AC coefficient, determine the predicted value of the second AC coefficient; The residual value of the second AC coefficient is determined based on the original value and the predicted value of the second AC coefficient.
31. The method according to claim 30, wherein, The method further includes: Write the residual value of the second AC coefficient into the bitstream.
32. The method according to claim 30 or 31, wherein, The method further includes: The reconstructed value of the second AC coefficient is determined based on the predicted value of the second AC coefficient and the residual value of the second AC coefficient. The reconstruction attribute value of the current node is determined by performing an inverse RAHT transformation based on the reconstructed value of the second AC coefficient and the DC coefficient.
33. A decoder, comprising: The unit is configured to parse the bitstream and determine the first reference node of the current node in the current frame. The prediction unit is configured to perform Region Adaptive Hierarchical Transformation (RAHT) inter-frame prediction on the attribute values of the current node if the first reference node satisfies a first condition, wherein the first condition is related to the difference between the reconstructed attribute values of the first reference node and the reconstructed attribute values of the second reference node in the reference frame.
34. A decoder, comprising: Memory, used to store computer programs; A processor, configured to perform the method as described in any one of claims 1 to 16 when running the computer program.
35. An encoder, comprising: The determination unit is configured to determine the first reference node of the current node in the current frame; The prediction unit is configured to perform Region Adaptive Hierarchical Transformation (RAHT) inter-frame prediction on the attribute values of the current node if the first reference node satisfies a first condition, wherein the first condition is related to the difference between the reconstructed attribute values of the first reference node and the reconstructed attribute values of the second reference node in the reference frame.
36. An encoder, comprising: Memory, used to store computer programs; A processor, configured to perform the method as described in any one of claims 17 to 32 when running the computer program.
37. A non-volatile computer-readable storage medium for storing a bitstream, said bitstream being generated by an encoding method using an encoder, or said bitstream being decoded by a decoding method using a decoder, wherein, The decoding method is the method as described in any one of claims 1 to 16, and the encoding method is the method as described in any one of claims 17 to 32.
38. A computer-readable storage medium storing a bitstream generated by the method of any one of claims 17 to 32.
39. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 16, or 17 to 32.