Coding method and apparatus, decoding method and apparatus, coder, decoder, bitstream, and storage medium
By determining the reference point in the point cloud encoder and using weighted prediction parameters to process the attribute values of the time domain reference point, the storage and transmission limitations of three-dimensional point cloud data and inter-frame reflectivity fluctuations are solved, and more efficient coding and better reconstruction quality are achieved.
Patent Information
- Application Number
- PCT/CN2023/142973
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, when processing three-dimensional point cloud data, there are storage space and transmission bandwidth limitations, and the average reflectivity fluctuations of the inter-frame point cloud data sets lead to inaccurate encoding prediction, resulting in degradation of reconstruction quality and redundancy of encoding bits.
By determining the reference point of the current point in the point cloud encoder and processing the attribute value of the time domain reference point based on the weighted prediction parameters, more accurate attribute residual values are obtained, which improves coding efficiency and reconstruction quality.
It effectively solves the problem of inaccurate prediction caused by reflectivity fluctuations in inter-frame point cloud datasets, improves coding efficiency and enhances point cloud reconstruction quality.
Smart Images

Figure CN2023142973_03072025_PF_FP_ABST
Abstract
Description
Coding and decoding method and device, codec, code stream, and storage medium Technical Field
[0001] The embodiments of the present application relate to point cloud encoding and decoding technology, including but not limited to encoding and decoding methods and devices, point cloud codecs, code streams, and storage media. Background Art
[0002] A point cloud is a collection of randomly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or scene. Point cloud data typically includes both geometric and attribute information about the sampling points. Geometric information includes the three-dimensional position of the sampling points, while attribute information includes their color and / or reflectivity.
[0003] Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes. Moreover, since point clouds are obtained by directly sampling real objects, they can provide a strong sense of reality while ensuring accuracy. Therefore, they are widely used, including virtual reality games, computer-aided design, geographic information systems, automatic navigation systems, digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive remote presentation, and three-dimensional reconstruction of biological tissues and organs.
[0004] As application demands grow, the processing of massive amounts of three-dimensional (3D) point cloud data is facing bottlenecks in storage space and transmission bandwidth. To better manage data, conserve server storage space, and reduce server-client transmission traffic and time, point cloud attribute encoding and decoding have become key issues in promoting the development of the point cloud industry.
[0005] Summary of the Invention
[0006] The encoding and decoding method and device, point cloud codec, code stream, and storage medium provided in the embodiments of the present application are implemented as follows:
[0007] According to a first aspect of an embodiment of the present application, a decoding method is provided, which is applied to a point cloud decoder, and the method includes: determining a reference point of a current point in the current point cloud based on geometric information of the current point cloud; wherein the reference point includes a time domain reference point, and the time domain reference point is one or more points in the reference point cloud; the reference point cloud is point cloud data that has been decoded before decoding the current point cloud; processing the attribute value of the time domain reference point according to a weighted prediction parameter to obtain a time domain reference attribute value; wherein the weighted prediction parameter is obtained by decoding a code stream; and determining an attribute reconstruction value of the current point based on the time domain reference attribute value of the time domain reference point.
[0008] It can be understood that in the embodiment of the present application, before determining the attribute residual value of the current point based on the reference point of the current point, the point cloud encoder first processes the attribute value of the time domain reference point in the reference point of the current point based on the weighted prediction parameter, and then determines the attribute residual value of the current point based on the time domain reference attribute value of the time domain reference point; in this way, since the weighted prediction parameter is determined based on the attribute value of at least part of the points of the current point cloud and the attribute reconstruction value of at least part of the points of the reference point cloud, the attribute value of the time domain reference point is first processed based on the weighted prediction parameter, and then the attribute residual value obtained based on the processing result is more accurate, which is beneficial to saving coding bit overhead.
[0009] According to a second aspect of an embodiment of the present application, a coding method is provided, which is applied to a point cloud encoder, and the method includes: determining a reference point of a current point in the current point cloud based on geometric information of the current point cloud; wherein the reference point includes a time domain reference point, and the time domain reference point is one or more points in the reference point cloud; the reference point cloud is point cloud data that has been encoded before encoding the current point cloud; obtaining weighted prediction parameters; wherein the weighted prediction parameters are determined based on attribute values of at least some points of the current point cloud and attribute reconstruction values of at least some points of the reference point cloud; processing the attribute values of the time domain reference points according to the weighted prediction parameters to obtain time domain reference attribute values; and determining the attribute residual value of the current point based on the time domain reference attribute values of the time domain reference points.
[0010] It can be understood that in an embodiment of the present application, before determining the attribute reconstruction value of the current point based on the attribute value of the time domain reference point, the attribute value of the time domain reference point is first processed based on the weighted prediction parameter, and then the attribute reconstruction value of the current point is determined based on the processing result (that is, the time domain reference attribute value of the time domain reference point); this is beneficial to improving the accuracy of the attribute reconstruction value of the current point, thereby improving the attribute reconstruction quality of the current point cloud.
[0011] According to a third aspect of an embodiment of the present application, a decoding device is provided, which is applied to a point cloud decoder, and the device includes: a first determination module, configured to determine a reference point of a current point in the current point cloud based on geometric information of the current point cloud; wherein the reference point includes a time domain reference point, and the time domain reference point is one or more points in the reference point cloud; the reference point cloud is point cloud data that has been decoded before decoding the current point cloud; a first processing module, configured to process the attribute value of the time domain reference point according to a weighted prediction parameter to obtain a time domain reference attribute value; wherein the weighted prediction parameter is obtained by decoding a code stream; a second determination module, configured to determine the attribute reconstruction value of the current point based on the time domain reference attribute value of the time domain reference point.
[0012] According to a fourth aspect of an embodiment of the present application, a point cloud decoder is provided, comprising a first memory and a first processor; wherein the first memory is used to store a computer program that can be run on the first processor; and the first processor is used to execute the decoding method described in the first aspect when running the computer program.
[0013] According to the fifth aspect of the embodiment of the present application, there is provided an encoding device, which is applied to a point cloud encoder, and the device includes: a third determination module, configured to determine the reference point of the current point in the current point cloud based on the geometric information of the current point cloud; wherein the reference point includes a time domain reference point, and the time domain reference point is one or more points in the reference point cloud; the reference point cloud is point cloud data that has been decoded before decoding the current point cloud; an acquisition module, configured to obtain weighted prediction parameters; wherein the weighted prediction parameters are determined based on the attribute values of at least some points of the current point cloud and the attribute reconstruction values of at least some points of the reference point cloud; a second processing module, configured to process the attribute values of the time domain reference points according to the weighted prediction parameters to obtain time domain reference attribute values; a fourth determination module, configured to determine the attribute residual value of the current point based on the time domain reference attribute values of the time domain reference points.
[0014] According to the sixth aspect of an embodiment of the present application, a point cloud encoder is provided, comprising a second memory and a second processor; wherein the second memory is used to store a computer program that can be run on the second processor; and the second processor is used to execute the encoding method described in the second aspect when running the computer program.
[0015] According to a seventh aspect of the embodiments of the present application, a code stream is provided, which is obtained by the encoding method described in the second aspect.
[0016] According to an eighth aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor adapted to execute a computer program; a computer-readable storage medium storing a computer program, wherein when the computer program is executed by the processor, the decoding method as described in the first aspect is implemented, or when the computer program is executed by the processor, the encoding method as described in the second aspect is implemented.
[0017] According to the ninth aspect of the embodiments of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, it implements the decoding method as described in the first aspect, or implements the encoding method as described in the second aspect.
[0018] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings herein are incorporated into and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, serve to illustrate the technical solutions of the present application. Obviously, the drawings described below are merely some embodiments of the present application. Those skilled in the art can, without inventive effort, derive other drawings from these drawings.
[0020] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0021] FIG1 is a schematic diagram of the structure of a point cloud encoder 1 in the MPEG G-PCC attribute codec framework;
[0022] FIG2 is a schematic diagram of the structure of the point cloud decoder 2 in the MPEG G-PCC attribute codec framework;
[0023] FIG3 is a schematic diagram of a PT encoding implementation process for point cloud attribute information;
[0024] Figure 4 is a schematic diagram of the LOD construction process of the point cloud;
[0025] FIG5 is a schematic diagram of a LT encoding implementation process for point cloud attribute information;
[0026] FIG6 is a schematic diagram of an implementation flow of the encoding method provided in an embodiment of the present application;
[0027] Figure 7 is a second schematic diagram of the LOD construction process of the point cloud;
[0028] FIG8 is a schematic diagram of an implementation flow of a method for determining an attribute residual value of a current point provided in an embodiment of the present application;
[0029] FIG9 is a schematic diagram of an implementation flow of a decoding method provided in an embodiment of the present application;
[0030] FIG10 is a schematic diagram of an implementation flow of a method for determining an attribute reconstruction value of a current point provided in an embodiment of the present application;
[0031] FIG11 is a schematic diagram of an implementation flow of updating the attribute reconstruction value of at least one point in the first refinement layer according to an embodiment of the present application;
[0032] FIG12 is a schematic diagram of the implementation flow of the core method provided in an embodiment of the present application;
[0033] FIG13 is a schematic diagram of test results of the encoding and decoding method provided in an embodiment of the present application under condition C1;
[0034] FIG14 is a schematic diagram of test results of the encoding and decoding method provided in an embodiment of the present application under condition C2;
[0035] FIG15 is a schematic diagram of test results of the encoding and decoding method provided in an embodiment of the present application under the condition CW;
[0036] FIG16 is a schematic diagram of test results of the encoding and decoding method provided in an embodiment of the present application under condition CY;
[0037] FIG17 is a schematic diagram of the structure of a decoding device provided in an embodiment of the present application;
[0038] FIG18 is a schematic diagram of the structure of an encoding device provided in an embodiment of the present application;
[0039] FIG19 is a schematic diagram of the structure of a point cloud decoder provided in an embodiment of the present application;
[0040] FIG20 is a schematic diagram of the structure of a point cloud encoder provided in an embodiment of the present application;
[0041] FIG21 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0044] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0045] The point cloud encoder and point cloud decoder frameworks and business scenarios described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Persons skilled in the art will appreciate that with the evolution of point cloud encoders and point cloud decoders and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application will also be applicable to similar technical problems.
[0046] Figure 1 is a schematic diagram of the point cloud encoder 1 within the MPEG G-PCC attribute codec framework. As shown in Figure 1, during the geometry encoding process, the geometric information is transformed so that the entire point cloud is contained within a bounding box. Quantization then occurs. This quantization step primarily serves a scaling purpose. Due to quantization rounding, the geometric information of some point clouds becomes identical. Parameters are then used to determine whether to remove duplicate points. This process of quantization and removing duplicate points is also known as voxelization. The bounding box is then partitioned into a multi-tree (such as an octree) or a prediction tree is constructed. During this process, arithmetic coding is performed on the points in the leaf nodes of the partition to generate a binary geometry bitstream. Alternatively, arithmetic coding is performed on the intersection points (vertices) generated by the partition (surface fitting is performed based on the intersection points) to generate a binary geometry bitstream. During the attribute encoding process, after the geometry encoding is completed and the geometric information is reconstructed, color conversion is required to convert the color information (an example of attribute information) from the RGB color space to the YUV color space. Then, the point cloud is recolored using the reconstructed geometric information so that the unencoded attribute information corresponds to the reconstructed geometric information. Attribute encoding is mainly performed on color information and / or reflectivity information. During the encoding of color information and / or reflectivity information, there are three main transformation methods: one is a distance-based transformation that relies on the level of detail (LOD) division, which is further divided into the predictive transform (PT) of point cloud attribute information and the lifting transform (LT) of point cloud attribute information; the other is to directly perform a region adaptive hierarchical transform (RAHT). In any of these three methods, the attribute information will be quantized and arithmetically encoded to generate a binary attribute bit stream.
[0047] Figure 2 is a schematic diagram of the structure of the point cloud decoder 2 in the MPEG G-PCC attribute codec framework. As shown in Figure 2, for the acquired binary bit stream, the geometric bit stream and attribute bit stream in the binary bit stream are first decoded independently. When decoding the geometric bit stream, the geometric information of the point cloud is obtained through arithmetic decoding-reconstruction of the multi-branch tree / reconstruction of the prediction tree-reconstruction of the geometry-coordinate inverse conversion; when decoding the attribute bit stream, the attribute information of the point cloud is obtained through arithmetic decoding-inverse quantization-LOD partitioning / RAHT and / or color inverse conversion, and the point cloud data to be encoded (i.e., the output point cloud) is restored based on the geometric information and attribute information. The attribute bit stream and geometric bit stream shown in Figures 1 and 2 can be understood as point cloud code streams.
[0048] It should be noted that, as shown in FIG1 or FIG2 , the geometric coding and decoding of MPEG G-PCC can be divided into geometric coding and decoding based on multi-tree (marked by a dotted box) and geometric coding and decoding based on prediction tree (marked by a dotted box).
[0049] Optionally, Figure 3 is a schematic diagram of a PT encoding implementation process for point cloud attribute information. As shown in Figure 3, the PT encoding process mainly includes: first, determining the three reconstructed nearest neighbor points of the current point based on the generation order of the level of detail (LOD) of the original point cloud; using the three reconstructed nearest neighbor points to predict the attributes of the current point to obtain the attribute prediction value of the current point; secondly, obtaining the attribute residual value of the current point (that is, the prediction residual) by subtracting the attribute prediction value from the original attribute value of the current point; finally, quantizing and arithmetically encoding the attribute residual value to generate a code stream of attribute information.
[0050] The following will provide a detailed introduction to the PT encoding process and PT decoding process according to the above steps; wherein, the PT encoding process includes the following steps 1 to 4.
[0051] Step 1, LOD generation:
[0052] The TMC13 software platform currently uses a distance-based LOD construction method. The specific steps for constructing LOD are as follows (1)-(5):
[0053] (1) Customize L Euclidean distances (d l ) l=0...L-1 , divided into L refinement layers (R l ) l=0...L-1 ;
[0054] (2) Mark all points in a frame of point cloud as unvisited, and set the visited point set V to an empty set;
[0055] (3) Refinement layer R l generate:
[0056] Traverse all points and ignore if the current point has been visited;
[0057] Otherwise, calculate the minimum distance D between the current point and the point set V;
[0058] If D is less than d l , then ignore the current point;
[0059] Otherwise, mark the current point as visited and add it to R l and V in;
[0060] Repeat this process until all points are traversed;
[0061] (4) l+1, based on d l+1 Repeat the process described in (3) to generate the refined layer R l+1 , until the refinement layer R is obtained L-1 .
[0062] (5) Take the refinement layers R0, R1, ..., R L-1 The union of the details level LOD (i.e. ); where (R l ) l=0...L-1 Represents the refinement level, and L is the number of LOD levels. As shown in Figure 4, L is equal to 3.
[0063] Step 2, optimal prediction value selection:
[0064] After the LOD is constructed, according to the LOD generation order, the three nearest neighboring points of the current point are first found from the encoded points in the current point cloud and the reference point cloud, and the attribute reconstruction values of these three nearest neighboring points are used as candidate prediction values for the current point; then, the best candidate prediction value is selected from them according to the rate-distortion optimization (RDO) as the attribute prediction value of the current point. For example, when encoding the attribute value of point P2 in Figure 4, the attribute reconstruction value of the nearest neighbor point P4 is used as a candidate prediction value, and its index is set to 1; the attribute reconstruction values of the second nearest neighbor point P5_ref (i.e., P5 comes from the reference point cloud) and the three nearest neighbor points P0 are used as candidate prediction values, and their indices are set to 2 and 3 respectively; the weighted average of points P0, P5_ref and P4 is used as a candidate prediction value, and its index is set to 0, as shown in Table 1; finally, the best prediction variable is selected using RDO. The formula for weighted average is shown in the following formula (1):
[0065] In the formula represents the spatial geometric weight from the neighboring point j to the current point i. The spatial geometric weight is determined as shown in the following formula (2):
[0066] represents the attribute prediction value of the current point i, j represents the index of any point of the three nearest neighboring points, represents the attribute value after reconstruction of the nearest neighbor point j (i.e., attribute reconstruction value), x i ,y i ,z i is the geometric position coordinate of the current point i, x ij ,y ij ,zij is the geometric coordinate of the nearest neighbor point j.
[0067] Table 1 Samples of candidate prediction items for attribute coding
[0068] Step 3: Attribute prediction residual (i.e. attribute residual value) and quantification:
[0069] The attribute prediction value of the current point i is obtained through the above prediction (k is the total number of points in the point cloud). Let (a i ) i∈0…k-1 is the original attribute value of the current point, then the attribute prediction residual of the current point (r i ) i∈0…k-1 The determination formula is as follows (3):
[0070] Furthermore, the attribute prediction residual is quantified according to the following formula (4):
[0071] Where Q i It represents the quantized attribute prediction residual of the current point i, Qs is the quantization step (Qs), which can be calculated by the quantization parameter (QP) specified by CTC.
[0072] Step 4: The encoder reconstructs the attribute value of the current point (i.e., obtains the attribute reconstruction value of the current point):
[0073] The purpose of reconstruction at the encoding end is to predict the attributes of subsequent points. Before reconstructing the attribute value, Q is first calculated according to the following formula (5): i Dequantize, remember The attribute prediction residual after dequantization is:
[0074] and attribute prediction values Add up to get the attribute reconstruction value of point i As shown in the following formula (6):
[0075] The lifting transform (LT) encoding process of point cloud attribute information is shown in Figure 5. The lifting transform LT also predicts the attributes of the point cloud based on LOD. The difference from the predictive transform PT is that the lifting transform LT first divides the LOD into high and low layers, predicts in the reverse order of the LOD generation layer, and introduces an update operator in the prediction process to update the quantized weights of the low-level LOD midpoints to improve the accuracy of the prediction. This is because the attribute values of the low-level LOD midpoints are frequently used to predict the attribute values of the high-level LOD midpoints, and the points in the low-level LOD should have greater influence.
[0076] The LT encoding process is as follows: Step 1 to Step 3:
[0077] Step 1, segmentation process:
[0078] The segmentation process is to divide the complete LOD layer into a low LOD layer L(N) and a high LOD layer H(N), where L(N) is generated based on the Euclidean distance d l Is greater than the Euclidean distance based on which H(N) is generated. If a point cloud has three levels of LOD, that is After segmentation, (R l ) l=1,2 It is the high LOD layer, denoted as H(N), (R l ) l=0 It is the low LOD layer, denoted as L(N).
[0079] Step 2, prediction process:
[0080] The points in the high LOD layer select the attribute information of the nearest neighbor points from the low LOD layers of the current point cloud and the reference point cloud, and determine the attribute prediction value P(N) of the current point based on this. The attribute prediction residual D(N) is shown in the following formula (7): D(N) = H(N) - P(N) (7);
[0081] LT has only one prediction mode, which is the weighted average of the three nearest neighbor points. The formula is shown in (8):
[0082] Where, represents the spatial geometric weight from the neighboring point j to the current point i, and is calculated as follows:
[0083] Where, Represents the attribute prediction value of the current point i, j represents the index of any point of the three nearest neighbor points, Represents the attribute value of the nearest neighbor point j after reconstruction, x i ,y i ,z iis the geometric position coordinate of the current point i, x ij ,y ij ,z ij is the geometric coordinate of the nearest neighbor point j.
[0084] Step 3, Update process:
[0085] Update the attribute prediction residual D(N) in the high LOD layer to obtain U(N), and use U(N) to improve the attribute value of the point in the low LOD layer (that is, the current attribute reconstruction value of the point belonging to the current point cloud among the three nearest neighboring points of the current point), as shown in formula (10): L′(N)=L(N)+U(N) (10);
[0086] The above process will iterate continuously until the lowest LOD according to the order of LOD from high to low.
[0087] As you can understand, point cloud inter-frame predictive coding is suitable for dynamic point cloud datasets acquired by lidar (LiDAR). LiDAR systems use laser beams to measure objects and terrain in the surrounding environment, then convert these measurements into point cloud data, where each point represents a discrete location in three-dimensional space. LiDAR point cloud data typically consists of geometric information and reflectivity information. Reflectivity, as a point attribute value, reflects the intensity of the echo captured by the laser scanner's receiver. It is related to the target's surface material, roughness, angle of incidence, as well as the instrument's emission energy, laser wavelength, and lighting conditions during the acquisition process. LiDAR point clouds have a wide range of uses in various applications, including autonomous driving, mapping, environmental perception, architectural surveying, geological exploration, remote sensing, virtual reality, and game development. This data can be used to create high-resolution 3D models to assist robots and autonomous vehicles in navigation, obstacle detection, and other tasks.
[0088] During the acquisition process, the average reflectivity of each frame in the point cloud dataset fluctuates due to factors such as daylight conditions, nighttime lighting conditions, traffic light on / off, vehicle lighting, and weather conditions. Therefore, when the point cloud codec performs inter-frame prediction, if the nearest neighbor point comes from the reference point cloud, the prediction will be inaccurate, resulting in reduced reconstruction quality and redundant encoding bits. The root cause is that related technologies do not compensate for the nearest neighbor points in the reference point cloud, but instead directly use their reconstructed values for prediction. Therefore, these technologies cannot achieve ideal compression effects when the average reflectivity of each frame in the point cloud dataset fluctuates.
[0089] The embodiment of the present application provides an encoding method, which can be applied to the point cloud encoder 1 shown in FIG1 . FIG6 is a schematic diagram of an implementation flow of the encoding method provided in the embodiment of the present application. As shown in FIG6 , the method includes the following steps 601 and 605:
[0090] Step 601: Determine a reference point for a current point in the current point cloud based on geometric information of the current point cloud; wherein the reference point includes a temporal reference point, which is one or more points in a reference point cloud; and the reference point cloud is point cloud data that has been encoded before encoding the current point cloud.
[0091] Step 602: Obtain weighted prediction parameters; wherein the weighted prediction parameters are determined based on attribute values of at least some points of the current point cloud and attribute reconstruction values of at least some points of the reference point cloud;
[0092] Step 603: Process the attribute value of the time domain reference point according to the weighted prediction parameter to obtain a time domain reference attribute value;
[0093] Step 604: Determine the attribute residual value of the current point according to the time domain reference attribute value of the time domain reference point.
[0094] In an embodiment of the present application, before determining the attribute residual value of the current point based on the reference point of the current point, the point cloud encoder first processes the attribute value of the time domain reference point in the reference point of the current point based on the weighted prediction parameter, and then determines the attribute residual value of the current point based on the time domain reference attribute value of the time domain reference point; in this way, since the weighted prediction parameter is determined based on the attribute value of at least part of the points of the current point cloud and the attribute reconstruction value of at least part of the points of the reference point cloud, the attribute value of the time domain reference point is first processed based on the weighted prediction parameter, and then a more accurate attribute residual value is obtained based on the processing result, which is beneficial to saving coding bit overhead.
[0095] The following describes further optional implementations and related terms of each of the above steps.
[0096] In step 601, based on the geometric information of the current point cloud, a reference point of the current point in the current point cloud is determined; wherein the reference point includes a time domain reference point, and the time domain reference point is one or more points in the reference point cloud; the reference point cloud is point cloud data that has been encoded before encoding the current point cloud.
[0097] For example, the current point cloud can be understood as the point cloud data of the current frame, and the reference point cloud can be understood as the point cloud data of the reference frame.
[0098] Optionally, in some embodiments, the point cloud encoder needs to construct the LOD of the current point cloud before determining the reference point of the current point. Of course, for a frame of point cloud, it is sufficient to construct a single LOD. That is, before determining the reference point of the first point in the current point cloud, that is, before encoding the attribute value of the first point in the current point cloud, the LOD of the current point cloud is first constructed based on the geometric information of the current point cloud. Subsequently, for each point in the current point cloud, its attribute value is encoded based on the constructed LOD.
[0099] Optionally, in some embodiments, the point cloud encoder can construct the LOD of the current point cloud based on the distance. As shown in FIG7 , the point cloud encoder can construct the LOD of the current point cloud through the following steps 701 to 708:
[0100] Step 701: Get L predefined distances (d l ) l=0...L-1 ; Among them, a d l Corresponding to a refinement layer, it can be divided into L refinement layers (R l ) l=0...L-1 ; Wherein, L is greater than 1;
[0101] For example, d l It can be understood as the lth predefined distance, which can be a parameter representing the geometric distance between two spatial coordinates, such as Euclidean distance or cosine distance.
[0102] Step 702: Mark all points in the current point cloud as unvisited, and set the visited point set V to an empty set; then execute the first process, which includes the following steps 703 to 706;
[0103] Step 703, traverse all points, if the current point has been visited, ignore it; otherwise, execute step 703;
[0104] Step 704, calculating the minimum distance D between the current point and the points in the point set V;
[0105] It can be understood that the minimum distance D and d l The properties are the same, for example, d l is the Euclidean distance, and the minimum distance D is also the Euclidean distance. l It is the cosine distance, and the minimum distance D is also the cosine distance.
[0106] Step 705: Determine whether the minimum distance D is less than d l ; If yes, ignore the current point; otherwise, execute step 706;
[0107] Step 706: Mark the current point as visited and add the current point to the lth refinement layer R l and point set V;
[0108] Continue to execute steps 703 to 706 for the next point until all points of the current point cloud are traversed. l+1 Execute steps 703 to 706 to obtain the refinement layer R l+1 , until L refinement layers (R l ) l=0...L-1 , proceed to step 707;
[0109] Step 707: Take the refinement layers R0, R1, ..., R L-1 The union of the details level LOD (i.e. ); where (R l ) l=0...L-1 Represents the refinement level, where L is the number of LOD levels.
[0110] After obtaining the LOD of the current point cloud, the point cloud encoder can encode the attribute values of the points in the current point cloud. In some embodiments, the point cloud encoder can implement step 601 as follows: determine the current refinement layer, the current refinement layer refers to a refinement layer of the current point in the LOD of the current point cloud; determine the reference point of the current point from at least one first refinement layer of the LOD of the current point cloud, the encoded points of the current refinement layer and / or at least one second refinement layer of the LOD of the reference point cloud; wherein the encoding order of the at least one first refinement layer is before the current refinement layer, and the predefined distance (such as Euclidean distance or cosine distance, etc.) based on which the second refinement layer is generated is equal to the predefined distance based on which the first refinement layer is generated.
[0111] Optionally, in some embodiments, the point cloud encoder can determine the reference point of the current point from the at least one first refinement layer, the encoded points of the current refinement layer and the at least one second refinement layer; in other embodiments, the point cloud encoder can also determine the reference point of the current point from the at least one first refinement layer and the at least one second refinement layer; in some other embodiments, the point cloud encoder can also determine the reference point of the current point from the encoded points of the current refinement layer and the at least one second refinement layer.
[0112] In an embodiment of the present application, there is no limitation on which point or points can be used as reference points for the current point. For example, a point whose distance to the current point (i.e., geometric distance) is less than or equal to a distance threshold can be used as a reference point for the current point; for another example, one or more points (i.e., nearest neighbor points) with the smallest distance to the current point can also be used as reference points for the current point.
[0113] That is, in some embodiments, the reference point of the current point is one or more first nearest neighbor points of the current point (that is, the first one or more points with the smallest distance to the current point), and the one or more first nearest neighbor points are from the current point cloud and / or the reference point cloud. In an embodiment of the present application, the one or more first nearest neighbor points can be N first nearest neighbor points, where N is greater than 1 and N can be a predefined value.
[0114] In the embodiment of the present application, there is no limitation on the relationship between the at least one first refinement layer and the current refinement layer, which is related to the encoding order of the current point cloud. For example, the encoding order of the current point cloud is R L-1 ,R L-2 ,…,R0, then the predefined distance based on which the at least one first refinement layer is generated is smaller than the predefined distance based on which the current refinement layer is generated; for example, the encoding order of the current point cloud is R0, R1,…, R L-1 , then the predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated. In short, the coding order of the at least one first refinement layer is before the current refinement layer.
[0115] In step 602, weighted prediction parameters are obtained; wherein the weighted prediction parameters are determined based on attribute values of at least some points of the current point cloud and attribute reconstruction values of at least some points of the reference point cloud.
[0116] In the embodiment of the present application, the weighted prediction parameter may be applicable to the attribute coding of each point in the current point cloud, or may be applicable only to the attribute coding of some points in the current point cloud.
[0117] For a scheme in which a weighted prediction parameter is only applicable to attribute encoding of a portion of the points in the current point cloud, in one possible implementation, if the weighted prediction parameter is determined based on the attribute values of the first portion of the points in the current point cloud and the attribute reconstructed values of the first portion of the points in the reference point cloud, then the determined weighted prediction parameter is used for encoding the attribute values of the first portion of the points in the current point cloud; if the weighted prediction parameter is determined based on the attribute values of the second portion of the points in the current point cloud and the attribute reconstructed values of the second portion of the points in the reference point cloud, then the determined weighted prediction parameter is used for encoding the attribute values of the second portion of the points in the current point cloud. Therefore, when writing the weighted prediction parameter into the bitstream, the point cloud encoder also needs to indicate to which portion of the points the weighted prediction parameter applies.
[0118] For the scheme in which weighted prediction parameters are applicable to the attribute encoding of each point in the current point cloud, the point cloud encoder may determine the weighted prediction parameters used by the current point cloud based on the attribute values of at least some of the points in the current point cloud and the attribute reconstruction values of at least some of the points in the reference point cloud before starting to encode the attributes of the current point cloud. The weighted prediction parameters can be directly obtained from the corresponding storage area and used when the encoding method is subsequently executed on each point in the current point cloud. Optionally, in some embodiments, the point cloud encoder may write the weighted prediction parameters into the bitstream before writing the attribute residual value of the first point in the current point cloud into the bitstream. The weighted prediction parameters may also be written into the bitstream after writing the attribute residual value of the first point in the current point cloud into the bitstream. The point cloud encoder may also write the weighted prediction parameters into the bitstream again before and / or before writing the attribute residual values of other points in the current point cloud into the bitstream.
[0119] In one possible implementation, the point cloud encoder may determine the weighted prediction parameters as follows: fitting the mapping relationship between the attribute values of at least some of the points in the current point cloud and the attribute reconstruction values of at least some of the points in the reference point cloud to obtain the weighted prediction parameters; wherein, the attribute values of at least some of the points in the current point cloud may be the true values of the points, that is, the collected attribute values, and the true values may be the collected original attribute values, or the values of the original attribute values after preprocessing (such as filtering, etc.).
[0120] Exemplarily, in some embodiments, the weighted prediction parameter includes a weighting factor and / or an offset.
[0121] For example, remember a i (i=0,1,…,n) is the attribute value of the i-th point in the current point cloud, where n is the number of points in the current point cloud; a_ref j (j=0,1,…,m) is the attribute reconstruction value of the jth point in the reference point cloud, where m is the number of points in the reference point cloud. The weighting factor q and the offset w can be obtained by fitting according to the following formula (11): i =q×a_ref j +w (11);
[0122] In step 603, the attribute value (such as the attribute reconstruction value) of the time domain reference point is processed according to the weighted prediction parameter to obtain a time domain reference attribute value.
[0123] Optionally, in some embodiments, the point cloud encoder may implement step 603 as follows: determining a first product of the weighting factor and the attribute reconstruction value of the time domain reference point; and offsetting the first product according to the offset to obtain the time domain reference attribute value.
[0124] For example, the time domain reference attribute value of the time domain reference point can be obtained according to the following formula (12):
[0125] Where temp j refers to the time domain reference attribute value of time domain reference point j, It refers to the attribute reconstruction value of the time domain reference point j.
[0126] In step 604, the attribute residual value of the current point is determined according to the time domain reference attribute value of the time domain reference point.
[0127] In some embodiments, as shown in FIG8 , the point cloud encoder may implement step 604 through the following steps 6041 and 6042 :
[0128] Step 6041: Determine the attribute prediction value of the current point according to the time domain reference attribute value of the time domain reference point.
[0129] In an embodiment of the present application, the point cloud encoder may use the following embodiment 1 or embodiment 2 to determine the attribute prediction value of the current point.
[0130] Optionally, in the first embodiment, the point cloud encoder may implement step 6041 as follows: determining a candidate prediction value for the current point based on the temporal reference attribute value of the temporal reference point; and determining an attribute prediction value for the current point based on a rate-distortion cost of the candidate prediction value for the current point. For example, in some embodiments, the attribute prediction value for the current point is equal to the candidate prediction value with the smallest rate-distortion cost.
[0131] It is understood that the reference points for the current point may include temporal reference points and / or spatial reference points. The spatial reference points are one or more points in the current point cloud, and the spatial reference points are points that were encoded before the current point was encoded. For the reference points of the current point obtained based on LOD search, i.e., the one or more first nearest neighbor points, the temporal reference points come from the second refinement layer, and the spatial reference points come from encoded points in the first refinement layer or the current refinement layer.
[0132] For the case where the reference point of the current point includes a time domain reference point and a spatial domain reference point, for embodiment one, further, in some embodiments, the point cloud encoder can determine the candidate prediction value of the current point as follows: based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point, determine the candidate prediction value of the current point.
[0133] Exemplarily, in some embodiments, the candidate prediction value of the current point includes the time domain reference attribute value of the time domain reference point, the weighted average of the time domain reference attribute value of the time domain reference point, the attribute reconstruction value of the spatial domain reference point and / or the weighted average of the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0134] In some embodiments, the method further includes: determining a prediction mode of the current point according to an index value of the candidate prediction value with the minimum rate-distortion cost in the candidate prediction values of the current point; and generating a bitstream according to the prediction mode.
[0135] In one possible implementation, the prediction mode of the current point is equal to the index value of the candidate prediction value with the minimum rate-distortion cost among the candidate prediction values of the current point, that is, the prediction mode of the current point is represented by the index value, and the prediction mode written into the bitstream is the index value.
[0136] In the embodiment of the present application, the index value of the candidate prediction value for the current point is determined based on the arrangement order of the candidate prediction values for the current point. The index value of the candidate prediction value can be equal to the arrangement position of the candidate prediction value among the candidate prediction values for the current point. In the embodiment of the present application, the arrangement rules for the candidate prediction values of the current point are not limited. Taking the reference point of the current point as the four first nearest neighbor points as an example, the arrangement order of the candidate prediction values of the current point is shown in Table 2 below.
[0137] Table 2
[0138] It should be noted that in Table 2, if the four first nearest neighbor points include both time domain reference points and spatial domain reference points, the candidate prediction value corresponding to the index value 1 is the weighted average of the time domain reference attribute values of all time domain reference points and the attribute reconstruction values of all spatial domain reference points; if the four first nearest neighbor points only include time domain reference points, the candidate prediction value corresponding to the index value 1 is the weighted average of the time domain reference attribute values of all time domain reference points; if the four first nearest neighbor points only include spatial domain reference points, the candidate prediction value corresponding to the index value 1 is the weighted average of the attribute reconstruction values of all spatial domain reference points.
[0139] In Table 2, the so-called first nearest neighbor refers to the point closest to the current point among the four first nearest neighbors, the so-called second nearest neighbor refers to the point with the second closest distance to the current point among the four first nearest neighbors, and so on. In Table 2, if the first (second, third, or fourth) nearest neighbor is a time domain reference point, the attribute value of the first (second, third, or fourth) nearest neighbor is the time domain reference attribute value of the point; if the first (second, third, or fourth) nearest neighbor is a spatial domain reference point, the attribute value of the first (second, third, or fourth) nearest neighbor is the attribute reconstructed value of the point.
[0140] Of course, the arrangement order of the candidate prediction values of the current point in the above Table 2 can also be any other order, and the embodiment of the present application is not limited to the arrangement rules shown in Table 2.
[0141] Optionally, in embodiment 2, the reference point also includes a spatial reference point, which is one or more encoded points in the current point cloud; the point cloud encoder can also implement step 6041 (determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point) in this way: determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial reference point.
[0142] Optionally, in embodiment 2, the point cloud encoder can use the method described in embodiment 1 above to determine the attribute prediction value of the current point, or it can determine the attribute prediction value of the current point as follows: determine the attribute prediction value of the current point based on the weighted average of the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0143] For example, in some embodiments, the attribute prediction value of the current point is equal to the weighted average value. In other embodiments, the attribute prediction value of the current point may also be equal to the weighted average value after optimization. In one possible implementation, the point cloud encoder may optimize the weighted average value by multiplying the weighted average value by a second value, or by multiplying the weighted average value by the second value and then adding the result to a third value. The second value and the third value may be predefined values.
[0144] Optionally, in embodiment 2, after determining the attribute residual value of the current point, the point cloud encoder can update the attribute reconstruction value of the spatial reference point (here refers to the current attribute reconstruction value of the spatial reference point) according to the attribute residual value of the current point.
[0145] In the embodiments of the present application, there is no limitation on the method for updating the attribute reconstruction value of the airspace reference point. In some embodiments, the attribute residual value of the current point can be converted into a weighting coefficient, and then the weighting coefficient is multiplied by the attribute reconstruction value of the airspace reference point to achieve the update of the attribute reconstruction value of the airspace reference point.
[0146] In other embodiments, the point cloud encoder may also update the attribute reconstruction value of the spatial reference point as follows: update the attribute residual value D(N) of the current point to obtain U(N); use U(N) to update the attribute reconstruction value of the spatial reference point; for example, update the attribute reconstruction value of the spatial reference point according to the following formula (13): L′(N)=L(N)+U(N) (13);
[0147] Where L′(N) is the attribute reconstruction value of the spatial reference point after the update, and L(N) is the attribute reconstruction value of the spatial reference point before the update.
[0148] Step 6042: Determine the attribute residual value of the current point based on the attribute value and the attribute prediction value of the current point.
[0149] In step 6042, the attribute value of the current point refers to the true value of the current point, that is, the collected attribute value. The true value can be the collected original attribute value or the value of the original attribute value after preprocessing (such as filtering, etc.).
[0150] In one possible implementation, the point cloud encoder can determine the attribute residual value D(N) of the current point according to the following formula (14): D(N)=H(N)-P(N) (14);
[0151] Where H(N) refers to the attribute value of the current point, and P(N) refers to the predicted attribute value of the current point.
[0152] In some embodiments, after obtaining the attribute residual value of the current point, the method further includes: generating a code stream according to the attribute residual value of the current point.
[0153] Optionally, in some embodiments, the point cloud encoder may quantize the attribute residual value of the current point, and then perform arithmetic coding on the quantized attribute residual value to generate a code stream.
[0154] Of course, in other embodiments, the point cloud encoder may not quantize the attribute residual value of the current point, but directly perform arithmetic coding on it to generate a code stream.
[0155] The present application provides a decoding method, which can be applied to the point cloud decoder 2 shown in FIG2 . FIG9 is a schematic diagram of an implementation flow of the decoding method provided in the present application embodiment. As shown in FIG9 , the method includes the following steps 901 and 903:
[0156] Step 901: Determine a reference point for a current point in the current point cloud based on geometric information of the current point cloud; wherein the reference point includes a temporal reference point, which is one or more points in a reference point cloud; and the reference point cloud is point cloud data that has been decoded before decoding the current point cloud.
[0157] Step 902: Process the attribute value of the time domain reference point according to the weighted prediction parameter to obtain a time domain reference attribute value; wherein the weighted prediction parameter is obtained by decoding the bitstream;
[0158] Step 903: Determine the attribute reconstruction value of the current point according to the time domain reference attribute value of the time domain reference point.
[0159] In an embodiment of the present application, before determining the attribute reconstruction value of the current point based on the attribute value of the time domain reference point, the attribute value of the time domain reference point is first processed based on the weighted prediction parameter, and then the attribute reconstruction value of the current point is determined based on the processing result (that is, the time domain reference attribute value of the time domain reference point); this is beneficial to improving the accuracy of the attribute reconstruction value of the current point, thereby improving the attribute reconstruction quality of the current point cloud.
[0160] It can be understood that the decoding method shown in Figure 9 corresponds to the encoding method shown in Figure 6, so the decoding method shown in Figure 9 and further or additional implementations of the method are similar to the description of the encoding method shown in Figure 6 and further or additional implementations of the method. For technical details not disclosed in the decoding method embodiment, please refer to the description of the above encoding method embodiment for understanding.
[0161] The following describes further optional implementations and related terms of each step shown in FIG9 .
[0162] In step 901, based on the geometric information of the current point cloud, a reference point of the current point in the current point cloud is determined; wherein the reference point includes a time domain reference point, and the time domain reference point is one or more points in the reference point cloud; the reference point cloud is point cloud data that has been decoded before decoding the current point cloud.
[0163] For example, the current point cloud can be understood as point cloud data of the current point cloud, and the reference point cloud can be understood as point cloud data of the reference point cloud.
[0164] Optionally, in some embodiments, the point cloud decoder needs to construct the LOD of the current point cloud before determining the reference point of the current point. Of course, for a frame of point cloud, constructing a single LOD is sufficient. That is, before determining the reference point of the first point in the current point cloud, or in other words, before decoding the attribute value of the first point in the current point cloud, the LOD of the current point cloud is constructed based on the geometric information of the current point cloud. Subsequently, for each point in the current point cloud, its attribute value is decoded based on the constructed LOD.
[0165] Optionally, in some embodiments, the method by which the point cloud decoder constructs the LOD of the current point cloud is the same as the method by which the above-mentioned point cloud encoder constructs the LOD of the current point cloud. Specifically, the construction of the LOD can be implemented with reference to the method shown in Figure 7, which will not be repeated here.
[0166] After obtaining the LOD of the current point cloud, the point cloud decoder can reconstruct the attribute values of the points in the current point cloud. In some embodiments, the point cloud decoder can implement step 901 as follows: determine the current refinement layer, the current refinement layer refers to a refinement layer of the current point in the LOD of the current point cloud; determine the reference point of the current point from at least one first refinement layer of the LOD of the current point cloud, the decoded points of the current refinement layer and / or at least one second refinement layer of the LOD of the reference point cloud; wherein the decoding order of the at least one first refinement layer is before the current refinement layer, and the predefined distance (such as Euclidean distance or cosine distance, etc.) based on which the second refinement layer is generated is equal to the predefined distance based on which the first refinement layer is generated.
[0167] In some embodiments, the point cloud decoder can determine the reference point of the current point from the at least one first refinement layer, the decoded points of the current refinement layer and the at least one second refinement layer; in other embodiments, the point cloud decoder can also determine the reference point of the current point from the at least one first refinement layer and the at least one second refinement layer; in yet other embodiments, the point cloud decoder can also determine the reference point of the current point from the decoded points of the current refinement layer and the at least one second refinement layer.
[0168] In an embodiment of the present application, there is no limitation on which point or points can be used as reference points for the current point. For example, a point whose distance to the current point (i.e., geometric distance) is less than or equal to a distance threshold can be used as a reference point for the current point; for another example, one or more points (i.e., nearest neighbor points) with the smallest distance to the current point can also be used as reference points for the current point.
[0169] That is, in some embodiments, the reference point is one or more first nearest neighbor points of the current point (i.e., the first one or more points with the smallest distance to the current point), and the one or more first nearest neighbor points are from the current point cloud and / or the reference point cloud. In an embodiment of the present application, the one or more first nearest neighbor points can be N first nearest neighbor points, where N is greater than 1 and N can be a predefined value.
[0170] In the embodiment of the present application, there is no limitation on the relationship between the at least one first refinement layer and the current refinement layer, which is related to the decoding order of the current point cloud. For example, the decoding order of the current point cloud is R L-1 ,R L-2 ,…,R0, then the predefined distance based on which the at least one first refinement layer is generated is smaller than the predefined distance based on which the current refinement layer is generated; for example, the decoding order of the current point cloud is R0, R1,…, R L-1, then the predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated. In short, the decoding order of the at least one first refinement layer is before the current refinement layer.
[0171] In step 902, the attribute value of the time domain reference point is processed according to the weighted prediction parameter to obtain a time domain reference attribute value; wherein the weighted prediction parameter is obtained by decoding the bitstream.
[0172] In the embodiment of the present application, the weighted prediction parameters may be applicable to the attribute reconstruction / attribute decoding of each point in the current point cloud, or may be applicable only to the attribute reconstruction / attribute decoding of some points in the current point cloud.
[0173] For a solution in which weighted prediction parameters are only applicable to attribute reconstruction of some points of the current point cloud, in one possible implementation, the point cloud decoder determines to which part of the points the weighted prediction parameters are applicable based on the decoded bitstream, and reconstructs attributes of the part of the points based on the weighted prediction parameters.
[0174] For attribute encoding schemes in which weighted prediction parameters are applied to each point in the current point cloud, optionally, in some embodiments, the weighted prediction parameters are determined by decoding the bitstream before decoding the attribute residual value of the first point in the current point cloud. The point cloud decoder can first decode the bitstream to determine the weighted prediction parameters and write them to a storage area before beginning attribute reconstruction for the current point cloud. The weighted prediction parameters can then be directly retrieved and used from the corresponding storage area when executing the encoding method for each point in the current point cloud.
[0175] In some embodiments, the weighted prediction parameter is used to characterize a mapping relationship between attribute values of at least some points in the current point cloud and attribute reconstructed values of at least some points in the reference point cloud.
[0176] In some embodiments, the weighted prediction parameter is a fitting coefficient obtained by fitting a mapping relationship between attribute values of at least some points in the current point cloud and attribute reconstruction values of at least some points in the reference point cloud.
[0177] In some embodiments, the weighted prediction parameters include weighting factors and / or offsets.
[0178] In some embodiments, the point cloud decoder may implement step 903 as follows: determining a first product of the weighting factor and the attribute reconstruction value of the time domain reference point; and offsetting the first product according to the offset to obtain the time domain reference attribute value.
[0179] The time domain reference attribute value for determining the time domain reference point here is the same as that of the encoding end, so it will not be described again here.
[0180] In step 903, the attribute reconstruction value of the current point is determined according to the time domain reference attribute value of the time domain reference point.
[0181] In some embodiments, the method further includes: decoding a code stream and determining a property residual value of the current point.
[0182] In one possible implementation, the point cloud decoder may perform arithmetic decoding on the bitstream carrying the attribute residual value, and then dequantize the arithmetically decoded value to obtain the attribute residual value of the current point.
[0183] As shown in FIG10 , the point cloud decoder can implement step 903 through the following steps 9031 and 9032:
[0184] Step 9031, determining the attribute prediction value of the current point according to the time domain reference attribute value of the time domain reference point;
[0185] In an embodiment of the present application, the point cloud decoder can adopt the following embodiment three or embodiment four to determine the attribute prediction value of the current point; wherein, embodiment three is a decoding operation corresponding to the above embodiment one. For the technical details not disclosed in embodiment three, they can be understood by referring to the description of the above embodiment one and further or additional implementation methods of embodiment one.
[0186] Optionally, in embodiment three, the point cloud decoder can implement step 9031 as follows: determine the candidate prediction value of the current point based on the time domain reference attribute value of the time domain reference point; decode the code stream to determine the prediction mode of the current point; determine the attribute prediction value of the current point based on the prediction mode and the candidate prediction value of the current point.
[0187] It is understood that the reference points for the current point may include temporal reference points and / or spatial reference points. The spatial reference points are one or more points in the current point cloud, and the spatial reference points are points that were encoded before decoding the current point. For the reference points of the current point obtained based on LOD search, i.e., the one or more first nearest neighbor points, the temporal reference points come from the second refinement layer, and the spatial reference points come from encoded points in the first refinement layer or the current refinement layer.
[0188] For the case where the reference point of the current point includes a time domain reference point and a spatial domain reference point, for embodiment three, further, in some embodiments, the point cloud decoder can determine the candidate prediction value of the current point as follows: based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point, determine the candidate prediction value of the current point.
[0189] Exemplarily, in some embodiments, the candidate prediction value of the current point includes the time domain reference attribute value of the time domain reference point, the weighted average of the time domain reference attribute value of the time domain reference point, the attribute reconstruction value of the spatial domain reference point and / or the weighted average of the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0190] Exemplarily, in some embodiments, the attribute prediction value of the current point is equal to the candidate prediction value corresponding to the prediction mode.
[0191] It can be understood that, at the encoder, the prediction mode written into the bitstream is the index value of the candidate prediction value with the lowest rate-distortion cost among the candidate prediction values at the current point. The index value of the candidate prediction value at the current point is determined based on the order in which the candidate prediction values are arranged at the current point. The order in which the candidate prediction values are arranged at the current point is the same at the decoder as at the encoder. Therefore, the order in which the candidate prediction values are arranged at the decoder and the determination of the index value of the candidate prediction value at the current point can be understood by referring to the description of the encoder.
[0192] In embodiment four, the reference point also includes a spatial reference point, which is one or more decoded points in the current point cloud; the point cloud decoder can implement step 9031 as follows: determine the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial reference point.
[0193] Optionally, in embodiment four, the point cloud decoder can use the method described in embodiment one above to determine the attribute prediction value of the current point, or it can determine the attribute prediction value of the current point in this way: determine the attribute prediction value of the current point based on the weighted average of the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0194] For example, in some embodiments, the attribute prediction value of the current point is equal to the weighted average value. In other embodiments, the attribute prediction value of the current point may also be equal to the weighted average value after optimization. In one possible implementation, the point cloud encoder may optimize the weighted average value by multiplying the weighted average value by a second value, or by multiplying the weighted average value by the second value and then adding the result to a third value. The second value and the third value may be predefined values.
[0195] Optionally, in embodiment four, before determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point, the point cloud decoder updates the attribute reconstruction value of the point of at least one first refinement layer based on the attribute residual value of the point of the current refinement layer.
[0196] It can be understood that, unlike the encoding end, the decoding end first updates the attribute reconstruction value of at least one point in the first refinement layer after obtaining the attribute residual value by decoding the code stream, and then performs attribute prediction based on the updated result.
[0197] Optionally, in some embodiments, as shown in FIG11 , the point cloud decoder may update the attribute reconstruction values of the points of the at least one first refinement layer through the following steps 1101 to 1102:
[0198] Step 1101: Execute a second process according to the i-th point of the current refinement layer; the second process includes the following steps 1101a and 1101b:
[0199] Step 1101a, determining one or more second nearest neighbor points of the i-th point of the current refinement layer from the at least one first refinement layer and the at least one second refinement layer of the LOD of the reference point cloud; wherein i is greater than 0 and less than or equal to the total number of points in the current refinement layer; the predefined distance based on which the at least one second refinement layer is generated corresponds to and is equal to the predefined distance based on which the at least one first refinement layer is generated; the at least one first refinement layer includes a refinement layer generated based on a maximum distance in the LOD of the current point cloud, and an initial attribute reconstruction value of a point in the refinement layer generated based on the maximum distance is equal to a first numerical value.
[0200] Step 1101b: updating the attribute reconstruction values of the points belonging to the first refinement layer among the one or more second nearest neighbor points according to the attribute residual value of the i-th point.
[0201] Exemplarily, in some embodiments, the point cloud decoder may implement step 1101b as follows: updating the attribute residual value D(N) of the i-th point (here, the attribute residual value after inverse quantization) to obtain U(N); using U(N) to update the attribute reconstruction value of the point belonging to the first refinement layer among the one or more second nearest neighbor points; for example, updating the attribute reconstruction value of the point belonging to the first refinement layer among the one or more second nearest neighbor points according to the following formula (15): L′(N)=L(N)+U(N) (15);
[0202] In formula (14), L′(N) is the updated attribute reconstruction value of the point belonging to the first refinement layer among the one or more second nearest neighbor points, and L(N) is the attribute reconstruction value of the point before the update.
[0203] Step 1102: Iterate the second process according to the i+1th point of the current refined layer until the second process is executed according to the last point of the current refined layer, and start determining the attribute prediction value of each point of the current refined layer based on this, for example, determine the attribute prediction value of the current point according to the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point; wherein the attribute reconstruction value of the spatial domain reference point is the updated attribute reconstruction value.
[0204] It can be understood that at the decoding end, the method for updating the attribute reconstruction value of the point of the first refinement layer as shown in Figure 11 is to first update the attribute reconstruction value of the point of the first refinement layer based on the attribute residual values of all points of the current refinement layer. Based on this, in the process of determining the attribute prediction value of each point of the current refinement layer, there is no need to update the attribute reconstruction value of the point of the first refinement layer. Instead, one or more first nearest neighbor points of the current point are directly searched from the at least one first refinement layer and the at least one second refinement layer, and the attribute prediction value of the current point is determined based on the current attribute reconstruction value of the spatial reference point in the one or more first nearest neighbor points and the time domain reference attribute value of the time domain reference point.
[0205] It can be understood that for the same point, the one or more second nearest neighbor points described in the update stage and the one or more first nearest neighbor points described in the attribute prediction stage are the same points. Therefore, there is no need to search for the one or more first nearest neighbor points of the current point again in the attribute prediction stage. Instead, the one or more second nearest neighbor points of the same point searched in the update stage can be directly used as the first nearest neighbor points.
[0206] Step 9032: Determine the attribute reconstruction value of the current point based on the attribute prediction value and the attribute residual value of the current point.
[0207] In step 9032, the attribute residual value is the inverse quantized attribute residual value, and the point cloud decoder can add the attribute prediction value of the current point to the attribute residual value to obtain the attribute reconstruction value of the current point.
[0208] This embodiment of the present application provides a PT encoding method, which includes the following steps 1201 to 1213:
[0209] Step 1201, constructing the LOD of the current point cloud;
[0210] Optionally, in some embodiments, the method by which the point cloud decoder constructs the LOD of the current point cloud is the same as the method by which the above-mentioned point cloud encoder constructs the LOD of the current point cloud. Specifically, the construction of the LOD can be implemented with reference to the method shown in Figure 7, which will not be repeated here.
[0211] Step 1202: determining a current refinement layer, where the current refinement layer refers to a refinement layer of the current point in the LOD of the current point cloud;
[0212] Step 1203: Determine N first nearest neighbor points of the current point from at least one first refinement level of the LOD of the current point cloud, the encoded points of the current refinement level, and at least one second refinement level of the LOD of the reference point cloud.
[0213] The decoding order of the at least one first refinement layer is before the current refinement layer, the predefined distance based on which the second refinement layer is generated is equal to the predefined distance based on which the first refinement layer is generated, and the predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated.
[0214] Step 1204 , fitting the mapping relationship between the attribute values of at least some points in the current point cloud and the attribute reconstruction values of at least some points in the reference point cloud to obtain weighted prediction parameters; wherein the weighted prediction parameters include weighting factors and / or offsets.
[0215] Step 1205: writing the weighted prediction parameters into the bitstream for transmission to the point cloud decoder;
[0216] Optionally, in some embodiments, the point cloud encoder may write the weighted prediction parameter into the bitstream before writing the attribute residual value of the first point of the current point cloud into the bitstream. Subsequently, for the current point cloud, the weighted prediction parameter may not be written into the bitstream again.
[0217] Step 1206 , compensating the attribute reconstruction value of the time domain reference point among the N first nearest neighboring points of the current point according to the weighted prediction parameter to obtain a time domain reference attribute value;
[0218] Step 1207: Determine a candidate prediction value for the current point based on the temporal reference attribute value of the temporal reference point among the N first nearest neighboring points and the attribute reconstruction value of the spatial reference point among the N first nearest neighboring points.
[0219] Step 1208 , determining the candidate prediction value with the minimum rate-distortion cost among the candidate prediction values of the current point as the attribute prediction value of the current point;
[0220] Step 1209: determining a prediction mode for the current point according to the index value of the candidate prediction value with the minimum rate-distortion cost among the candidate prediction values for the current point;
[0221] Step 1210, generating a bitstream according to the prediction mode of the current point;
[0222] Step 1211, determining the attribute residual value of the current point based on the attribute value and the attribute prediction value of the current point;
[0223] Step 1212, quantizing the attribute residual value of the current point to obtain a quantized residual;
[0224] Step 1213: performing arithmetic coding on the quantized residual of the current point to generate a bit stream;
[0225] Step 1214: Dequantize the quantized residual of the current point to obtain a dequantized residual.
[0226] Step 1215 : Determine the attribute reconstruction value of the current point based on the inverse quantized residual of the current point and the attribute prediction value of the current point.
[0227] This embodiment of the present application provides a PT decoding method, which includes the following steps 1301 to 1311:
[0228] Step 1301, constructing the LOD of the current point cloud;
[0229] Step 1302: determining a current refinement layer, where the current refinement layer refers to a refinement layer of the current point in the LOD of the current point cloud;
[0230] Step 1303: Determine N first nearest neighbor points of the current point from at least one first refinement level of the LOD of the current point cloud, the encoded points of the current refinement level, and at least one second refinement level of the LOD of the reference point cloud.
[0231] The decoding order of the at least one first refinement layer is before the current refinement layer, the predefined distance based on which the second refinement layer is generated is equal to the predefined distance based on which the first refinement layer is generated, and the predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated.
[0232] Step 1304: decode the code stream and determine weighted prediction parameters; wherein the weighted prediction parameters include weighting factors and / or offsets.
[0233] Optionally, in some embodiments, the weighted prediction parameter is determined by decoding the code stream before decoding the attribute residual value of the first point of the current point cloud.
[0234] Step 1305 , compensating the attribute reconstruction value of the time domain reference point among the N first nearest neighboring points of the current point according to the weighted prediction parameter to obtain a time domain reference attribute value;
[0235] Step 1306: Determine a candidate prediction value for the current point based on the temporal reference attribute value of the temporal reference point among the N first nearest neighboring points and the attribute reconstruction value of the spatial reference point among the N first nearest neighboring points.
[0236] Step 1307: Decode the code stream and determine the prediction mode of the current point;
[0237] Step 1308: determining the attribute prediction value of the current point according to the prediction mode and the candidate prediction value of the current point;
[0238] Step 1309: Decode the code stream and determine the attribute residual value of the current point;
[0239] Step 1310: Dequantize the quantized residual of the current point to obtain a dequantized residual.
[0240] Step 1311 : Determine the attribute reconstruction value of the current point based on the inverse quantized residual of the current point and the attribute prediction value of the current point.
[0241] This embodiment of the present application provides an LT encoding method, which includes the following steps 1401 to 1413:
[0242] Step 1401, constructing the LOD of the current point cloud;
[0243] Step 1402: determining a current refinement layer, where the current refinement layer refers to a refinement layer of the current point in the LOD of the current point cloud;
[0244] Step 1403: Determine N first nearest neighbor points of the current point from at least one first refinement layer of the current point cloud, the encoded points of the current refinement layer, and at least one second refinement layer of the LOD of the reference point cloud;
[0245] The decoding order of the at least one first refinement layer is before the current refinement layer, the predefined distance based on which the second refinement layer is generated is equal to the predefined distance based on which the first refinement layer is generated, and the predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated.
[0246] Step 1404: Fit the mapping relationship between the attribute values of at least some points in the current point cloud and the attribute reconstruction values of at least some points in the reference point cloud to obtain weighted prediction parameters; wherein the weighted prediction parameters include weighting factors and / or offsets.
[0247] Step 1405: writing the weighted prediction parameters into the bitstream for transmission to the point cloud decoder;
[0248] Optionally, in some embodiments, the point cloud encoder may write the weighted prediction parameter into the bitstream before writing the attribute residual value of the first point of the current point cloud into the bitstream. Subsequently, for the current point cloud, the weighted prediction parameter may not be written into the bitstream again.
[0249] Step 1406 , compensating the attribute reconstruction value of the time domain reference point among the N first nearest neighboring points of the current point according to the weighted prediction parameter to obtain a time domain reference attribute value;
[0250] Step 1407: Determine an attribute prediction value of the current point based on a weighted average of the temporal reference attribute values of the temporal reference points among the N first nearest neighboring points and the attribute reconstruction values of the spatial reference points among the N first nearest neighboring points.
[0251] Step 1408, determining the attribute residual value of the current point based on the attribute value and the attribute prediction value of the current point;
[0252] Step 1409, quantizing the attribute residual value of the current point to obtain a quantized residual;
[0253] Step 1410: performing arithmetic coding on the quantized residual of the current point to generate a bitstream;
[0254] Step 1411, dequantize the quantized residual of the current point to obtain a dequantized residual;
[0255] Step 1412: Determine the attribute reconstruction value of the current point based on the inverse quantized residual of the current point and the attribute prediction value of the current point.
[0256] Step 1413: Update the attribute reconstruction value of the spatial reference point according to the attribute residual value of the current point.
[0257] This embodiment of the present application provides an LT decoding method, which includes the following steps 1501 to 1513:
[0258] Step 1501, constructing the LOD of the current point cloud;
[0259] Step 1502: determining a current refinement layer, where the current refinement layer refers to a refinement layer of the current point in the LOD of the current point cloud;
[0260] Step 1503: Determine at least one first refinement layer whose decoding order is before the current refinement layer; wherein the at least one first refinement layer belongs to the LOD of the current point cloud, and the predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated.
[0261] Step 1504: Decode the code stream to determine the attribute residual value of the point in the current refinement layer;
[0262] Step 1505: Determine N second nearest neighbor points of the i-th point of the current refinement layer from the at least one first refinement layer and the at least one second refinement layer of the LOD of the reference point cloud; wherein i is greater than 0 and less than or equal to the total number of points in the current refinement layer; the predefined distance based on which the at least one second refinement layer is generated is equal to the predefined distance based on which the at least one first refinement layer is generated; and the predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated.
[0263] Step 1506: Update the attribute reconstruction value of the point belonging to the second refinement layer among the one or more second nearest neighbor points based on the attribute residual value of the i-th point; i+1, return to execute steps 1505 and 1506 until step 1506 is completed based on the attribute residual value of the last point in the current refinement layer, and then proceed to step 1507;
[0264] Step 1507 , determining N first nearest neighbor points of the current point from the at least one first refinement layer and the at least one second refinement layer of the LOD of the reference point cloud;
[0265] Step 1508: Decode the bitstream and determine weighted prediction parameters;
[0266] Optionally, in some embodiments, the weighted prediction parameter is determined by decoding the code stream before decoding the attribute residual value of the first point of the current point cloud.
[0267] Step 1509 , compensating the attribute reconstruction values of the points belonging to the reference point cloud (i.e., the time-domain reference points) among the N first nearest neighbor points according to the weighted prediction parameters to obtain time-domain reference attribute values;
[0268] Step 1510: Determine an attribute prediction value of the current point based on a weighted average of the temporal reference attribute values of the temporal reference points among the N first nearest neighboring points and the attribute reconstruction values of the spatial reference points among the N first nearest neighboring points.
[0269] Step 1511: decode the code stream and determine the attribute residual value of the current point;
[0270] Step 1512: Dequantize the quantized residual of the current point to obtain a dequantized residual.
[0271] Step 1513, determine the attribute reconstruction value of the current point based on the inverse quantized residual of the current point and the attribute prediction value of the current point; return to execute steps 1507 to step 1513 to determine the attribute reconstruction value of the next point of the current refined layer, until the attribute reconstruction value of each point in the current refined layer is obtained, and return to execute step 1502 to determine the attribute reconstruction value of each point in the next refined layer.
[0272] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0273] The embodiment of the present application provides a GPCC inter-frame attribute prediction optimization and quality enhancement technology based on weighted prediction.
[0274] It is understandable that in the current GPCC inter-frame attribute prediction method, there is no compensation operation for the attribute reconstruction values of neighboring points from the reference point cloud, and the fluctuation of attribute mean values between different frames is not taken into account. However, the ratio of attribute mean values between adjacent frames of the current CTC dataset multi-frame point cloud is close to 1, which indicates that the multi-frame point cloud in the CTC dataset may have been collected in a relatively ideal environment, and the impact of factors such as the collection environment and changes in lighting conditions during the collection process on point cloud attributes is not fully considered. In order to improve the existing technology, it is necessary to consider how to process point cloud data under different lighting conditions to improve the accuracy and robustness of attribute prediction.
[0275] Based on this, in the embodiment of the present application, the attribute reconstruction values of the neighboring points from the reference point cloud are weighted, and the weighted attribute reconstruction values are used for subsequent operations. The implementation process of the core method is shown in Figure 12, where A i (i=1,2,3) are the attribute reconstruction values of the three nearest neighboring points of the current point. The coefficients q and w are weighted prediction parameters obtained by fitting the attribute values of each point in the current point cloud and the attribute reconstruction values of each point in the reference point cloud. The attribute reconstruction values of the nearest neighboring points from the reference point cloud are weighted according to q and w to compensate for the prediction error caused by the fluctuation of the attribute mean between frames.
[0276] The implementation details of the encoding end are described below.
[0277] Embodiment 5: Predicting Transform (PT) encoding of point cloud attribute information specifically includes the following steps 1 to 4.
[0278] Step 1, LOD generation, is the same as the LOD construction process corresponding to Figure 4 above, so it will not be repeated here.
[0279] Step 2, optimal prediction value selection:
[0280] After the LOD is constructed, according to the LOD generation order, the three nearest neighbor points of the current point are first found from the encoded data points of the current point cloud and the reference point cloud (i.e., an example of the N first nearest neighbor points). This technical solution compensates the nearest neighbor points from the reference point cloud, and the specific operations are as follows:
[0281] 1) Remember a i (i=0,1,…,n) is the attribute value of the i-th point in the current point cloud, where n is the number of points in the current point cloud; a_ref j (j=0,1,…,m) is the attribute reconstruction value of the jth point in the reference point cloud, where m is the number of points in the reference point cloud. q and w are obtained by fitting according to formula (16): i =q×a_ref i +w (16);
[0282] 2) Write q and w into the code stream and transmit it to the decoding end.
[0283] 3) Reconstruct the attribute values of the three nearest neighbor points of each point in the current point cloud Compensation is performed, j represents the index of any point of the three nearest neighbor points, Represents the attribute reconstruction value of the nearest neighbor point j: Determine whether the nearest neighbor point j comes from the reference point cloud. If so, Weighted compensation to obtain the reconstructed attribute value If not, then
[0284] Reconstruct the attributes of these three nearest neighbors into temp j , as the candidate prediction value of the current point; then, the best candidate prediction value is selected from them according to the rate-distortion optimization (RDO) as the attribute prediction value of the current point. For example, when encoding the attribute value of point P2 in Figure 4, the attribute reconstruction value temp of the nearest neighbor point P4 is used. P4 As a candidate prediction value, its index is set to 1; the attribute reconstruction value temp of the next nearest neighbor point P5_ref (ie P5_ref comes from the reference point cloud) and the three nearest neighbor points P0 P5_ref and temp P0 As candidate prediction values, their indexes are set to 2 and 3 respectively; the attribute reconstruction values temp of points P0, P5_ref and P4 are P0 , temp P5_ref and temp P4 The weighted average of is taken as a candidate prediction value, and its index is set to 0, as shown in Table 3; finally, RDO is used to select the best candidate prediction value as the attribute prediction value of the current point. The weighted average formula is shown in the following formula (17):
[0285] Where, represents the spatial geometric weight from the nearest neighbor point j to the current point i, and is determined by the following formula (18):
[0286] Represents the attribute prediction value of the current point i, j represents the index of any point of the three nearest neighbor points, temp j Represents the attribute reconstruction value of the nearest neighbor point j, x i ,y i ,z i is the geometric position coordinate of the current point i, x ij ,y ij ,z ij is the geometric coordinate of the nearest neighbor point j.
[0287] Table 3 Samples of candidate prediction items for attribute coding
[0288] In Table 3, the so-called first nearest neighbor point refers to the point closest to the current point among the three nearest neighbor points, the so-called second nearest neighbor point refers to the point second closest to the current point among the three first nearest neighbors, and so on.
[0289] Step 3, attribute prediction residual and quantification:
[0290] The attribute prediction value of the current point i is obtained through the above prediction (k is the total number of points in the point cloud). Let (a i ) i∈0…k-1 is the original attribute value of the current point, then the attribute residual value of the current point i (r i ) i∈0…k-1 It is expressed as the following formula (19):
[0291] Furthermore, the attribute residual value is quantified according to the following formula (20):
[0292] Where Q i It represents the quantized attribute residual value of the current point i, and Qs is the quantization step (Qs), which can be calculated by the quantization parameter QP (QP) specified by CTC.
[0293] Step 4: The encoding end reconstructs the attribute value:
[0294] The purpose of reconstruction at the encoding end is to predict the subsequent points. Before reconstructing the attribute value, the residual Q is calculated according to the following formula (21): iDequantize, remember is the residual after inverse quantization:
[0295] and attribute prediction values Add up to get the attribute reconstruction value of point i For example, formula (22):
[0296] Example 6: The implementation details of the decoding end corresponding to the PT encoding end include the following steps 1 to 4, which are described as follows:
[0297] Step 1, LOD generation, is the same as the LOD construction process corresponding to Figure 4 above, so it will not be repeated here.
[0298] Step 2, optimal prediction value selection:
[0299] After the LOD is constructed, according to the LOD generation order, the three nearest neighbor points of the current point are first found from the encoded data points of the current point cloud and the reference point cloud. This technical solution compensates the nearest neighbor points from the reference point cloud. The specific operations are as follows:
[0300] 1) parse q and w from the code stream;
[0301] 2) Reconstruct the attributes of the three nearest neighbor points of each point in the current point cloud Compensation is performed, j represents the index of any point of the three nearest neighbor points, Represents the attribute reconstruction value of the nearest neighbor point j: Determine the nearest neighbor point Is it from the reference point cloud? If so, weight it to get the compensated attribute reconstruction value If not, then
[0302] Reconstruct the attributes of these three nearest neighbors into temp j and their weighted average as the candidate prediction value of the current point; then, the best prediction mode is parsed from the bitstream. The attribute prediction value of the current point is obtained based on the best prediction mode, where the weighted average is determined as shown in the following formula (23):
[0303] In the formula represents the spatial geometric weight from the nearest neighbor point j to the current point i, and is determined by the following formula (24):
[0304] Represents the attribute prediction value of the current point i, j represents the index of any point of the three nearest neighbor points, temp jRepresents the attribute reconstruction value of the nearest neighbor point j, x i ,y i ,z i is the geometric position coordinate of the current point i, x ij ,y ij ,z ij is the geometric coordinate of the neighboring point j.
[0305] Step 3, residual decoding:
[0306] Parse the attribute residual value Q from the code stream i , and dequantize the attribute residual value obtained by the analysis according to the following formula (25), record is the attribute residual value after inverse quantization:
[0307] Step 4: The decoder reconstructs the attribute value:
[0308] and attribute prediction values Add up to get the attribute reconstruction value of point i As shown in the following formula (26):
[0309] Example 7: Lifting Transform (LT) encoding of point cloud attribute information
[0310] Figure 5 shows the encoding process of the lifting transform. The lifting transform also predicts and encodes point cloud attributes based on LOD. The difference from the predictive transform is that the lifting transform first divides the LOD into high and low layers, predicts in the reverse order of the LOD generation layer, and introduces an update operator during the prediction process to update the quantized weights of the points in the low-level LOD to improve the accuracy of the prediction. This is because the attribute values of the points in the low-level LOD are frequently used to predict the attribute values of the points in the high-level LOD, and the points in the low-level LOD should have a greater influence.
[0311] In the seventh embodiment, the LT encoding process includes the following steps 1 to 3:
[0312] Step 1, segmentation process:
[0313] The segmentation process is to divide the complete LOD layer into a low LOD layer L(N) and a high LOD layer H(N), where L(N) is generated based on the Euclidean distance d l Is greater than the Euclidean distance based on which H(N) is generated. If a point cloud has three levels of LOD, that is After segmentation, (R l ) l=1,2 It is the high LOD layer, denoted as H(N), (R l )l=0 It is the low LOD layer, denoted as L(N).
[0314] Step 2, prediction process:
[0315] The points in the high LOD layer select the attribute information of the nearest neighbor points from the low LOD layers of the current point cloud and the reference point cloud, and determine the attribute prediction value P(N) of the current point based on this. The attribute prediction residual D(N) is shown in the following formula (27): D(N) = H(N) - P(N) (27);
[0316] This technical solution compensates for neighboring points from the reference point cloud. The specific operations are as follows:
[0317] 1) Remember a i (i=0,1,…,n) is the attribute value of the i-th point in the current point cloud, where n is the number of points in the current point cloud; a_ref j (j=0,1,…,m) is the attribute reconstruction value of the jth point in the reference point cloud, where m is the number of points in the reference point cloud. q and w are obtained by fitting according to formula (28): i =q×a_ref i +w (28);
[0318] 2) Write q (i.e., weighting factor) and w (i.e., offset) into the bitstream and transmit it to the decoding end.
[0319] 3) Reconstruct the attribute values of the three nearest neighbor points of each point in the current point cloud Compensation is performed, j represents the index of any point of the three nearest neighbor points, Represents the attribute reconstruction value of the nearest neighbor point j: Determine whether the nearest neighbor point j comes from the reference point cloud. If so, Weighted compensation to obtain the reconstructed attribute value If not, then
[0320] 4) Update the prediction process. LT coding is different from PT coding. LT coding has only one prediction mode, that is, the attribute prediction value of the current point is equal to the weighted average of the attribute reconstruction values of the three nearest neighboring points. The formula is as follows (29):
[0321] In the formula represents the spatial geometric weight from the nearest neighbor point j to the current point i, and is calculated as shown in the following formula (30):
[0322] Step 3, Update process:
[0323] Update the attribute prediction residual D(N) in the high-level LOD to obtain U(N), and use U(N) to improve the attribute value of the midpoint of the low-level LOD, as shown in formula (31): L′(N)=L(N)+U(N) (31);
[0324] As shown in FIG5 , the above process will iterate continuously until the lowest LOD according to the order of LOD from high to low.
[0325] Example 8: The implementation details of the decoding end corresponding to the LT encoding end include the following steps 1 to 3:
[0326] The lifting transformation at the decoding end is the inverse process of the encoding end. The LOD layer is divided into high and low layers, and updated and predicted from low to high.
[0327] Step 1, segmentation process:
[0328] The segmentation process is to divide the complete LOD layer into a low LOD layer L(N) and a high LOD layer H(N), where L(N) is generated based on the Euclidean distance d l Is greater than the Euclidean distance based on which H(N) is generated. If a point cloud has three levels of LOD, that is After segmentation, (R l ) l=1,1 It is the high LOD layer, denoted as H(N), (R l ) l=0 It is the low LOD layer, denoted as L(N).
[0329] Step 2, Update process:
[0330] First, the residual D(N) is parsed from the bitstream and dequantized. The dequantized D(N) is updated to obtain U(N). The attribute reconstruction value of the midpoint of the lower layer LOD is updated with U(N), as shown in formula (32): R'0(N) = R0(N) - U(N) (32);
[0331] Step 3, prediction process:
[0332] The point in R1(N) selects the attribute information of the nearest neighbor point from the current point cloud and the reference point cloud R′0(N) as the attribute prediction value P(N) of the current point, and obtains the attribute reconstruction value R1(N), as shown in formula (33): R1(N)=D(N)+P(N) (33);
[0333] This technical solution compensates the nearest neighbor points from the reference point cloud. The specific operations are as follows:
[0334] 1) parse q and w from the code stream;
[0335] 2) Reconstruct the attribute values of the three nearest neighbor points of each point in the current point cloud Compensation is performed, j represents the index of any point of the three nearest neighbor points, Represents the attribute reconstruction value of the nearest neighbor point j: Determine whether the nearest neighbor point j comes from the reference point cloud. If so, Weighted compensation to obtain the reconstructed attribute value If not, then
[0336] Reconstruct the attributes of these three nearest neighbors into temp j The attribute prediction value is obtained by weighting, and the weighted average formula is shown in (34):
[0337] In the formula represents the spatial geometric weight from the nearest neighbor point j to the current point i, and is calculated as shown in the following formula (35):
[0338] Represents the attribute prediction value of the current point i, j represents the index of any point of the three nearest neighbor points, temp j Represents the attribute reconstruction value of the nearest neighbor point j, x i ,y i ,z i is the geometric position coordinate of the current point i, x ij ,y ij ,z ij is the geometric coordinate of the nearest neighbor point j.
[0339] The above process will iterate continuously until the highest LOD according to the order of LOD from low to high.
[0340] Since the average properties of all seven point cloud sequences in the Cat3 dataset specified by CTC are almost identical between each sequence frame, it is impossible to verify the effectiveness of this technical solution, nor can it reflect the shortcomings of existing technical solutions. Therefore, based on the Cat3 dataset, this technical solution randomly linearly weights the reflectivity of each frame of the point cloud to obtain a modified dataset. The specific operation is as follows: Taking the ford_01 sequence as an example, 32 consecutive frames of point cloud frames are randomly selected from the original point cloud sequence. A random number is added to the original reflectivity of each frame of the point cloud, and no other information is changed. This is the modified point cloud. The same process is performed for the other six point cloud sequences.
[0341] The proposed technical solution was implemented on the G-PCC reference software TMC13V23-rc1 and tested under the CTC-C1 / C2 / CW / CY test conditions. The test results are shown in Figures 13 to 16. It can be seen that the technical solution has gains under all CTC test conditions, and the time complexity remains unchanged or slightly reduced. Under the C1 / C2 / CW / CY test conditions, gains of -1.1%, -12.1%, 94.7%, and -1.3% can be obtained, respectively.
[0342] Condition C1 refers to lossless geometry and lossy attribute coding, condition C2 refers to lossy geometry and lossy attribute coding, condition CW refers to lossless geometry and lossless attribute coding, and condition CY refers to lossless geometry and near-lossless attribute coding. In the figure, End-to-End BD-AttrRate represents the BD-Rate of the end-to-end attribute values for the attribute bitstream, End-to-End BD-TotalRate represents the BD-Rate of the end-to-end attribute values for the total bitstream, and Geom.End-to-End BD-TotalRate represents the BD-Rate of the end-to-end geometry for the total bitstream. The BD-Rate reflects the difference in PSNR curves between the two scenarios. A decrease in BD-Rate indicates improved performance with a reduced bitrate while maintaining the same PSNR; conversely, performance decreases. In other words, the greater the decrease in BD-Rate, the better the compression effect. BPIP represents the number of bits consumed per input point. D1 and D2 represent geometric distortion, with the former representing point-to-point distortion and the latter representing point-to-plane distortion. Cat3-frame average represents the average of the test results for each point cloud sequence in the dataset.
[0343] It can be understood that this scheme proposes a GPCC inter-frame attribute prediction optimization and quality enhancement technology based on weighted prediction, which effectively solves the prediction inaccuracy problem caused by the reflectivity fluctuation of multi-frame point cloud data sets. By fitting the linear relationship between the point cloud attributes of the current point cloud and the reference point cloud, the attribute reconstruction values of the nearest neighbor points in the reference point cloud are weighted according to this linear relationship, which improves the coding efficiency and realizes the quality enhancement of the reconstructed point cloud at the decoding end.
[0344] It should be noted that although the steps of the method of the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all steps must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps; or steps in different embodiments may be combined to form a new technical solution.
[0345] Based on the above embodiments, the present application provides a decoding device for use in a point cloud decoder. FIG17 is a schematic structural diagram of the decoding device provided in the present application. As shown in FIG17 , the decoding device 170 includes:
[0346] A first determining module 1701 is configured to determine a reference point of a current point in the current point cloud based on geometric information of the current point cloud; wherein the reference point includes a temporal reference point, which is one or more points in a reference point cloud; and the reference point cloud is point cloud data that has been decoded before decoding the current point cloud.
[0347] The first processing module 1702 is configured to process the attribute value of the time domain reference point according to the weighted prediction parameter to obtain a time domain reference attribute value; wherein the weighted prediction parameter is obtained by decoding the bitstream;
[0348] The second determining module 1703 is configured to determine the attribute reconstruction value of the current point according to the time domain reference attribute value of the time domain reference point.
[0349] In some embodiments, the weighted prediction parameter is used to characterize a mapping relationship between attribute values of at least some points in the current point cloud and attribute reconstructed values of at least some points in the reference point cloud.
[0350] In some embodiments, the weighted prediction parameter is a fitting coefficient obtained by fitting a mapping relationship between attribute values of at least some points in the current point cloud and attribute reconstruction values of at least some points in the reference point cloud.
[0351] In some embodiments, the weighted prediction parameters include weighting factors and / or offsets.
[0352] In some embodiments, the weighted prediction parameter is determined by decoding the code stream before decoding the attribute residual value of the first point of the current point cloud.
[0353] In some embodiments, the attribute value of the time domain reference point is processed according to the weighted prediction parameter to obtain the time domain reference attribute value, including: determining a first product of the weighting factor and the attribute reconstruction value of the time domain reference point; offsetting the first product according to the offset to obtain the time domain reference attribute value.
[0354] In some embodiments, determining the reference point of the current point in the current point cloud based on the geometric information of the current point cloud includes: determining a current refinement layer, where the current refinement layer refers to a refinement layer of the current point in the LOD of the current point cloud; determining the reference point of the current point from at least one first refinement layer of the LOD of the current point cloud, the decoded points of the current refinement layer and / or at least one second refinement layer of the LOD of the reference point cloud; wherein the decoding order of the at least one first refinement layer is before the current refinement layer, and the predefined distance based on which the second refinement layer is generated is equal to the predefined distance based on which the first refinement layer is generated.
[0355] In some embodiments, the predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated.
[0356] In some embodiments, the reference point is one or more first nearest neighbor points of the current point, and the one or more first nearest neighbor points are from the current point cloud and / or the reference point cloud.
[0357] In some embodiments, the decoding device 170 also includes a decoding module, which is configured to: decode the code stream to obtain the attribute residual value of the current point; determine the attribute reconstruction value of the current point based on the time domain reference attribute value of the time domain reference point, including: determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point; determine the attribute reconstruction value of the current point based on the attribute prediction value and attribute residual value of the current point.
[0358] In some embodiments, the construction process of the LOD of the current point cloud includes: obtaining L predefined distances, L is greater than 1; marking all points in the current point cloud as unvisited, and setting the visited point set to an empty set; executing a first process based on the lth predefined distance, l is greater than or equal to 0 and less than or equal to L; the first process includes: traversing all points, and ignoring if the current point has been visited; if the current point has not been visited, determining the minimum distance between the current point and the points in the visited point set; if the minimum distance is less than the lth predefined distance, ignoring the current point; if the minimum distance is greater than or equal to the lth predefined distance, marking the current point as visited, and adding the current point to the lth refinement layer and the visited point set; executing the first process based on the l+1th predefined distance until the Lth refinement layer corresponding to the Lth predefined distance is obtained, taking the union of the refinement layers corresponding to the L predefined distances respectively, and obtaining the LOD of the current point cloud.
[0359] In some embodiments, determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point includes: determining the candidate prediction value of the current point based on the time domain reference attribute value of the time domain reference point; decoding the code stream to determine the prediction mode of the current point; and determining the attribute prediction value of the current point based on the prediction mode and the candidate prediction value of the current point.
[0360] In some embodiments, the reference point also includes a spatial domain reference point, which is one or more decoded points in the current point cloud; determining the candidate prediction value of the current point based on the time domain reference attribute value of the time domain reference point includes: determining the candidate prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0361] In some embodiments, the candidate prediction value of the current point includes the time domain reference attribute value of the time domain reference point, the weighted average of the time domain reference attribute value of the time domain reference point, the attribute reconstruction value of the spatial domain reference point and / or the weighted average of the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0362] In some embodiments, the attribute prediction value of the current point is equal to the candidate prediction value corresponding to the prediction mode.
[0363] In some embodiments, the reference point also includes a spatial domain reference point, which is one or more decoded points in the current point cloud; determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point includes: determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0364] In some embodiments, determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point includes: determining the attribute prediction value of the current point based on the weighted average of the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0365] In some embodiments, the attribute prediction value of the current point is equal to the weighted average value.
[0366] In some embodiments, the decoding device 170 also includes a first update module, configured to: before determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point, update the attribute reconstruction value of the point of at least one first refined layer according to the attribute residual value of the point of the current refined layer.
[0367] In some embodiments, updating the attribute reconstruction values of the points of the at least one first refinement layer based on the attribute residual values of the points of the current refinement layer includes: performing a second process based on the i-th point of the current refinement layer; the second process includes: determining one or more second nearest neighboring points of the i-th point of the current refinement layer from the at least one first refinement layer and at least one second refinement layer of the LOD of the reference point cloud; wherein i is greater than 0 and less than or equal to the total number of points of the current refinement layer; updating the attribute reconstruction values of the points belonging to the first refinement layer among the one or more second nearest neighboring points based on the attribute residual value of the i-th point; and iterating the second process based on the i+1-th point of the current refinement layer until the second process is performed based on the last point of the current refinement layer.
[0368] In some embodiments, the at least one first refinement layer includes a refinement layer generated based on a maximum distance in the LOD of the current point cloud, and an initial attribute reconstruction value of a point in the refinement layer generated based on the maximum distance is equal to a first value.
[0369] The description of the above decoding device embodiment is similar to the description of the above decoding method embodiment, and has similar beneficial effects as the decoding method embodiment. For technical details not disclosed in the decoding device embodiment of this application, please refer to the description of the decoding method embodiment of this application for understanding.
[0370] Based on the aforementioned embodiments, an embodiment of the present application provides an encoding device, which is applied to a point cloud encoder. Figure 18 is a structural schematic diagram of the encoding device provided by the embodiment of the present application. As shown in Figure 18, the encoding device 180 includes: a third determination module 1801, configured to determine the reference point of the current point in the current point cloud based on the geometric information of the current point cloud; wherein the reference point includes a time domain reference point, and the time domain reference point is one or more points in the reference point cloud; the reference point cloud is point cloud data that has been decoded before decoding the current point cloud; an acquisition module 1802, configured to obtain weighted prediction parameters; wherein the weighted prediction parameters are determined based on the attribute values of at least some points of the current point cloud and the attribute reconstruction values of at least some points of the reference point cloud; a second processing module 1803, configured to process the attribute values of the time domain reference points according to the weighted prediction parameters to obtain time domain reference attribute values; a fourth determination module 1804, configured to determine the attribute residual value of the current point based on the time domain reference attribute values of the time domain reference points.
[0371] In some embodiments, the encoding device 180 further includes an encoding module configured to generate a bit stream according to the weighted prediction parameters.
[0372] In some embodiments, generating a bitstream based on the weighted prediction parameters includes: writing the weighted prediction parameters into the bitstream before writing the attribute residual value of the first point of the current point cloud into the bitstream.
[0373] In some embodiments, the weighted prediction parameter is used to characterize a mapping relationship between attribute values of at least some points in the current point cloud and attribute reconstructed values of at least some points in the reference point cloud.
[0374] In some embodiments, the weighted prediction parameter is a fitting coefficient obtained by fitting a mapping relationship between attribute values of at least some points in the current point cloud and attribute reconstruction values of at least some points in the reference point cloud.
[0375] In some embodiments, the weighted prediction parameters include weighting factors and / or offsets.
[0376] In some embodiments, the attribute value of the time domain reference point is processed according to the weighted prediction parameter to obtain the time domain reference attribute value, including: determining a first product of the weighting factor and the attribute reconstruction value of the time domain reference point; offsetting the first product according to the offset to obtain the time domain reference attribute value.
[0377] In some embodiments, determining the reference point of the current point in the current point cloud based on the geometric information of the current point cloud includes: determining a current refinement layer, where the current refinement layer refers to a refinement layer of the current point in the LOD of the current point cloud; determining the reference point of the current point from at least one first refinement layer of the LOD of the current point cloud, the encoded points of the current refinement layer and / or at least one second refinement layer of the LOD of the reference point cloud; wherein the encoding order of the at least one first refinement layer is before the current refinement layer, and the predefined distance based on which the second refinement layer is generated is equal to the predefined distance based on which the first refinement layer is generated.
[0378] In some embodiments, the predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated.
[0379] In some embodiments, the reference point is one or more first nearest neighbor points of the current point, and the one or more first nearest neighbor points are from the current point cloud and / or the reference point cloud.
[0380] In some embodiments, determining the attribute residual value of the current point based on the time domain reference attribute value of the time domain reference point includes: determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point; determining the attribute residual value of the current point based on the attribute value and attribute prediction value of the current point.
[0381] In some embodiments, the encoding module is further configured to: generate a code stream according to the attribute residual value of the current point.
[0382] In some embodiments, determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point includes: determining the candidate prediction value of the current point based on the time domain reference attribute value of the time domain reference point; and determining the attribute prediction value of the current point based on the rate-distortion cost of the candidate prediction value of the current point.
[0383] In some embodiments, the reference point also includes a spatial domain reference point, which is one or more encoded points in the current point cloud; determining the candidate prediction value of the current point based on the time domain reference attribute value of the time domain reference point includes: determining the candidate prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0384] In some embodiments, the candidate prediction value of the current point includes the time domain reference attribute value of the time domain reference point, the weighted average of the time domain reference attribute value of the time domain reference point, the attribute reconstruction value of the spatial domain reference point and / or the weighted average of the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0385] In some embodiments, the attribute prediction value of the current point is equal to the candidate prediction value with the minimum rate-distortion cost.
[0386] In some embodiments, the fourth determination module 1804 is further configured to determine the prediction mode of the current point based on the index value of the candidate prediction value with the smallest rate-distortion cost in the candidate prediction values of the current point; the encoding module is further configured to generate a code stream based on the prediction mode.
[0387] In some embodiments, the reference point also includes a spatial reference point, which is one or more encoded points in the current point cloud; determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point includes: determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial reference point.
[0388] In some embodiments, determining the attribute prediction value of the current point based on the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point includes: determining the attribute prediction value of the current point based on the weighted average of the time domain reference attribute value of the time domain reference point and the attribute reconstruction value of the spatial domain reference point.
[0389] In some embodiments, the attribute prediction value of the current point is equal to the weighted average value.
[0390] In some embodiments, the encoding device 180 further includes a second updating module configured to update the attribute reconstruction value of the spatial reference point according to the attribute residual value of the current point after determining the attribute residual value of the current point.
[0391] The description of the above encoding device embodiment is similar to the description of the above encoding method embodiment, and has similar beneficial effects as the encoding method embodiment. For technical details not disclosed in the encoding device embodiment of this application, please refer to the description of the encoding method embodiment of this application for understanding.
[0392] It should be noted that the division of modules in the device described in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation. In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or they can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. It can also be implemented in the form of a combination of software and hardware.
[0393] It should be noted that, in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0394] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, it implements an encoding method such as a point cloud encoder side, or implements a decoding method such as a point cloud decoder side.
[0395] An embodiment of the present application provides a point cloud decoder. As shown in FIG19 , the point cloud decoder 190 includes: a first communication interface 1901, a first memory 1902, and a first processor 1903; each component is coupled together through a first bus system 1904. It can be understood that the first bus system 1904 is used to achieve connection and communication between these components. In addition to the data bus, the first bus system 1904 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as the first bus system 1904 in FIG19. Among them,
[0396] The first communication interface 1901 is used to receive and send signals when sending and receiving information with other external network elements;
[0397] A first memory 1902 is used to store computer programs that can be run on the first processor 1903;
[0398] The first processor 1903 is configured to execute the decoding method described in the embodiment of the present application when running the computer program.
[0399] It is understood that the first memory 1902 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The first memory 1902 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0400] The first processor 1903 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the first processor 1903. The above-mentioned first processor 1903 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the first memory 1902 , and the first processor 1903 reads the information in the first memory 1902 and completes the steps of the above method in combination with its hardware.
[0401] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0402] Optionally, as another embodiment, the first processor 1903 is further configured to execute any of the aforementioned decoding method embodiments when running the computer program.
[0403] The present application implements a point cloud encoder, as shown in FIG20 , the point cloud encoder 200 includes: a second communication interface 2001, a second memory 2002, and a second processor 2003; each component is coupled together through a second bus system 2004. It can be understood that the second bus system 2004 is used to achieve connection and communication between these components. In addition to the data bus, the second bus system 2004 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as the second bus system 2004 in FIG20. Among them,
[0404] The second communication interface 2001 is used for sending and receiving signals during the process of sending and receiving information with other external network elements;
[0405] The second memory 2002 is used to store computer programs that can be run on the second processor 2003;
[0406] The second processor 2003 is configured to execute the encoding method described in the embodiment of the present application when running the computer program.
[0407] Optionally, as another embodiment, the second processor 2003 is further configured to execute any of the aforementioned encoding method embodiments when running the computer program.
[0408] It can be understood that the hardware functions of the second memory 2002 are similar to those of the first memory 1902, and the hardware functions of the second processor 2003 are similar to those of the first processor 1903; they will not be described in detail here.
[0409] An embodiment of the present application provides an electronic device. Figure 21 is a structural diagram of the electronic device provided by the embodiment of the present application. As shown in Figure 21, the electronic device 210 includes a memory 2101 and a processor 2102. The memory 2101 stores a computer program that can be run on the processor 2102. When the processor 2102 executes the program, the steps in the method provided in the above embodiment are implemented.
[0410] It should be noted that the memory 2101 is configured to store instructions and applications executable by the processor 2102, and can also cache data to be processed or processed by the processor 2102 and various modules in the electronic device 210 (for example, image data, audio data, voice communication data and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0411] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method provided in the above embodiment are implemented.
[0412] An embodiment of the present application provides a computer program product containing instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.
[0413] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0414] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, they will not be repeated here.
[0415] The term "and / or" in this article is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, object A and / or object B can mean: object A exists alone, object A and object B exist at the same time, and object B exists alone.
[0416] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0417] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.
[0418] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed across multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of this embodiment.
[0419] In addition, all functional modules in the embodiments of the present application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the above-mentioned integrated modules can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0420] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.
[0421] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks or optical disks.
[0422] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined, if they do not conflict, to obtain new method embodiments. The features disclosed in the several product embodiments provided in this application can be arbitrarily combined, if they do not conflict, to obtain new product embodiments. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined, if they do not conflict, to obtain new method embodiments or device embodiments.
[0423] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A decoding method, which is applied to a point cloud decoder, and the method includes: Determining a reference point of a current point in the current point cloud according to geometric information of the current point cloud; wherein, the reference point includes a time-domain reference point, and the time-domain reference point is one or more points in a reference point cloud; the reference point cloud is point cloud data that has been decoded before decoding the current point cloud; Processing attribute values of the time-domain reference points according to weighted prediction parameters to obtain time-domain reference attribute values; wherein, the weighted prediction parameters are obtained by decoding a code stream; Determining an attribute reconstruction value of the current point according to the time-domain reference attribute values of the time-domain reference points.
2. The method according to claim 1, wherein, The weighted prediction parameters are used to represent a mapping relationship between attribute values of at least some points in the current point cloud and attribute reconstruction values of at least some points in the reference point cloud.
3. The method according to claim 2, wherein, The weighted prediction parameters are fitting coefficients obtained by fitting a mapping relationship between attribute values of at least some points in the current point cloud and attribute reconstruction values of at least some points in the reference point cloud.
4. The method according to claim 3, wherein The weighted prediction parameters include a weighting factor and / or an offset.
5. The method according to any one of claims 1-4, wherein, The weighted prediction parameters are determined by decoding the code stream before decoding an attribute residual value of the first point of the current point cloud.
6. The method according to claim 4, wherein, The processing the attribute values of the time-domain reference points according to the weighted prediction parameters to obtain time-domain reference attribute values includes: Determining a first product of the weighting factor and the attribute reconstruction value of the time-domain reference point; Offsetting the first product according to the offset to obtain the time-domain reference attribute value.
7. The method according to any one of claims 1-6, wherein, The determining the reference point of the current point in the current point cloud according to the geometric information of the current point cloud includes: Determining a current refinement layer, where the current refinement layer refers to a refinement layer of the current point in the LOD of the current point cloud; Determining the reference point of the current point from at least one first refinement layer of the LOD of the current point cloud, decoded points of the current refinement layer, and / or at least one second refinement layer of the LOD of the reference point cloud; wherein, the decoding order of the at least one first refinement layer is before the current refinement layer, and a predefined distance based on which the second refinement layer is generated is correspondingly equal to a predefined distance based on which the first refinement layer is generated.
8. The method according to claim 7, wherein, The predefined distance based on which the at least one first refinement layer is generated is greater than the predefined distance based on which the current refinement layer is generated.
9. The method according to claim 7, wherein The reference point is one or more first nearest neighbor points of the current point, and the one or more first nearest neighbor points come from the current point cloud and / or the reference point cloud.
10. The method according to any one of claims 7-9, wherein, The method further includes: decoding the code stream to obtain an attribute residual value of the current point; The determining the attribute reconstruction value of the current point according to the time-domain reference attribute values of the time-domain reference points includes: Determining an attribute prediction value of the current point according to the time-domain reference attribute values of the time-domain reference points; Determining the attribute reconstruction value of the current point according to the attribute prediction value and the attribute residual value of the current point.
11. The method according to claim 7, wherein, The construction process of the LOD of the current point cloud includes: Obtaining L predefined distances, where L>1; Mark all points in the current point cloud as unvisited, and set the set of visited points to an empty set; Execute a first process based on the l-th predefined distance, where l is greater than or equal to 0 and less than or equal to L; the first process includes: traversing all points, if the current point has been visited, ignore it; if the current point has not been visited, determine the minimum distance between the current point and the points in the set of visited points; if the minimum distance is less than the l-th predefined distance, ignore the current point; if the minimum distance is greater than or equal to the l-th predefined distance, mark the current point as visited, and add the current point to the l-th refinement layer and the set of visited points; Execute the first process based on the (l + 1)-th predefined distance until the L-th refinement layer corresponding to the L-th predefined distance is obtained, and take the union of the refinement layers corresponding to the L predefined distances respectively to obtain the LOD of the current point cloud.
12. The method according to claim 10, wherein The determining the attribute prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point includes: Determine the candidate prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point; Decode the bitstream to determine the prediction mode of the current point; Determine the attribute prediction value of the current point according to the prediction mode and the candidate prediction value of the current point.
13. The method according to claim 12, wherein, The reference point further includes a spatial-domain reference point, and the spatial-domain reference point is one or more decoded points in the current point cloud; The determining the candidate prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point includes: Determine the candidate prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point and the attribute reconstruction value of the spatial-domain reference point.
14. The method according to claim 12 or 13, wherein The candidate prediction value of the current point includes the time-domain reference attribute value of the time-domain reference point, the weighted average value of the time-domain reference attribute values of the time-domain reference point, the attribute reconstruction value of the spatial-domain reference point, and / or the weighted average value of the time-domain reference attribute value of the time-domain reference point and the attribute reconstruction value of the spatial-domain reference point.
15. The method according to any one of claims 12 - 14, wherein, The attribute prediction value of the current point is equal to the candidate prediction value corresponding to the prediction mode.
16. The method according to claim 10, wherein, The reference point further includes a spatial-domain reference point, and the spatial-domain reference point is one or more decoded points in the current point cloud; The determining the attribute prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point includes: Determine the attribute prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point and the attribute reconstruction value of the spatial-domain reference point.
17. The method according to claim 16, wherein The determining the attribute prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point and the attribute reconstruction value of the spatial-domain reference point includes: Determine the attribute prediction value of the current point according to the weighted average value of the time-domain reference attribute value of the time-domain reference point and the attribute reconstruction value of the spatial-domain reference point.
18. The method according to claim 17, wherein, The attribute prediction value of the current point is equal to the weighted average value.
19. The method according to claim 16, wherein Before determining the attribute prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point and the attribute reconstruction value of the spatial-domain reference point, the method further includes: Update the attribute reconstruction values of the points in the at least one first refinement layer according to the attribute residual values of the points in the current refinement layer.
20. The method according to claim 19, wherein The updating the attribute reconstruction values of the points in the at least one first refinement layer according to the attribute residual values of the points in the current refinement layer includes: Perform a second process according to the i-th point in the current refinement layer; the second process includes: Determine one or more second nearest neighbor points of the i-th point in the current refinement layer from the at least one first refinement layer and at least one second refinement layer of the LOD of the reference point cloud; where i is greater than 0 and less than or equal to the total number of points in the current refinement layer; Update the attribute reconstruction values of the points belonging to the first refinement layer among the one or more second nearest neighbor points according to the attribute residual value of the i-th point; Iterate the second process according to the (i + 1)-th point in the current refinement layer until the second process is performed according to the last point in the current refinement layer.
21. The method according to claim 20, wherein, The at least one first refinement layer includes a refinement layer generated based on the maximum distance in the LOD of the current point cloud, and the initial attribute reconstruction value of the points in the refinement layer generated based on the maximum distance is equal to the first value.
22. An encoding method, which is applied to a point cloud encoder, and the method includes: Determine a reference point of the current point in the current point cloud according to the geometric information of the current point cloud; where the reference point includes a time-domain reference point, and the time-domain reference point is one or more points in the reference point cloud; the reference point cloud is point cloud data that has been encoded before encoding the current point cloud; Obtain a weighted prediction parameter; where the weighted prediction parameter is determined based on the attribute values of at least some points in the current point cloud and the attribute reconstruction values of at least some points in the reference point cloud; Process the attribute value of the time-domain reference point according to the weighted prediction parameter to obtain a time-domain reference attribute value; Determine the attribute residual value of the current point according to the time-domain reference attribute value of the time-domain reference point.
23. The method according to claim 22, wherein The method further includes: Generate a bitstream according to the weighted prediction parameter.
24. The method according to claim 23, wherein, The generating a bitstream according to the weighted prediction parameter includes: Before writing the attribute residual value of the first point of the current point cloud into the bitstream, write the weighted prediction parameter into the bitstream.
25. The method according to any one of claims 22-24, wherein, The weighted prediction parameter is used to characterize the mapping relationship between the attribute values of at least some points in the current point cloud and the attribute reconstruction values of at least some points in the reference point cloud.
26. The method according to claim 25, wherein The weighted prediction parameter is a fitting coefficient obtained by fitting the mapping relationship between the attribute values of at least some points in the current point cloud and the attribute reconstruction values of at least some points in the reference point cloud.
27. The method according to claim 25 or 26, wherein, The weighted prediction parameter includes a weighting factor and / or an offset.
28. The method according to claim 27, wherein The processing the attribute value of the time-domain reference point according to the weighted prediction parameter to obtain a time-domain reference attribute value includes: Determine a first product of the weighting factor and the attribute reconstruction value of the time-domain reference point; Offset the first product according to the offset to obtain the time-domain reference attribute value.
29. The method according to any one of claims 22-28, wherein, The determining a reference point of the current point in the current point cloud according to the geometric information of the current point cloud includes: Determine the current refinement level, where the current refinement level refers to a refinement level of the current point in the LOD of the current point cloud; Determine the reference point of the current point from at least one first refinement level of the LOD of the current point cloud, the encoded points of the current refinement level, and / or at least one second refinement level of the LOD of the reference point cloud; wherein the encoding order of the at least one first refinement level is before the current refinement level, and the predefined distance based on which the second refinement level is generated is correspondingly equal to the predefined distance based on which the first refinement level is generated.
30. The method according to claim 29, wherein The predefined distance based on which the at least one first refinement level is generated is greater than the predefined distance based on which the current refinement level is generated.
31. The method according to claim 29 or 30, wherein, The construction process of the LOD of the current point cloud includes: Obtain L predefined distances, where L is greater than 1; Mark all points in the current point cloud as unvisited, and set the set of visited points to be an empty set; Execute a first process based on the l-th predefined distance, where l is greater than or equal to 0 and less than or equal to L; the first process includes: traverse all points, if the current point has been visited, ignore it; if the current point has not been visited, determine the minimum distance between the current point and the points in the set of visited points; if the minimum distance is less than the l-th predefined distance, ignore the current point; if the minimum distance is greater than or equal to the l-th predefined distance, mark the current point as visited, and add the current point to the l-th refinement level and the set of visited points; Execute the first process based on the (l + 1)-th predefined distance until the L-th refinement level corresponding to the L-th predefined distance is obtained, and take the union of the refinement levels corresponding to the L predefined distances to obtain the LOD of the current point cloud.
32. The method according to claim 29, wherein The reference point is one or more first nearest neighbor points of the current point, and the one or more first nearest neighbor points come from the current point cloud and / or the reference point cloud.
33. The method according to any one of claims 22-32, wherein, The determining the attribute residual value of the current point according to the time-domain reference attribute value of the time-domain reference point includes: Determine the attribute prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point; Determine the attribute residual value of the current point according to the attribute value and the attribute prediction value of the current point.
34. The method according to claim 33, wherein, The method further includes: Generate a bitstream according to the attribute residual value of the current point.
35. The method according to claim 33, wherein, The determining the attribute prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point includes: Determine the candidate prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point; Determine the attribute prediction value of the current point according to the rate-distortion cost of the candidate prediction value of the current point.
36. The method according to claim 35, wherein, The reference point further includes a spatial-domain reference point, and the spatial-domain reference point is one or more encoded points in the current point cloud; The determining the candidate prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point includes: Determine the candidate prediction value of the current point according to the time-domain reference attribute value of the time-domain reference point and the attribute reconstruction value of the spatial-domain reference point.
37. The method according to claim 35 or 36, wherein, The candidate prediction value of the current point includes the temporal reference attribute value of the temporal reference point, the weighted average value of the temporal reference attribute values of the temporal reference points, the attribute reconstruction value of the spatial reference point, and / or the weighted average value of the temporal reference attribute value of the temporal reference point and the attribute reconstruction value of the spatial reference point.
38. The method according to claim 35 or 36, wherein The attribute prediction value of the current point is equal to the candidate prediction value with the minimum rate-distortion cost.
39. The method according to claim 38, wherein, The method further includes: Determining the prediction mode of the current point according to the index value of the candidate prediction value with the minimum rate-distortion cost among the candidate prediction values of the current point; Generating a bitstream according to the prediction mode.
40. The method according to claim 33, wherein, The reference point further includes a spatial reference point, and the spatial reference point is one or more encoded points in the current point cloud; The determining the attribute prediction value of the current point according to the temporal reference attribute value of the temporal reference point includes: Determining the attribute prediction value of the current point according to the temporal reference attribute value of the temporal reference point and the attribute reconstruction value of the spatial reference point.
41. The method according to claim 40, wherein, The determining the attribute prediction value of the current point according to the temporal reference attribute value of the temporal reference point and the attribute reconstruction value of the spatial reference point includes: Determining the attribute prediction value of the current point according to the weighted average value of the temporal reference attribute value of the temporal reference point and the attribute reconstruction value of the spatial reference point.
42. The method according to claim 41, wherein, The attribute prediction value of the current point is equal to the weighted average value.
43. The method according to claim 40, wherein, After determining the attribute residual value of the current point, the method further includes: updating the attribute reconstruction value of the spatial reference point according to the attribute residual value of the current point.
44. A decoding device, applied to a point cloud decoder, the device includes: A first determination module, configured to determine the reference point of the current point in the current point cloud according to the geometric information of the current point cloud; wherein, the reference point includes a temporal reference point, and the temporal reference point is one or more points in the reference point cloud; the reference point cloud is the point cloud data that has been decoded before decoding the current point cloud; A first processing module, configured to process the attribute value of the temporal reference point according to a weighted prediction parameter to obtain a temporal reference attribute value; wherein, the weighted prediction parameter is obtained by decoding the bitstream; A second determination module, configured to determine the attribute reconstruction value of the current point according to the temporal reference attribute value of the temporal reference point.
45. A point cloud decoder, including a first memory and a first processor; wherein, The first memory is used to store a computer program that can run on the first processor; The first processor is used to execute the method according to any one of claims 1 to 21 when running the computer program.
46. An encoding device, applied to a point cloud encoder, the device includes: A third determination module, configured to determine the reference point of the current point in the current point cloud according to the geometric information of the current point cloud; wherein, the reference point includes a temporal reference point, and the temporal reference point is one or more points in the reference point cloud; the reference point cloud is the point cloud data that has been decoded before decoding the current point cloud; An acquisition module configured to acquire weighted prediction parameters; wherein the weighted prediction parameters are determined based on the attribute values of at least some points of the current point cloud and the attribute reconstruction values of at least some points of the reference point cloud; A second processing module configured to process the attribute values of the time-domain reference points according to the weighted prediction parameters to obtain time-domain reference attribute values; A fourth determination module configured to determine the attribute residual value of the current point according to the time-domain reference attribute values of the time-domain reference points.
47. A point cloud encoder, comprising a second memory and a second processor; wherein, The second memory is used to store a computer program that can run on the second processor; The second processor is configured to execute the method according to any one of claims 22 to 43 when running the computer program.
48. A bitstream obtained by any one of claims 22 to 43.
49. An electronic device, comprising: A processor adapted to execute a computer program; A computer-readable storage medium storing a computer program, which when executed by the processor, implements the method according to any one of claims 1 to 21, or when the computer program is executed by the processor, implements the method according to any one of claims 22 to 43.
50. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, which when executed, implements the method according to any one of claims 1 to 21 or implements the method according to any one of claims 22 to 43.
Citation Information
Patent Citations
Point cloud geometrical information inter-frame encoding and decoding method
CN112565764A
Point cloud attribute encoding method, device and system
CN113179410A
Coding method, device, decoding method, device, equipment and readable storage medium
CN113766229A
Point cloud sequence coding and decoding method and device based on two-dimensional regularization plane projection
CN114915791A
Point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device
US20240062428A1