Point cloud coding and decoding method, device, equipment and storage medium

CN120303940APending Publication Date: 2025-07-11GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380078227.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-01-06
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing technology fails to effectively utilize inter-frame information during the point cloud encoding process, resulting in low encoding and decoding performance of the point cloud.

Method used

In the point cloud encoding and decoding process, N prediction nodes of the current node are determined in the prediction reference frame of the current frame to be decoded, and the coordinate information of the current node is predicted and decoded based on the geometric encoding and decoding information of these prediction nodes, taking into account the relative Temporal correlation between adjacent frames.

Benefits of technology

It improves the efficiency of encoding and decoding point cloud geometric information and improves the transmission performance of point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303940A_ABST
    Figure CN120303940A_ABST
Patent Text Reader

Abstract

The invention provides a point cloud coding and decoding method and device, equipment and a storage medium, and the method comprises the steps: determining N prediction nodes of a current node in a prediction reference frame of a current to-be-coded and decoded frame, and carrying out the prediction coding and decoding of the coordinate information of a midpoint of the current node based on the geometric coding and decoding information of the midpoint of the N prediction nodes. That is, according to the embodiment of the invention, optimization is carried out when DCM direct coding and decoding are carried out on the nodes, the correlation between adjacent frames on the time domain is considered, and the geometric information of the prediction nodes in the prediction reference frame is utilized to carry out prediction coding and decoding on the geometric information of the points in the IDCM nodes (namely the current nodes) of the to-be-processed nodes. And the geometric information coding and decoding efficiency of the point cloud is further improved by considering the time domain correlation between adjacent frames.
Need to check novelty before this filing date? Find Prior Art

Description

Point cloud encoding and decoding method, device, equipment and storage medium Technical Field

[0001] The present application relates to the field of point cloud technology, and in particular to a point cloud encoding and decoding method, apparatus, device and storage medium. Background Art

[0002] Capturing the surface of an object using a capture device creates point cloud data, which can contain hundreds of thousands or even more points. During video production, this point cloud data is transmitted between the point cloud encoding device and the point cloud decoding device in the form of point cloud media files. However, such a large number of points poses a challenge to transmission, so the point cloud encoding device must compress the point cloud data before transmission.

[0003] Point cloud compression is also known as point cloud encoding. During the point cloud encoding process, using infer direct mode coding (IDCM) can significantly reduce complexity for points that are isolated in geometric space. When using direct coding to encode and decode the current node, the geometric information of the point in the current node is directly encoded. However, this current encoding of the geometric information of the point in the current node does not take into account inter-frame information, which can reduce point cloud encoding and decoding performance.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a point cloud encoding and decoding method, apparatus, device, and storage medium, which take into account inter-frame information when encoding the geometric information of the node midpoint, thereby improving the encoding and decoding performance of the point cloud.

[0006] In a first aspect, an embodiment of the present application provides a point cloud decoding method, comprising:

[0007] Determine N prediction nodes of a current node in a prediction reference frame of a current frame to be decoded, where the current node is a node to be decoded in the current frame to be decoded, and N is a positive integer;

[0008] Based on the geometric decoding information of the N predicted nodes, the coordinate information of the midpoint of the current node is predicted and decoded.

[0009] In a second aspect, the present application provides a point cloud encoding method, comprising:

[0010] Determine N prediction nodes of a current node in a prediction reference frame of a current frame to be encoded, where the current node is a node to be encoded in the current frame to be encoded, and N is a positive integer;

[0011] Based on the geometric coding information of the N predicted nodes, the coordinate information of the midpoint of the current node is predictively coded.

[0012] In a third aspect, the present application provides a point cloud decoding device for executing the method of the first aspect or its respective implementations. Specifically, the device includes a functional unit for executing the method of the first aspect or its respective implementations.

[0013] In a fourth aspect, the present application provides a point cloud encoding device for executing the method of the second aspect or its respective implementations. Specifically, the device includes a functional unit for executing the method of the second aspect or its respective implementations.

[0014] In a fifth aspect, a point cloud decoder is provided, comprising a processor and a memory. The memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the first aspect or its respective implementations.

[0015] In a sixth aspect, a point cloud encoder is provided, comprising a processor and a memory. The memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the second aspect or its respective implementations.

[0016] In a seventh aspect, a point cloud encoding and decoding system is provided, comprising a point cloud encoder and a point cloud decoder. The point cloud decoder is configured to execute the method of the first aspect or its respective implementations, and the point cloud encoder is configured to execute the method of the second aspect or its respective implementations.

[0017] In an eighth aspect, a chip is provided for implementing the method described in any one of the first and second aspects above, or their respective implementations. Specifically, the chip includes a processor configured to load and execute a computer program from a memory, causing a device equipped with the chip to perform the method described in any one of the first and second aspects above, or their respective implementations.

[0018] In a ninth aspect, a computer-readable storage medium is provided for storing a computer program, which enables a computer to execute the method of any one of the first to second aspects or their respective implementations.

[0019] In a tenth aspect, a computer program product is provided, comprising computer program instructions, which enable a computer to execute the method of any one of the first to second aspects or their respective implementations.

[0020] In an eleventh aspect, a computer program is provided, which, when executed on a computer, enables the computer to execute the method in any one of the first to second aspects or their respective implementations.

[0021] In a twelfth aspect, a code stream is provided. The code stream is generated based on the method of the second aspect. Optionally, the code stream includes at least one of a first parameter and a second parameter.

[0022] Based on the above technical solution, when decoding the current node in the current codec frame, in the predicted reference frame of the current frame to be coded, N predicted nodes of the current node are determined, and based on the geometric coding and decoding information of the midpoints of these N predicted nodes, the coordinate information of the midpoint of the current node is predicted and coded. In other words, the embodiment of the present application optimizes the direct DCM coding and decoding of the node, and by considering the correlation in the time domain between adjacent frames, the geometric information of the predicted node in the predicted reference frame is used to predict and code the geometric information of the midpoint of the IDCM node (i.e., the current node) of the to-be-coded node, and by considering the time domain correlation between adjacent frames, the efficiency of coding and decoding the geometric information of the point cloud is further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1A is a schematic diagram of a point cloud;

[0024] Figure 1B is a partial enlarged view of the point cloud;

[0025] FIG2 is a schematic diagram of six viewing angles of a point cloud image;

[0026] FIG3 is a schematic block diagram of a point cloud encoding and decoding system according to an embodiment of the present application;

[0027] FIG4A is a schematic block diagram of a point cloud encoder provided in an embodiment of the present application;

[0028] FIG4B is a schematic block diagram of a point cloud decoder provided in an embodiment of the present application;

[0029] FIG5A is a schematic plan view;

[0030] FIG5B is a schematic diagram of node coding sequence;

[0031] FIG5C is a schematic diagram of a plane mark;

[0032] Figure 5D is a schematic diagram of sibling nodes;

[0033] Figure 5E is a schematic diagram of the intersection of the laser radar and the node;

[0034] FIG5F is a schematic diagram of neighborhood nodes at the same partition depth and the same coordinates;

[0035] FIG5G is a schematic diagram of neighboring nodes when the node is located at a lower plane position of the parent node;

[0036] FIG5H is a schematic diagram of neighboring nodes when the node is located at a high plane position of the parent node;

[0037] FIG5I is a schematic diagram of predictive coding of planar position information of a laser radar point cloud;

[0038] FIG6A is a schematic diagram of IDCM encoding;

[0039] FIG6B is a schematic diagram of coordinate transformation of a point cloud acquired by a rotating laser radar;

[0040] FIG6C is a schematic diagram of predictive coding in the X or Y axis direction;

[0041] FIG6D is a schematic diagram showing the angle of the X or Y plane predicted by the horizontal azimuth angle;

[0042] FIG6E is a schematic diagram of predictive coding of the X or Y axis;

[0043] 7A to 7C are schematic diagrams of geometric information encoding based on triangular facets;

[0044] FIG8A is a schematic diagram of an AVS coding framework;

[0045] FIG8B is a schematic diagram of the decoding framework of AVS;

[0046] FIG9A is a schematic diagram of reference nodes selected by each sub-node;

[0047] FIG9B is a schematic diagram of four groups of reference neighbor nodes of the current node;

[0048] FIG9C is a schematic diagram showing that sub-blocks correspond to six adjacent parent blocks;

[0049] FIG9D is a schematic diagram of 18 adjacent blocks used by the current block to be encoded and their Morton sequence numbers;

[0050] FIG9E is a simplified diagram of a prediction tree;

[0051] FIG10 is a schematic diagram of a point cloud decoding method according to an embodiment of the present application;

[0052] FIG11 is a schematic diagram of an octree partition;

[0053] FIG12 is a schematic diagram of a prediction node;

[0054] FIG13 is a schematic diagram of a domain node;

[0055] Figure 14 is a schematic diagram of the corresponding nodes of the domain node;

[0056] FIG15A is a schematic diagram of a predicted node of a current node in a predicted reference frame;

[0057] FIG15B is a schematic diagram of the predicted nodes of the current node in two predicted reference frames;

[0058] FIG16A is a schematic diagram of IDCM encoding;

[0059] FIG16B is a schematic diagram of IDCM decoding;

[0060] FIG17 is a schematic diagram of a point cloud encoding method according to an embodiment of the present application;

[0061] FIG18 is a schematic block diagram of a point cloud decoding device provided in an embodiment of the present application;

[0062] FIG19 is a schematic block diagram of a point cloud encoding device provided in an embodiment of the present application;

[0063] FIG20 is a schematic block diagram of an electronic device provided in an embodiment of the present application;

[0064] Figure 21 is a schematic block diagram of the point cloud encoding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] The present application can be applied to the field of point cloud upsampling technology, for example, it can be applied to the field of point cloud compression technology.

[0066] To facilitate understanding of the embodiments of the present application, the following briefly introduces the relevant concepts involved in the embodiments of the present application:

[0067] A point cloud is a set of irregularly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or scene. Figure 1A is a schematic diagram of a 3D point cloud image, and Figure 1B is a zoomed-in view of Figure 1A. As can be seen from Figures 1A and 1B, the point cloud surface is composed of densely distributed points.

[0068] 2D images contain information at every pixel, and their distribution is regular, so there's no need to record their location. However, the distribution of points in a point cloud in 3D space is random and irregular, so recording the location of every point in space is necessary to fully represent a point cloud. Similar to 2D images, each location in the data collection process has corresponding attribute information.

[0069] Point cloud data is a specific record format for point clouds. Points in a point cloud can include both their location information and attribute information. For example, the location information of a point can be its 3D coordinate information. This information can also be referred to as its geometric information. For example, the attribute information of a point can include color information, reflectance information, normal vector information, and so on. Color information reflects the color of an object, while reflectance information reflects the surface material of the object. The color information can be information in any color space. For example, the color information can be in RGB. Another example is luminance and chrominance (YCbCr, YUV) information. For example, Y represents luminance (Luma), Cb (U) represents blue color difference, Cr (V) represents red, and U and V represent chroma (Chroma) to describe color difference information. For example, a point cloud obtained using laser measurement principles can include both its 3D coordinate information and its laser reflection intensity (reflectance). Another example is a point cloud obtained using photogrammetry principles, which can include both its 3D coordinate information and its color information. For example, a point cloud is obtained by combining the principles of laser measurement and photogrammetry. The points in the point cloud may include the three-dimensional coordinate information of the point, the laser reflection intensity (reflectance) of the point, and the color information of the point. Figure 2 shows a point cloud image, where Figure 2 shows six viewing angles of the point cloud image. Table 1 shows the point cloud data storage format consisting of a file header information part and a data part:

[0070] Table 1

[0071]

[0072] In Table 1, the header information includes the data format, data representation type, the total number of point cloud points, and the content represented by the point cloud. For example, the point cloud in this example is in the ".ply" format, represented by ASCII code, with a total number of 207242 points. Each point has three-dimensional position information XYZ and three-dimensional color information RGB.

[0073] Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes. Moreover, since point clouds are obtained by directly sampling real objects, they can provide a strong sense of reality while ensuring accuracy. Therefore, they are widely used, including virtual reality games, computer-aided design, geographic information systems, automatic navigation systems, digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive remote presentation, and three-dimensional reconstruction of biological tissues and organs.

[0074] The ways to obtain point cloud data may include but are not limited to at least one of the following: (1) generation by computer equipment. Computer equipment can generate point cloud data based on virtual three-dimensional objects and virtual three-dimensional scenes. (2) 3D (3-Dimension) laser scanning acquisition. 3D laser scanning can obtain point cloud data of static real-world three-dimensional objects or three-dimensional scenes, and millions of point cloud data can be obtained per second; (3) 3D photogrammetry acquisition. 3D photography equipment (i.e., a group of cameras or camera equipment with multiple lenses and sensors) is used to collect real-world visual scenes to obtain point cloud data of real-world visual scenes. 3D photography can obtain point cloud data of dynamic real-world three-dimensional objects or three-dimensional scenes. (4) Point cloud data of biological tissues and organs can be obtained through medical equipment. In the medical field, point cloud data of biological tissues and organs can be obtained through medical equipment such as magnetic resonance imaging (MRI), computed tomography (CT), and electromagnetic positioning information.

[0075] Point clouds can be divided into dense point clouds and sparse point clouds according to the acquisition method.

[0076] Point clouds are divided into the following types according to the time series of the data:

[0077] The first type of static point cloud: the object is stationary and the device used to obtain the point cloud is also stationary;

[0078] The second type of dynamic point cloud: the object is moving, but the device that obtains the point cloud is stationary;

[0079] The third type of dynamic point cloud acquisition: the device that acquires the point cloud is moving.

[0080] Point clouds are divided into two categories according to their uses:

[0081] Category 1: Machine perception point cloud, which can be used in scenarios such as autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, and disaster relief robots;

[0082] Category 2: Human eye perception point cloud, which can be used in point cloud application scenarios such as digital cultural heritage, free viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.

[0083] The aforementioned point cloud acquisition technologies reduce the cost and time required to acquire point cloud data, while improving data accuracy. This evolution in point cloud data acquisition has made it possible to acquire large amounts of point cloud data. However, as application demands grow, the processing of massive amounts of 3D point cloud data is facing bottlenecks due to storage space and transmission bandwidth limitations.

[0084] Taking a point cloud video with a frame rate of 30 fps (frames per second) as an example, each frame contains 700,000 points, each with coordinate information (xyz, float) and color information (RGB, uchar). Therefore, the data volume of a 10-second point cloud video is approximately 0.7 million points x (4 bytes x 3 + 1 byte x 3) x 30 fps x 10 seconds = 3.15 GB. For a 1280 x 720 two-dimensional video with a YUV sampling format of 4:2:0 and a frame rate of 24 fps, the data volume for 10 seconds is approximately 1280 x 720 x 12 bits x 24 frames x 10 seconds, which is approximately 0.33 GB. A 10-second two-view 3D video has a data volume of approximately 0.33 x 2 = 0.66 GB. Therefore, the data volume of a point cloud video far exceeds that of a 2D or 3D video of the same length. Therefore, point cloud compression has become a key issue in promoting the development of the point cloud industry to better manage data, save server storage space, and reduce the transmission traffic and time between the server and client.

[0085] The following introduces the relevant knowledge of point cloud encoding and decoding.

[0086] Figure 3 is a schematic block diagram of a point cloud encoding and decoding system involved in an embodiment of the present application. It should be noted that Figure 3 is only an example, and the point cloud encoding and decoding system of the embodiment of the present application includes but is not limited to that shown in Figure 3. As shown in Figure 3, the point cloud encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compression) the point cloud data to generate a code stream, and transmit the code stream to the decoding device. The decoding device decodes the code stream generated by the encoding device to obtain decoded point cloud data.

[0087] The encoding device 110 of the embodiment of the present application can be understood as a device with a point cloud encoding function, and the decoding device 120 can be understood as a device with a point cloud decoding function, that is, the embodiment of the present application includes a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, point cloud game consoles, vehicle-mounted computers, etc.

[0088] In some embodiments, the encoding device 110 may transmit the encoded point cloud data (such as a code stream) to the decoding device 120 via the channel 130. The channel 130 may include one or more media and / or devices capable of transmitting the encoded point cloud data from the encoding device 110 to the decoding device 120.

[0089] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded point cloud data directly to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded point cloud data according to a communication standard and transmit the modulated point cloud data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media can also include wired communication media, such as one or more physical transmission lines.

[0090] In another example, channel 130 includes a storage medium that can store the point cloud data encoded by encoding device 110. The storage medium includes various locally accessible data storage media, such as optical disks, DVDs, and flash memory. In this example, decoding device 120 can retrieve the encoded point cloud data from the storage medium.

[0091] In another example, the channel 130 may include a storage server that can store the point cloud data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded point cloud data from the storage server. Alternatively, the storage server can store the encoded point cloud data and transmit the encoded point cloud data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.

[0092] In some embodiments, the encoding device 110 includes a point cloud encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.

[0093] In some embodiments, the encoding device 110 may further include a point cloud source 111 in addition to the point cloud encoder 112 and the input interface 113 .

[0094] The point cloud source 111 may include at least one of a point cloud acquisition device (e.g., a scanner), a point cloud archive, a point cloud input interface, and a computer graphics system, wherein the point cloud input interface is used to receive point cloud data from a point cloud content provider, and the computer graphics system is used to generate point cloud data.

[0095] The point cloud encoder 112 encodes the point cloud data from the point cloud source 111 to generate a code stream. The point cloud encoder 112 transmits the encoded point cloud data directly to the decoding device 120 via the output interface 113. The encoded point cloud data can also be stored on a storage medium or storage server for subsequent reading by the decoding device 120.

[0096] In some embodiments, the decoding device 120 includes an input interface 121 and a point cloud decoder 122 .

[0097] In some embodiments, the decoding device 120 may further include a display device 123 in addition to the input interface 121 and the point cloud decoder 122 .

[0098] The input interface 121 includes a receiver and / or a modem and can receive the encoded point cloud data via the channel 130 .

[0099] The point cloud decoder 122 is used to decode the encoded point cloud data to obtain decoded point cloud data, and transmit the decoded point cloud data to the display device 123.

[0100] The decoded point cloud data is displayed on the display device 123. The display device 123 may be integrated with the decoding device 120 or external to the decoding device 120. The display device 123 may include various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0101] In addition, Figure 3 is only an example, and the technical solution of the embodiment of the present application is not limited to Figure 3. For example, the technology of the present application can also be applied to unilateral point cloud encoding or unilateral point cloud decoding.

[0102] Current point cloud encoders can use two point cloud compression coding technology routes proposed by the Moving Picture Experts Group (MPEG) of the International Organization for Standardization: Video-based Point Cloud Compression (VPCC) and Geometry-based Point Cloud Compression (GPCC). VPCC projects a 3D point cloud onto a 2D image and uses existing 2D coding tools to encode the projected 2D image. GPCC uses a hierarchical structure to divide the point cloud into multiple units, encoding the entire point cloud by recording the division process.

[0103] The following uses the GPCC encoding and decoding framework as an example to illustrate the point cloud encoder and point cloud decoder applicable to the embodiments of the present application.

[0104] Figure 4A is a schematic block diagram of the point cloud encoder provided in an embodiment of the present application.

[0105] As can be seen from the above, points in a point cloud can include both their location information and their attribute information. Therefore, the encoding of points in a point cloud mainly includes location encoding and attribute encoding. In some examples, the location information of points in a point cloud is also called geometric information, and the corresponding location encoding of points in the point cloud can also be called geometric encoding.

[0106] In the GPCC coding framework, the geometric information of the point cloud and the corresponding attribute information are encoded separately.

[0107] As shown in Figure 4A below, the current geometric coding and decoding of G-PCC can be divided into octree-based geometric coding and decoding and prediction tree-based geometric coding and decoding.

[0108] The position encoding process involves preprocessing the points in the point cloud, such as coordinate transformation, quantization, and duplicate point removal. Next, geometric encoding is performed on the preprocessed point cloud, such as constructing an octree or prediction tree. Based on the constructed octree or prediction tree, geometric encoding is performed to form a geometric bitstream. Simultaneously, the position information of each point in the point cloud data is reconstructed based on the position information output by the constructed octree or prediction tree, resulting in a reconstructed value for each point's position information.

[0109] The attribute encoding process includes: given the reconstruction information of the input point cloud position information and the original value of the attribute information, selecting one of the three prediction modes for point cloud prediction, quantizing the predicted result, and performing arithmetic coding to form an attribute code stream.

[0110] As shown in Figure 4A, position encoding can be achieved through the following units:

[0111] Coordinate conversion (Tanmsform coordinates) unit 201, voxel (Voxelize) unit 202, octree partition (Analyze octree) unit 203, geometry reconstruction (Reconstruct geometry) unit 204, arithmetic encoding (Arithmetic enconde) unit 205, surface fitting unit (Analyze surface approximation) 206 and prediction tree construction unit 207.

[0112] The coordinate conversion unit 201 can be used to convert the world coordinates of a point in the point cloud into relative coordinates. For example, the geometric coordinates of the point are subtracted from the minimum value of the x, y, and z coordinate axes, which is equivalent to a DC removal operation, to convert the coordinates of the point in the point cloud from world coordinates to relative coordinates.

[0113] Voxelize unit 202, also known as the quantize and remove points unit, reduces the number of coordinates through quantization. After quantization, previously different points may be assigned the same coordinates. Based on this, duplicate points can be removed through deduplication. For example, multiple clouds with the same quantized position but different attribute information can be merged into a single cloud through attribute conversion. In some embodiments of the present application, voxel unit 202 is an optional unit module.

[0114] The octree partitioning unit 203 may encode the quantized point position information using an octree encoding scheme. For example, the point cloud may be partitioned using an octree, so that point positions correspond one-to-one with octree positions. Geometric encoding is performed by counting the point positions in the octree and setting their flags to 1.

[0115] In some embodiments, in the geometric information encoding process based on a triangle soup (trisoup), the point cloud is also octree-partitioned by the octree partitioning unit 203. However, unlike the geometric information encoding based on the octree, the trisoup does not need to divide the point cloud into unit cubes with a side length of 1X1X1 step by step. Instead, the division is stopped when the block (sub-block) has a side length of W. Based on the surface formed by the distribution of the point cloud in each block, at most twelve vertices (intersections) generated by the surface and the twelve edges of the block are obtained. The intersections are surface fitted by the surface fitting unit 206, and the fitted intersections are geometrically encoded.

[0116] The prediction tree construction unit 207 can encode the quantized point position information using a prediction tree encoding method. For example, the point cloud is divided into a prediction tree, so that the point positions correspond one-to-one with the positions of the nodes in the prediction tree. By counting the positions of the points in the prediction tree, different prediction modes are selected to predict the geometric position information of the nodes to obtain prediction residuals, and the geometric prediction residuals are quantized using quantization parameters. Finally, through continuous iteration, the prediction residuals of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary bitstream.

[0117] The geometric reconstruction unit 204 can perform position reconstruction based on the position information output by the octree partitioning unit 203 or the intersection points fitted by the surface fitting unit 206 to obtain a reconstructed value of the position information of each point in the point cloud data. Alternatively, the geometric reconstruction unit 204 can perform position reconstruction based on the position information output by the prediction tree construction unit 207 to obtain a reconstructed value of the position information of each point in the point cloud data.

[0118] The arithmetic coding unit 205 may perform entropy coding on the position information output by the octree analysis unit 203 or the intersection points fitted by the surface fitting unit 206, or the geometric prediction residual values ​​output by the prediction tree construction unit 207 to generate a geometric code stream; the geometric code stream may also be referred to as a geometry bitstream.

[0119] Attribute encoding can be achieved through the following units:

[0120] A color conversion unit 210 , a transfer attributes unit 211 , a region adaptive hierarchical transform (RAHT) unit 212 , a generate LOD unit 213 , a lifting transform unit 214 , a quantize coefficients unit 215 , and an arithmetic coding unit 216 .

[0121] It should be noted that the point cloud encoder 200 may include more, fewer, or different functional components than those shown in FIG. 4A .

[0122] The color conversion unit 210 may be configured to convert the RGB color space of a point in the point cloud into a YCbCr format or other formats.

[0123] The recoloring unit 211 recolors the color information using the reconstructed geometric information so that the uncoded attribute information corresponds to the reconstructed geometric information.

[0124] After the original value of the point attribute information is converted by the recoloring unit 211, any transformation unit can be selected to transform the points in the point cloud. The transformation units may include: RAHT transformation 212 and lifting transformation unit 214. The lifting transformation relies on generating the level of detail (LOD).

[0125] Either the RAHT transform or the lifting transform can be understood as being used to predict the attribute information of a point in a point cloud to obtain a predicted value of the attribute information of the point, and then to obtain a residual value of the attribute information of the point based on the predicted value of the attribute information of the point. For example, the residual value of the attribute information of the point can be the original value of the attribute information of the point minus the predicted value of the attribute information of the point.

[0126] In one embodiment of the present application, the process of generating LOD by the LOD generation unit includes: obtaining the Euclidean distance between points based on the position information of the points in the point cloud; and dividing the points into different detail expression layers based on the Euclidean distance. In one embodiment, the Euclidean distances can be sorted and then Euclidean distances in different ranges can be divided into different detail expression layers. For example, a point can be randomly selected as the first detail expression layer. The Euclidean distances between the remaining points and the point are then calculated, and the points whose Euclidean distances meet the first threshold requirement are classified as the second detail expression layer. The centroid of the points in the second detail expression layer is obtained, and the Euclidean distances between the points other than the first and second detail expression layers and the centroid are calculated, and the points whose Euclidean distances meet the second threshold requirement are classified as the third detail expression layer. And so on, all points are classified into the detail expression layer. By adjusting the threshold of the Euclidean distance, the number of points in each LOD layer can be increased. It should be understood that the LOD division method can also be adopted in other ways, and this application is not limited to this.

[0127] It should be noted that the point cloud can be directly divided into one or more detail expression layers, or the point cloud can be first divided into multiple point cloud slices, and then each point cloud slice can be divided into one or more LOD layers.

[0128] For example, a point cloud can be divided into multiple point cloud tiles, each containing between 550,000 and 1.1 million points. Each point cloud tile can be considered a separate point cloud. Each point cloud tile can be further divided into multiple detail expression layers, each containing multiple points. In one embodiment, the detail expression layers can be divided based on the Euclidean distance between points.

[0129] The quantization unit 215 may be used to quantize the residual value of the attribute information of the point. For example, if the quantization unit 215 is connected to the RAHT transformation unit 212, the quantization unit 215 may be used to quantize the residual value of the attribute information of the point output by the RAHT transformation unit 212.

[0130] The arithmetic coding unit 216 may perform entropy coding on the residual value of the attribute information of the point using zero run length coding to obtain an attribute code stream. The attribute code stream may be bit stream information.

[0131] Figure 4B is a schematic block diagram of the point cloud decoder provided in an embodiment of the present application.

[0132] As shown in Figure 4B, the decoder 300 can obtain the point cloud code stream from the encoding device and obtain the position information and attribute information of the points in the point cloud by parsing the code. The decoding of the point cloud includes position decoding and attribute decoding.

[0133] The position decoding process includes: performing arithmetic decoding on the geometric code stream; constructing an octree and then merging it to reconstruct the point position information to obtain the reconstructed position information of the point; and performing coordinate transformation on the reconstructed position information of the point to obtain the point position information. The point position information can also be called the point's geometric information.

[0134] The attribute decoding process includes: obtaining the residual value of the attribute information of the point in the point cloud by parsing the attribute code stream; obtaining the residual value of the attribute information of the point after dequantization by dequantizing the residual value of the attribute information of the point; based on the reconstruction information of the point position information obtained in the position decoding process, selecting one of the following RAHT inverse transform and lifting inverse transform to perform point cloud prediction to obtain the predicted value, and adding the predicted value to the residual value to obtain the reconstructed value of the attribute information of the point; performing inverse color space conversion on the reconstructed value of the attribute information of the point to obtain the decoded point cloud.

[0135] As shown in Figure 4B, position decoding can be achieved by the following units:

[0136] Arithmetic decoding unit 301, octree reconstruction unit 302, surface reconstruction unit 303, geometry reconstruction unit 304, inverse transform coordinates unit 305 and prediction tree reconstruction unit 306.

[0137] Attribute encoding can be achieved through the following units:

[0138] an arithmetic decoding unit 310 , an inverse quantization unit 311 , an inverse RAHT transform unit 312 , a LOD generation unit 313 , an inverse lifting transform unit 314 , and an inverse color transform unit 315 .

[0139] It should be noted that decompression is the inverse process of compression. Similarly, the functions of each unit in the decoder 300 can refer to the functions of the corresponding units in the encoder 200. In addition, the point cloud decoder 300 may include more, fewer, or different functional components than those in Figure 4B.

[0140] For example, the decoder 300 can divide the point cloud into multiple LODs based on the Euclidean distance between points in the point cloud. The decoder 300 then decodes the attribute information of the points in the LODs in sequence. For example, the number of zeros (zero_cnt) in the zero-run encoding technique is calculated to decode the residual based on zero_cnt. The decoding framework 200 then dequantizes the decoded residual value and adds the dequantized residual value to the predicted value of the current point to obtain the reconstructed value of the point cloud until all point clouds are decoded. The current point will be used as the nearest neighbor of the subsequent LOD point, and the reconstructed value of the current point will be used to predict the attribute information of the subsequent point.

[0141] The above is the basic process of the point cloud codec based on the GPCC codec framework. With the development of technology, some modules or steps of the framework or process may be optimized. This application is applicable to the basic process of the point cloud codec based on the GPCC codec framework, but is not limited to this framework and process.

[0142] The following introduces octree-based geometric coding and prediction tree-based geometric coding.

[0143] The geometric encoding based on octree includes: first, coordinate transformation of the geometric information so that all point clouds are contained in a bounding box. Then quantization is performed. This step of quantization mainly plays a role of scaling. Due to the quantization rounding, the geometric information of some points is the same. The parameters are used to decide whether to remove duplicate points. The process of quantization and removal of duplicate points is also called voxelization. Next, the bounding box is continuously divided into trees (octree / quadtree / binary tree) in the order of breadth-first traversal, and the placeholder code of each node is encoded. In an implicit geometric division method, the bounding box of the point cloud is first calculated. Assume that the d x >d y >d z The bounding box corresponds to a cuboid. When geometrically partitioning, the binary tree partitioning is first performed based on the x-axis to obtain two child nodes; until d is satisfied x =d y >d z When the conditions are met, the quadtree partitioning will be performed based on the x and y axes to obtain four child nodes; when d is finally satisfied x =d y =d zWhen the conditions are met, the octree partitioning will continue until the leaf node obtained by the partitioning is a 1x1x1 unit cube. The partitioning will stop and the points in the leaf node will be encoded to generate a binary code stream. In the process of binary tree / quadtree / octree partitioning, two parameters are introduced: K and M. Parameter K indicates the maximum number of binary tree / quadtree partitions before octree partitioning; parameter M is used to indicate that the minimum block side length corresponding to binary tree / quadtree partitioning is 2 M . At the same time, K and M must meet the following conditions: Assume d max =max(d x ,d y ,d z ),d min =min(d x ,d y ,d z ), parameter K satisfies: K>=d max -d min ; Parameter M satisfies: M>=d min The parameters K and M meet the above conditions because the priority of the partitioning method in the current G-PCC implicit geometric partitioning process is binary tree, quadtree and octree. When the node block size does not meet the conditions of binary tree / quadtree, the node will be partitioned into octree until the minimum unit of leaf node 1X1X1 is reached.

[0144] The octree-based geometric information coding mode can effectively encode the geometric information of the point cloud by utilizing the correlation between adjacent points in space. However, for some relatively flat nodes or nodes with planar characteristics, the coding efficiency of the point cloud geometric information can be further improved by using plane coding.

[0145] For example, as shown in Figure 5A, the (a) series belongs to the low plane position in the Z-axis direction, and the (b) series belongs to the high plane position in the Z-axis direction. Taking (a) as an example, it can be seen that the four occupied child nodes of the current node are all located in the low plane position of the current node in the Z-axis direction. Therefore, it can be considered that the current node belongs to a Z plane and is a low plane in the Z-axis direction. Similarly, (b) shows that the occupied child nodes of the current node are located in the high plane position of the current node in the Z-axis direction.

[0146] Taking (a) as an example, the efficiency of octree coding and plane coding is compared. As shown in Figure 5B, if the octree coding method is used for (a) in Figure 1, the placeholder information of the current node is represented as: 11001100. However, if the plane coding method is used, first, an identifier needs to be encoded to indicate that the current node is a plane in the Z-axis direction. Secondly, if the current node is a plane in the Z-axis direction, the plane position of the current node needs to be represented. Secondly, only the placeholder information of the low plane node in the Z-axis direction needs to be encoded (that is, the placeholder information of the four child nodes 0246). Therefore, encoding the current node based on the plane coding method only requires encoding 6 bits, which can reduce the representation of 2 bits compared to the original octree coding. Based on this analysis, plane coding has more obvious coding efficiency than octree coding. Therefore, for an occupied node, if the plane coding method is used in a certain dimension, as shown in Figure 5C, first, the plane identification (planarMode) and plane position (PlanePos) information of the current node in the dimension need to be represented, and then the occupancy information of the current node is encoded based on the plane information of the current node. It should be noted that: PlaneMode i (i=0,1,2): 0 means the current node is not a plane in the direction of i axis. When the node is a plane in the direction of i axis, PlanePosition i :0 means the current node is a plane in the direction of the i-axis and the plane position is a low plane, 1 means the current node is a high plane in the direction of the i-axis. For example, i=0 represents the X-axis, i=1 represents the Y-axis, and i=2 represents the Z-axis.

[0147] The following details how to determine whether a node meets the plane coding conditions in the current G-PCC standard and predictively encode the node plane identifier and plane position information when the node meets the plane coding conditions.

[0148] Currently, there are three types of conditions in G-PCC to determine whether a node meets the conditions for plane coding. The following describes them one by one:

[0149] The first method is to judge based on the plane probability of the node in each dimension.

[0150] First, determine the local area density (local_node_density) of the current node and the probability Prob(i) of the current node in each dimension.

[0151] When the local area density of a node is less than the threshold Th (Th = 3), the plane probability Prob(i) of the current node in three dimensions is compared with the thresholds Th0, Th1, and Th2, where Th0 < Th1 < Th2 (Th0 = 0.6, Th1 = 0.77, Th2 = 0.88). Next, Eligible i (i = 0, 1, 2) is used to indicate whether plane coding is started in each dimension, where Eligible i The judgment process is shown in formula (1). For example, if Eligible i >= threshold, it means that plane coding is started in the i-th dimension:

[0152] Eligible i = Prob(i) >= threshold (1)

[0153] It should be noted that threshold changes adaptively. For example, when Prob(0) > Prob(1) > Prob(2), the value of threshold is as shown in formula (2):

[0154] Eligible0 = Prob(0) >= Th0

[0155] Eligible1 = Prob(1) >= Th1

[0156] Eligible2 = Prob(2) >= Th2 (2)

[0157] Next, the update process of local_node_density and the update of Prob(i) are introduced.

[0158] In one example, Prob(i) is updated by the following formula (3):

[0159] Prob(i) new = (Lx Prob(i) + δ(coded node)) / L + 1 (3)

[0160] where L = 255. When the coded node is a plane, it is 1; otherwise, it is 0.

[0161] In one example, local_node_density is updated by the following formula (4):

[0162] local_node_density new = local_node_density + 4 * numSiblings (4)

[0163] Among them, local_node_density is initialized to 4, numSiblings is the number of sibling nodes of the node, as shown in Figure 5D, the current node is the left node, the right node is the sibling node of the current node, and the number of sibling nodes of the current node is 5 (including itself).

[0164] The second method is to determine whether the nodes in the current layer meet the requirements of plane coding based on the point cloud density of the current layer.

[0165] The density of the points in the current layer is used to determine whether to perform plane coding on the nodes in the current layer. Assuming that the number of points in the current point cloud to be coded is pointCount, the number of points reconstructed after IDCM coding is numPointCountRecon, and because the octree is coded in the order of breadth-first traversal, the number of nodes to be coded in the current layer can be obtained as nodeCount. It is assumed that planarEligibleKOctreeDepth is used to indicate whether the current layer starts plane coding. The judgment process of planarEligibleKOctreeDepth is shown in formula (5):

[0166] planarEligibleKOctreeDepth=(pointCount-numPointCountRecon) <nodeCount*1.3 (5)

[0167] When planarEligibleKOctreeDepth is true, all nodes in the current layer are plane coded; otherwise, no plane coding is performed and only octree coding is used.

[0168] The third method is to determine whether the current node meets the requirements of plane coding based on the acquisition parameters of the lidar point cloud.

[0169] As shown in Figure 5E, the large cube node at the top is simultaneously traversed by two lasers, so the current node is not a plane in the Z-axis direction. The small cube node at the bottom is small enough that it cannot be traversed by both nodes simultaneously, so it is likely a plane. Therefore, based on the number of lasers corresponding to the current node, we can determine whether the current node meets the requirements for plane coding.

[0170] The following describes the predictive coding of plane identification information and plane position information for nodes that currently meet the plane coding conditions.

[0171] 1. Predictive Coding of Plane Marking Information

[0172] Currently, three contexts are used to encode plane identification information, that is, the plane representation in each dimension is designed separately.

[0173] The following introduces the encoding of planar position information of non-lidar point clouds and lidar point clouds respectively.

[0174] 1) Encoding of non-lidar point cloud planar position information

[0175] 1. Predictive coding of planar position information.

[0176] The plane position information is predictively coded based on the following information:

[0177] (1) Using the occupancy information of neighboring nodes, the plane position information of the current node is predicted to be three elements: predicted as low plane, predicted as high plane, and unpredictable;

[0178] (2) The spatial distance between the nodes at the same partition depth and the same coordinates as the current node and the current node is “close” or “far”;

[0179] (3) The plane position of the node at the same partition depth and the same coordinates as the current node;

[0180] (4) Coordinate dimension (i=0, 1, 2).

[0181] As shown in Figure 5F, the current node to be encoded is the left node, then the neighboring node is searched for as the right node at the same octree partition depth level and the same vertical coordinate, and the distance between the two nodes is judged as "near" and "far", and the plane position of the reference node is used.

[0182] In one example, as shown in FIG5G , the black node is the current node. If the current node is located on the lower plane of the parent node, the plane position of the current node is determined as follows:

[0183] a) If any of the child nodes 4 to 7 of the dashed node is occupied, and all the dot nodes are unoccupied, it is very likely that there is a plane in the current node, and the plane is at a lower position.

[0184] b) If the child nodes 4 to 7 of the dashed node are not occupied, and any dotted node is occupied, it is very likely that there is a plane in the current node, and the plane is at a higher position.

[0185] c) If the child nodes 4 to 7 of the dashed node are all empty nodes and the dotted nodes are all empty nodes, the plane position cannot be inferred and is therefore marked as unknown.

[0186] If any of the child nodes 4 to 7 of the dashed node are occupied and any of the dotted nodes are occupied, the plane position cannot be inferred and is therefore marked as unknown.

[0187] In another example, as shown in FIG5H , the black node is the current node. If the node is at a high plane position of the parent node, the plane position of the current node is determined as follows:

[0188] a) If any of the dot node's child nodes 4 to 7 is occupied, and the dashed node is not occupied, it is very likely that there is a plane in the current node, and the plane is at a lower position.

[0189] b) If the child nodes 4 to 7 of the dot node are not occupied, but the node with the dashed line is occupied, it is very likely that a plane exists in the current node, and the plane is located at a higher position.

[0190] c) If the child nodes 4 to 7 of the dot node are all unoccupied, and the dashed node is unoccupied, the plane position cannot be inferred and is therefore marked as unknown.

[0191] d) If one of the child nodes 4-7 of the dotted node is occupied and the dashed node is occupied, the plane position cannot be inferred and is therefore marked as unknown.

[0192] 2) Coding of LiDAR point cloud plane position information

[0193] Figure 5I shows the predictive coding of the plane position information of the laser radar point cloud. The plane position of the current node is predicted by using the laser radar acquisition parameters. The position is quantized into four intervals by using the intersection position of the current node and the laser ray, and finally used as the context of the plane position of the current node. The specific calculation process is as follows: Assume that the coordinates of the laser radar are (x Lidar ,y Lidar ,z Lidar ), the geometric coordinates of the current point are (x, y, z), then first calculate the vertical tangent value tanθ of the current point relative to the lidar. The calculation process is shown in formula (6):

[0194]

[0195] Because each laser has a certain offset angle relative to the laser radar, the relative tangent value tanθ of the current node relative to the laser is calculated. corr,L , the specific calculation process is shown in formula (7):

[0196]

[0197] Finally, the corrected tangent value of the current node is used to predict the plane position of the current node. Specifically, assuming that the tangent value of the lower boundary of the current node is tan(θ bottom), and the tangent value of the upper boundary is tan(θ top), according to tanθ corr,L The plane position is quantized into 4 quantization intervals, which are the contexts of the plane position.

[0198] However, the octree-based geometric information coding mode only has an efficient compression rate for points with correlation in space. For points that are isolated in the geometric space, the use of the Direct Coding Model (DCM) can greatly reduce the complexity. For all nodes in the octree, the use of DCM is not indicated by flag information, but is inferred from the parent node and neighbor information of the current node. There are three ways to determine whether the current node is eligible for DCM coding, as shown in Figure 6A:

[0199] (1) The current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.

[0200] (2) The parent node of the current node has only one child node, the current node, and the six neighbor nodes that share a face with the current node are all empty nodes.

[0201] (3) The number of sibling nodes of the current node is greater than 1.

[0202] If the current node does not meet the DCM coding qualifications, it will be divided into octrees. If it meets the DCM coding qualifications, the number of points contained in the node will be further determined. When the number of points is less than the threshold 2, the node will be DCM-encoded, otherwise the octree division will continue. When the DCM coding mode is applied, it is first necessary to encode whether the current node is a true isolated point, that is, IDCM_flag. When IDCM_flag is true, the current node uses DCM coding, otherwise octree coding is still used. When the current node meets the DCM coding requirements, the DCM coding mode of the current node needs to be encoded. There are currently two DCM modes: 1: There is only one point (or multiple points, but they are duplicate points); 2: Contains two points. Finally, the geometric information of each point needs to be encoded. Assume that the side length of the node is 2 d When encoding each component of the node's geometric coordinates, d bits are required, and these bits are directly encoded into the bitstream. It is important to note that when encoding LiDAR point clouds, the efficiency of geometric information coding can be further improved by predictively encoding the three-dimensional coordinate information using LiDAR acquisition parameters.

[0203] Next, the IDCM encoding process is introduced in detail:

[0204] When the current node meets the direct coding mode (DCM), the number of points of the current node, numPoints, is first encoded. The number of points of the current node is encoded according to different DirectModes, specifically including the following methods:

[0205] 1. If the current node does not meet the requirements of the DCM node, exit directly (that is, the number of points is greater than 2 points and is not a duplicate point).

[0206] 2. If the number of points numPonts in the current node is less than or equal to 2, the encoding process is as follows:

[0207] 1) First, encode whether the numPonts of the current node is greater than 1;

[0208] 2) If the current node has only one point and the geometry coding environment is geometry lossless coding, it is necessary to encode the second point of the current node to ensure that it is not a duplicate point.

[0209] 3. If the number of points numPonts in the current node is greater than 2, the encoding process is as follows:

[0210] 1) First, encode the numPonts of the current node to be less than or equal to 1;

[0211] 2) Secondly, encode whether the second point of the current node is a repeated point, and then encode whether the number of repeated points of the current node is greater than 1. When the number of repeated points is greater than 1, it is necessary to perform exponential Golomb decoding on the remaining number of repeated points.

[0212] After encoding the number of points in the current node, the coordinate information of the points contained in the current node is encoded. The following will introduce the lidar point cloud and the human eye point cloud separately.

[0213] Human eye point cloud

[0214] 1) If the current node contains only one point, the geometric information of the point in three dimensions will be directly encoded (Bypass coding).

[0215] 2) If the current node contains two points, the priority coding coordinate axis dirextAxis will be obtained by using the geometric coordinates of the points. It should be noted that the coordinate axes currently compared only include the x and y axes, not the z axis. Assuming that the geometric coordinates of the current node are nodePos, the priority coding coordinate axis is determined using the method shown in formula (8):

[0216] dirextAxis=!(nodePos[0] <nodePos[1]) (8)

[0217] That is, the axis with the smaller node coordinate geometric position is used as the coordinate axis dirextAxis for priority encoding.

[0218] Secondly, first encode the geometry information of the priority-encoded coordinate axis dirextAxis as follows, assuming that the bit depth of the geometry to be encoded corresponding to the priority-encoded axis is nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1] respectively:

[0219]

[0220] After encoding the priority axis dirextAxis, the geometric coordinates of the current point are directly encoded. Assuming that the remaining encoding bit depth of each point is nodeSizeLog2, the specific encoding process is as follows:

[0221] for(int axisIdx=0; axisIdx<3; ++axisIdx)

[0222] for(int mask=(1<<nodeSizeLog2[axisIdx])> >1;mask;mask>>1)

[0223] encodePosBit(!!(pointPos[axisIdx]&mask));

[0224] For LiDAR point clouds

[0225] 1) If the current node contains two points, the priority coding coordinate axis dirextAxis will be obtained by using the geometric coordinates of the points. Assuming that the geometric coordinates of the current node are nodePos, the priority coding coordinate axis is determined by the method shown in formula (9):

[0226] dirextAxis=!(nodePos[0] <nodePos[1]) (9)

[0227] That is, the axis with the smaller node coordinate geometric position is used as the coordinate axis dirextAxis for priority encoding. It should be noted here that the currently compared coordinate axes only include the x and y axes, but not the z axis.

[0228] Secondly, first encode the geometry information of the priority-encoded coordinate axis dirextAxis as follows, assuming that the bit depth of the geometry to be encoded corresponding to the priority-encoded axis is nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1] respectively:

[0229]

[0230]

[0231] After encoding the priority encoding axis dirextAxis, the geometric coordinates of the current point are encoded.

[0232] Since the LiDAR point cloud can obtain the acquisition parameters of the LiDAR point cloud, the geometric coordinate information of the current node can be predicted, thereby further improving the efficiency of the geometric information encoding of the point cloud. Similarly, the geometric information nodePos of the current node is first used to obtain a directly encoded main axis direction, and then the geometric information of the already encoded direction is used to predict the geometric information of the other dimension. Assuming that the axis direction of the direct encoding is directAxis and the bit depth to be encoded in the direct encoding is nodeSizeLog2, the encoding method is as follows:

[0233] for(int mask=(1<<nodeSizeLog2)> >1;mask;mask>>1)

[0234] encodePosBit(!!(pointPos[directAxis]&mask));

[0235] It should be noted here that all geometric accuracy information in the directAxis direction will be encoded here.

[0236] After encoding all the precision of the directAxis coordinate direction, the LaserIdx corresponding to the current point will be calculated first, that is, pointLaserIdx in Figure 6B, and the LaserIdx of the current node will be calculated, that is, nodeLaserIdx. Then, the LaserIdx of the node, that is, nodeLaserIdx, will be used to predict the LaserIdx of the point, that is, pointLaserIdx. The LaserIdx of the node or point is calculated as follows:

[0237] Assume that the geometric coordinates of the point are pointPos, the starting coordinates of the laser ray are LidarOrigin, and the number of lasers is LaserNum, and the tangent value of each laser is tanθ i , the vertical offset position of each Laser is Z i ,but:

[0238]

[0239] After calculating the current point's LaserIdx, the pointLaserIdx of the point is first predictively encoded using the current node's LaserIdx. After encoding the current point's LaserIdx, the three-dimensional geometric information of the current point is predictively encoded using the LiDAR acquisition parameters.

[0240] The specific algorithm is shown in Figure 6C. First, the LaserIdx corresponding to the current point is used to obtain the corresponding horizontal azimuth prediction value, that is, Secondly, the node geometry information corresponding to the current point is used to obtain the horizontal azimuth angle corresponding to the node Among them, the horizontal azimuth The calculation method between the node geometry information is shown in formula (10), assuming that the geometric coordinates of the node are nodePos:

[0241]

[0242] By using the acquisition parameters of the laser radar, the number of rotation points of each laser, numPoints, can be obtained, which represents the number of points obtained by each laser ray rotating one circle. The rotation angular velocity deltaPhi of each laser can then be calculated using the number of rotation points of each laser, as shown in formula (11):

[0243]

[0244] As shown in FIG6D , the horizontal azimuth angle of the node is used And the horizontal azimuth of the previous Laser code point corresponding to the current point Calculate the predicted horizontal azimuth angle corresponding to the current point The specific calculation formula is shown in formula (12):

[0245]

[0246] Finally, as shown in FIG6E , by using the predicted value of the horizontal azimuth angle and the low plane horizontal azimuth of the current node and the horizontal azimuth of the high plane To predict the geometric information of the current node. The details are as follows:

[0247]

[0248] After encoding the LaserIdx of the point, the Z-axis direction of the current point will be predicted and encoded using the LaserIdx corresponding to the current point. That is, the depth information radius of the cylindrical coordinate system is calculated by using the x and y information of the current point. Then, the tangent value of the current point and the vertical direction are obtained using the laser LaserIdx of the current point. The predicted value of the Z-axis direction of the current point, namely Z_pred, can be obtained:

[0249]

[0250]

[0251] Finally, Z_pred is used to predict the geometric information of the current point in the Z-axis direction to obtain the prediction residual Z_res, and Z_res is finally encoded.

[0252] It is important to note that when partitioning nodes into leaf nodes, the number of duplicate points in the leaf nodes must be encoded in the case of lossless geometric coding. Ultimately, the placeholder information for all nodes is encoded to generate a binary bitstream. Furthermore, G-PCC currently introduces a plane coding mode. During the geometric partitioning process, it determines whether the child nodes of the current node are in the same plane. If the child nodes of the current node meet the condition of being in the same plane, the child nodes of the current node are represented by that plane.

[0253] In octree-based geometric decoding, the decoder follows a breadth-first traversal. Before decoding each node's occupancy information, it first uses the reconstructed geometric information to determine whether the current node is for plane decoding or IDCM decoding. If the current node meets the requirements for plane decoding, it first decodes the plane identifier and plane position information of the current node. Then, based on the plane information, it decodes the current node's occupancy information. If the current node meets the requirements for IDCM decoding, it first decodes whether the current node is a true IDCM node. If so, it continues to parse the DCM decoding mode of the current node, then obtains the number of points in the current DCM node, and finally decodes the geometric information of each point. For nodes that do not meet either plane decoding or DCM decoding requirements, the current node's occupancy information is decoded. By continuously parsing in this way, the placeholder code of each node is obtained, and the node is continuously partitioned until a 1x1x1 unit cube is obtained. The number of points contained in each leaf node is parsed, and the geometrically reconstructed point cloud information is finally recovered.

[0254] The following is a detailed introduction to the IDCM decoding process:

[0255] The same process as encoding is used. First, a priori information is used to determine whether the node should start IDCM. The starting conditions of IDCM are as follows:

[0256] (1) The current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.

[0257] (2) The parent node of the current node has only one child node, the current node, and the six neighbor nodes that share a face with the current node are all empty nodes.

[0258] (3) The number of sibling nodes of the current node is greater than 1.

[0259] When a node meets the conditions for DCM encoding, it first decodes whether the current node is a real DCM node, that is, IDCM_flag. When IDCM_flag is true, the current node adopts DCM encoding, otherwise it still adopts octree encoding.

[0260] Next, decode the number of points numPoints of the current node. The specific decoding method is as follows:

[0261] 1) First decode whether the numPonts of the current node is greater than 1;

[0262] 2) If the numPonts of the current node is greater than 1, continue decoding to see if the second point is a duplicate point. If the second point is not a duplicate point, it can be implicitly inferred that the second type of DCM mode contains only two points.

[0263] 3) If the numPonts of the current node obtained by decoding is less than or equal to 1, continue decoding whether the second point is a repeated point. If the second point is not a repeated point, it can be implicitly inferred that the second type of DCM mode is satisfied, which contains only one point; if the second point obtained by decoding is a repeated point, it can be inferred that the third type of DCM mode is satisfied, which contains multiple points, but they are all repeated points. Then continue decoding whether the number of repeated points is greater than 1 (entropy decoding). If it is greater than 1, continue decoding the number of remaining repeated points (using exponential Columbus decoding).

[0264] If the current node does not meet the requirements of the DCM node, that is, the number of points is greater than 2 points and it is not a duplicate point, exit directly.

[0265] After decoding the number of points in the current node, the coordinate information of the points contained in the current node is decoded. The following will introduce the lidar point cloud and the human eye point cloud separately.

[0266] Human eye point cloud

[0267] 1) If the current node contains only one point, the geometric information of the point in three dimensions will be directly decoded (Bypass coding);

[0268] 2) If the current node contains two points, the priority decoding axis dirextAxis will be obtained by using the geometric coordinates of the points. It should be noted that the coordinate axes currently compared only include the x and y axes, not the z axis. Assuming that the geometric coordinates of the current node are nodePos, the method shown in formula (13) is used to determine the priority encoding axis:

[0269] dirextAxis=!(nodePos[0] <nodePos[1]) (13)

[0270] That is to say, the axis with the smaller node coordinate geometric position is used as the coordinate axis dirextAxis for priority decoding.

[0271] Secondly, decode the geometry information of the priority decoded coordinate axis dirextAxis as follows, assuming that the bit depth of the geometry to be decoded corresponding to the priority decoded axis is nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1] respectively:

[0272]

[0273]

[0274] After decoding the priority decoding axis dirextAxis, the geometric coordinates of the current point are directly decoded. Assuming that the remaining encoding bit depth of each point is nodeSizeLog2, the specific decoding process is as follows, assuming that the coordinate information of the point is pointPos:

[0275]

[0276] For LiDAR point clouds

[0277] 1) If the current node contains two points, the priority decoding coordinate axis dirextAxis will be obtained by using the geometric coordinates of the points. Assuming that the geometric coordinates of the current node are nodePos, the priority encoding coordinate axis is determined by the method shown in formula (14):

[0278] dirextAxis=!(nodePos[0] <nodePos[1]) (14)

[0279] That is to say, the axis with the smaller node coordinate geometric position is used as the coordinate axis dirextAxis for priority decoding. It should be noted here that the currently compared coordinate axes only include the x and y axes, and do not include the z axis.

[0280] Secondly, first decode the geometry information of the priority-encoded coordinate axis dirextAxis as follows, assuming that the bit depth of the geometry to be encoded corresponding to the priority-encoded axis is nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1] respectively:

[0281]

[0282] After decoding the priority decoding axis dirextAxis, the geometric coordinates of the current point are decoded.

[0283] Similarly, we first use the current node's geometry information nodePos to get a direct decoding main axis direction, and then use the geometry information of the decoded direction to decode the geometry information of the other dimension. Assuming that the axis direction of direct decoding is directAxis and the bit depth to be decoded in direct decoding is nodeSizeLog2, the decoding method is as follows:

[0284]

[0285] It should be noted here that all geometric accuracy information in the directAxis direction will be decoded here.

[0286] After decoding all the precision of the directAxis coordinate direction, the LaserIdx of the current node, i.e., nodeLaserIdx, is first calculated. Then, the LaserIdx of the node, i.e., nodeLaserIdx, is used to predict and decode the LaserIdx of the point, i.e., pointLaserIdx. The calculation method of the LaserIdx of the node or point is the same as that of the encoder. Finally, the LaserIdx of the current point and the predicted residual information of the LaserIdx of the node are decoded to obtain ResLaserIdx. The calculation formula is shown in Formula 15:

[0287] PointLaserIdx=nodeLaserIdx+ResLaserIdx (15)

[0288] After decoding the LaserIdx of the current point, the three-dimensional geometric information of the current point is predicted and decoded using the acquisition parameters of the laser radar.

[0289] Specifically, as shown in FIG6B , the LaserIdx corresponding to the current point is first used to obtain the corresponding predicted value of the horizontal azimuth angle, that is, Secondly, the node geometry information corresponding to the current point is used to obtain the horizontal azimuth angle corresponding to the node Assume that the geometric coordinates of the node are nodePos, and the horizontal azimuth is The calculation method between the node geometry information is shown in formula (16):

[0290]

[0291] By using the acquisition parameters of the laser radar, the number of rotation points of each laser, numPoints, can be obtained, which represents the number of points obtained by each laser ray rotating one circle. The rotation angular velocity deltaPhi of each laser can then be calculated using the number of rotation points of each laser, as shown in formula (17):

[0292]

[0293] Next, as shown in FIG6D , the horizontal azimuth angle of the node is used And the horizontal azimuth of the previous Laser code point corresponding to the current point Calculate the predicted horizontal azimuth angle corresponding to the current point The predicted value of the horizontal azimuth angle is calculated as shown in formula (18):

[0294]

[0295] Finally, by using the predicted value of the horizontal azimuth and the low plane horizontal azimuth of the current node and the horizontal azimuth of the high plane To predict the geometric information of the current node. The details are as follows:

[0296]

[0297] After decoding the LaserIdx of the point, the Z-axis direction of the current point will be predicted and decoded using the LaserIdx corresponding to the current point. That is, the depth information radius of the cylindrical coordinate system is calculated by using the x and y information of the current point. Then, the tangent value of the current point and the vertical offset are obtained using the laser LaserIdx of the current point. Then, the predicted value of the Z-axis direction of the current point, namely Z_pred, can be obtained:

[0298]

[0299] Finally, the decoded Z_res and Z_pred are used to reconstruct and restore the geometric information of the current point in the Z-axis direction.

[0300] In the trisoup (triangle soup)-based geometric information coding framework, geometric partitioning is also performed first. However, unlike geometric information coding based on binary trees, quad trees, and octrees, this method does not need to gradually partition the point cloud into unit cubes with side lengths of 1x1x1. Instead, the partitioning stops when the block (sub-block) has a side length of W. Based on the surface formed by the distribution of the point cloud in each block, the surface and the twelve edges of the block are obtained. The vertex coordinates of each block are encoded in sequence to generate a binary code stream.

[0301] When reconstructing point cloud geometry based on trisoup, the decoding end first decodes vertex coordinates to complete triangle reconstruction. This process is shown in Figures 7A to 7C. The block shown in Figure 7A contains three vertices (v1, v2, v3). The set of triangles formed by these three vertices in a certain order is called triangle soup, or trisoup, as shown in Figure 7B. Afterwards, sampling is performed on this set of triangles, and the resulting sampling points are used as the reconstructed point cloud within the block, as shown in Figure 7C.

[0302] The geometric coding based on the prediction tree includes: first, sorting the input point cloud. The currently used sorting methods include unordered, Morton order, azimuth order, and radial distance order. At the encoding end, the prediction tree structure is established by using two different methods, including: KD-Tree (high-latency slow mode) and using the lidar calibration information to divide each point into different Lasers and establish a prediction structure according to different Lasers (low-latency fast mode). Next, based on the structure of the prediction tree, each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter. Finally, through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream.

[0303] Based on the geometric decoding of the prediction tree, the decoding end reconstructs the prediction tree structure by continuously parsing the bit stream. Secondly, the geometric position prediction residual information and quantization parameters of each prediction node are obtained through parsing, and the prediction residual is dequantized to restore the reconstructed geometric position information of each node, finally completing the geometric reconstruction at the decoding end.

[0304] After the geometric encoding is completed, the geometric information is reconstructed. At present, attribute encoding is mainly performed on color information. First, the color information is converted from the RGB color space to the YUV color space. Then, the point cloud is recolored using the reconstructed geometric information so that the unencoded attribute information corresponds to the reconstructed geometric information. In color information encoding, there are two main transformation methods. One is the distance-based lifting transformation that relies on LOD (Level of Detail) division, and the other is to directly perform RAHT (Region Adaptive Hierarchal Transform) transformation. Both methods will convert the color information from the spatial domain to the frequency domain, and obtain high-frequency coefficients and low-frequency coefficients through transformation. Finally, the coefficients are quantized and encoded to generate a binary code stream.

[0305] When using geometric information to predict attribute information, Morton codes can be used to perform nearest neighbor search. The Morton code corresponding to each point in the point cloud can be obtained from the geometric coordinates of the point. The specific method for calculating the Morton code is described as follows. For each component of the three-dimensional coordinate represented by a d-bit binary number, its three components can be expressed as formula (19):

[0306]

[0307] in, The highest bits of x, y, and z are To the lowest position The corresponding binary value. The Morton code M is x, y, z starting from the highest bit, arranged in sequence To the lowest bit, the calculation formula of M is shown in the following formula (20):

[0308]

[0309] in, The highest bit of M To the lowest position After obtaining the Morton code M of each point in the point cloud, the points in the point cloud are arranged in ascending order of the Morton code, and the weight w of each point is set to 1.

[0310] There are 4 general test conditions for GPCC:

[0311] Condition 1: The geometric position is limited and the attributes are lost;

[0312] Condition 2: Geometric position lossless, attribute lossy;

[0313] Condition 3: Geometric position lossless, attribute loss limited;

[0314] Condition 4: Geometric position and attributes are lossless.

[0315] The general test sequences include Cat1A, Cat1B, Cat3-fused, and Cat3-frame, a total of four categories. Among them, Cat2-frame point cloud only contains reflectance attribute information, Cat1A and Cat1B point clouds only contain color attribute information, and Cat3-fused point cloud contains both color and reflectance attribute information.

[0316] There are two technical routes of GPCC, which are distinguished by the algorithm used for geometric compression, and are divided into octree coding branch and prediction tree coding branch.

[0317] Among them, in the octree coding branch, at the encoding end, the bounding box is divided into sub-cubes in sequence, and the non-empty sub-cubes (containing points in the point cloud) are divided until the leaf node obtained by division is a 1X1X1 unit cube. The division stops when the division is completed. In the case of geometric lossless coding, the number of points contained in the leaf node needs to be encoded, and finally the geometric octree encoding is completed to generate a binary code stream. At the decoding end, the decoding end obtains the placeholder code of each node by continuous parsing in the order of breadth-first traversal, and continuously divides the nodes in sequence until the division is a 1x1x1 unit cube. In the case of geometric lossless decoding, the number of points contained in each leaf node needs to be parsed to finally recover the geometric reconstructed point cloud information.

[0318] In the prediction tree coding branch, the encoder establishes the prediction tree structure using two different approaches: a KD-Tree (high-latency, slow mode) and a low-latency, fast mode, where each point is assigned to a different laser using lidar calibration information and the prediction structure is established accordingly. Next, based on the prediction tree structure, each node in the tree is traversed, and the geometric position information of the node is predicted using different prediction modes to obtain a prediction residual. This geometric prediction residual is then quantized using a quantization parameter. Finally, through continuous iteration, the prediction residuals of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary bitstream. On the decoder side, the decoder continuously parses the bitstream to reconstruct the prediction tree structure. The geometric position prediction residual information and quantization parameters for each prediction node are then parsed and dequantized to recover the reconstructed geometric position information for each node, completing the geometric reconstruction at the decoder.

[0319] The following is an introduction to the AVS codec framework.

[0320] In the point cloud AVS encoder framework, the geometric information of the point cloud and the attribute information corresponding to each point are encoded separately.

[0321] Figure 8A is a schematic diagram of the encoding framework of AVS, and Figure 8B is a schematic diagram of the decoding framework of AVS. As shown in Figure 8A, the geometric information is first transformed so that all point clouds are contained in a bounding box. Before the preprocessing process, it is decided whether to divide the entire point cloud sequence into multiple slices based on the parameter configuration, and each divided slice is treated as a single independent point cloud for serial processing. The preprocessing process includes quantization and removal of duplicate points. Quantization mainly plays a role in scaling. Due to the rounding of quantization, the geometric information of some points is the same, and whether to remove duplicate points is determined based on the parameters. Next, the bounding box is divided in the order of breadth-first traversal (octree / quadtree / binary tree), and the placeholder code of each node is encoded.

[0322] As shown in Figure 8B, in the octree-based geometric coding framework, the bounding box is divided into sub-cubes in sequence, and the non-empty (containing points in the point cloud) sub-cubes are divided until the leaf node obtained by division is a 1x1x1 unit cube. The division is stopped. Then, in the case of geometric lossless coding, the number of points contained in the leaf node is encoded, and finally the geometric octree encoding is completed to generate a binary code stream. In the octree-based geometric decoding process, the decoding end obtains the placeholder code of each node by continuously parsing in the order of breadth-first traversal, and continuously divides the nodes in sequence until the division is a 1x1x1 unit cube. The number of points contained in each leaf node is parsed, and finally the geometric reconstructed point cloud information is restored.

[0323] There are two encoding methods in the current AVS geometric coding, one is octree coding and the other is prediction tree coding.

[0324] Octree encoding: If octree encoding is used, there are two context encoding models. Context model 1 is used for cat1-A and cat2 point cloud sequences; context model 2 is used for cat1-B and cat3 sequences.

[0325] The following is an introduction to Model 1.

[0326] The following model 1 includes the sub-layer neighbor prediction of the current point and the neighbor prediction of the current point layer.

[0327] 1) Sub-layer neighbor prediction of the current point

[0328] Under the octree breadth-first traversal partitioning method, the neighbor information that can be obtained when encoding the child node of the current point includes the neighbor child nodes in the three directions of left, front, and bottom. The context model of the child node layer is designed as follows: for the child node layer to be encoded, find the occupancy of 3 coplanar, 3 colinear, and 1 co-point nodes in the left, front, and bottom directions of the same layer as the child node to be encoded, as well as the node in the negative direction of the dimension with the shortest node side length, which is two node side lengths away from the current child node to be encoded. Taking the node with the shortest side length in the X dimension as an example, the reference node selected by each child node is shown in Figure 9A. The dotted box node is the current node, the gray node is the current child node to be encoded, and the solid box node is the reference node selected by each child node.

[0329] Among them, the occupancy of the 3 coplanar nodes, 3 collinear nodes and the node with the shortest side length in the negative direction away from the current sub-node to be encoded is considered in detail. The occupancy of these 7 nodes is 2 7 = 128 cases. If not all are unoccupied, there are 2 7 -1 = 127 possible cases, with one context allocated for each. If all seven nodes are unoccupied, the occupied position of the common neighbor node is considered. This common neighbor has two possibilities: occupied or unoccupied. A separate context is allocated for the occupied case of the common neighbor node. If this common neighbor is also unoccupied, the occupied position of the current node's neighbors, described below, is considered. Thus, the neighbors at the subnode level to be encoded correspond to a total of 127 + 2 - 1 = 128 contexts.

[0330] 2) Neighbor prediction of the current node layer

[0331] If the eight reference nodes in the same layer of the child node to be encoded are not occupied, the occupancy of the four groups of neighbors in the current node layer is considered as shown in Figure 9B, where the dotted frame node is the current node and the solid frame node is the neighbor node.

[0332] For the current node layer, the context is determined as follows:

[0333] 1. First consider the three coplanar neighbors to the upper right of the current node. The occupancy of the three coplanar neighbors to the upper right of the current node is 2 3 = 8 possibilities. For the cases where all nodes are not occupied, one context is assigned to each node. Considering that the child node to be encoded is located at the position of the current node, the group of neighboring nodes provides a total of (8-1)×8=56 contexts. If the three coplanar neighbors to the upper right of the current point are not occupied, then continue to consider the remaining three groups of neighbors at the current node layer.

[0334] 2. Consider the distance between the most recently occupied node and the current node.

[0335] The specific correspondence between neighbor node distribution and distance is shown in Table 2.

[0336] Table 2 The corresponding relationship between the current node layer occupancy and distance

[0337] The current node layer occupancy situation is 1: the distance between the left front and lower coplanar neighbors occupied or the upper right and rear collinear neighbors occupied. The left front and lower coplanar neighbors and the upper right and rear collinear neighbors are not occupied, and the left front and lower collinear neighbors are occupied. 2: none of the four groups of neighbors of the current node layer are occupied. 3:

[0338] As shown in Table 2, the distance has three possible values. One context is assigned to each of these three values. Considering the position of the child node to be encoded within the current node, there are a total of 3 × 8 = 24 contexts.

[0339] So far, this set of context models has allocated a total of 128+56+24=208 contexts.

[0340] The following is an introduction to context model 2.

[0341] This method uses a two-layer context reference relationship configuration, as shown in formula (21). The first layer is the occupancy of the encoded adjacent blocks of the parent node of the current sub-block to be encoded (i.e., ctxIdxParent), and the second layer is the occupancy of the adjacent encoded blocks at the same depth as the current sub-block to be encoded (i.e., ctxIdxChild).

[0342] First, for each sub-block to be coded, the ctxIdxChild of the second layer is as shown in formula (22), C i 1 Indicates that the current sub-block Occupancy of the three coded sub-blocks with a distance of 1.

[0343] idx=LUT[ctxIdxParent][ctxIdxChild] (21)

[0344]

[0345]

[0346] Secondly, for the relative positions of different sub-blocks, the first layer’s ctxIdxParent is used to find the adjacent parent blocks that are coplanar and colinear with them by looking up the table, and the ctxIdxParent is calculated according to the occupancy of the adjacent parent blocks according to formula (23). Figure 9C is a schematic diagram of the six adjacent parent blocks corresponding to the sub-blocks. As shown in Figure 9C, each sub-graph shows the relative position relationship of the six adjacent parent blocks found by the i-th sub-block, including three coplanar parent blocks (P i,0 ,P i,1 ,P i,2 ) and 3 collinear parent blocks (P i,3,P i,4 ,P i,5 ). The positional relationship between each sub-block and the adjacent parent block is obtained through the method in Table 2. The numbers in Table 2 correspond to the Morton sequence numbers in Figure 6. This method takes into account the positions of different sub-blocks and the geometric center rotation symmetry. Figure 9D is a schematic diagram of the 18 adjacent blocks and their Morton sequence numbers used by the current block to be encoded. It can be seen from Figure 9D that with the current block as the center, this method has a larger receptive field and can utilize up to 18 adjacent parent blocks that have been encoded in the surrounding area. The method used in formula (3) is the permutation and combination of the occupancy of the three coplanar parent blocks and the sum of the number of occupancy of the three collinear parent blocks.

[0347] Therefore, the number of contexts used in this method is at most 2 3 ×2 5 =256.

[0348] Table 3 shows the relationship between child block i and its adjacent parent block j. The numbers in the table correspond to the Morton sequence numbers in FIG9D .

[0349] Table 3

[0350]

[0351] Prediction Tree Coding: If prediction tree coding is used, the encoder first uses the point cloud's geometric information to perform Morton code sorting. Then, a KD-Tree is used to predictively encode the point cloud's geometric information. This is similar to a single-chain structure, where the parent node predicts the geometry of the child nodes. As shown in Figure 9E, the prediction tree uses a single-chain structure: except for the sole leaf node, each tree node has only one child node. Except for the root node, which is predicted by default, all other nodes receive geometric predictions from their parent nodes.

[0352] During the multitree geometry coding process, the isolated point direct coding mode is effective when the current block satisfies the following three conditions at the same time:

[0353] 1. The isolated point direct coding mode identifier in the geometry header information is 1;

[0354] 2. The current block contains only one point cloud data point;

[0355] 3. The sum of the number of Morton code bits to be encoded for the points in the current block is greater than twice the number of directions that do not reach the minimum side length.

[0356] This branch is entered when all three of the above conditions are met. A flag is introduced to indicate whether the current node uses isolated point encoding mode. This flag uses a context for entropy encoding. If the flag is True, isolated point mode is used, directly encoding the geometric coordinates of the point, and octree partitioning is terminated. If the flag is False, occupancy code is encoded and octree partitioning continues.

[0357] In certain cases, this flag can be inferred to be False and not encoded. If the parent block of the current block already allows the use of isolated point coding mode, and the current block is the only child node of the parent block, then the current block must not contain isolated points. Therefore, in this case, the bits for encoding the flag can be omitted.

[0358] After encoding the flag bit, since the current block contains only one point cloud point, the uncoded bits of the Morton code corresponding to the geometric coordinates of the point cloud point are directly encoded. The specific encoding process is as follows:

[0359] Assuming that the remaining encoding bit depth of the point is nodeSizeLog2, the specific encoding process is as follows:

[0360] for(int axisIdx=0; axisIdx<3; ++axisIdx)

[0361] for(int mask=(1<<nodeSizeLog2[axisIdx])> >1;mask;mask>>1)

[0362] encodePosBit(!!(pointPos[axisIdx]&mask));

[0363] After geometric encoding is completed, the geometric information is reconstructed. Currently, attribute encoding is mainly performed on color and reflectance information. As shown in Figure 8A, the encoding end first determines whether to perform color space conversion. If color space conversion is performed, the color information is converted from RGB color space to YUV color space. Then, the reconstructed point cloud is recolored using the original point cloud so that the unencoded attribute information corresponds to the reconstructed geometric information. Color information encoding is divided into two modules: attribute prediction and attribute transformation. The attribute prediction process is as follows: first, the point cloud is reordered, and then differential prediction is performed. There are two reordering methods: Morton reordering and Hilbert reordering. For cat1A and cat2 sequences, Hilbert reordering is performed; for cat1B and cat3 sequences, Morton reordering is performed. The sorted point cloud is used to perform attribute prediction using a differential method. Finally, the prediction residual is quantized and entropy coded to generate a binary code stream. The attribute transformation process is as follows: first, wavelet transform is performed on the point cloud attributes and the transform coefficients are quantized; secondly, the attribute reconstruction value is obtained through inverse quantization and inverse wavelet transform; then the difference between the original attribute and the attribute reconstruction value is calculated to obtain the attribute residual and quantize it; finally, the quantized transform coefficients and attribute residuals are entropy coded to generate a binary code stream.

[0364] The following introduces the general test conditions of AVS PCC.

[0365] There are four general test conditions for AVS:

[0366] Condition 1: The geometric position is limited and the attributes are lost;

[0367] Condition 2: Geometric position lossless, attribute lossy;

[0368] Condition 3: Geometric position lossless, attribute loss limited;

[0369] Condition 4: Geometric position and attributes are lossless.

[0370] The general test sequences include five categories: Cat1A, Cat1B, Cat1C, Cat2-frame and Cat3. Among them, Cat1A and Cat2-frame point clouds only contain reflectance attribute information, Cat1B and Cat3 point clouds only contain color attribute information, and Cat1B point cloud contains both color and reflectance attribute information.

[0371] Technical routes: There are 2 types in total, distinguished by the algorithm used for attribute compression.

[0372] Technical Route 1: Prediction branch, attribute compression uses an intra-frame prediction-based method:

[0373] At the encoding end, the points in the point cloud are processed in a certain order (the original acquisition order of the point cloud, the Morton order, the Hilbert order, etc.). First, the prediction algorithm is used to obtain the attribute prediction value. The attribute residual is obtained based on the attribute value and the attribute prediction value. Then, the attribute residual is quantized to generate the quantized residual. Finally, the quantized residual is encoded.

[0374] At the decoding end, the points in the point cloud are processed in a certain order (the original acquisition order of the point cloud, Morton order, Hilbert order, etc.). First, the prediction algorithm is used to obtain the attribute prediction value, then the decoding is performed to obtain the quantized residual, and then the quantized residual is dequantized. Finally, the attribute reconstruction value is obtained based on the attribute prediction value and the dequantized residual.

[0375] Technical Route 2: Prediction Transform Branch - Resources are limited. Attribute compression uses a method based on intra-frame prediction and DCT transform. When encoding the quantized transform coefficients, there is a maximum point number X (such as 4096), that is, at most every X points can be encoded as a group:

[0376] At the encoding end, the points in the point cloud are processed in a certain order (the original acquisition order of the point cloud, Morton order, Hilbert order, etc.). First, the entire point cloud is divided into several small groups with a maximum length of Y (such as 2). These small groups are then combined into several large groups (the number of points in each large group does not exceed X, such as 4096). Then, a prediction algorithm is used to obtain attribute prediction values. Based on the attribute values ​​and attribute prediction values, attribute residuals are obtained. The attribute residuals are transformed by DCT in small groups to generate transform coefficients. The transform coefficients are then quantized to generate quantized transform coefficients. Finally, the quantized transform coefficients are encoded in large groups.

[0377] At the decoding end, the points in the point cloud are processed in a certain order (the original acquisition order of the point cloud, Morton order, Hilbert order, etc.). First, the entire point cloud is divided into several small groups with a maximum length of Y (such as 2). Then these small groups are combined into several large groups (the number of points in each large group does not exceed X, such as 4096). The quantized transform coefficients are decoded in large groups, and then the prediction algorithm is used to obtain the attribute prediction value. The quantized transform coefficients are then dequantized and inversely transformed in small groups. Finally, the attribute reconstruction value is obtained based on the attribute prediction value and the dequantized and inversely transformed coefficients.

[0378] Technical Route 3: Prediction Transform Branch - Resources are not limited. Attribute compression uses a method based on intra-frame prediction and DCT transform. When encoding the quantized transform coefficients, there is no limit on the maximum number of points X, that is, all coefficients are encoded together:

[0379] At the encoding end, the points in the point cloud are processed in a certain order (the original acquisition order of the point cloud, Morton order, Hilbert order, etc.). First, the entire point cloud is divided into several small groups with a maximum length of Y (such as 2). Then, a prediction algorithm is used to obtain attribute prediction values. Based on the attribute values ​​and attribute prediction values, attribute residuals are obtained. The attribute residuals are transformed by DCT in groups to generate transformation coefficients. The transformation coefficients are then quantized to generate quantized transformation coefficients. Finally, the quantized transformation coefficients of the entire point cloud are encoded.

[0380] At the decoding end, the points in the point cloud are processed in a certain order (the original acquisition order of the point cloud, Morton order, Hilbert order, etc.). First, the entire point cloud is divided into several small groups with a maximum length of Y (such as 2). The quantized transformation coefficients of the entire point cloud are obtained by decoding, and then the prediction algorithm is used to obtain the attribute prediction value. The quantized transformation coefficients are then dequantized and inversely transformed in groups. Finally, the attribute reconstruction value is obtained based on the attribute prediction value and the dequantized and inversely transformed coefficients.

[0381] Technical Route 4: Multi-layer transformation branch, attribute compression adopts a method based on multi-layer wavelet transform:

[0382] At the encoding end, the entire point cloud is subjected to multi-layer wavelet transform to generate transform coefficients, which are then quantized to generate quantized transform coefficients. Finally, the quantized transform coefficients of the entire point cloud are encoded.

[0383] At the decoding end, decoding obtains the quantized transform coefficients of the entire point cloud, and then dequantizes and inversely transforms the quantized transform coefficients to obtain attribute reconstruction values.

[0384] When directly encoding the current node, the encoder encodes the position information of the current node's midpoint after determining that the current node is eligible for direct encoding and decoding. However, when encoding the geometric information of the current node's midpoint, inter-frame information is not considered, which reduces the encoding and decoding performance of the point cloud.

[0385] In order to solve the above technical problems, when encoding and decoding the current node in the current decoding frame, the embodiment of the present application determines N predicted nodes of the current node in the predicted reference frame of the current frame to be encoded and decoded, and predicts and decodes the coordinate information of the midpoint of the current node based on the geometric encoding and decoding information of the midpoints of these N predicted nodes. In other words, the embodiment of the present application optimizes the direct DCM encoding and decoding of the node, and uses the geometric information of the predicted node in the predicted reference frame to predict and decode the geometric information of the midpoint of the IDCM node (i.e., the current node) of the to-be-encoded node by considering the time domain correlation between adjacent frames. The efficiency of geometric information encoding and decoding of the point cloud is further improved by considering the time domain correlation between adjacent frames.

[0386] The following describes the point cloud encoding and decoding method involved in the embodiments of the present application in conjunction with specific embodiments.

[0387] First, taking the decoding end as an example, the point cloud decoding method provided in the embodiment of the present application is introduced.

[0388] Figure 10 is a schematic diagram of a point cloud decoding method according to an embodiment of the present application. The point cloud decoding method according to an embodiment of the present application can be implemented by the point cloud decoding device or point cloud decoder shown in Figure 3 or Figure 4B or Figure 8B above.

[0389] As shown in FIG10 , the point cloud decoding method of the embodiment of the present application includes:

[0390] S101 : Determine N prediction nodes of a current node in a prediction reference frame of a current frame to be decoded.

[0391] The current node is the node to be decoded in the current frame to be decoded.

[0392] As can be seen from the above, a point cloud includes geometric information and attribute information, and decoding of a point cloud includes geometric decoding and attribute decoding. The embodiments of the present application relate to geometric decoding of a point cloud.

[0393] In some embodiments, the geometric information of the point cloud is also referred to as the position information of the point cloud. Therefore, the geometric decoding of the point cloud is also referred to as the position decoding of the point cloud.

[0394] In the octree-based encoding method, the encoding end constructs the octree structure of the point cloud based on the geometric information of the point cloud. As shown in Figure 11, the point cloud is enclosed by the smallest rectangular block. The bounding box is first divided into octrees to obtain 8 nodes. The occupied nodes among these 8 nodes, that is, the nodes including the points, are further divided into octrees, and so on, until the division is to the voxel level, for example, to a 1X1X1 cube. The point cloud octree structure obtained by such division includes multiple layers of nodes, for example, N layers. During encoding, the occupancy information of each layer is encoded layer by layer until the voxel-level leaf nodes of the last layer are encoded. That is to say, in octree encoding, the point cloud is divided into octrees, and finally the points in the point cloud are divided into the voxel-level leaf nodes of the octree. The encoding of the point cloud is achieved by encoding the entire octree.

[0395] Correspondingly, the decoder first decodes the point cloud geometry stream to obtain the occupancy information of the root node of the point cloud's octree. Based on this occupancy information, it determines the child nodes of the root node, that is, the nodes in the second layer of the octree. Next, it decodes the geometry stream to obtain the occupancy information of each node in the second layer. Based on this occupancy information, it determines the nodes in the third layer of the octree, and so on.

[0396] However, the octree-based geometric information encoding mode has an efficient compression rate for points with correlation in space, and for points in isolated positions in the geometric space, the use of direct encoding can greatly reduce the complexity and improve the encoding and decoding efficiency.

[0397] Since direct encoding directly encodes the geometric information of the points included in a node, if the node contains a large number of points, the compression effect of direct encoding is poor. Therefore, before performing direct encoding on a node in the octree, it is first determined whether the node can be encoded using direct encoding. If it is determined that the node can be encoded using direct encoding, the geometric information of the points included in the node is directly encoded using direct encoding. If it is determined that the node cannot be encoded using direct encoding, the node is further divided using the octree method.

[0398] Specifically, the encoder first determines whether the node is eligible for direct encoding. If so, it then determines whether the node's point count is less than or equal to a preset threshold. If so, the node is determined to be eligible for direct encoding. Next, the number of points in the node and the geometric information of each point are encoded into the bitstream. Correspondingly, after determining that the node is eligible for direct decoding, the decoder decodes the bitstream, obtains the node's point count and geometric information of each point, and performs geometric decoding of the node.

[0399] Currently, when predicting the position information of the midpoint of the current node, inter-frame information is not considered, resulting in low coding performance of the point cloud.

[0400] In order to solve the above problems, in an embodiment of the present application, the decoding end predicts and decodes the position information of the midpoint of the current node based on the inter-frame information corresponding to the current node, thereby improving the decoding efficiency and decoding performance of the point cloud.

[0401] Specifically, the decoding end first determines N prediction nodes of the current node in the prediction reference frame of the current frame to be decoded.

[0402] It should be noted that the current frame to be decoded is a point cloud frame. In some embodiments, the current frame to be decoded is also referred to as the current frame, the current point cloud frame, or the point cloud frame to be decoded. The current node can be understood as any non-leaf node in the current frame to be decoded, which is a non-empty node. In other words, the current node is not a leaf node in the octree corresponding to the current frame to be decoded, that is, the current node is any middle node in the octree, and the current node is not a non-empty node, that is, it includes at least one point.

[0403] In an embodiment of the present application, when decoding a current node in a current frame to be decoded, the decoder first determines a prediction reference frame of the current frame to be decoded, and then determines N prediction nodes of the current node in the prediction reference frame. For example, FIG12 shows a prediction node of the current node in the prediction reference frame.

[0404] It should be noted that the embodiments of the present application do not limit the number of prediction reference frames for the current frame to be decoded. For example, the current frame to be decoded may have one prediction reference frame, or the current frame to be decoded may have multiple prediction reference frames. Furthermore, the embodiments of the present application do not limit the number N of prediction nodes for the current node, and this number is determined based on actual needs.

[0405] The embodiment of the present application does not limit the specific method of determining the prediction reference frame of the current frame to be decoded.

[0406] In some embodiments, one or several decoded frames before the current frame to be decoded are determined as prediction reference frames for the current frame to be decoded.

[0407] For example, if the current frame to be decoded is a P frame, the inter-frame reference frame of the P frame includes the previous frame of the P frame (i.e., the forward frame). Therefore, the previous frame of the current frame to be decoded (i.e., the forward frame) can be determined as the predicted reference frame of the current frame to be decoded.

[0408] For another example, if the current frame to be decoded is a B frame, the inter-frame reference frames of the B frame include the previous frame of the P frame (i.e., the forward frame) and the next frame of the P frame (i.e., the backward frame). Therefore, the previous frame of the current frame to be decoded (i.e., the forward frame) can be determined as the predicted reference frame of the current frame to be decoded.

[0409] In some embodiments, one or several decoded frames following the current frame to be decoded are determined as prediction reference frames for the current frame to be decoded.

[0410] For example, if the current frame to be decoded is a B frame, the next frame after the current frame to be decoded may be determined as a prediction reference frame for the current frame to be decoded.

[0411] In some embodiments, one or several decoded frames before the current frame to be decoded, and one or several decoded frames after the current frame to be decoded, are determined as prediction reference frames for the current frame to be decoded.

[0412] For example, if the current frame to be decoded is a B frame, the previous frame and the next frame of the current frame to be decoded can be determined as prediction reference frames of the current frame to be decoded. In this case, the current frame to be decoded has two prediction reference frames.

[0413] The following takes the current frame to be decoded including K prediction reference frames as an example to introduce the specific process of determining N prediction nodes of the current node in the prediction reference frames of the current frame to be decoded in S101-A.

[0414] In some embodiments, the decoding end selects at least one prediction reference frame from the K prediction reference frames based on the placeholder information of the node in the current frame to be decoded and the placeholder information of the node in each of the K prediction reference frames, and then searches for the predicted node of the current node in the at least one prediction reference frame. For example, at least one prediction reference frame whose placeholder information of the node is closest to the placeholder information of the node in the current frame to be decoded is selected from the K prediction reference frames, and then searches for the predicted node of the current node in the at least one prediction reference frame.

[0415] In some embodiments, the decoding end may determine N predicted nodes of the current node through the following steps S101-A1 and S101-A2:

[0416] S101-A1. For a k-th prediction reference frame among K prediction reference frames, determine at least one prediction node of a current node in the k-th prediction reference frame, where k is a positive integer less than or equal to K, and K is a positive integer;

[0417] S101-A2: Determine N prediction nodes of the current node based on at least one prediction node of the current node in K prediction reference frames.

[0418] In this embodiment, the decoding end determines at least one prediction node of the current node from each of the K prediction reference frames, and finally aggregates at least one prediction node in each of the K prediction reference frames to obtain N prediction nodes of the current node.

[0419] Among them, the process of the decoding end determining at least one prediction point of the current node in each of the K prediction reference frames is the same. For the sake of convenience of description, the kth prediction reference frame among the K prediction reference frames is used as an example for explanation.

[0420] The specific process of determining at least one prediction node of the current node in the kth prediction reference frame in the above S101-A1 is introduced below.

[0421] The embodiment of the present application does not limit the specific manner in which the decoding end determines at least one prediction node of the current node in the kth prediction reference frame.

[0422] Method 1: In the kth prediction reference frame, a prediction node of the current node is determined. For example, a node in the kth prediction reference frame that has the same partition depth as the current node is determined as the prediction node of the current node.

[0423] For example, assuming that the current node is located at the third level of the octree of the current frame to be decoded, the nodes at the third level of the octree in the k-th predicted reference frame can be obtained, and then the prediction node of the current node can be determined from these nodes.

[0424] In one example, if the number of prediction nodes of the current node in the kth prediction reference frame is 1, then among the points at which the kth prediction reference frame and the current node are at the same division depth, a node whose occupancy information is the smallest different from that of the current node can be selected, recorded as node 1, and node 1 is determined as a prediction node of the current node in the kth prediction reference frame.

[0425] In another example, if the number of prediction nodes of the current node in the kth prediction reference frame is greater than 1, the node 1 determined above and at least one domain node of node 1 in the kth prediction reference frame, such as at least one domain node that is coplanar, colinear, or co-point with node 1, are determined as the prediction nodes of the current node in the kth prediction reference frame.

[0426] Method 2, in the above S101-A1, determining at least one prediction node of the current node in the k-th prediction reference frame includes the following steps S101-A11 to S101-A13:

[0427] S101-A11. In a current frame to be decoded, determine M domain nodes of a current node, where the M domain nodes include the current node, and M is a positive integer.

[0428] S101-A12, for the i-th domain node among the M domain nodes, determine the corresponding node of the i-th domain node in the k-th prediction reference frame, where i is a positive integer less than or equal to M;

[0429] S101-A13. Determine at least one prediction node of the current node in the kth prediction reference frame based on the corresponding nodes of the M domain nodes in the kth prediction reference frame.

[0430] In this implementation, before determining at least one prediction node of the current node in the kth prediction reference frame, the decoding end first determines M domain nodes of the current node in the current frame to be decoded, and the M domain nodes include the current node itself.

[0431] It should be noted that in the embodiment of the present application, there is no restriction on the specific method of determining the M domain nodes of the current node.

[0432] In one example, the M domain nodes of the current node include at least one domain node among the domain nodes that are coplanar, colinear, and co-point with the current node in the current frame to be decoded. As shown in Figure 13, the current node includes 6 coplanar nodes, 12 colinear nodes, and 8 co-point nodes.

[0433] In another example, the M domain nodes of the current node may include not only at least one domain node in the current frame to be decoded that is coplanar, colinear, and co-point with the current node, but also other nodes within the reference neighborhood range. This embodiment of the present application does not impose any restrictions on this.

[0434] Based on the above steps, the decoding end determines the M domain nodes of the current node in the current frame to be decoded, determines the corresponding node of each of the M domain nodes in the k-th prediction reference frame, and then determines at least one prediction node of the current node in the k-th prediction reference frame based on the corresponding nodes of the M domain nodes in the k-th prediction reference frame.

[0435] The embodiment of the present application does not limit the specific implementation method of S101-A13.

[0436] In one possible implementation, at least one corresponding node is selected from the corresponding nodes of the M domain nodes in the k-th prediction reference frame as the at least one prediction node of the current node in the k-th prediction reference frame. For example, at least one corresponding node whose placeholder information has the smallest difference between the placeholder information of the M domain nodes in the k-th prediction reference frame and the placeholder information of the current node is selected from the corresponding nodes of the M domain nodes in the k-th prediction reference frame as the at least one prediction node of the current node in the k-th prediction reference frame. The method for determining the difference between the placeholder information of the corresponding node and the placeholder information of the current node can refer to the above-mentioned process for determining the difference in placeholder information, for example, performing an XOR operation on the placeholder information of the corresponding node and the placeholder information of the current node, and using the XOR operation result as the difference between the placeholder information of the corresponding node and the placeholder information of the current node.

[0437] In another possible implementation, the decoding end determines the corresponding nodes of the M domain nodes in the kth prediction reference frame as at least one prediction node for the current node in the kth prediction reference frame. For example, each of the M domain nodes has a corresponding node in the kth prediction reference frame, resulting in M ​​corresponding nodes. These M corresponding nodes are determined as the prediction nodes for the current node in the kth prediction reference frame, for a total of M prediction nodes.

[0438] The above describes the process of determining at least one prediction node for the current node in the kth prediction reference frame. Thus, the decoder can use the same method as above to determine at least one prediction node for the current node in each of the K prediction reference frames.

[0439] For example, if the current frame to be decoded is a P frame, the K predicted reference frames include the forward frame of the current frame to be decoded. At this point, the decoding end can determine at least one prediction node of the current node in the forward frame based on the above steps. Exemplarily, as shown in Figure 15A, it is assumed that the current node includes three domain nodes, which are respectively recorded as node 11, node 12 (current node) and node 13. These three domain nodes correspond to a corresponding node in the forward frame, respectively recorded as node 21, node 22 and node 23, and then node 21, node 22 and node 23 are determined as the three prediction nodes of the current node in the forward frame, or 1 or 2 nodes are selected from node 21, node 22 and node 23 to be determined as 1 or 2 prediction nodes of the current node in the forward frame.

[0440] For another example, if the current frame to be decoded is a B frame, the K prediction reference frames include the forward frame and the backward frame of the current frame to be decoded. At this time, based on the above steps, the decoding end can determine at least one prediction node of the current node in the forward frame, and at least one prediction node of the current node in the backward frame. For example, as shown in Figure 15B, it is assumed that the current node includes three domain nodes, respectively recorded as node 11, node 12, and node 13. These three domain nodes correspond to one corresponding node in the forward frame, respectively, recorded as node 21, node 22, and node 23. These three domain nodes correspond to one corresponding node in the backward frame, respectively, recorded as node 41, node 42, and node 43. In this way, the decoding end can determine node 21, node 22, and node 23 as the three prediction nodes of the current node in the forward frame, or select one or two nodes from node 21, node 22, and node 23 to determine as one or two prediction nodes of the current node in the forward frame. Similarly, the decoding end can determine node 41, node 42 and node 43 as the three prediction nodes of the current node in the backward frame, or select one or two nodes from node 41, node 42 and node 43 as one or two prediction nodes of the current node in the backward frame.

[0441] After the decoding end determines at least one prediction node of the current node in each of the K prediction reference frames, it performs the above step S101-B, that is, determines N prediction nodes of the current node based on at least one prediction node of the current node in the K prediction reference frames.

[0442] In one example, at least one prediction node of the current node in K prediction reference frames is determined as N prediction nodes of the current node.

[0443] For example, K=2, that is, the K prediction reference frames include the first prediction reference frame and the second prediction reference frame. Assume that the current node has 2 prediction nodes in the first prediction reference frame and 3 prediction nodes in the second prediction reference frame. In this way, it can be determined that the current node has 5 prediction nodes, and N=5.

[0444] In another example, N prediction nodes of the current node are screened out from at least one prediction node of the current node in K prediction reference frames.

[0445] Continuing with the above example, assume K = 2, meaning the K prediction reference frames include the first prediction reference frame and the second prediction reference frame. Assume the current node has two prediction nodes in the first prediction reference frame and three prediction nodes in the second prediction reference frame. From these five prediction nodes, select three prediction nodes as the final prediction nodes for the current node. For example, from these five prediction nodes, select the three prediction nodes whose placeholder information differs minimally from the placeholder information of the current node and determine them as the final prediction nodes for the current node.

[0446] In the second method, after the decoding end determines the M domain nodes of the current node in the current frame to be decoded, it determines the corresponding node of each of the M domain nodes in the kth prediction reference frame, and then determines at least one prediction point of the current node in the kth prediction reference frame based on the corresponding node of each of the M domain nodes.

[0447] Mode 3, in the above S101-A1, determining at least one prediction node of the current node in the k-th prediction reference frame includes the following steps S101-B11 to S101-B13:

[0448] S101-B11, determining the corresponding node of the current node in the kth prediction reference frame;

[0449] S101-B12, determining at least one domain node of the corresponding node;

[0450] S101-B13. Determine at least one domain node as at least one prediction node of the current node in the k-th prediction reference frame.

[0451] In this method 3, for each of the K predicted reference frames, the decoding end first determines the corresponding node of the current node in each predicted reference frame. For example, the corresponding node 1 of the current node in the predicted reference frame 1 is determined, and the corresponding node 2 of the current node in the predicted reference frame 2 is determined. Then, the decoding end determines at least one domain node of each corresponding node. For example, at least one domain node of the corresponding node 1 is determined in the predicted reference frame 1, and at least one domain node of the corresponding node 2 is determined in the predicted reference frame 2. In this way, at least one domain node of the corresponding node 1 in the predicted reference frame 1 can be determined as at least one predicted node of the current node in the predicted reference frame 1, and at least one domain node of the corresponding node 2 in the predicted reference frame 2 can be determined as at least one predicted node of the current node in the predicted reference frame 2.

[0452] Determining the corresponding node of the i-th domain node in the k-th prediction reference frame in S101-A12 of the second method is essentially the same as determining the corresponding node of the current node in the k-th prediction reference frame in S101-B11 of the third method described above. For ease of description, the i-th domain node and the current node are referred to as the i-th node. The specific process of determining the corresponding node of the i-th node in the k-th prediction reference frame is described below.

[0453] The decoding end determines the corresponding node of the i-th node in the k-th prediction reference frame in at least the following ways:

[0454] In method 1, a node in the k-th prediction reference frame that has the same division depth as the i-th node is determined as the corresponding node of the i-th node.

[0455] For example, assuming that the i-th node is located at the third level of the octree of the current frame to be decoded, the nodes at the third level of the octree in the k-th prediction reference frame can be obtained, and the corresponding node of the i-th node can be determined from these nodes. For example, among the points in the k-th prediction reference frame that are at the same partition depth as the i-th node, the node whose placeholder information differs the least from that of the i-th node is selected and determined as the corresponding node of the i-th node in the k-th prediction reference frame.

[0456] Mode 2: The above-mentioned S101-A12 and S101-B11 include the following steps:

[0457] S101-A121, in the current frame to be decoded, determine the parent node of the i-th node as the i-th parent node;

[0458] S101-A122, determine the matching node of the i-th parent node in the k-th prediction reference frame as the i-th matching node;

[0459] S101-A123: Determine one of the child nodes of the i matching nodes as the corresponding node of the i-th node in the k-th prediction reference frame.

[0460] In this method 2, for the i-th node, the decoding end determines the parent node of the i-th node in the current frame to be decoded, and then determines the matching node of the parent node of the i-th prediction domain node in the k-th prediction reference frame. For ease of description, the parent node of the i-th node is recorded as the i-th parent node, and the matching node of the parent node of the i-th node in the k-th prediction reference frame is determined as the i-th matching node. Then, a child node of the child node of the i-th matching node is determined as the corresponding node of the i-th node in the k-th prediction reference frame, thereby accurately determining the corresponding node of the i-th node in the k-th prediction reference frame.

[0461] The specific process of determining the matching node of the i-th parent node in the k-th prediction reference frame in the above S101-A122 is introduced below.

[0462] The embodiment of the present application does not limit the specific method by which the decoding end determines the matching node of the i-th parent node in the k-th prediction reference frame.

[0463] In some embodiments, the partition depth of the i-th parent node in the current frame to be decoded is determined, for example, the i-th parent node is at the second level of the octree of the current frame to be decoded. In this way, the decoding end can determine one of the nodes in the k-th prediction reference frame that have the same partition depth as the i-th parent node as the matching node of the i-th parent node in the k-th prediction reference frame. For example, one of the nodes in the second level of the k-th prediction reference frame can be determined as the matching node of the i-th parent node in the k-th prediction reference frame.

[0464] In some embodiments, the decoding end determines a matching node for the i-th parent node in the k-th predicted reference frame based on the placeholder information of the i-th parent node. Specifically, since the placeholder information for the i-th parent node in the current frame to be decoded has been decoded, and the placeholder information for each node in the k-th predicted reference frame has also been decoded, the decoding end can search for a matching node for the i-th parent node in the k-th predicted reference frame based on the placeholder information of the i-th parent node.

[0465] For example, the node with the smallest difference between the placeholder information of the k-th prediction reference frame and the placeholder information of the i-th parent node is determined as the matching node of the i-th parent node in the k-th prediction reference frame.

[0466] For example, assuming the placeholder information of the i-th parent node is 11001101, the k-th predicted reference frame is searched for the node whose placeholder information has the smallest difference from the placeholder information 11001101. Specifically, the decoder performs an XOR operation on the placeholder information of the i-th parent node and the placeholder information of each node in the k-th predicted reference frame. The node with the smallest XOR result in the k-th predicted reference frame is determined as the matching node of the i-th parent node in the k-th predicted reference frame.

[0467] For example, assuming that the occupancy information of node 1 in the k-th predicted reference frame is 10001101, 11001101 and 10001101 are XORed, where the first bit of 11001101 and the first bit of 10001101 are both 1. Therefore, the XOR result of the first bit of the two is 0, the second bit of 11001101 is different from the second bit of 10001111, so the XOR result of the second bit of the two is 1, and so on. The XOR result of 11001101 and 10001111 is 0+1+0+0+0+0+1+0=2. According to this method, the decoding end can determine the XOR operation result of the occupancy information of the i-th parent node and the occupancy information of each node in the k-th predicted reference frame, and then determine the node in the k-th predicted reference frame with the smallest XOR operation with the occupancy information of the i-th parent node as the matching node of the i-th parent node in the k-th predicted reference frame.

[0468] Based on the above steps, the decoding end can determine the matching node of the i-th parent node in the k-th prediction reference frame. For ease of description, this matching node is recorded as the i-th matching node.

[0469] Next, the decoding end determines one of the child nodes of the i-th matching node as the corresponding node of the i-th domain node in the k-th prediction reference frame.

[0470] For example, the decoding end determines a default child node among the child nodes included in the i-th matching node as the corresponding node of the i-th node in the k-th prediction reference frame. Assume that the first child node of the i-th matching node is determined as the corresponding node of the i-th node in the k-th prediction reference frame.

[0471] For another example, the decoding end determines the first sequence number of the i-th node among the child nodes included in the parent node; and determines the child node with the first sequence number among the child nodes of the i-th matching node as the corresponding node of the i-th node in the k-th prediction reference frame. For example, as shown in FIG14 , the i-th node is the second child node of the i-th parent node, and in this case, the first sequence number is 2. In this way, the second child node of the i-th matching node can be determined as the corresponding node of the i-th node.

[0472] The above describes the process of determining the corresponding node of the i-th domain node among M domain nodes in the k-th prediction reference frame, and the corresponding node of the current node in the k-th prediction reference frame. Thus, the decoder can use Method 2 or Method 3 to determine the N prediction nodes for the current node in the prediction reference frame.

[0473] Based on the above steps, the decoding end determines N prediction nodes of the current node in the prediction reference frame of the current frame to be decoded, and then performs the following step S102.

[0474] S102 : Based on the geometric decoding information of the midpoints of the N predicted nodes, predictively decode the position information of the midpoint of the current node.

[0475] Due to the correlation between adjacent frames of the point cloud, the embodiment of the present application refers to the relevant information between frames when predicting the position information of the current node midpoint based on the correlation between adjacent frames of the point cloud. Specifically, the WeChat information of the current node midpoint is predictively encoded based on the geometric decoding information of the N predicted node midpoints of the current node, thereby improving the encoding and decoding efficiency and performance of the point cloud.

[0476] In one example, as shown in Figure 16A , the encoder's process for directly encoding the current node includes determining whether the current node is eligible for direct encoding. If so, setting IDCMEligible to true. Next, determining whether the number of points in the current node is less than a preset threshold. If so, direct encoding is used for the current node, specifically directly encoding the number of points in the current node and the geometric information of the midpoints in the current node.

[0477] Correspondingly, when decoding the current node, as shown in Figure 16B , the decoder first determines whether the current node is eligible for direct decoding. If so, it sets IDCMEligible to true. Next, it decodes the geometric information of the midpoint of the current node.

[0478] It should be noted that, in the embodiments of the present application, predictive decoding of the coordinate information of the midpoint of the current node based on the geometric decoding information of the N predicted nodes can be understood as using the geometric decoding information of the N predicted nodes as context to predictively decode the coordinate information of the midpoint of the current node. For example, the decoding end determines the index of the context model based on the geometric decoding information of the N predicted nodes, and then, based on the index of the context model, determines the target context model from multiple preset context models, and uses the context model to predictively decode the coordinate information of the midpoint of the current node.

[0479] In an embodiment of the present application, the process of predicting and decoding the coordinate information of each point in the current node based on the geometric decoding information of N predicted nodes is basically the same. For the sake of convenience of description, the predicting and decoding of the coordinate information of the current point in the current node is taken as an example to illustrate.

[0480] In some embodiments, the above S102 includes the following steps:

[0481] S102-A, determining an index of a context model based on geometric decoding information of N prediction nodes;

[0482] S102-B, determining a context model based on the context model index;

[0483] S102-C: Use the context model to predict and decode the coordinate information of the current point in the current node.

[0484] In an embodiment of the present application, multiple context models, for example, Q context models, are set for the decoding process of the coordinate information. The embodiment of the present application does not limit the specific number of context models corresponding to the coordinate information, as long as Q is greater than 1. That is, in an embodiment of the present application, an optimal context model is selected from at least two context models to perform predictive decoding on the coordinate information of the current point in the current node, so as to improve the decoding efficiency of the coordinate information of the current point.

[0485] For example, the coordinate information corresponds to multiple context models as shown in Table 4:

[0486] Index context model 0 context model A1 context model B…………

[0487] In this way, the decoder determines the index of the context model based on the geometric decoding information of the N prediction nodes. Then, based on the index of the context model, a context model is selected from the corresponding context models in Table 4 to perform predictive decoding on the coordinate information of the current point in the current node.

[0488] In the embodiments of the present application, the geometric decoding information of the prediction node can be understood as any information involved in the geometric decoding process of the prediction node, including, for example, the number of points included in the prediction node, the placeholder information of the prediction node, the decoding method of the prediction node, the geometric information of the midpoint of the prediction node, etc.

[0489] In some embodiments, the geometric decoding information of the prediction node includes direct decoding information of the prediction node and / or coordinate information of a midpoint of the prediction node, wherein the direct decoding information of the prediction node is used to indicate whether the prediction node meets the conditions for decoding in a direct decoding manner.

[0490] Based on this, the above S102-A includes the following steps S102-A1:

[0491] S102-A1. Determine a first context index based on direct decoding information of the N prediction nodes, and / or determine a second context index based on coordinate information of midpoints of the N prediction nodes.

[0492] Correspondingly, the above S102-B includes the following steps S102-B:

[0493] S102-B1. Select a context model from a plurality of preset context models based on the first context index and / or the second context index.

[0494] In this embodiment, if the geometric decoding information of the prediction node includes the direct decoding information of the prediction node and / or the coordinate information of the midpoint of the prediction node, the decoding end can determine the first context index based on the direct decoding information of the N prediction nodes, and / or determine the second context index based on the position information of the midpoint of the N prediction nodes, and then select the final context model from the preset multiple context models based on the first context index and / or the second context index.

[0495] It can be seen that in this embodiment, the decoding end determines the context model in the following ways, but is not limited to:

[0496] In one possible implementation, if the geometric decoding information of the prediction node includes the direct decoding information of the prediction node, the process of determining the context model may be to determine the first context index based on the direct decoding information of N prediction nodes, and then based on the first context index, select the final context model from the preset multiple context models to decode the coordinate information of the current point.

[0497] For example, the decoding end selects the final context model from the context models shown in Table 4 based on the first context index.

[0498] In another possible implementation, if the geometric decoding information of the prediction node includes the coordinate information of the midpoint of the prediction node, the process of determining the context model may be to determine the second context index based on the coordinate information of the midpoints of N prediction nodes, and then based on the second context index, select the final context model from the preset multiple context models to decode the coordinate information of the current point.

[0499] For example, the decoding end selects the final context model from the context models shown in Table 4 based on the second context index.

[0500] In another possible implementation, if the geometric decoding information of the prediction node includes the direct decoding information of the prediction node and the coordinate information of the midpoint of the prediction node, the process of determining the context model may be to determine the first context index based on the direct decoding information of N prediction nodes, and then determine the second context index based on the first context index and the coordinate information of the midpoint of the N prediction nodes, and then select the final context model from the preset multiple context models based on the second context index, the first context index and the second context index to decode the coordinate information of the current point.

[0501] Exemplarily, the correspondence between the first context index, the second context index, and the context model is shown in Table 5:

[0502] Table 5

[0503] Second context index 1, second context index 2, second context index 3, …, first context index 1, context model 11, context model 12, context model 13, …, first context index 2, context model 21, context model 22, context model 23, …, first context index 3, context model 31, context model 32, context model 33, ……………………………

[0504] In this approach 3, the decoder determines the first context index based on the directly decoded information of the N prediction nodes, and the second context index based on the coordinate information of the midpoint of the N prediction nodes. Then, the decoder consults Table 5 to obtain the final context model. For example, the decoder determines the first context index to be first context index 2 based on the directly decoded information of the N prediction nodes, and determines the second context index to be second context index 3 based on the coordinate information of the midpoint of the N prediction nodes. By consulting Table 5, the final context model obtained is context model 23. The decoder then uses context model 23 to decode the coordinate information of the current point.

[0505] The following describes a specific process of determining the first context index based on the direct decoding information of the N prediction nodes in S102-A1.

[0506] In the embodiment of the present application, the decoding end determines the first context index in the following ways but is not limited to:

[0507] Method 1: The above S102-A1 includes the following steps S102-A1-11 and S102-A1-12:

[0508] S102-A1-11. For any prediction node among the N prediction nodes, determine a first numerical value corresponding to the prediction node based on direct decoding information of the prediction node.

[0509] In this manner, for each of the N prediction nodes, a first value corresponding to the prediction node is determined based on direct decoding information of the prediction node, and finally a first context index is determined based on the first values ​​corresponding to the N prediction nodes.

[0510] The following describes the process of determining the first value corresponding to the prediction node.

[0511] As can be seen from the above, the direct decoding information of the prediction node is used to indicate whether the prediction node meets the conditions for decoding in the direct decoding mode. The embodiment of the present application does not limit the specific content of the direct decoding information.

[0512] In some embodiments, the direct decoding information includes the number of points included in the prediction node. In this way, the first numerical value corresponding to the prediction node can be determined based on the number of points included in the prediction node.

[0513] In one example, under the GPCC framework, when the number of points included in the prediction node is greater than or equal to 2, the first value corresponding to the prediction node is determined to be 1; if the number of points included in the prediction node is less than 2, the first value corresponding to the prediction node is determined to be 0. Under the AVS framework, when the number of points included in the prediction node is greater than or equal to 1, the first value corresponding to the prediction node is determined to be 1; if the number of points included in the prediction node is less than 1, the first value corresponding to the prediction node is determined to be 0.

[0514] In another example, the number of points included in the prediction node is determined as the first value corresponding to the prediction node. For example, when the prediction node includes 2 points, the first value corresponding to the prediction node is determined to be 2.

[0515] In some embodiments, the direct decoding information of the prediction node includes a direct decoding mode of the prediction node. In this case, the above S102-A1-11 includes: numbering the direct decoding mode of the prediction node and determining a first value corresponding to the prediction node.

[0516] For example, in the GPCC framework, if the direct decoding mode of the prediction node is mode 0, the first value corresponding to the prediction node is determined to be 0. If the direct decoding mode of the prediction node is mode 1, the first value corresponding to the prediction node is determined to be 1. If the direct decoding mode of the prediction node is mode 2, the first value corresponding to the prediction node is determined to be 2.

[0517] For another example, in the AVS framework, if the direct decoding mode of the prediction node is mode 0, the first value corresponding to the prediction node is determined to be 0. If the direct decoding mode of the prediction node is mode 1, the first value corresponding to the prediction node is determined to be 1.

[0518] Based on the above steps, after the decoding end determines the first value corresponding to each of the N prediction nodes, it executes the following step S102-A1-12.

[0519] S102-A1-12. Determine a first context index based on first values ​​corresponding to the N prediction nodes.

[0520] After determining the first values ​​corresponding to the N prediction nodes based on the above steps, the decoding end determines a first context index based on the first values ​​corresponding to the N prediction nodes.

[0521] Determining the first context index based on the first values ​​corresponding to the N prediction nodes includes at least the following implementations:

[0522] In method 1, an average value of the sum of the first numerical values ​​corresponding to the N prediction nodes is determined as the first context index.

[0523] Mode 2, S102-A1-12 includes the following steps S102-A1-121 to S102-A1-123:

[0524] S102-A1-121, determine a first weight corresponding to the prediction node;

[0525] S102-A1-122. Perform weighted processing on the first values ​​corresponding to the N prediction nodes based on the first weight to obtain a first weighted prediction value;

[0526] S102-A1-123. Determine a first context index based on the first weighted prediction value.

[0527] In method 2, if the current node includes multiple prediction nodes, i.e., N prediction nodes, when determining the first context index based on the first numerical values ​​corresponding to the N prediction nodes, a weight, i.e., the first weight, can be determined for each of the N prediction nodes. In this way, the first numerical values ​​corresponding to each prediction node can be weighted based on the first weight of each prediction node, and then the first context index can be determined based on the final weighted result, thereby improving the accuracy of determining the first context index based on the geometric decoding information of the N prediction nodes.

[0528] The embodiment of the present application does not limit the determination of the first weights corresponding to the N prediction nodes.

[0529] In some embodiments, the first weight corresponding to each of the N prediction nodes is a preset value. As can be seen from the above, the N prediction nodes are determined based on the M domain nodes of the current node. Assuming that prediction node 1 is the prediction node corresponding to domain node 1, if domain node 1 is a coplanar node with the current node, then the first weight of prediction node 1 is the preset weight 1. If domain node 1 is a colinear node with the current node, then the first weight of prediction node 1 is the preset weight 2. If domain node 1 is a co-point node with the current node, then the first weight of prediction node 1 is the preset weight 3.

[0530] In some embodiments, for each of the N prediction nodes, a first weight corresponding to the prediction node is determined based on the distance between the domain node corresponding to the prediction node and the current node. For example, the smaller the distance between the domain node and the current node, the stronger the inter-frame correlation between the prediction node corresponding to the domain node and the current node, and thus the greater the first weight of the prediction node.

[0531] For example, taking prediction node 1 among N prediction nodes as an example, assuming that prediction node 1 is the corresponding point of domain node 1 among the M domain nodes of the current node in the prediction reference frame, the first weight of prediction node 1 can be determined based on the distance between domain node 1 and the current node. For example, the inverse of the distance between domain node 1 and the current node is determined as the first weight of prediction node 1.

[0532] In one example, if domain node 1 is a coplanar node of the current node, the first weight of the predicted node 1 is 1; if domain node 1 is a colinear node of the current node, the first weight of the predicted node 1 is a preset weight. If domain node 1 is a common node of the current node, the first weight of predicted node 1 is the preset weight

[0533] In one example, if domain node 1 is a coplanar node of the current node, the first weight of predicted node 1 is If domain node 1 is a collinear node of the current node, the first weight of prediction node 1 is the preset weight If domain node 1 is a common node of the current node, the first weight of predicted node 1 is the preset weight

[0534] In some embodiments, based on the above steps, after determining the weight corresponding to each prediction node in the N prediction nodes, the weight is normalized, and the normalized weight is used as the final first weight of the prediction node.

[0535] The embodiment of the present application does not limit the specific method of obtaining the first weighted prediction value by weighting the first numerical values ​​corresponding to N prediction nodes based on the first weight.

[0536] In one example, based on the first weight, a weighted average is performed on the first values ​​corresponding to the N prediction nodes to obtain a first weighted prediction value.

[0537] In another example, based on the first weight, a weighted sum is performed on the first numerical values ​​corresponding to the N prediction nodes to obtain a first weighted prediction value.

[0538] After determining the first weighted prediction value based on the method steps, the first context index is determined based on the first weighted prediction value. That is, the above S102-A1-123 includes at least the following examples:

[0539] Example 1: Determine the first weighted prediction value as the first context index.

[0540] Example 2: Determine the weighted prediction value range in which the first weighted prediction value is located, and determine the index corresponding to the range as the first context index.

[0541] The above describes the process of obtaining the first context index by performing weighted processing on N prediction nodes at the decoding end.

[0542] In some embodiments, the decoding end may also adopt the following second method to determine the first context index.

[0543] Method 2: If K is greater than 1, determine the second weighted prediction value corresponding to each of the K prediction reference frames, and then determine the first context index based on the second weighted prediction values ​​corresponding to the K prediction reference frames. In this case, the above S102-A1 includes the following steps S102-A1-21 to S102-A1-23:

[0544] S102-A1-21. For a j-th prediction reference frame among the K prediction reference frames, determine, based on direct decoding information of a prediction node of a current node in the j-th prediction reference frame, a first value corresponding to the prediction node in the j-th prediction reference frame, where j is a positive integer less than or equal to K.

[0545] S102-A1-22, determining a first weight corresponding to the prediction node, and performing weighted processing on the first value corresponding to the prediction node in the j-th prediction reference frame based on the first weight to obtain a second weighted prediction value corresponding to the j-th prediction reference frame;

[0546] S102-A1-23. Determine a first context index based on second weighted prediction values ​​corresponding to the K predicted reference frames.

[0547] In this second approach, when determining the first context index, each of the K predicted reference frames is considered separately as separate context information. Specifically, direct decoding information of the prediction nodes included in each of the K predicted reference frames is determined, and a second weighted prediction value corresponding to each predicted reference frame is determined. Furthermore, based on the second weighted prediction value corresponding to each predicted reference frame, the first context index is determined, thereby achieving accurate selection of the first context index and improving point cloud decoding efficiency.

[0548] In the embodiment of the present application, the specific method in which the decoding end determines the second weighted prediction value corresponding to each of the K prediction reference frames is the same. For the sake of convenience of description, the j-th prediction reference frame among the K prediction reference frames is used as an example for illustration.

[0549] In an embodiment of the present application, the current node includes at least one prediction node in the j-th prediction reference frame, so the first value of the at least one prediction node is determined based on the direct decoding information of the at least one prediction node in the j-th prediction reference frame.

[0550] For example, the j-th prediction reference frame includes two prediction nodes of the current node, which are respectively recorded as prediction node 1 and prediction node 2. Then, based on the direct decoding information of prediction node 1, the first value of prediction node 1 is determined, and based on the direct decoding information of prediction node 2, the first value of prediction node 2 is determined. The process of determining the first value corresponding to the prediction node based on the direct decoding information of the prediction node can refer to the description of the above embodiment. Exemplarily, the first value corresponding to the prediction node is determined based on the direct decoding mode of the prediction node. For example, the number of the direct decoding mode of the prediction node (0, 1, or 2) is determined as the first value corresponding to the prediction node.

[0551] After the decoding end determines the first value of at least one prediction node included in the j-th prediction reference frame, it determines the first weight corresponding to each of the at least one prediction node, and performs weighted processing on the first value corresponding to the at least one prediction node based on the first weight to obtain a second weighted prediction value corresponding to the j-th prediction reference frame.

[0552] In one example, based on the first weight, a weighted average is performed on the first values ​​corresponding to the prediction nodes in the j-th prediction reference frame to obtain a second weighted prediction value corresponding to the j-th prediction reference frame.

[0553] In another example, based on the first weight, a weighted sum is performed on the first values ​​corresponding to the prediction nodes in the j-th prediction reference frame to obtain a second weighted prediction value corresponding to the j-th prediction reference frame.

[0554] The process of determining the first weight may refer to the description of the above embodiment and will not be repeated here.

[0555] The above introduces the process of determining the second weighted prediction value corresponding to the j-th prediction reference frame among the K prediction reference frames. The second weighted prediction values ​​corresponding to other prediction reference frames among the K prediction reference frames are determined in accordance with the method corresponding to the j-th prediction reference frame.

[0556] After the decoding end determines the second weighted prediction value corresponding to each of the K predicted reference frames, it executes the above step S102-A1-23.

[0557] The present application does not limit the specific method of determining the first context index based on the second weighted prediction value corresponding to K prediction reference frames.

[0558] In some embodiments, the decoding end determines an average value of the second weighted prediction values ​​corresponding to the K prediction reference frames as the first context index.

[0559] In some embodiments, the decoding end determines second weights corresponding to the K predicted reference frames, and performs weighted processing on second weighted prediction values ​​corresponding to the K predicted reference frames based on the second weights to obtain a first context index.

[0560] In this embodiment, the decoding end first determines the second weight corresponding to each of the K prediction reference frames. The embodiment of the present application does not limit the determination of the second weight corresponding to each of the K prediction reference frames.

[0561] In some embodiments, the second weight corresponding to each of the K predicted reference frames is a preset value. As can be seen from the above, the K predicted reference frames are forward frames and / or backward frames of the current frame to be decoded. Assuming that predicted reference frame 1 is the forward frame of the current frame to be decoded, the second weight corresponding to predicted reference frame 1 is the preset weight 1. If predicted reference frame 1 is the backward frame of the current frame to be decoded, the second weight corresponding to predicted reference frame 1 is the preset weight 2.

[0562] In some embodiments, the second weight corresponding to the predicted reference frame is determined based on the time difference between the predicted reference frame and the current frame to be decoded. In an embodiment of the present application, each point cloud includes time information, and the time information can be the time when the point cloud acquisition device acquires the point cloud of the frame. Based on this, the smaller the time difference between the predicted reference frame and the current frame to be decoded, the stronger the inter-frame correlation between the predicted reference frame and the current frame to be decoded, and thus the larger the second weight corresponding to the predicted reference frame. For example, the inverse of the time difference between the predicted reference frame and the current frame to be decoded can be determined as the second weight corresponding to the predicted reference frame.

[0563] After determining the second weight corresponding to each of the K prediction reference frames, weighted processing is performed on the second weighted prediction values ​​corresponding to the K prediction reference frames based on the second weight to obtain a first context index.

[0564] For example, assuming K=2, for example, the current frame to be decoded includes 2 prediction reference frames, and these 2 prediction reference frames include the forward frame and backward frame of the current frame to be decoded. Assuming that the second weight corresponding to the forward frame is W1 and the second weight corresponding to the backward frame is W2, based on W1 and W2, the second weighted prediction value corresponding to the forward frame and the second weighted prediction value corresponding to the backward frame are weighted to obtain the first context index.

[0565] In one example, based on the second weight, weighted averaging is performed on the second weighted prediction values ​​corresponding to the K prediction reference frames to obtain the first context index.

[0566] In another example, based on the second weight, the second weighted prediction values ​​corresponding to the K prediction reference frames are weightedly summed to obtain the first context index.

[0567] The above describes the process of determining the first context index at the decoding end.

[0568] The following describes the process of determining the second context index at the decoding end.

[0569] From the above process of determining the prediction node, it can be seen that each of the N prediction nodes includes one point or multiple points. If each of the N prediction nodes includes one point, the one point included in each prediction node is used to determine the second context index.

[0570] In some embodiments, if the predicted node includes multiple points, one point is selected from the multiple points to determine the second context index. In this case, the above S102-A1 includes the following steps S102-A1-31 and S102-A1-32:

[0571] S102-A1-31. For any prediction node among the N prediction nodes, select the first point corresponding to the current point of the current node from the points included in the prediction node;

[0572] S102-A1-32. Determine a second context index based on the coordinate information of the first point included in the N prediction nodes.

[0573] For example, assume that N prediction nodes include prediction node 1 and prediction node 2, where prediction node 1 includes points 1 and 2, and prediction node 2 includes points 3, 4, and 5. Then, a point is selected from points 1 and 2 included in prediction node 1 as the first point, and a point is selected from points 3, 4, and 5 included in prediction node 2 as the first point. In this way, the geometric information of the current point can be determined based on the geometric information of the first point in prediction node 1 and the first point in prediction node 2.

[0574] The embodiment of the present application does not limit the specific method of selecting the first point corresponding to the current point of the current node from the points included in the prediction node.

[0575] In one possible implementation, the points in the predicted node that are in the same order as the current point are determined as the first point corresponding to the current point. For example, assuming the current point is the second point in the current node, point 2 in predicted node 1 can be determined as the first point corresponding to the current point, and point 4 in predicted node 2 can be determined as the first point corresponding to the current point. For another example, if the predicted node includes only one point, the point included in the predicted node is determined as the first point corresponding to the current point.

[0576] In one possible implementation, if the encoder selects the first point corresponding to the current point from the points included in the prediction node based on the rate-distortion cost (or approximate cost), the encoder writes the identification information of the first point in the prediction node into the bitstream, so that the decoder obtains the first point in the prediction node by decoding the bitstream.

[0577] The decoding end determines, for each of the N prediction nodes, the first point corresponding to the current point in each prediction node based on the above method, and then executes the above step S102-A1-32.

[0578] In the embodiment of the present application, the decoding end encodes the coordinate information of the current point on different coordinate axes separately. Based on, the above S102-A1-32 includes the following step S102-A1-321:

[0579] S102-A1-321. Determine a second context index corresponding to the i-th coordinate axis based on coordinate information of the first point included in the N prediction nodes on the i-th coordinate axis.

[0580] The above-mentioned i-th coordinate axis can be the X-axis, Y-axis or Z-axis, and this embodiment of the application does not limit this.

[0581] In some embodiments, if the point cloud is a lidar point cloud, then it can be seen from the above that the i-th coordinate axis is the X-axis or the Y-axis.

[0582] In some embodiments, if the point cloud is a point cloud facing the human eye, it can be seen from the above that the i-th coordinate axis can be any one of the X-axis, Y-axis or Z-axis.

[0583] In an embodiment of the present application, when the decoding end decodes the coordinate information of the current point on the i-th coordinate axis, it determines the second context index corresponding to the i-th coordinate axis based on the coordinate information of the first point on the i-th coordinate axis included in the N prediction nodes. In this way, based on the first context index and / or the second context index corresponding to the i-th coordinate axis, the context model corresponding to the i-th coordinate axis can be selected from multiple context models, and then the context model corresponding to the i-th coordinate axis can be used to predict and decode the coordinate information of the current point on the i-th coordinate axis. For example, the decoding end determines the second context index corresponding to the X-axis based on the coordinate information of the first point on the X-axis included in the N prediction nodes, and based on the first context index and / or the second context index corresponding to the X-axis, selects the context model corresponding to the X-axis from multiple context models, and then uses the context model corresponding to the X-axis to predict and decode the coordinate information of the current point on the X-axis to obtain the X-coordinate value of the current point. For another example, the decoding end determines the second context index corresponding to the Y-axis based on the coordinate information of the first point on the Y-axis included in the N prediction nodes, and based on the first context index and / or the second context index corresponding to the Y-axis, selects the context model corresponding to the Y-axis from multiple context models, and then uses the context model corresponding to the Y-axis to predict and decode the coordinate information of the current point on the Y-axis to obtain the Y coordinate value of the current point.

[0584] The following describes a process in which the decoding end determines the second context index corresponding to the i-th coordinate axis based on the coordinate information of the first point included in the N prediction nodes on the i-th coordinate axis.

[0585] In the embodiment of the present application, the implementation methods of the above S102-A1-321 include but are not limited to the following:

[0586] Method 1: Weight the first points included in the N prediction nodes, and determine the second context index corresponding to the i-th coordinate axis based on the weighted coordinate information. In this case, the above S102-A1-321 includes the following steps S102-A1-321-11 to S102-A1-321-13:

[0587] S102-A1-321-11. Determine a first weight corresponding to the prediction node;

[0588] S102-A1-321-12. Based on the first weight, weighted processing is performed on the coordinate information of the first point included in the N prediction nodes to obtain a first weighted point;

[0589] S102-A1-321-13. Determine a second context index corresponding to the i-th coordinate axis based on the coordinate information of the first weighted point on the i-th coordinate axis.

[0590] In this first approach, if the current node includes multiple prediction nodes, i.e., N prediction nodes, when determining the second context index corresponding to the i-th coordinate axis based on the coordinate information of the first point included in the N prediction nodes, a weight, i.e., a first weight, can be determined for each of the N prediction nodes. In this way, based on the first weight of each prediction node, the coordinate information of the first point included in each prediction node can be weighted to obtain a first weighted point. Furthermore, based on the coordinate information of the first weighted point on the i-th coordinate axis, the second context index corresponding to the i-th coordinate axis can be determined, thereby improving the accuracy of decoding the current point based on the geometric decoding information of the N prediction nodes.

[0591] The process of determining the first weights corresponding to the N prediction nodes in the embodiment of the present application can refer to the description of the above embodiment and will not be repeated here.

[0592] After determining the first weight corresponding to each of the N prediction nodes, the decoding end performs weighted processing on the coordinate information of the first point included in the N prediction nodes based on the first weight to obtain a first weighted point.

[0593] The embodiment of the present application does not limit the specific method of obtaining the first weighted point by weighting the coordinate information of the first point included in the N prediction nodes based on the first weight.

[0594] In one example, based on the first weight, weighted averaging is performed on the coordinate information of the first point included in the N prediction nodes to obtain a first weighted point.

[0595] Based on the method steps, after the first weighted point is determined, the second context index corresponding to the i-th coordinate axis is determined based on the coordinate information of the first weighted point on the i-th coordinate axis.

[0596] As can be seen from the above, the first weighted point is obtained by weighting the first point in the N prediction nodes, wherein the values ​​of each bit of the first point after the prediction node have only two results, 0 or 1. Therefore, in some embodiments, the values ​​of each bit of the first weighted point obtained by weighting the first point in the N prediction nodes are also 0 or 1. In this way, when decoding the i-th bit of the current point on the i-th coordinate axis, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined based on the value of the i-th bit of the first weighted point on the i-th coordinate axis. For example, if the value of the i-th bit of the first weighted point on the i-th coordinate axis is 0, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined to be 0. For another example, if the value of the i-th bit of the first weighted point on the i-th coordinate axis is 1, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined to be 1. Finally, the decoding end determines the context model corresponding to the i-th bit on the i-th coordinate axis based on the first context index and / or the second context index corresponding to the i-th bit on the i-th coordinate axis, and uses the context model to predict and decode the value of the i-th bit of the current point on the i-th coordinate axis.

[0597] In addition to determining the second context index based on the above-mentioned method 1, the decoding end may also determine the second context index through the following method 2.

[0598] Method 2: If K is greater than 1, the first point included in the prediction node in each of the K prediction reference frames is weighted, and the second context index corresponding to the i-th coordinate axis is determined based on the weighted coordinate information. In this case, the above S102-B includes the following steps S102-B21 to S102-B23:

[0599] S102-A1-321-21. For a j-th prediction reference frame among the K prediction reference frames, determine a first weight corresponding to a prediction node in the j-th prediction reference frame;

[0600] S102-A1-321-22. Perform weighted processing on the coordinate information of the first point included in the prediction node in the j-th prediction reference frame based on the first weight to obtain a second weighted point corresponding to the j-th prediction reference frame, where j is a positive integer less than or equal to K.

[0601] S102-A1-321-23. Determine a second context index corresponding to the i-th coordinate axis based on the second weighted points corresponding to the K predicted reference frames.

[0602] In this second approach, when determining the geometric information of the current point, each of the K predicted reference frames is considered separately. Specifically, the coordinate information of the first point in the prediction node of each of the K predicted reference frames is determined, and the second weighted point corresponding to each predicted reference frame is determined. Then, based on the coordinate information of the second weighted point corresponding to each predicted reference frame, the second context index corresponding to the i-th coordinate axis is determined, thereby achieving accurate prediction of the second context index and improving the decoding efficiency of the point cloud.

[0603] In the embodiment of the present application, the specific method in which the decoding end determines the second weighted point corresponding to each of the K prediction reference frames is the same. For ease of description, the j-th prediction reference frame among the K prediction reference frames is used as an example for illustration.

[0604] In an embodiment of the present application, the current node includes at least one prediction node in the j-th prediction reference frame, so that based on the coordinate information of the first point included in the at least one prediction node in the j-th prediction reference frame, the second weighted point corresponding to the j-th prediction reference frame is determined.

[0605] For example, the j-th prediction reference frame includes two prediction nodes of the current node, which are respectively recorded as prediction node 1 and prediction node 2, and then the geometric information of the first point included in the prediction node 1 and the coordinate information of the first point included in the prediction node 2 are weighted to obtain the second weighted point corresponding to the j-th prediction reference frame.

[0606] Before the decoder performs weighted processing on the geometric information of the first point included in the prediction node in the j-th prediction reference frame, it is necessary to first determine the first weight corresponding to each prediction node in the j-th prediction reference frame. The process of determining the first weight can be referred to the description of the above embodiment and is not repeated here.

[0607] Next, the decoding end performs weighted processing on the coordinate information of the first point included in the prediction node in the j-th prediction reference frame based on the first weight to obtain a second weighted point corresponding to the j-th prediction reference frame.

[0608] In one example, based on the first weight, weighted averaging is performed on the coordinate information of the first point included in the prediction node in the j-th prediction reference frame to obtain a second weighted point corresponding to the j-th prediction reference frame.

[0609] The above describes the process of determining the second weighted point corresponding to the j-th prediction reference frame among the K prediction reference frames. The second weighted points corresponding to other prediction reference frames among the K prediction reference frames can be determined in the same manner as that of the j-th prediction reference frame.

[0610] After the decoding end determines the second weighted point corresponding to each of the K predicted reference frames, it executes the above-mentioned step S102-A1-321-23.

[0611] The present application does not impose any limitation on the specific method of determining the second context index corresponding to the i-th coordinate axis based on the second weighted points corresponding to the K predicted reference frames.

[0612] In some embodiments, the decoding end determines an average value of coordinate information of the second weighted points corresponding to the K predicted reference frames on the i-th coordinate axis, and determines a second context index corresponding to the i-th coordinate axis based on the average value.

[0613] In some embodiments, the above S102-A1-321-23 includes the following steps S102-A1-321-231 to S102-A1-321-233:

[0614] S102-A1-321-231, determine second weights corresponding to K prediction reference frames;

[0615] S102-B232, weighting the geometric information of the second weighted points corresponding to the K prediction reference frames based on the second weight to obtain a third weighted point;

[0616] S102-B233. Determine a second context index corresponding to the i-th coordinate axis based on the coordinate information of the third weighted point on the i-th coordinate axis.

[0617] In this embodiment, the decoding end may refer to the method of the above embodiment to determine the second weight corresponding to each of the K prediction reference frames. Then, based on the second weight, the coordinate information of the second weighted points corresponding to the K prediction reference frames is weighted to obtain a third weighted point.

[0618] For example, assuming K=2, for example, the current frame to be decoded includes 2 predicted reference frames, and these 2 predicted reference frames include the forward frame and backward frame of the current frame to be decoded. Assuming that the second weight corresponding to the forward frame is W1 and the second weight corresponding to the backward frame is W2, based on W1 and W2, the geometric information of the second weighted point corresponding to the forward frame and the geometric information of the second weighted point corresponding to the backward frame are weighted to obtain the geometric information of the third weighted point.

[0619] In one example, based on the second weight, weighted averaging is performed on the geometric information of the second weighted points corresponding to the K prediction reference frames to obtain coordinate information of the third weighted point.

[0620] After determining the geometric information of the third weighted point based on the above steps, the decoding end determines the second context index corresponding to the i-th coordinate axis based on the coordinate information of the third weighted point on the i-th coordinate axis.

[0621] As can be seen from the above, the third weighted point is obtained by weighting the first point in the prediction node in each prediction reference frame, wherein the value of each bit of the first point after the prediction node has only two possible results, 0 or 1. Therefore, in some embodiments, the value of each bit of the first weighted point obtained by weighting the first point in the N prediction nodes is also 0 or 1. In this way, when decoding the i-th bit of the current point on the i-th coordinate axis, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined based on the value of the i-th bit of the third weighted point on the i-th coordinate axis. For example, if the value of the i-th bit of the third weighted point on the i-th coordinate axis is 0, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined to be 0. For another example, if the value of the i-th bit of the third weighted point on the i-th coordinate axis is 1, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined to be 1. Finally, the decoding end determines the context model corresponding to the i-th bit on the i-th coordinate axis based on the first context index and / or the second context index corresponding to the i-th bit on the i-th coordinate axis, and uses the context model to predict and decode the value of the i-th bit of the current point on the i-th coordinate axis.

[0622] After the decoding end determines the first context index and / or the second context index based on the above steps, it determines the context model based on the first context index and / or the second context index, and uses the context model to decode the coordinate information of the current point.

[0623] In one example, assuming that the DCM mode information of the prediction node is PredDCMode, the number of points contained in the prediction node is PredNumPoints, and the geometric information of the first point in the prediction node is predPointPos. Assume that the decoder uses the IDCM mode of the prediction node and the geometric information of the points in the prediction node to predict and decode the geometric information of the current point. That is, the geometric decoding information of the prediction node used by the decoder includes the following two types:

[0624] 1) Predict the IDCM mode of the node;

[0625] 2) The geometric information of the midpoint of the predicted node (i.e., the first point), that is, the bit information (0 or 1) corresponding to the accuracy of the midpoint of the predicted node.

[0626] For example, in the GPCC framework, the IDCM mode of the prediction node includes PredDCMode(0,1,2). In the AVS framework, the IDCM of the prediction node

[0627] Mode includes PredDCMode(0,1).

[0628] Assuming that the number of points in the current node is numPoints, and the geometric information of each point is PointPos, and the bit precision depth to be encoded is nodeSizeLog2, the geometric information decoding process of each point in the current node is as follows:

[0629]

[0630]

[0631] Through the above decoding process, the geometric information of each point in the current node can be obtained. Wherein, ctx1 is the first context index and ctx2 is the second context index.

[0632] The point cloud decoding method provided by the embodiment of the present application, when decoding the current node in the current decoding frame, determines N predicted nodes of the current node in the predicted reference frame of the current frame to be decoded, and predictively decodes the coordinate information of the midpoint of the current node based on the geometric decoding information of the midpoints of these N predicted nodes. In other words, the embodiment of the present application optimizes the direct DCM decoding of the node, and by considering the correlation in the time domain between adjacent frames, uses the geometric information of the predicted nodes in the predicted reference frame to predictively decode the geometric information of the midpoint of the IDCM node (i.e., the current node) of the to-be-decoded node, and further improves the efficiency of geometric information decoding of the point cloud by considering the correlation in the time domain between adjacent frames.

[0633] The above takes the decoding end as an example to introduce in detail the point cloud decoding method provided in the embodiment of the present application. The following takes the encoding end as an example to introduce the point cloud encoding method provided in the embodiment of the present application.

[0634] Figure 17 is a schematic diagram of a point cloud encoding method according to an embodiment of the present application. The point cloud encoding method according to the embodiment of the present application can be implemented by the point cloud encoding device shown in Figure 3, Figure 4A, or Figure 8A.

[0635] As shown in FIG17 , the point cloud encoding method of the embodiment of the present application includes:

[0636] S201 : Determine N prediction nodes of a current node in a prediction reference frame of a current frame to be encoded.

[0637] The current node is the node to be encoded in the current frame to be encoded.

[0638] As can be seen from the above, a point cloud includes geometric information and attribute information, and encoding of a point cloud includes geometric encoding and attribute encoding. The embodiments of the present application relate to geometric encoding of a point cloud.

[0639] In some embodiments, the geometric information of the point cloud is also referred to as the position information of the point cloud. Therefore, the geometric encoding of the point cloud is also referred to as the position encoding of the point cloud.

[0640] In the octree-based encoding method, the encoding end constructs the octree structure of the point cloud based on the geometric information of the point cloud. As shown in Figure 11, the point cloud is enclosed by the smallest rectangular block. The bounding box is first divided into octrees to obtain 8 nodes. The occupied nodes among these 8 nodes, that is, the nodes including the points, are further divided into octrees, and so on, until the division is to the voxel level, for example, to a 1X1X1 cube. The point cloud octree structure obtained by such division includes multiple layers of nodes, for example, N layers. During encoding, the occupancy information of each layer is encoded layer by layer until the voxel-level leaf nodes of the last layer are encoded. That is to say, in octree encoding, the point cloud is divided into octrees, and finally the points in the point cloud are divided into the voxel-level leaf nodes of the octree. The encoding of the point cloud is achieved by encoding the entire octree.

[0641] However, the octree-based geometric information coding mode has an efficient compression rate for points with correlation in space, and for points in isolated positions in the geometric space, the use of direct coding can greatly reduce the complexity and improve the coding efficiency.

[0642] Since direct encoding directly encodes the geometric information of the points included in a node, if the node contains a large number of points, the compression effect of direct encoding is poor. Therefore, before performing direct encoding on a node in the octree, it is first determined whether the node can be encoded using direct encoding. If it is determined that the node can be encoded using direct encoding, the geometric information of the points included in the node is directly encoded using direct encoding. If it is determined that the node cannot be encoded using direct encoding, the node is further divided using the octree method.

[0643] Specifically, the encoder first determines whether the node is eligible for direct encoding. If so, it then determines whether the node's point count is less than or equal to a preset threshold. If so, the node is considered eligible for direct encoding. The number of points in the node and the geometric information of each point are then encoded into the bitstream.

[0644] Currently, when predicting the position information of the midpoint of the current node, inter-frame information is not considered, resulting in low coding performance of the point cloud.

[0645] In order to solve the above problems, in an embodiment of the present application, the encoding end predictively encodes the position information of the midpoint of the current node based on the inter-frame information corresponding to the current node, thereby improving the encoding efficiency and encoding performance of the point cloud.

[0646] Specifically, the encoder first determines N prediction nodes of the current node in the prediction reference frame of the current frame to be encoded.

[0647] It should be noted that the current frame to be encoded is a point cloud frame. In some embodiments, the current frame to be encoded is also referred to as the current frame, the current point cloud frame, or the point cloud frame to be encoded. The current node can be understood as any non-leaf node in the current frame to be encoded that is not a non-empty node. In other words, the current node is not a leaf node in the octree corresponding to the current frame to be encoded, that is, the current node is any middle node in the octree, and the current node is not a non-empty node, that is, it includes at least one point.

[0648] In an embodiment of the present application, when encoding a current node in a current frame to be encoded, the encoder first determines a prediction reference frame for the current frame to be encoded, and then determines N prediction nodes for the current node in the prediction reference frame. For example, FIG12 shows a prediction node for the current node in the prediction reference frame.

[0649] It should be noted that the embodiment of the present application does not limit the number of prediction reference frames of the current frame to be encoded. For example, the current frame to be encoded may have one prediction reference frame, or the current frame to be encoded may have multiple prediction reference frames. Furthermore, the embodiment of the present application does not limit the number N of prediction nodes of the current node, and this number is determined based on actual needs.

[0650] The embodiment of the present application does not limit the specific method of determining the prediction reference frame of the current frame to be encoded.

[0651] In some embodiments, one or several encoded frames before the current frame to be encoded are determined as prediction reference frames for the current frame to be encoded.

[0652] For example, if the current frame to be encoded is a P frame, the inter-frame reference frame of the P frame includes the previous frame of the P frame (i.e., the forward frame). Therefore, the previous frame of the current frame to be encoded (i.e., the forward frame) can be determined as the predicted reference frame of the current frame to be encoded.

[0653] For another example, if the current frame to be encoded is a B frame, the inter-frame reference frames of the B frame include the previous frame of the P frame (i.e., the forward frame) and the next frame of the P frame (i.e., the backward frame). Therefore, the previous frame of the current frame to be encoded (i.e., the forward frame) can be determined as the predicted reference frame of the current frame to be encoded.

[0654] In some embodiments, one or several encoded frames following the current frame to be encoded are determined as prediction reference frames for the current frame to be encoded.

[0655] For example, if the current frame to be encoded is a B frame, the next frame after the current frame to be encoded may be determined as a prediction reference frame for the current frame to be encoded.

[0656] In some embodiments, one or several encoded frames before the current frame to be encoded, and one or several encoded frames after the current frame to be encoded, are determined as prediction reference frames for the current frame to be encoded.

[0657] For example, if the current frame to be encoded is a B frame, the previous frame and the next frame of the current frame to be encoded may be determined as prediction reference frames of the current frame to be encoded. In this case, the current frame to be encoded has two prediction reference frames.

[0658] The following takes the current frame to be encoded including K prediction reference frames as an example to introduce the specific process of determining N prediction nodes of the current node in the prediction reference frames of the current frame to be encoded in S201-A.

[0659] In some embodiments, the encoder selects at least one prediction reference frame from the K prediction reference frames based on the placeholder information of the node in the current frame to be encoded and the placeholder information of the node in each of the K prediction reference frames, and then searches for a predicted node for the current node in the at least one prediction reference frame. For example, at least one prediction reference frame whose placeholder information of the node is closest to the placeholder information of the node in the current frame to be encoded is selected from the K prediction reference frames, and then searches for a predicted node for the current node in the at least one prediction reference frame.

[0660] In some embodiments, the encoder may determine N prediction nodes of the current node through the following steps S201-A1 and S201-A2:

[0661] S201-A1. For a k-th prediction reference frame among K prediction reference frames, determine at least one prediction node of a current node in the k-th prediction reference frame, where k is a positive integer less than or equal to K, and K is a positive integer;

[0662] S201-A2: Determine N prediction nodes of the current node based on at least one prediction node of the current node in K prediction reference frames.

[0663] In this embodiment, the encoding end determines at least one prediction node of the current node from each of the K prediction reference frames, and finally aggregates at least one prediction node in each of the K prediction reference frames to obtain N prediction nodes of the current node.

[0664] Among them, the process of the encoding end determining at least one prediction point of the current node in each of the K prediction reference frames is the same. For the sake of convenience of description, the kth prediction reference frame among the K prediction reference frames is used as an example for explanation.

[0665] The specific process of determining at least one prediction node of the current node in the kth prediction reference frame in the above S201-A1 is introduced below.

[0666] The embodiment of the present application does not limit the specific manner in which the encoder determines at least one prediction node of the current node in the kth prediction reference frame.

[0667] Method 1: In the kth prediction reference frame, a prediction node of the current node is determined. For example, a node in the kth prediction reference frame that has the same partition depth as the current node is determined as the prediction node of the current node.

[0668] In one example, if the number of prediction nodes of the current node in the kth prediction reference frame is 1, then among the points at which the kth prediction reference frame and the current node are at the same division depth, a node whose occupancy information is the smallest different from that of the current node can be selected, recorded as node 1, and node 1 is determined as a prediction node of the current node in the kth prediction reference frame.

[0669] In another example, if the number of prediction nodes of the current node in the kth prediction reference frame is greater than 1, the node 1 determined above and at least one domain node of node 1 in the kth prediction reference frame, such as at least one domain node that is coplanar, colinear, or co-point with node 1, are determined as the prediction nodes of the current node in the kth prediction reference frame.

[0670] Method 2, in the above S201-A1, determining at least one prediction node of the current node in the k-th prediction reference frame includes the following steps S201-A11 to S201-A13:

[0671] S201-A11. In a current frame to be encoded, determine M domain nodes of a current node, where the M domain nodes include the current node, and M is a positive integer.

[0672] S201-A12: for the i-th domain node among the M domain nodes, determine the corresponding node of the i-th domain node in the k-th prediction reference frame, where i is a positive integer less than or equal to M;

[0673] S201-A13. Determine at least one prediction node of the current node in the kth prediction reference frame based on the corresponding nodes of the M domain nodes in the kth prediction reference frame.

[0674] In this implementation, before determining at least one prediction node of the current node in the kth prediction reference frame, the encoder first determines M domain nodes of the current node in the current frame to be encoded, and the M domain nodes include the current node itself.

[0675] It should be noted that in the embodiment of the present application, there is no restriction on the specific method of determining the M domain nodes of the current node.

[0676] In one example, the M domain nodes of the current node include at least one domain node among the domain nodes that are coplanar, colinear, and co-point with the current node in the current frame to be encoded. As shown in Figure 13, the current node includes 6 coplanar nodes, 12 colinear nodes, and 8 co-point nodes.

[0677] In another example, the M domain nodes of the current node may include other nodes within the reference neighborhood in addition to at least one domain node in the current frame to be encoded that is coplanar, colinear, and co-point with the current node. This embodiment of the present application does not impose any restrictions on this.

[0678] Based on the above steps, the encoding end determines the M domain nodes of the current node in the current frame to be encoded, determines the corresponding node of each of the M domain nodes in the k-th prediction reference frame, and then determines at least one prediction node of the current node in the k-th prediction reference frame based on the corresponding nodes of the M domain nodes in the k-th prediction reference frame.

[0679] The embodiment of the present application does not limit the specific implementation method of S201-A13.

[0680] In one possible implementation, at least one corresponding node is selected from the corresponding nodes of the M domain nodes in the k-th prediction reference frame as the at least one prediction node of the current node in the k-th prediction reference frame. For example, at least one corresponding node whose placeholder information has the smallest difference between the placeholder information of the M domain nodes in the k-th prediction reference frame and the placeholder information of the current node is selected from the corresponding nodes of the M domain nodes in the k-th prediction reference frame as the at least one prediction node of the current node in the k-th prediction reference frame. The method for determining the difference between the placeholder information of the corresponding node and the placeholder information of the current node can refer to the above-mentioned process for determining the difference in placeholder information, for example, performing an XOR operation on the placeholder information of the corresponding node and the placeholder information of the current node, and using the XOR operation result as the difference between the placeholder information of the corresponding node and the placeholder information of the current node.

[0681] In another possible implementation, the encoder determines the corresponding nodes of the M domain nodes in the kth prediction reference frame as at least one prediction node for the current node in the kth prediction reference frame. For example, each of the M domain nodes has a corresponding node in the kth prediction reference frame, resulting in M ​​corresponding nodes. These M corresponding nodes are determined as the prediction nodes for the current node in the kth prediction reference frame, for a total of M prediction nodes.

[0682] The above describes the process of determining at least one prediction node for the current node in the kth prediction reference frame. Thus, the encoder can use the same method as above to determine at least one prediction node for the current node in each of the K prediction reference frames.

[0683] After the encoder determines at least one prediction node of the current node in each of the K prediction reference frames, it performs the above step S201-B, that is, determines N prediction nodes of the current node based on at least one prediction node of the current node in the K prediction reference frames.

[0684] In one example, at least one prediction node of the current node in K prediction reference frames is determined as N prediction nodes of the current node.

[0685] For example, K=2, that is, the K prediction reference frames include the first prediction reference frame and the second prediction reference frame. Assume that the current node has 2 prediction nodes in the first prediction reference frame and 3 prediction nodes in the second prediction reference frame. In this way, it can be determined that the current node has 5 prediction nodes, and N=5.

[0686] In another example, N prediction nodes of the current node are screened out from at least one prediction node of the current node in K prediction reference frames.

[0687] Continuing with the above example, assume K = 2, meaning the K prediction reference frames include the first prediction reference frame and the second prediction reference frame. Assume the current node has two prediction nodes in the first prediction reference frame and three prediction nodes in the second prediction reference frame. From these five prediction nodes, select three prediction nodes as the final prediction nodes for the current node. For example, from these five prediction nodes, select the three prediction nodes whose placeholder information differs minimally from the placeholder information of the current node and determine them as the final prediction nodes for the current node.

[0688] In the second method, after the encoding end determines the M domain nodes of the current node in the current frame to be encoded, it determines the corresponding node of each of the M domain nodes in the kth prediction reference frame, and then determines at least one prediction point of the current node in the kth prediction reference frame based on the corresponding node of each of the M domain nodes.

[0689] Mode 3, in the above S201-A1, determining at least one prediction node of the current node in the k-th prediction reference frame includes the following steps S201-B11 to S201-B13:

[0690] S201-B11, determining the corresponding node of the current node in the kth prediction reference frame;

[0691] S201-B12, determining at least one domain node of the corresponding node;

[0692] S201-B13. Determine at least one domain node as at least one prediction node of the current node in the k-th prediction reference frame.

[0693] In this method 3, for each of the K predicted reference frames, the encoding end first determines the corresponding node of the current node in each predicted reference frame. For example, the corresponding node 1 of the current node in the predicted reference frame 1 is determined, and the corresponding node 2 of the current node in the predicted reference frame 2 is determined. Next, the encoding end determines at least one domain node of each corresponding node. For example, at least one domain node of the corresponding node 1 is determined in the predicted reference frame 1, and at least one domain node of the corresponding node 2 is determined in the predicted reference frame 2. In this way, at least one domain node of the corresponding node 1 in the predicted reference frame 1 can be determined as at least one predicted node of the current node in the predicted reference frame 1, and at least one domain node of the corresponding node 2 in the predicted reference frame 2 can be determined as at least one predicted node of the current node in the predicted reference frame 2.

[0694] Determining the corresponding node of the i-th domain node in the k-th prediction reference frame in S201-A12 of the second method is essentially the same as determining the corresponding node of the current node in the k-th prediction reference frame in S201-B11 of the third method described above. For ease of description, the i-th domain node and the current node are referred to as the i-th node. The specific process of determining the corresponding node of the i-th node in the k-th prediction reference frame is described below.

[0695] The encoder determines the corresponding node of the i-th node in the k-th prediction reference frame in at least the following ways:

[0696] In method 1, a node in the k-th prediction reference frame that has the same division depth as the i-th node is determined as the corresponding node of the i-th node.

[0697] Mode 2: The above-mentioned S201-A12 and S201-B11 include the following steps:

[0698] S201-A121, in the current frame to be encoded, determine the parent node of the i-th node as the i-th parent node;

[0699] S201-A122, determine the matching node of the i-th parent node in the k-th prediction reference frame as the i-th matching node;

[0700] S201-A123: Determine one of the child nodes of the i matching nodes as the corresponding node of the i-th node in the k-th prediction reference frame.

[0701] In this method 2, for the i-th node, the encoding end determines the parent node of the i-th node in the current frame to be encoded, and then determines the matching node of the parent node of the i-th prediction domain node in the k-th prediction reference frame. For ease of description, the parent node of the i-th node is recorded as the i-th parent node, and the matching node of the parent node of the i-th node in the k-th prediction reference frame is determined as the i-th matching node. Then, a child node of the child node of the i-th matching node is determined as the corresponding node of the i-th node in the k-th prediction reference frame, thereby accurately determining the corresponding node of the i-th node in the k-th prediction reference frame.

[0702] The specific process of determining the matching node of the i-th parent node in the k-th prediction reference frame in the above S201-A122 is introduced below.

[0703] The embodiment of the present application does not limit the specific method by which the encoder determines the matching node of the i-th parent node in the k-th prediction reference frame.

[0704] In some embodiments, the partition depth of the i-th parent node in the current frame to be encoded is determined, for example, the i-th parent node is at the second level of the octree of the current frame to be encoded. In this way, the encoder can determine one of the nodes in the k-th prediction reference frame that have the same partition depth as the i-th parent node as the matching node of the i-th parent node in the k-th prediction reference frame. For example, one of the nodes in the second level of the k-th prediction reference frame can be determined as the matching node of the i-th parent node in the k-th prediction reference frame.

[0705] In some embodiments, the encoder determines a matching node for the i-th parent node in the k-th predicted reference frame based on the placeholder information of the i-th parent node. Specifically, since the placeholder information for the i-th parent node in the current frame to be encoded has been encoded, and the placeholder information for each node in the k-th predicted reference frame has also been encoded, the encoder can search for a matching node for the i-th parent node in the k-th predicted reference frame based on the placeholder information of the i-th parent node.

[0706] For example, the node with the smallest difference between the placeholder information of the k-th prediction reference frame and the placeholder information of the i-th parent node is determined as the matching node of the i-th parent node in the k-th prediction reference frame.

[0707] For example, assuming the placeholder information of the i-th parent node is 11001101, the k-th predicted reference frame is searched for the node whose placeholder information has the smallest difference from the placeholder information 11001101. Specifically, the encoder performs an XOR operation on the placeholder information of the i-th parent node and the placeholder information of each node in the k-th predicted reference frame. The node with the smallest XOR result in the k-th predicted reference frame is determined as the matching node of the i-th parent node in the k-th predicted reference frame.

[0708] Based on the above steps, the encoder can determine the matching node of the i-th parent node in the k-th prediction reference frame. For ease of description, this matching node is recorded as the i-th matching node.

[0709] Next, the encoder determines one of the child nodes of the i-th matching node as the corresponding node of the i-th domain node in the k-th prediction reference frame.

[0710] For example, the encoder determines a default child node among the child nodes included in the i-th matching node as the corresponding node of the i-th node in the k-th prediction reference frame. Assume that the first child node of the i-th matching node is determined as the corresponding node of the i-th node in the k-th prediction reference frame.

[0711] For another example, the encoder determines the first sequence number of the i-th node among the child nodes included in the parent node; and determines the child node with the first sequence number among the child nodes of the i-th matching node as the corresponding node of the i-th node in the k-th prediction reference frame. For example, as shown in FIG14 , the i-th node is the second child node of the i-th parent node, and in this case, the first sequence number is 2. In this way, the second child node of the i-th matching node can be determined as the corresponding node of the i-th node.

[0712] The above describes the process of determining the corresponding node of the i-th domain node among M domain nodes in the k-th prediction reference frame, and the corresponding node of the current node in the k-th prediction reference frame. Thus, the encoder can use Method 2 or Method 3 to determine the N prediction nodes for the current node in the prediction reference frame.

[0713] Based on the above steps, the encoder determines N prediction nodes of the current node in the prediction reference frame of the current frame to be encoded, and then performs the following step S202.

[0714] S202 . Based on the geometric coding information of the midpoints of the N predicted nodes, predictively encode the position information of the midpoint of the current node.

[0715] Due to the correlation between adjacent frames of the point cloud, the present embodiment references the relevant information between frames when predicting the position information of the current node midpoint based on the correlation between adjacent frames of the point cloud. Specifically, the WeChat information of the current node midpoint is predictively encoded based on the geometric encoding information of the N predicted node midpoints of the current node, thereby improving the encoding efficiency and encoding performance of the point cloud.

[0716] It should be noted that, in the embodiment of the present application, predictive encoding of the coordinate information of the midpoint of the current node based on the geometric coding information of N predicted nodes can be understood as using the geometric coding information of the N predicted nodes as context to predictively encode the coordinate information of the midpoint of the current node. For example, the encoding end determines the index of the context model based on the geometric coding information of the N predicted nodes, and then, based on the index of the context model, determines the target context model from multiple preset context models, and uses the context model to predictively encode the coordinate information of the midpoint of the current node.

[0717] In an embodiment of the present application, the process of predictive encoding the coordinate information of each point in the current node based on the geometric coding information of N prediction nodes is basically the same. For the sake of convenience of description, the predictive encoding of the coordinate information of the current point in the current node is used as an example to illustrate.

[0718] In some embodiments, the above S202 includes the following steps:

[0719] S202-A, determining an index of a context model based on geometric coding information of N prediction nodes;

[0720] S202-B, determining a context model based on an index of the context model;

[0721] S202-C. Use the context model to perform predictive coding on the coordinate information of the current point in the current node.

[0722] In an embodiment of the present application, multiple context models, for example, Q context models, are set for the encoding process of the coordinate information. The embodiment of the present application does not limit the specific number of context models corresponding to the coordinate information, as long as Q is greater than 1. That is, in an embodiment of the present application, an optimal context model is selected from at least two context models to perform predictive encoding on the coordinate information of the current point in the current node, so as to improve the encoding efficiency of the coordinate information of the current point.

[0723] In the embodiments of the present application, the geometric coding information of the prediction node can be understood as any information involved in the geometric coding process of the prediction node, including, for example, the number of points included in the prediction node, the placeholder information of the prediction node, the coding method of the prediction node, the geometric information of the midpoint of the prediction node, etc.

[0724] In some embodiments, the geometric coding information of the prediction node includes direct coding information of the prediction node and / or coordinate information of a midpoint of the prediction node, wherein the direct coding information of the prediction node is used to indicate whether the prediction node meets the conditions for encoding in a direct coding manner.

[0725] Based on this, the above S202-A includes the following step S202-A1:

[0726] S202-A1. Determine a first context index based on direct encoding information of the N prediction nodes, and / or determine a second context index based on coordinate information of midpoints of the N prediction nodes.

[0727] Correspondingly, the above-mentioned S202-B includes the following steps S202-B:

[0728] S202-B1. Select a context model from a plurality of preset context models based on the first context index and / or the second context index.

[0729] In this embodiment, if the geometric coding information of the prediction node includes the direct coding information of the prediction node and / or the coordinate information of the midpoint of the prediction node, the encoding end can determine the first context index based on the direct coding information of the N prediction nodes, and / or determine the second context index based on the position information of the midpoint of the N prediction nodes, and then select the final context model from the preset multiple context models based on the first context index and / or the second context index.

[0730] It can be seen that in this embodiment, the encoding end determines the context model in the following ways, but is not limited to:

[0731] In one possible implementation, if the geometric coding information of the prediction node includes the direct coding information of the prediction node, the process of determining the context model may be to determine the first context index based on the direct coding information of N prediction nodes, and then based on the first context index, select the final context model from the preset multiple context models to encode the coordinate information of the current point.

[0732] For example, the encoder selects the final context model from the context models shown in Table 4 based on the first context index.

[0733] In another possible implementation, if the geometric coding information of the prediction node includes the coordinate information of the midpoint of the prediction node, the process of determining the context model may be to determine the second context index based on the coordinate information of the midpoints of N prediction nodes, and then based on the second context index, select the final context model from the preset multiple context models to encode the coordinate information of the current point.

[0734] For example, the encoder selects the final context model from the context models shown in Table 4 based on the second context index.

[0735] In another possible implementation, if the geometric coding information of the prediction node includes the direct coding information of the prediction node and the coordinate information of the midpoint of the prediction node, the process of determining the context model may be to determine the first context index based on the direct coding information of N prediction nodes, and then determine the second context index based on the first context index and the coordinate information of the midpoint of the N prediction nodes, and then select the final context model from the preset multiple context models based on the second context index, the first context index and the second context index to encode the coordinate information of the current point.

[0736] Exemplarily, the correspondence between the first context index, the second context index and the context model is shown in Table 5.

[0737] In this approach 3, the encoder determines the first context index based on the directly coded information of the N prediction nodes, and the second context index based on the coordinate information of the midpoint of the N prediction nodes. Then, the encoder consults Table 5 to obtain the final context model. For example, the encoder determines the first context index to be first context index 2 based on the directly coded information of the N prediction nodes, and determines the second context index to be second context index 3 based on the coordinate information of the midpoint of the N prediction nodes. By consulting Table 5, the final context model obtained is context model 23. The encoder then uses context model 23 to encode the coordinate information of the current point.

[0738] The specific process of determining the first context index based on the direct encoding information of the N prediction nodes in S202-A1 is introduced below.

[0739] In the embodiment of the present application, the encoding end determines the first context index in the following ways but is not limited to:

[0740] Method 1: the above S202-A1 includes the following steps S202-A1-11 and S202-A1-12:

[0741] S202-A1-11. For any prediction node among the N prediction nodes, determine a first numerical value corresponding to the prediction node based on direct encoding information of the prediction node.

[0742] In this manner, for each of the N prediction nodes, a first value corresponding to the prediction node is determined based on direct coding information of the prediction node, and finally a first context index is determined based on the first values ​​corresponding to the N prediction nodes.

[0743] The following describes the process of determining the first value corresponding to the prediction node.

[0744] As can be seen from the above, the direct coding information of the prediction node is used to indicate whether the prediction node meets the conditions for encoding in the direct coding mode. The embodiment of the present application does not limit the specific content of the direct coding information.

[0745] In some embodiments, the direct encoding information includes the number of points included in the prediction node. In this way, the first numerical value corresponding to the prediction node can be determined based on the number of points included in the prediction node.

[0746] In one example, under the GPCC framework, when the number of points included in the prediction node is greater than or equal to 2, the first value corresponding to the prediction node is determined to be 1; if the number of points included in the prediction node is less than 2, the first value corresponding to the prediction node is determined to be 0. Under the AVS framework, when the number of points included in the prediction node is greater than or equal to 1, the first value corresponding to the prediction node is determined to be 1; if the number of points included in the prediction node is less than 1, the first value corresponding to the prediction node is determined to be 0.

[0747] In another example, the number of points included in the prediction node is determined as the first value corresponding to the prediction node. For example, when the prediction node includes 2 points, the first value corresponding to the prediction node is determined to be 2.

[0748] In some embodiments, the direct encoding information of the prediction node includes the direct encoding mode of the prediction node. In this case, the above S202-A1-11 includes: numbering the direct encoding mode of the prediction node and determining a first value corresponding to the prediction node.

[0749] For example, in the GPCC framework, if the direct encoding mode of the prediction node is mode 0, the first value corresponding to the prediction node is determined to be 0. If the direct encoding mode of the prediction node is mode 1, the first value corresponding to the prediction node is determined to be 1. If the direct encoding mode of the prediction node is mode 2, the first value corresponding to the prediction node is determined to be 2.

[0750] For another example, in the AVS framework, if the direct decoding mode of the prediction node is mode 0, the first value corresponding to the prediction node is determined to be 0. If the direct decoding mode of the prediction node is mode 1, the first value corresponding to the prediction node is determined to be 1.

[0751] Based on the above steps, after the encoder determines the first value corresponding to each of the N prediction nodes, it executes the following step S202-A1-12.

[0752] S202-A1-12. Determine a first context index based on first values ​​corresponding to the N prediction nodes.

[0753] After determining the first values ​​corresponding to the N prediction nodes based on the above steps, the encoder determines a first context index based on the first values ​​corresponding to the N prediction nodes.

[0754] Determining the first context index based on the first values ​​corresponding to the N prediction nodes includes at least the following implementations:

[0755] In method 1, an average value of the sum of the first numerical values ​​corresponding to the N prediction nodes is determined as the first context index.

[0756] Mode 2, S202-A1-12 includes the following steps S202-A1-121 to S202-A1-123:

[0757] S202-A1-121, determine a first weight corresponding to the prediction node;

[0758] S202-A1-122, based on the first weight, weighting the first values ​​corresponding to the N prediction nodes to obtain a first weighted prediction value;

[0759] S202-A1-123. Determine a first context index based on the first weighted prediction value.

[0760] In method 2, if the current node includes multiple prediction nodes, i.e., N prediction nodes, when determining the first context index based on the first numerical values ​​corresponding to the N prediction nodes, a weight, i.e., the first weight, can be determined for each of the N prediction nodes. In this way, the first numerical values ​​corresponding to each prediction node can be weighted based on the first weight of each prediction node, and then the first context index can be determined based on the final weighted result, thereby improving the accuracy of determining the first context index based on the geometric coding information of the N prediction nodes.

[0761] The embodiment of the present application does not limit the determination of the first weights corresponding to the N prediction nodes.

[0762] In some embodiments, the first weight corresponding to each of the N prediction nodes is a preset value. As can be seen from the above, the N prediction nodes are determined based on the M domain nodes of the current node. Assuming that prediction node 1 is the prediction node corresponding to domain node 1, if domain node 1 is a coplanar node with the current node, then the first weight of prediction node 1 is the preset weight 1. If domain node 1 is a colinear node with the current node, then the first weight of prediction node 1 is the preset weight 2. If domain node 1 is a co-point node with the current node, then the first weight of prediction node 1 is the preset weight 3.

[0763] In some embodiments, for each of the N prediction nodes, a first weight corresponding to the prediction node is determined based on the distance between the domain node corresponding to the prediction node and the current node. For example, the smaller the distance between the domain node and the current node, the stronger the inter-frame correlation between the prediction node corresponding to the domain node and the current node, and thus the greater the first weight of the prediction node.

[0764] In some embodiments, based on the above steps, after determining the weight corresponding to each prediction node in the N prediction nodes, the weight is normalized, and the normalized weight is used as the final first weight of the prediction node.

[0765] The embodiment of the present application does not limit the specific method of obtaining the first weighted prediction value by weighting the first numerical values ​​corresponding to N prediction nodes based on the first weight.

[0766] In one example, based on the first weight, a weighted average is performed on the first values ​​corresponding to the N prediction nodes to obtain a first weighted prediction value.

[0767] In another example, based on the first weight, a weighted sum is performed on the first numerical values ​​corresponding to the N prediction nodes to obtain a first weighted prediction value.

[0768] After determining the first weighted prediction value based on the method steps, the first context index is determined based on the first weighted prediction value. That is, the above S202-A1-123 includes at least the following examples:

[0769] Example 1: Determine the first weighted prediction value as the first context index.

[0770] Example 2: Determine the weighted prediction value range in which the first weighted prediction value is located, and determine the index corresponding to the range as the first context index.

[0771] The above describes the process of obtaining the first context index by performing weighted processing on N prediction nodes by the encoder.

[0772] In some embodiments, the encoding end may also adopt the following second method to determine the first context index.

[0773] Method 2: If K is greater than 1, determine the second weighted prediction value corresponding to each of the K prediction reference frames, and then determine the first context index based on the second weighted prediction values ​​corresponding to the K prediction reference frames. In this case, the above S202-A1 includes the following steps S202-A1-21 to S202-A1-23:

[0774] S202-A1-21. For a j-th prediction reference frame among the K prediction reference frames, determine a first value corresponding to the prediction node in the j-th prediction reference frame based on direct encoding information of the prediction node of the current node in the j-th prediction reference frame, where j is a positive integer less than or equal to K.

[0775] S202-A1-22, determining a first weight corresponding to the prediction node, and performing weighted processing on the first value corresponding to the prediction node in the j-th prediction reference frame based on the first weight to obtain a second weighted prediction value corresponding to the j-th prediction reference frame;

[0776] S202-A1-23. Determine a first context index based on the second weighted prediction values ​​corresponding to the K predicted reference frames.

[0777] In this second approach, when determining the first context index, each of the K predicted reference frames is considered separately as separate context information. Specifically, direct coding information of the prediction nodes included in each of the K predicted reference frames is determined, and a second weighted prediction value corresponding to each predicted reference frame is determined. Furthermore, based on the second weighted prediction value corresponding to each predicted reference frame, the first context index is determined, thereby achieving accurate selection of the first context index and improving the coding efficiency of the point cloud.

[0778] In the embodiment of the present application, the specific method in which the encoding end determines the second weighted prediction value corresponding to each of the K prediction reference frames is the same. For the sake of convenience of description, the j-th prediction reference frame among the K prediction reference frames is used as an example for illustration.

[0779] In an embodiment of the present application, the current node includes at least one prediction node in the j-th prediction reference frame, so that the first value of the at least one prediction node is determined based on the direct encoding information of the at least one prediction node in the j-th prediction reference frame.

[0780] For example, the j-th prediction reference frame includes two prediction nodes of the current node, which are respectively recorded as prediction node 1 and prediction node 2. Then, based on the direct coding information of prediction node 1, the first value of prediction node 1 is determined, and based on the direct coding information of prediction node 2, the first value of prediction node 2 is determined. The process of determining the first value corresponding to the prediction node based on the direct coding information of the prediction node can refer to the description of the above embodiment. Exemplarily, the first value corresponding to the prediction node is determined based on the direct coding mode of the prediction node. For example, the number of the direct coding mode of the prediction node (0, 1, or 2) is determined as the first value corresponding to the prediction node.

[0781] After the encoding end determines the first value of at least one prediction node included in the j-th prediction reference frame, it determines the first weight corresponding to the at least one prediction node, and performs weighted processing on the first value corresponding to the at least one prediction node based on the first weight to obtain a second weighted prediction value corresponding to the j-th prediction reference frame.

[0782] In one example, based on the first weight, a weighted average is performed on the first values ​​corresponding to the prediction nodes in the j-th prediction reference frame to obtain a second weighted prediction value corresponding to the j-th prediction reference frame.

[0783] In another example, based on the first weight, a weighted sum is performed on the first values ​​corresponding to the prediction nodes in the j-th prediction reference frame to obtain a second weighted prediction value corresponding to the j-th prediction reference frame.

[0784] The process of determining the first weight may refer to the description of the above embodiment and will not be repeated here.

[0785] The above introduces the process of determining the second weighted prediction value corresponding to the j-th prediction reference frame among the K prediction reference frames. The second weighted prediction values ​​corresponding to other prediction reference frames among the K prediction reference frames are determined in accordance with the method corresponding to the j-th prediction reference frame.

[0786] After the encoding end determines the second weighted prediction value corresponding to each of the K prediction reference frames, it executes the above step S202-A1-23.

[0787] The present application does not limit the specific method of determining the first context index based on the second weighted prediction value corresponding to K prediction reference frames.

[0788] In some embodiments, the encoding end determines an average value of the second weighted prediction values ​​corresponding to the K prediction reference frames as the first context index.

[0789] In some embodiments, the encoding end determines second weights corresponding to the K predicted reference frames, and performs weighted processing on second weighted prediction values ​​corresponding to the K predicted reference frames based on the second weights to obtain a first context index.

[0790] In this embodiment, the encoder first determines the second weight corresponding to each of the K prediction reference frames. This embodiment of the application does not limit the determination of the second weight corresponding to each of the K prediction reference frames.

[0791] In some embodiments, the second weight corresponding to each of the K predicted reference frames is a preset value. As can be seen from the above, the K predicted reference frames are forward frames and / or backward frames of the current frame to be encoded. Assuming that predicted reference frame 1 is the forward frame of the current frame to be encoded, the second weight corresponding to predicted reference frame 1 is the preset weight 1. If predicted reference frame 1 is the backward frame of the current frame to be encoded, the second weight corresponding to predicted reference frame 1 is the preset weight 2.

[0792] In some embodiments, the second weight corresponding to the predicted reference frame is determined based on the time difference between the predicted reference frame and the current frame to be encoded. In an embodiment of the present application, each point cloud includes time information, and the time information can be the time when the point cloud acquisition device acquires the point cloud of the frame. Based on this, if the time difference between the predicted reference frame and the current frame to be encoded is smaller, the inter-frame correlation between the predicted reference frame and the current frame to be encoded is stronger, and thus the second weight corresponding to the predicted reference frame is larger. For example, the inverse of the time difference between the predicted reference frame and the current frame to be encoded can be determined as the second weight corresponding to the predicted reference frame.

[0793] After determining the second weight corresponding to each of the K prediction reference frames, weighted processing is performed on the second weighted prediction values ​​corresponding to the K prediction reference frames based on the second weight to obtain a first context index.

[0794] For example, assuming K=2, for example, the current frame to be encoded includes 2 prediction reference frames, and these 2 prediction reference frames include the forward frame and backward frame of the current frame to be encoded. Assuming that the second weight corresponding to the forward frame is W1 and the second weight corresponding to the backward frame is W2, based on W1 and W2, the second weighted prediction value corresponding to the forward frame and the second weighted prediction value corresponding to the backward frame are weighted to obtain the first context index.

[0795] In one example, based on the second weight, weighted averaging is performed on the second weighted prediction values ​​corresponding to the K prediction reference frames to obtain the first context index.

[0796] In another example, based on the second weight, the second weighted prediction values ​​corresponding to the K prediction reference frames are weightedly summed to obtain the first context index.

[0797] The above describes the process of determining the first context index at the encoder.

[0798] The following describes the process of determining the second context index at the encoding end.

[0799] From the above process of determining the prediction node, it can be seen that each of the N prediction nodes includes one point or multiple points. If each of the N prediction nodes includes one point, the one point included in each prediction node is used to determine the second context index.

[0800] In some embodiments, if the predicted node includes multiple points, one point is selected from the multiple points to determine the second context index. In this case, the above S202-A1 includes the following steps S202-A1-31 and S202-A1-32:

[0801] S202-A1-31. For any prediction node among the N prediction nodes, select the first point corresponding to the current point of the current node from the points included in the prediction node;

[0802] S202-A1-32. Determine a second context index based on the coordinate information of the first point included in the N prediction nodes.

[0803] For example, assume that N prediction nodes include prediction node 1 and prediction node 2, where prediction node 1 includes points 1 and 2, and prediction node 2 includes points 3, 4, and 5. Then, a point is selected from points 1 and 2 included in prediction node 1 as the first point, and a point is selected from points 3, 4, and 5 included in prediction node 2 as the first point. In this way, the geometric information of the current point can be determined based on the geometric information of the first point in prediction node 1 and the first point in prediction node 2.

[0804] The embodiment of the present application does not limit the specific method of selecting the first point corresponding to the current point of the current node from the points included in the prediction node.

[0805] In one possible implementation, the points in the predicted node that are in the same order as the current point are determined as the first point corresponding to the current point. For example, assuming the current point is the second point in the current node, point 2 in predicted node 1 can be determined as the first point corresponding to the current point, and point 4 in predicted node 2 can be determined as the first point corresponding to the current point. For another example, if the predicted node includes only one point, the point included in the predicted node is determined as the first point corresponding to the current point.

[0806] In one possible implementation, if the encoder selects the first point corresponding to the current point from the points included in the prediction node based on the rate-distortion cost (or approximate cost), the encoder writes the identification information of the first point in the prediction node into the bitstream. In this way, the encoder obtains the first point in the prediction node by encoding the bitstream.

[0807] For each of the N prediction nodes, the encoder determines the first point corresponding to the current point in each prediction node based on the above method, and then performs the above step S202-A1-32.

[0808] In the embodiment of the present application, the encoding end encodes the coordinate information of the current point on different coordinate axes separately. Based on, the above S202-A1-32 includes the following step S202-A1-321:

[0809] S202-A1-321. Determine a second context index corresponding to the i-th coordinate axis based on coordinate information of the first point included in the N prediction nodes on the i-th coordinate axis.

[0810] The above-mentioned i-th coordinate axis can be the X-axis, Y-axis or Z-axis, and this embodiment of the application does not limit this.

[0811] In an embodiment of the present application, when the encoding end encodes the coordinate information of the current point on the i-th coordinate axis, the second context index corresponding to the i-th coordinate axis is determined based on the coordinate information of the first point on the i-th coordinate axis included in the N prediction nodes. In this way, based on the first context index and / or the second context index corresponding to the i-th coordinate axis, the context model corresponding to the i-th coordinate axis can be selected from multiple context models, and then the context model corresponding to the i-th coordinate axis is used to predictively encode the coordinate information of the current point on the i-th coordinate axis. For example, the encoding end determines the second context index corresponding to the X-axis based on the coordinate information of the first point on the X-axis included in the N prediction nodes, and based on the first context index and / or the second context index corresponding to the X-axis, selects the context model corresponding to the X-axis from multiple context models, and then uses the context model corresponding to the X-axis to predictively encode the coordinate information of the current point on the X-axis to obtain the X-coordinate value of the current point. For another example, the encoding end determines the second context index corresponding to the Y-axis based on the coordinate information of the first point on the Y-axis included in the N prediction nodes, and based on the first context index and / or the second context index corresponding to the Y-axis, selects the context model corresponding to the Y-axis from multiple context models, and then uses the context model corresponding to the Y-axis to predict and encode the coordinate information of the current point on the Y-axis to obtain the Y coordinate value of the current point.

[0812] The following describes a process in which the encoder determines the second context index corresponding to the i-th coordinate axis based on the coordinate information of the first point included in the N prediction nodes on the i-th coordinate axis.

[0813] In the embodiment of the present application, the implementation of the above S202-A1-321 includes but is not limited to the following:

[0814] Method 1: Weight the first points included in the N prediction nodes, and determine the second context index corresponding to the i-th coordinate axis based on the weighted coordinate information. In this case, the above S202-A1-321 includes the following steps S202-A1-321-11 to S202-A1-321-13:

[0815] S202-A1-321-11. Determine a first weight corresponding to the prediction node;

[0816] S202-A1-321-12. Based on the first weight, weighted processing is performed on the coordinate information of the first point included in the N prediction nodes to obtain a first weighted point;

[0817] S202-A1-321-13. Determine a second context index corresponding to the i-th coordinate axis based on the coordinate information of the first weighted point on the i-th coordinate axis.

[0818] In this first approach, if the current node includes multiple prediction nodes, i.e., N prediction nodes, when determining the second context index corresponding to the i-th coordinate axis based on the coordinate information of the first point included in the N prediction nodes, a weight, i.e., a first weight, can be determined for each of the N prediction nodes. In this way, based on the first weight of each prediction node, the coordinate information of the first point included in each prediction node can be weighted to obtain a first weighted point. Furthermore, based on the coordinate information of the first weighted point on the i-th coordinate axis, the second context index corresponding to the i-th coordinate axis can be determined, thereby improving the accuracy of encoding the current point based on the geometric encoding information of the N prediction nodes.

[0819] The process of determining the first weights corresponding to the N prediction nodes in the embodiment of the present application can refer to the description of the above embodiment and will not be repeated here.

[0820] After determining the first weight corresponding to each of the N prediction nodes, the encoder performs weighted processing on the coordinate information of the first point included in the N prediction nodes based on the first weight to obtain a first weighted point.

[0821] The embodiment of the present application does not limit the specific method of obtaining the first weighted point by weighting the coordinate information of the first point included in the N prediction nodes based on the first weight.

[0822] In one example, based on the first weight, weighted averaging is performed on the coordinate information of the first point included in the N prediction nodes to obtain a first weighted point.

[0823] Based on the method steps, after the first weighted point is determined, the second context index corresponding to the i-th coordinate axis is determined based on the coordinate information of the first weighted point on the i-th coordinate axis.

[0824] As can be seen from the above, the first weighted point is obtained by weighting the first point in the N prediction nodes, wherein the values ​​of each bit of the first point after the prediction node can only be two results, 0 or 1. Therefore, in some embodiments, the values ​​of each bit of the first weighted point obtained by weighting the first point in the N prediction nodes are also 0 or 1. In this way, when encoding the i-th bit of the current point on the i-th coordinate axis, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined based on the value of the i-th bit of the first weighted point on the i-th coordinate axis. For example, if the value of the i-th bit of the first weighted point on the i-th coordinate axis is 0, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined to be 0. For another example, if the value of the i-th bit of the first weighted point on the i-th coordinate axis is 1, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined to be 1. Finally, the encoding end determines the context model corresponding to the i-th bit on the i-th coordinate axis based on the first context index and / or the second context index corresponding to the i-th bit on the i-th coordinate axis, and uses the context model to predict the value of the i-th bit of the current point on the i-th coordinate axis.

[0825] In addition to determining the second context index based on the above-mentioned method 1, the encoder can also determine the second context index through the following method 2.

[0826] Method 2: If K is greater than 1, the first point included in the prediction node in each of the K prediction reference frames is weighted, and the second context index corresponding to the i-th coordinate axis is determined based on the weighted coordinate information. In this case, the above S202-B includes the following steps S202-B21 to S202-B23:

[0827] S202-A1-321-21. For a j-th prediction reference frame among the K prediction reference frames, determine a first weight corresponding to a prediction node in the j-th prediction reference frame;

[0828] S202-A1-321-22. Based on the first weight, weight the coordinate information of the first point included in the prediction node in the j-th prediction reference frame to obtain a second weighted point corresponding to the j-th prediction reference frame, where j is a positive integer less than or equal to K.

[0829] S202-A1-321-23. Determine a second context index corresponding to the i-th coordinate axis based on the second weighted points corresponding to the K predicted reference frames.

[0830] In this second approach, when determining the geometric information of the current point, each of the K predicted reference frames is considered separately. Specifically, the coordinate information of the first point in the prediction node of each of the K predicted reference frames is determined, and the second weighted point corresponding to each predicted reference frame is determined. Then, based on the coordinate information of the second weighted point corresponding to each predicted reference frame, the second context index corresponding to the i-th coordinate axis is determined, thereby achieving accurate prediction of the second context index and improving the coding efficiency of the point cloud.

[0831] In the embodiment of the present application, the specific method in which the encoding end determines the second weighted point corresponding to each of the K prediction reference frames is the same. For ease of description, the j-th prediction reference frame among the K prediction reference frames is used as an example for illustration.

[0832] In an embodiment of the present application, the current node includes at least one prediction node in the j-th prediction reference frame, so that based on the coordinate information of the first point included in the at least one prediction node in the j-th prediction reference frame, the second weighted point corresponding to the j-th prediction reference frame is determined.

[0833] For example, the j-th prediction reference frame includes two prediction nodes of the current node, which are respectively recorded as prediction node 1 and prediction node 2, and then the geometric information of the first point included in the prediction node 1 and the coordinate information of the first point included in the prediction node 2 are weighted to obtain the second weighted point corresponding to the j-th prediction reference frame.

[0834] Before the encoder performs weighted processing on the geometric information of the first point included in the prediction node in the j-th prediction reference frame, it first needs to determine the first weight corresponding to each prediction node in the j-th prediction reference frame. The process of determining the first weight can refer to the description of the above embodiment and is not repeated here.

[0835] Next, the encoder performs weighted processing on the coordinate information of the first point included in the prediction node in the j-th prediction reference frame based on the first weight to obtain a second weighted point corresponding to the j-th prediction reference frame.

[0836] In one example, based on the first weight, weighted averaging is performed on the coordinate information of the first point included in the prediction node in the j-th prediction reference frame to obtain a second weighted point corresponding to the j-th prediction reference frame.

[0837] The above describes the process of determining the second weighted point corresponding to the j-th prediction reference frame among the K prediction reference frames. The second weighted points corresponding to other prediction reference frames among the K prediction reference frames can be determined in the same manner as that of the j-th prediction reference frame.

[0838] After the encoding end determines the second weighted point corresponding to each of the K predicted reference frames, it executes the above-mentioned step S202-A1-321-23.

[0839] The present application does not impose any limitation on the specific method of determining the second context index corresponding to the i-th coordinate axis based on the second weighted points corresponding to the K predicted reference frames.

[0840] In some embodiments, the encoding end determines an average value of coordinate information of the second weighted points corresponding to the K predicted reference frames on the i-th coordinate axis, and determines a second context index corresponding to the i-th coordinate axis based on the average value.

[0841] In some embodiments, the above S202-A1-321-23 includes the following steps S202-A1-321-231 to S202-A1-321-233:

[0842] S202-A1-321-231, determine second weights corresponding to K prediction reference frames;

[0843] S202-B232, weighting the geometric information of the second weighted points corresponding to the K prediction reference frames based on the second weight to obtain a third weighted point;

[0844] S202-B233. Determine a second context index corresponding to the i-th coordinate axis based on the coordinate information of the third weighted point on the i-th coordinate axis.

[0845] In this embodiment, the encoder can refer to the method of the above embodiment to determine the second weight corresponding to each of the K prediction reference frames. Then, based on the second weight, the coordinate information of the second weighted points corresponding to the K prediction reference frames is weighted to obtain a third weighted point.

[0846] For example, assuming K=2, for example, the current frame to be encoded includes 2 prediction reference frames, and these 2 prediction reference frames include the forward frame and backward frame of the current frame to be encoded. Assuming that the second weight corresponding to the forward frame is W1 and the second weight corresponding to the backward frame is W2, based on W1 and W2, the geometric information of the second weighted point corresponding to the forward frame and the geometric information of the second weighted point corresponding to the backward frame are weighted to obtain the geometric information of the third weighted point.

[0847] In one example, based on the second weight, weighted averaging is performed on the geometric information of the second weighted points corresponding to the K prediction reference frames to obtain coordinate information of the third weighted point.

[0848] After determining the geometric information of the third weighted point based on the above steps, the encoder determines the second context index corresponding to the i-th coordinate axis based on the coordinate information of the third weighted point on the i-th coordinate axis.

[0849] As can be seen from the above, the third weighted point is obtained by weighting the first point in the prediction node in each prediction reference frame, wherein the value of each bit of the first point after the prediction node has only two possible results, 0 or 1. Therefore, in some embodiments, the value of each bit of the first weighted point obtained by weighting the first point in the N prediction nodes is also 0 or 1. In this way, when encoding the i-th bit of the current point on the i-th coordinate axis, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined based on the value of the i-th bit of the third weighted point on the i-th coordinate axis. For example, if the value of the i-th bit of the third weighted point on the i-th coordinate axis is 0, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined to be 0. For another example, if the value of the i-th bit of the third weighted point on the i-th coordinate axis is 1, the second context index corresponding to the i-th bit on the i-th coordinate axis is determined to be 1. Finally, the encoding end determines the context model corresponding to the i-th bit on the i-th coordinate axis based on the first context index and / or the second context index corresponding to the i-th bit on the i-th coordinate axis, and uses the context model to predict the value of the i-th bit of the current point on the i-th coordinate axis.

[0850] After the encoding end determines the first context index and / or the second context index based on the above steps, it determines the context model based on the first context index and / or the second context index, and uses the context model to encode the coordinate information of the current point.

[0851] In one example, assuming that the DCM mode information of the prediction node is PredDCMode, the number of points in the prediction node is PredNumPoints, and the geometric information of the first point in the prediction node is predPointPos. Assume that the encoder uses the IDCM mode of the prediction node and the geometric information of the points in the prediction node to predictively encode the geometric information of the current point. That is, the geometric coding information of the prediction node used by the encoder includes the following two types:

[0852] 1) Predict the IDCM mode of the node;

[0853] 2) The geometric information of the midpoint of the predicted node (i.e., the first point), that is, the bit information (0 or 1) corresponding to the accuracy of the midpoint of the predicted node.

[0854] For example, in the GPCC framework, the IDCM mode of a prediction node includes PredDCMode(0,1,2). In the AVS framework, the IDCM mode of a prediction node includes PredDCMode(0,1).

[0855] Assuming that the number of points in the current node is numPoints, and the geometric information of each point is PointPos, and the bit precision depth to be encoded is nodeSizeLog2, the geometric information encoding process of each point in the current node is as follows:

[0856]

[0857] Through the above encoding process, the geometric information of each point in the current node can be obtained. Wherein, ctx1 is the first context index and ctx2 is the second context index.

[0858] The point cloud coding method provided by the embodiment of the present application, when encoding the current node in the current coding frame, determines N predicted nodes of the current node in the predicted reference frame of the current frame to be coded, and predictively codes the coordinate information of the midpoint of the current node based on the geometric coding information of the midpoints of these N predicted nodes. In other words, the embodiment of the present application optimizes the direct DCM coding of the node, and by considering the temporal correlation between adjacent frames, uses the geometric information of the predicted nodes in the predicted reference frame to predictively code the geometric information of the midpoint of the IDCM node (i.e., the current node) of the to-be-coded node, thereby further improving the efficiency of the geometric information coding of the point cloud by considering the temporal correlation between adjacent frames.

[0859] It should be understood that Figures 10 to 17 are merely examples of the present application and should not be understood as limiting the present application.

[0860] The preferred embodiments of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, a variety of simple modifications can be made to the technical solution of the present application, and these simple modifications all fall within the scope of protection of the present application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different embodiments of the present application can also be arbitrarily combined, and as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.

[0861] It should also be understood that in the various method embodiments of the present application, the size of the sequence numbers of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in the embodiments of the present application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three types of relationships can exist. Specifically, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the related objects before and after are in an "or" relationship.

[0862] The above describes in detail the method embodiment of the present application in conjunction with Figures 10 to 17, and the following describes in detail the device embodiment of the present application in conjunction with Figures 18 to 19.

[0863] Figure 18 is a schematic block diagram of the point cloud decoding device provided in an embodiment of the present application.

[0864] As shown in FIG18 , the point cloud decoding device 10 may include:

[0865] A determining unit 11 is configured to determine N prediction nodes of a current node in a prediction reference frame of a current frame to be decoded, where the current node is a node to be decoded in the current frame to be decoded, and N is a positive integer;

[0866] The decoding unit 12 is configured to perform predictive decoding on the coordinate information of the midpoint of the current node based on the geometric decoding information of the N predicted nodes.

[0867] In some embodiments, the current frame to be decoded includes K prediction reference frames, and the determination unit 11 is specifically used to determine, for the kth prediction reference frame among the K prediction reference frames, at least one prediction node of the current node in the kth prediction reference frame, where k is a positive integer less than or equal to K, and K is a positive integer; and determine N prediction nodes of the current node based on at least one prediction node of the current node in the K prediction reference frames.

[0868] In some embodiments, the determination unit 11 is specifically used to determine the M domain nodes of the current node in the current frame to be decoded, where the M domain nodes include the current node, and M is a positive integer; for the i-th domain node among the M domain nodes, determine the corresponding node of the i-th domain node in the k-th prediction reference frame, where i is a positive integer less than or equal to M; based on the corresponding nodes of the M domain nodes in the k-th prediction reference frame, determine at least one prediction node of the current node in the k-th prediction reference frame.

[0869] In some embodiments, the determination unit 11 is specifically used to determine the corresponding node of the current node in the kth prediction reference frame; determine at least one domain node of the corresponding node; and determine the at least one domain node as at least one prediction node of the current node in the kth prediction reference frame.

[0870] In some embodiments, the determination unit 11 is specifically used to determine the parent node of the i-th node in the current frame to be decoded as the i-th parent node, the i-th node being the i-th domain node or the current node; determine the matching node of the i-th parent node in the k-th prediction reference frame as the i-th matching node; and determine one of the child nodes of the i-th matching node as the corresponding node of the i-th node in the k-th prediction reference frame.

[0871] In some embodiments, the determining unit 11 is specifically configured to determine a matching node of the i-th parent node in the k-th prediction reference frame based on the occupancy information of the i-th parent node.

[0872] In some embodiments, the determination unit 11 is specifically configured to determine the node whose placeholder information in the k-th prediction reference frame has the smallest difference with the placeholder information of the i-th parent node as the matching node of the i-th parent node in the k-th prediction reference frame.

[0873] In some embodiments, the determination unit 11 is specifically used to determine the first serial number of the i-th node among the child nodes included in the parent node; and determine the child node with the first serial number among the child nodes of the i-th matching node as the corresponding node of the i-th node in the k-th predicted reference frame.

[0874] In some embodiments, the determining unit 11 is specifically configured to determine corresponding nodes of the M domain nodes in the k-th prediction reference frame as at least one prediction node of the current node in the k-th prediction reference frame.

[0875] In some embodiments, the determining unit 11 is specifically configured to determine at least one prediction node of the current node in the K prediction reference frames as the N prediction nodes of the current node.

[0876] In some embodiments, if the current frame to be decoded is a P frame, the K prediction reference frames include a forward frame of the current frame to be decoded.

[0877] In some embodiments, if the current frame to be decoded is a B frame, the K prediction reference frames include a forward frame and a backward frame of the current frame to be decoded.

[0878] In some embodiments, the prediction unit 12 is specifically used to determine the index of the context model based on the geometric decoding information of the N prediction nodes; determine the context model based on the index of the context model; and use the context model to predict and decode the coordinate information of the current point in the current node.

[0879] In some embodiments, the geometric decoding information of the prediction node includes direct decoding information of the prediction node and / or position information of the midpoint of the prediction node, and the direct decoding information is used to indicate whether the prediction node meets the conditions for decoding by direct decoding. The prediction unit 12 is specifically used to determine the first context index based on the direct decoding information of the N prediction nodes, and / or determine the second context index based on the coordinate information of the midpoint of the N prediction nodes; based on the first context index and / or the second context index, select the context model from a plurality of preset context models.

[0880] In some embodiments, the prediction unit 12 is specifically configured to determine, for any one of the N prediction nodes, a first numerical value corresponding to the prediction node based on direct decoding information of the prediction node; and determine the first context index based on the first numerical values ​​corresponding to the N prediction nodes.

[0881] In some embodiments, the prediction unit 12 is specifically configured to determine the first numerical value corresponding to the prediction node using the number of the direct decoding mode of the prediction node.

[0882] In some embodiments, the prediction unit 12 is specifically used to determine a first weight corresponding to the prediction node; based on the first weight, weighted processing is performed on the first numerical values ​​corresponding to the N prediction nodes to obtain a first weighted prediction value; based on the first weighted prediction value, the first context index is determined.

[0883] In some embodiments, the prediction unit 12 is specifically used to determine, for the j-th prediction reference frame among the K prediction reference frames, a first numerical value corresponding to the prediction node in the j-th prediction reference frame based on the direct decoding information of the prediction node of the current node in the j-th prediction reference frame, where j is a positive integer less than or equal to K; determine a first weight corresponding to the prediction node, and perform weighted processing on the first numerical value corresponding to the prediction node in the j-th prediction reference frame based on the first weight to obtain a second weighted prediction value corresponding to the j-th prediction reference frame; and determine the first context index based on the second weighted prediction values ​​corresponding to the K prediction reference frames.

[0884] In some embodiments, the prediction unit 12 is specifically used to determine the second weights corresponding to the K prediction reference frames; and perform weighted processing on the second weighted prediction values ​​corresponding to the K prediction reference frames based on the second weights to obtain the first context index.

[0885] In some embodiments, the prediction unit 12 is specifically used to select, for any prediction node among the N prediction nodes, a first point corresponding to the current point in the current node from the points included in the prediction node; and determine the second context index based on the coordinate information of the first point included in the N prediction nodes.

[0886] In some embodiments, the prediction unit 12 is specifically used to determine the second context index corresponding to the i-th coordinate axis based on the coordinate information of the first point included in the N prediction nodes on the i-th coordinate axis, where the i-th coordinate axis is the X-coordinate axis, the Y-coordinate axis or the Z-coordinate axis; based on the first context index and / or the second context index corresponding to the i-th coordinate axis, select the context model corresponding to the i-th coordinate axis from the multiple context models; and use the context model corresponding to the i-th coordinate axis to predict and decode the coordinate information of the current point on the i-th coordinate axis.

[0887] In some embodiments, the prediction unit 12 is specifically used to determine a first weight corresponding to the prediction node; based on the first weight, weighted processing is performed on the coordinate information of the first point included in the N prediction nodes to obtain a first weighted point; based on the coordinate information of the first weighted point on the i-th coordinate axis, the second context index corresponding to the i-th coordinate axis is determined.

[0888] In some embodiments, if K is greater than 1, the prediction unit 12 is specifically used to determine the first weight corresponding to the prediction node in the j-th prediction reference frame based on the j-th prediction reference frame among the K prediction reference frames; based on the first weight, weighted processing is performed on the coordinate information of the first point included in the prediction node in the j-th prediction reference frame to obtain the second weighted point corresponding to the j-th prediction reference frame, where j is a positive integer less than or equal to K; based on the second weighted points corresponding to the K prediction reference frames, the second context index corresponding to the i-th coordinate axis is determined.

[0889] In some embodiments, the prediction unit 12 is specifically used to determine the second weight corresponding to the K prediction reference frames; perform weighted processing on the coordinate information of the second weighted point corresponding to the K prediction reference frames based on the second weight to obtain a third weighted point; and determine the second context index corresponding to the i-th coordinate axis based on the coordinate information of the third weighted point on the i-th coordinate axis.

[0890] In some embodiments, the prediction unit 12 is specifically configured to determine a first weight corresponding to the prediction node based on a distance between a domain node corresponding to the prediction node and the current node.

[0891] In some embodiments, the prediction unit 12 is specifically configured to determine a second weight corresponding to the prediction reference frame based on a time difference between the prediction reference frame and the current frame to be decoded.

[0892] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions can refer to the method embodiment. To avoid repetition, no further description is given here. Specifically, the point cloud decoding device 10 shown in FIG18 may correspond to the corresponding subject in the point cloud decoding method of the embodiment of the present application, and the aforementioned and other operations and / or functions of each unit in the point cloud decoding device 10 are respectively for implementing the corresponding processes in the point cloud decoding method. For the sake of brevity, no further description is given here.

[0893] Figure 19 is a schematic block diagram of the point cloud encoding device provided in an embodiment of the present application.

[0894] As shown in FIG19 , the point cloud encoding device 20 includes:

[0895] The determining unit 21 is specifically configured to determine N prediction nodes of a current node in a prediction reference frame of a current frame to be encoded, where the current node is a node to be encoded in the current frame to be encoded, and N is a positive integer;

[0896] The encoding unit 22 is configured to perform predictive encoding on the coordinate information of the midpoint of the current node based on the geometric encoding information of the N predicted nodes.

[0897] In some embodiments, the current frame to be encoded includes K prediction reference frames, and the determination unit 21 is specifically used to determine, for the kth prediction reference frame among the K prediction reference frames, at least one prediction node of the current node in the kth prediction reference frame, where k is a positive integer less than or equal to K, and K is a positive integer; and determine N prediction nodes of the current node based on at least one prediction node of the current node in the K prediction reference frames.

[0898] In some embodiments, the determination unit 21 is specifically used to determine the M domain nodes of the current node in the current frame to be encoded, where the M domain nodes include the current node, and M is a positive integer; for the i-th domain node among the M domain nodes, determine the corresponding node of the i-th domain node in the k-th prediction reference frame, where i is a positive integer less than or equal to M; based on the corresponding nodes of the M domain nodes in the k-th prediction reference frame, determine at least one prediction node of the current node in the k-th prediction reference frame.

[0899] In some embodiments, the determination unit 21 is specifically used to determine the corresponding node of the current node in the kth prediction reference frame; determine at least one domain node of the corresponding node; and determine the at least one domain node as at least one prediction node of the current node in the kth prediction reference frame.

[0900] In some embodiments, the determination unit 21 is specifically used to determine the parent node of the i-th node in the current frame to be encoded as the i-th parent node, the i-th node being the i-th domain node or the current node; determine the matching node of the i-th parent node in the k-th prediction reference frame as the i-th matching node; and determine one of the child nodes of the i-matching node as the corresponding node of the i-th node in the k-th prediction reference frame.

[0901] In some embodiments, the determining unit 21 is specifically configured to determine a matching node of the i-th parent node in the k-th prediction reference frame based on the occupancy information of the i-th parent node.

[0902] In some embodiments, the determination unit 21 is specifically configured to determine the node whose placeholder information in the k-th prediction reference frame has the smallest difference with the placeholder information of the i-th parent node as the matching node of the i-th parent node in the k-th prediction reference frame.

[0903] In some embodiments, the determination unit 21 is specifically used to determine the first serial number of the i-th node among the child nodes included in the parent node; and determine the child node with the first serial number among the child nodes of the i-th matching node as the corresponding node of the i-th node in the k-th predicted reference frame.

[0904] In some embodiments, the determining unit 21 is specifically configured to determine corresponding nodes of the M domain nodes in the k-th prediction reference frame as at least one prediction node of the current node in the k-th prediction reference frame.

[0905] In some embodiments, the determining unit 21 is specifically configured to determine at least one prediction node of the current node in the K prediction reference frames as the N prediction nodes of the current node.

[0906] In some embodiments, if the current frame to be encoded is a P frame, the K prediction reference frames include a forward frame of the current frame to be encoded.

[0907] In some embodiments, if the current frame to be encoded is a B frame, the K prediction reference frames include a forward frame and a backward frame of the current frame to be encoded.

[0908] In some embodiments, the encoding unit 22 is specifically used to determine the index of the context model based on the geometric encoding information of the N prediction nodes; determine the context model based on the index of the context model; and use the context model to predict and encode the coordinate information of the current point in the current node.

[0909] In some embodiments, the geometric coding information of the prediction node includes direct coding information of the prediction node and / or position information of the midpoint of the prediction node, and the direct coding information is used to indicate whether the prediction node meets the conditions for encoding in a direct coding manner. The encoding unit 22 is specifically used to determine a first context index based on the direct coding information of the N prediction nodes, and / or determine a second context index based on the coordinate information of the midpoint of the N prediction nodes; based on the first context index and / or the second context index, select the context model from a plurality of preset context models.

[0910] In some embodiments, the encoding unit 22 is specifically configured to determine, for any one of the N prediction nodes, a first numerical value corresponding to the prediction node based on direct encoding information of the prediction node; and determine the first context index based on the first numerical values ​​corresponding to the N prediction nodes.

[0911] In some embodiments, the direct encoding information includes a direct encoding mode of the prediction node, and the encoding unit 22 is specifically configured to determine a first value corresponding to the prediction node using the number of the direct encoding mode of the prediction node.

[0912] In some embodiments, the encoding unit 22 is specifically used to determine a first weight corresponding to the prediction node; based on the first weight, weighted processing is performed on the first numerical values ​​corresponding to the N prediction nodes to obtain a first weighted prediction value; based on the first weighted prediction value, the first context index is determined.

[0913] In some embodiments, the encoding unit 22 is specifically used to determine, for the j-th prediction reference frame among the K prediction reference frames, a first numerical value corresponding to the prediction node in the j-th prediction reference frame based on the direct encoding information of the prediction node of the current node in the j-th prediction reference frame, where j is a positive integer less than or equal to K; determine a first weight corresponding to the prediction node, and weightedly process the first numerical value corresponding to the prediction node in the j-th prediction reference frame based on the first weight to obtain a second weighted prediction value corresponding to the j-th prediction reference frame; and determine the first context index based on the second weighted prediction values ​​corresponding to the K prediction reference frames.

[0914] In some embodiments, the encoding unit 22 is specifically used to determine the second weights corresponding to the K prediction reference frames; and perform weighted processing on the second weighted prediction values ​​corresponding to the K prediction reference frames based on the second weights to obtain the first context index.

[0915] In some embodiments, the encoding unit 22 is specifically configured to, for any prediction node among the N prediction nodes, select a first point corresponding to the current point in the current node from the points included in the prediction node; and determine the second context index based on the coordinate information of the first point included in the N prediction nodes.

[0916] In some embodiments, the encoding unit 22 is specifically used to determine the second context index corresponding to the i-th coordinate axis based on the coordinate information of the first point included in the N prediction nodes on the i-th coordinate axis, where the i-th coordinate axis is the X-coordinate axis, the Y-coordinate axis or the Z-coordinate axis; based on the first context index and / or the second context index corresponding to the i-th coordinate axis, select the context model corresponding to the i-th coordinate axis from the multiple context models; and use the context model corresponding to the i-th coordinate axis to predict and encode the coordinate information of the current point on the i-th coordinate axis.

[0917] In some embodiments, the encoding unit 22 is specifically used to determine a first weight corresponding to the prediction node; based on the first weight, weighted processing is performed on the coordinate information of the first point included in the N prediction nodes to obtain a first weighted point; based on the coordinate information of the first weighted point on the i-th coordinate axis, determining a second context index corresponding to the i-th coordinate axis.

[0918] In some embodiments, the encoding unit 22 is specifically used to determine, for the j-th prediction reference frame among the K prediction reference frames, a first weight corresponding to the prediction node in the j-th prediction reference frame; based on the first weight, weighted processing is performed on the coordinate information of the first point included in the prediction node in the j-th prediction reference frame to obtain a second weighted point corresponding to the j-th prediction reference frame, where j is a positive integer less than or equal to K; based on the second weighted points corresponding to the K prediction reference frames, determine the second context index corresponding to the i-th coordinate axis.

[0919] In some embodiments, the encoding unit 22 is specifically used to determine the second weights corresponding to the K predicted reference frames; weight the coordinate information of the second weighted points corresponding to the K predicted reference frames based on the second weights to obtain a third weighted point; and determine the second context index corresponding to the i-th coordinate axis based on the coordinate information of the third weighted point on the i-th coordinate axis.

[0920] In some embodiments, the encoding unit 22 is specifically configured to determine a first weight corresponding to the prediction node based on a distance between a domain node corresponding to the prediction node and the current node.

[0921] In some embodiments, the encoding unit 22 is specifically configured to determine a second weight corresponding to the predicted reference frame based on a time difference between the predicted reference frame and the current frame to be encoded.

[0922] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, no further description is given here. Specifically, the point cloud coding device 20 shown in FIG19 may correspond to the corresponding subject in the point cloud coding method of the embodiment of the present application, and the aforementioned and other operations and / or functions of each unit in the point cloud coding device 20 are respectively for implementing the corresponding processes in the point cloud coding method. For the sake of brevity, no further description is given here.

[0923] The above describes the apparatus and system of the embodiment of the present application from the perspective of functional units in conjunction with the accompanying drawings. It should be understood that the functional unit can be implemented in the form of hardware, can be implemented by instructions in the form of software, or can be implemented by a combination of hardware and software units. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software units in the decoding processor. Optionally, the software unit can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.

[0924] Figure 20 is a schematic block diagram of an electronic device provided in an embodiment of the present application.

[0925] As shown in FIG20 , the electronic device 30 may be a point cloud decoding device or a point cloud encoding device as described in an embodiment of the present application. The electronic device 30 may include:

[0926] The memory 33 and the processor 32 are configured to store a computer program 34 and transmit the program code 34 to the processor 32. In other words, the processor 32 can call and run the computer program 34 from the memory 33 to implement the method in the embodiment of the present application.

[0927] For example, the processor 32 may be configured to execute the steps of the method 200 according to the instructions in the computer program 34 .

[0928] In some embodiments of the present application, the processor 32 may include but is not limited to:

[0929] General-purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0930] In some embodiments of the present application, the memory 33 includes but is not limited to:

[0931] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0932] In some embodiments of the present application, the computer program 34 may be divided into one or more units, which are stored in the memory 33 and executed by the processor 32 to implement the method provided by the present application. The one or more units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program 34 in the electronic device 30.

[0933] As shown in FIG20 , the electronic device 30 may further include:

[0934] The transceiver 33 may be connected to the processor 32 or the memory 33 .

[0935] The processor 32 can control the transceiver 33 to communicate with other devices. Specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 33 may include a transmitter and a receiver. The transceiver 33 may further include an antenna, and the number of antennas may be one or more.

[0936] It should be understood that the various components in the electronic device 30 are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0937] Figure 21 is a schematic block diagram of the point cloud encoding and decoding system provided in an embodiment of the present application.

[0938] As shown in Figure 21, the point cloud encoding and decoding system 40 may include: a point cloud encoder 41 and a point cloud decoder 42, wherein the point cloud encoder 41 is used to execute the point cloud encoding method involved in the embodiment of the present application, and the point cloud decoder 42 is used to execute the point cloud decoding method involved in the embodiment of the present application.

[0939] The present application also provides a code stream, which is generated according to the above encoding method.

[0940] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a computer, enables the computer to perform the method of the above-mentioned method embodiment. In other words, the present application also provides a computer program product containing instructions, which, when executed by a computer, enables the computer to perform the meth...

Claims

1. A point cloud decoding method, It is characterized in that include: In a prediction reference frame of a current frame to be decoded, determine N prediction nodes of a current node, where the current node is a node to be decoded in the current frame to be decoded, and N is a positive integer; Based on the geometric decoding information of the N predicted nodes, the coordinate information of the midpoint of the current node is predicted and decoded.

2. The method according to claim 1, It is characterized in that The current frame to be decoded includes K prediction reference frames, and determining N prediction nodes of the current node in the prediction reference frames of the current frame to be decoded includes: For a k-th prediction reference frame among the K prediction reference frames, determining at least one prediction node of the current node in the k-th prediction reference frame, where k is a positive integer less than or equal to K, and K is a positive integer; Based on at least one prediction node of the current node in the K prediction reference frames, N prediction nodes of the current node are determined.

3. The method according to claim 2, It is characterized in that The determining at least one prediction node of the current node in the k-th prediction reference frame comprises: In the current frame to be decoded, determine M domain nodes of the current node, the M domain nodes include the current node, and M is a positive integer; For an i-th domain node among the M domain nodes, determine a corresponding node of the i-th domain node in the k-th prediction reference frame, where i is a positive integer less than or equal to M; Based on the corresponding nodes of the M domain nodes in the k-th prediction reference frame, at least one prediction node of the current node in the k-th prediction reference frame is determined.

4. The method according to claim 2, It is characterized in that The determining at least one prediction node of the current node in the k-th prediction reference frame comprises: Determine a corresponding node of the current node in the k-th prediction reference frame; Determining at least one domain node of the corresponding node; The at least one domain node is determined as at least one prediction node of the current node in the k-th prediction reference frame.

5. The method according to claim 3 or 4, It is characterized in that The method further comprises: In the current frame to be decoded, determine the parent node of the ith node as the ith parent node, the ith node being the ith domain node or the current node; Determine a matching node of the i-th parent node in the k-th prediction reference frame as the i-th matching node; One of the child nodes of the i matching nodes is determined as the corresponding node of the i-th node in the k-th prediction reference frame.

6. The method according to claim 5, It is characterized in that The determining a matching node of the i-th parent node in the k-th prediction reference frame includes: Based on the placeholder information of the i-th parent node, a matching node of the i-th parent node in the k-th prediction reference frame is determined.

7. The method according to claim 6, It is characterized in that The determining, based on the placeholder information of the i-th parent node, a matching node of the i-th parent node in the k-th prediction reference frame includes: A node whose placeholder information in the k-th prediction reference frame has the smallest difference with the placeholder information of the i-th parent node is determined as a matching node of the i-th parent node in the k-th prediction reference frame.

8. The method according to claim 5, It is characterized in that The step of determining one of the child nodes of the i matching nodes as a corresponding node of the i-th node in the k-th prediction reference frame includes: Determine the first sequence number of the i-th node among the child nodes included in the parent node; The child node with the first sequence number among the child nodes of the i-th matching node is determined as the corresponding node of the i-th node in the k-th prediction reference frame.

9. The method according to claim 3, It is characterized in that The determining, based on corresponding nodes of the M domain nodes in the kth prediction reference frame, at least one prediction node of the current node in the kth prediction reference frame comprises: The corresponding nodes of the M domain nodes in the k-th prediction reference frame are determined as at least one prediction node of the current node in the k-th prediction reference frame.

10. The method according to claim 2, It is characterized in that The determining, based on at least one prediction node of the current node in the K prediction reference frames, N prediction nodes of the current node comprises: At least one prediction node of the current node in the K prediction reference frames is determined as N prediction nodes of the current node.

11. The method according to claim 2, It is characterized in that If the current frame to be decoded is a P frame, the K prediction reference frames include a forward frame of the current frame to be decoded.

12. The method according to claim 2, It is characterized in that If the current frame to be decoded is a B frame, the K prediction reference frames include a forward frame and a backward frame of the current frame to be decoded.

13. The method according to any one of claims 2 to 12, It is characterized in that The predictive decoding of the coordinate information of the midpoint of the current node based on the geometric decoding information of the N predicted nodes includes: Determining an index of a context model based on the geometric decoding information of the N prediction nodes; Determining the context model based on the index of the context model; Using the context model, the coordinate information of the current point in the current node is predicted and decoded.

14. The method according to claim 13, It is characterized in that The geometric decoding information of the prediction node includes direct decoding information of the prediction node and / or position information of a midpoint of the prediction node, the direct decoding information is used to indicate whether the prediction node satisfies a condition for decoding in a direct decoding manner, and determining an index of a context model based on the geometric decoding information of the N prediction nodes includes: Determine a first context index based on direct decoding information of the N prediction nodes, and / or determine a second context index based on coordinate information of midpoints of the N prediction nodes; The selecting the context model based on the index of the context model comprises: Based on the first context index and / or the second context index, the context model is selected from a plurality of preset context models.

15. The method according to claim 14, It is characterized in that The determining, based on the direct decoding information of the N prediction nodes, a first context index comprises: For any prediction node among the N prediction nodes, determining a first value corresponding to the prediction node based on direct decoding information of the prediction node; The first context index is determined based on first values ​​corresponding to the N prediction nodes.

16. The method according to claim 15, It is characterized in that The direct decoding information includes a direct decoding mode of the prediction node, and determining a first value corresponding to the prediction node based on the direct decoding information of the prediction node includes: The direct decoding mode number of the prediction node is used to determine a first value corresponding to the prediction node.

17. The method according to claim 15, It is characterized in that The determining the first context index based on the first values ​​corresponding to the N prediction nodes includes: Determining a first weight corresponding to the prediction node; Based on the first weight, weighting the first values ​​corresponding to the N prediction nodes to obtain a first weighted prediction value; Based on the first weighted prediction value, the first context index is determined.

18. The method according to claim 14, It is characterized in that If K is greater than 1, determining the first context index based on the direct decoding information of the N prediction nodes includes: For a j-th prediction reference frame among the K prediction reference frames, determining a first value corresponding to the prediction node in the j-th prediction reference frame based on direct decoding information of the prediction node of the current node in the j-th prediction reference frame, where j is a positive integer less than or equal to K; Determine a first weight corresponding to the prediction node, and perform weighted processing on a first value corresponding to the prediction node in the j-th prediction reference frame based on the first weight to obtain a second weighted prediction value corresponding to the j-th prediction reference frame; The first context index is determined based on second weighted prediction values ​​corresponding to the K prediction reference frames.

19. The method according to claim 18, It is characterized in that The determining the first context index based on the second weighted prediction values ​​corresponding to the K prediction reference frames includes: Determine second weights corresponding to the K prediction reference frames; Based on the second weight, weighted processing is performed on the second weighted prediction values ​​respectively corresponding to the K prediction reference frames to obtain the first context index.

20. The method according to claim 14, It is characterized in that The determining the second context index based on the coordinate information of the midpoints of the N prediction nodes includes: For any prediction node among the N prediction nodes, selecting a first point corresponding to a current point in the current node from the points included in the prediction node; The second context index is determined based on coordinate information of a first point included in the N prediction nodes.

21. The method according to claim 20, It is characterized in that The determining the second context index based on the coordinate information of the first point included in the N prediction nodes includes: Determine, based on coordinate information of a first point included in the N prediction nodes on an i-th coordinate axis, a second context index corresponding to the i-th coordinate axis, where the i-th coordinate axis is an X-coordinate axis, a Y-coordinate axis, or a Z-coordinate axis; The selecting the context model from a plurality of preset context models based on the first context index and / or the second context index includes: Based on the first context index and / or the second context index corresponding to the i-th coordinate axis, selecting a context model corresponding to the i-th coordinate axis from the multiple context models; The using the context model to predict and decode the coordinate information of the current point in the current node includes: Use the context model corresponding to the i-th coordinate axis to predict and decode the coordinate information of the current point on the i-th coordinate axis.

22. The method according to claim 21, It is characterized in that The determining, based on coordinate information of a first point included in the N prediction nodes on an i-th coordinate axis, a second context index corresponding to the i-th coordinate axis includes: Determining a first weight corresponding to the prediction node; Based on the first weight, weighted processing is performed on the coordinate information of the first point included in the N prediction nodes to obtain a first weighted point; Based on the coordinate information of the first weighted point on the i-th coordinate axis, a second context index corresponding to the i-th coordinate axis is determined.

23. The method according to claim 21, It is characterized in that If K is greater than 1, determining the second context index corresponding to the i-th coordinate axis based on the coordinate information of the first point included in the N prediction nodes on the i-th coordinate axis includes: For a j-th prediction reference frame among the K prediction reference frames, determining a first weight corresponding to a prediction node in the j-th prediction reference frame; Based on the first weight, weighted processing is performed on the coordinate information of the first point included in the prediction node in the j-th prediction reference frame to obtain a second weighted point corresponding to the j-th prediction reference frame, where j is a positive integer less than or equal to K; Based on the second weighted points corresponding to the K prediction reference frames, a second context index corresponding to the i-th coordinate axis is determined.

24. The method according to claim 23, It is characterized in that The determining, based on the coordinate information of the second weighted point corresponding to the K prediction reference frames, the second context index corresponding to the i-th coordinate axis includes: Determine second weights corresponding to the K prediction reference frames; performing weighted processing on the coordinate information of the second weighted points corresponding to the K prediction reference frames based on the second weight to obtain a third weighted point; Based on the coordinate information of the third weighted point on the i-th coordinate axis, a second context index corresponding to the i-th coordinate axis is determined.

25. The method according to claim 17, 18, 22 or 23, It is characterized in that Determining a first weight corresponding to the prediction node includes: Based on the distance between the domain node corresponding to the prediction node and the current node, a first weight corresponding to the prediction node is determined.

26. The method according to claim 19 or 24, It is characterized in that The determining the second weights corresponding to the K prediction reference frames includes: Based on the time difference between the predicted reference frame and the current frame to be decoded, a second weight corresponding to the predicted reference frame is determined.

27. A point cloud encoding method, It is characterized in that include: In a prediction reference frame of a current frame to be encoded, determining N prediction nodes of a current node, wherein the current node is a node to be encoded in the current frame to be encoded, and N is a positive integer; Based on the geometric coding information of the N predicted nodes, the coordinate information of the midpoint of the current node is predicted and coded.

28. The method according to claim 27, It is characterized in that The current frame to be encoded includes K prediction reference frames, and determining N prediction nodes of the current node in the prediction reference frame of the current frame to be encoded includes: For a k-th prediction reference frame among the K prediction reference frames, determining at least one prediction node of the current node in the k-th prediction reference frame, where k is a positive integer less than or equal to K, and K is a positive integer; Based on at least one prediction node of the current node in the K prediction reference frames, N prediction nodes of the current node are determined.

29. The method according to claim 28, It is characterized in that The determining at least one prediction node of the current node in the k-th prediction reference frame comprises: In the current frame to be encoded, determine M domain nodes of the current node, the M domain nodes include the current node, and M is a positive integer; For an i-th domain node among the M domain nodes, determine a corresponding node of the i-th domain node in the k-th prediction reference frame, where i is a positive integer less than or equal to M; Based on the corresponding nodes of the M domain nodes in the k-th prediction reference frame, at least one prediction node of the current node in the k-th prediction reference frame is determined.

30. The method according to claim 28, It is characterized in that The determining at least one prediction node of the current node in the k-th prediction reference frame comprises: Determine a corresponding node of the current node in the k-th prediction reference frame; Determining at least one domain node of the corresponding node; The at least one domain node is determined as at least one prediction node of the current node in the k-th prediction reference frame.

31. The method according to claim 29 or 30, It is characterized in that The method further comprises: In the current frame to be encoded, determine the parent node of the ith node as the ith parent node, the ith node being the ith domain node or the current node; Determine a matching node of the i-th parent node in the k-th prediction reference frame as the i-th matching node; One of the child nodes of the i matching nodes is determined as the corresponding node of the i-th node in the k-th prediction reference frame.

32. The method according to claim 31, It is characterized in that The determining a matching node of the i-th parent node in the k-th prediction reference frame includes: Based on the placeholder information of the i-th parent node, a matching node of the i-th parent node in the k-th prediction reference frame is determined.

33. The method according to claim 32, It is characterized in that The determining, based on the placeholder information of the i-th parent node, a matching node of the i-th parent node in the k-th prediction reference frame includes: A node whose placeholder information in the k-th prediction reference frame has the smallest difference with the placeholder information of the i-th parent node is determined as a matching node of the i-th parent node in the k-th prediction reference frame.

34. The method according to claim 31, It is characterized in that The step of determining one of the child nodes of the i matching nodes as a corresponding node of the i-th node in the k-th prediction reference frame includes: Determine the first sequence number of the i-th node among the child nodes included in the parent node; The child node with the first sequence number among the child nodes of the i-th matching node is determined as the corresponding node of the i-th node in the k-th prediction reference frame.

35. The method according to claim 29, It is characterized in that The determining, based on corresponding nodes of the M domain nodes in the kth prediction reference frame, at least one prediction node of the current node in the kth prediction reference frame comprises: The corresponding nodes of the M domain nodes in the k-th prediction reference frame are determined as at least one prediction node of the current node in the k-th prediction reference frame.

36. The method according to claim 28, It is characterized in that The determining, based on at least one prediction node of the current node in the K prediction reference frames, N prediction nodes of the current node comprises: At least one prediction node of the current node in the K prediction reference frames is determined as N prediction nodes of the current node.

37. The method according to claim 28, It is characterized in that If the current frame to be encoded is a P frame, the K prediction reference frames include a forward frame of the current frame to be encoded.

38. The method according to claim 28, It is characterized in that If the current frame to be encoded is a B frame, the K prediction reference frames include a forward frame and a backward frame of the current frame to be encoded.

39. The method according to any one of claims 28 to 38, It is characterized in that The predictive coding of the coordinate information of the midpoint of the current node based on the geometric coding information of the N predicted nodes includes: Determining an index of a context model based on the geometric coding information of the N prediction nodes; Determining the context model based on the index of the context model; Using the context model, the coordinate information of the current point in the current node is predictively encoded.

40. The method according to claim 39, It is characterized in that The geometric coding information of the prediction node includes direct coding information of the prediction node and / or position information of a midpoint of the prediction node, the direct coding information is used to indicate whether the prediction node satisfies a condition for encoding in a direct coding manner, and determining an index of a context model based on the geometric coding information of the N prediction nodes includes: Determine a first context index based on direct encoding information of the N prediction nodes, and / or determine a second context index based on coordinate information of midpoints of the N prediction nodes; The selecting the context model based on the index of the context model comprises: Based on the first context index and / or the second context index, the context model is selected from a plurality of preset context models.

41. The method according to claim 40, It is characterized in that The determining a first context index based on the direct encoding information of the N prediction nodes includes: For any prediction node among the N prediction nodes, determining a first value corresponding to the prediction node based on direct encoding information of the prediction node; The first context index is determined based on first values ​​corresponding to the N prediction nodes.

42. The method according to claim 41, It is characterized in that The direct encoding information includes a direct encoding mode of the prediction node, and determining a first value corresponding to the prediction node based on the direct encoding information of the prediction node includes: The direct coding mode number of the prediction node is used to determine a first value corresponding to the prediction node.

43. The method according to claim 41, It is characterized in that The determining the first context index based on the first values ​​corresponding to the N prediction nodes includes: Determining a first weight corresponding to the prediction node; Based on the first weight, weighting the first values ​​corresponding to the N prediction nodes to obtain a first weighted prediction value; Based on the first weighted prediction value, the first context index is determined.

44. The method according to claim 40, It is characterized in that If K is greater than 1, determining the first context index based on the direct encoding information of the N prediction nodes includes: For a j-th prediction reference frame among the K prediction reference frames, determining a first value corresponding to the prediction node in the j-th prediction reference frame based on direct encoding information of the prediction node of the current node in the j-th prediction reference frame, where j is a positive integer less than or equal to K; Determine a first weight corresponding to the prediction node, and perform weighted processing on a first value corresponding to the prediction node in the j-th prediction reference frame based on the first weight to obtain a second weighted prediction value corresponding to the j-th prediction reference frame; The first context index is determined based on second weighted prediction values ​​corresponding to the K prediction reference frames.

45. The method according to claim 44, It is characterized in that The determining the first context index based on the second weighted prediction values ​​corresponding to the K prediction reference frames includes: Determine second weights corresponding to the K prediction reference frames; Based on the second weight, weighted processing is performed on the second weighted prediction values ​​respectively corresponding to the K prediction reference frames to obtain the first context index.

46. ​​The method according to claim 40, It is characterized in that The determining the second context index based on the coordinate information of the midpoints of the N prediction nodes includes: For any prediction node among the N prediction nodes, selecting a first point corresponding to a current point in the current node from the points included in the prediction node; The second context index is determined based on coordinate information of a first point included in the N prediction nodes.

47. The method according to claim 46, It is characterized in that The determining the second context index based on the coordinate information of the first point included in the N prediction nodes includes: Determine, based on coordinate information of a first point included in the N prediction nodes on an i-th coordinate axis, a second context index corresponding to the i-th coordinate axis, where the i-th coordinate axis is an X-coordinate axis, a Y-coordinate axis, or a Z-coordinate axis; The selecting the context model from a plurality of preset context models based on the first context index and / or the second context index includes: Based on the first context index and / or the second context index corresponding to the i-th coordinate axis, selecting a context model corresponding to the i-th coordinate axis from the multiple context models; The using the context model to predictively encode the coordinate information of the current point in the current node includes: The context model corresponding to the i-th coordinate axis is used to predictively encode the coordinate information of the current point on the i-th coordinate axis.

48. The method according to claim 47, It is characterized in that The determining, based on coordinate information of a first point included in the N prediction nodes on an i-th coordinate axis, a second context index corresponding to the i-th coordinate axis includes: Determining a first weight corresponding to the prediction node; Based on the first weight, weighted processing is performed on the coordinate information of the first point included in the N prediction nodes to obtain a first weighted point; Based on the coordinate information of the first weighted point on the i-th coordinate axis, a second context index corresponding to the i-th coordinate axis is determined.

49. The method according to claim 47, It is characterized in that If K is greater than 1, determining the second context index corresponding to the i-th coordinate axis based on the coordinate information of the first point included in the N prediction nodes on the i-th coordinate axis includes: For a j-th prediction reference frame among the K prediction reference frames, determining a first weight corresponding to a prediction node in the j-th prediction reference frame; Based on the first weight, weighted processing is performed on the coordinate information of the first point included in the prediction node in the j-th prediction reference frame to obtain a second weighted point corresponding to the j-th prediction reference frame, where j is a positive integer less than or equal to K; Based on the second weighted points corresponding to the K prediction reference frames, a second context index corresponding to the i-th coordinate axis is determined.

50. The method according to claim 49, It is characterized in that The determining, based on the coordinate information of the second weighted point corresponding to the K prediction reference frames, the second context index corresponding to the i-th coordinate axis includes: Determine second weights corresponding to the K prediction reference frames; performing weighted processing on the coordinate information of the second weighted points corresponding to the K prediction reference frames based on the second weight to obtain a third weighted point; Based on the coordinate information of the third weighted point on the i-th coordinate axis, a second context index corresponding to the i-th coordinate axis is determined.

51. The method of claim 43, 44, 48 or 49, It is characterized in that Determining a first weight corresponding to the prediction node includes: Based on the distance between the domain node corresponding to the prediction node and the current node, a first weight corresponding to the prediction node is determined.

52. The method according to claim 45 or 50, It is characterized in that The determining the second weights corresponding to the K prediction reference frames includes: Based on the time difference between the predicted reference frame and the current frame to be encoded, a second weight corresponding to the predicted reference frame is determined.

53. A point cloud decoding device, It is characterized in that include: A determination unit, configured to determine N prediction nodes of a current node in a prediction reference frame of a current frame to be decoded, wherein the current node is a node to be decoded in the current frame to be decoded, and N is a positive integer; A decoding unit is used to predict and decode the coordinate information of the midpoint of the current node based on the geometric decoding information of the N predicted nodes.

54. A point cloud encoding device, It is characterized in that include: A determination unit, specifically configured to determine N prediction nodes of a current node in a prediction reference frame of a current frame to be encoded, wherein the current node is a node to be encoded in the current frame to be encoded, and N is a positive integer; The encoding unit is used to predict the coordinate information of the midpoint of the current node based on the geometric encoding information of the N prediction nodes.

55. An electronic device, It is characterized in that include: Processor and memory; The memory is used to store computer programs; The processor is used to call and run the computer program stored in the memory to perform the method according to any one of claims 1 to 26 or 27 to 52.

56. A computer readable storage medium, It is characterized in that Used to store a computer program, the computer program causing a computer to execute the method according to any one of claims 1 to 26 or 27 to 52.