Point cloud coding and decoding method, device, equipment and storage medium

CN120345252APending Publication Date: 2025-07-18GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085781.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-01-06
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the process of point cloud coding, the existing technology only relies on a priori reference information for predictive coding for flat nodes or nodes with planar characteristics, resulting in poor predictive coding performance of planar structure information and affecting the coding efficiency of point cloud geometric information.

Method used

By determining the N domain nodes of the current node, and based on the occupancy information of these domain nodes, the plane structure information of the current node is predictively encoded and decoded, which improves the predictive encoding and decoding performance of the plane structure information and improves the encoding and decoding efficiency of the point cloud. and performance.

Benefits of technology

It effectively improves the predictive encoding and decoding performance of planar structure information, improves the coding efficiency and performance of point cloud data, and improves the transmission and storage efficiency of point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345252A_ABST
    Figure CN120345252A_ABST
Patent Text Reader

Abstract

The invention provides a point cloud coding and decoding method and device, equipment and a storage medium, and the method comprises the steps: determining N domain nodes of a current node when the plane structure information of the current node is coded and decoded, and carrying out the coding and decoding of the plane structure information of the current node based on the occupation information of the N domain nodes. In other words, when the plane structure information of the current node is subjected to prediction coding and decoding, the correlation between the plane structure information of the adjacent nodes is considered, so that the geometric information coding and decoding efficiency of the point cloud can be effectively improved, the prediction coding and decoding performance of the plane structure information is improved, and the coding and decoding efficiency and performance of the point cloud are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Point cloud encoding and decoding method, device, equipment and storage medium Technical Field

[0001] The present application relates to the field of point cloud technology, and in particular to a point cloud encoding and decoding method, apparatus, device and storage medium. Background Art

[0002] Capturing the surface of an object using a capture device creates point cloud data, which can contain hundreds of thousands or even more points. During video production, this point cloud data is transmitted between the point cloud encoding device and the point cloud decoding device in the form of point cloud media files. However, such a large number of points poses a challenge to transmission, so the point cloud encoding device must compress the point cloud data before transmission.

[0003] Point cloud compression, also known as point cloud encoding, can further improve the efficiency of encoding point cloud geometry information for relatively flat or planar nodes by utilizing plane coding. However, current methods only predict the plane structure of the current node based on some prior reference information, resulting in poor performance.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a point cloud encoding and decoding method, apparatus, device and storage medium, which predictively encode and decode the plane structure information of the current node based on the plane structure information of the domain node, thereby improving the predictive encoding and decoding performance of the plane structure information of the current node.

[0006] In a first aspect, an embodiment of the present application provides a point cloud decoding method, comprising:

[0007] Determine N domain nodes of the current node, where N is a positive integer;

[0008] Based on the occupancy information of the N domain nodes, the plane structure information of the current node is predicted and decoded.

[0009] In a second aspect, the present application provides a point cloud encoding method, comprising:

[0010] Determine N domain nodes of the current node, where N is a positive integer;

[0011] Based on the placeholder information of the N domain nodes, the plane structure information of the current node is predictively encoded.

[0012] In a third aspect, the present application provides a point cloud decoding device for executing the method of the first aspect or its respective implementations. Specifically, the device includes a functional unit for executing the method of the first aspect or its respective implementations.

[0013] In a fourth aspect, the present application provides a point cloud encoding device for executing the method of the second aspect or its respective implementations. Specifically, the device includes a functional unit for executing the method of the second aspect or its respective implementations.

[0014] In a fifth aspect, a point cloud decoder is provided, comprising a processor and a memory. The memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the first aspect or its respective implementations.

[0015] In a sixth aspect, a point cloud encoder is provided, comprising a processor and a memory. The memory is configured to store a computer program, and the processor is configured to call and execute the computer program stored in the memory to perform the method of the second aspect or its respective implementations.

[0016] In a seventh aspect, a point cloud encoding and decoding system is provided, comprising a point cloud encoder and a point cloud decoder. The point cloud decoder is configured to execute the method of the first aspect or its respective implementations, and the point cloud encoder is configured to execute the method of the second aspect or its respective implementations.

[0017] In an eighth aspect, a chip is provided for implementing the method described in any one of the first and second aspects above, or their respective implementations. Specifically, the chip includes a processor configured to load and execute a computer program from a memory, causing a device equipped with the chip to perform the method described in any one of the first and second aspects above, or their respective implementations.

[0018] In a ninth aspect, a computer-readable storage medium is provided for storing a computer program, which enables a computer to execute the method of any one of the first to second aspects or their respective implementations.

[0019] In a tenth aspect, a computer program product is provided, comprising computer program instructions, which enable a computer to execute the method of any one of the first to second aspects or their respective implementations.

[0020] In an eleventh aspect, a computer program is provided, which, when executed on a computer, enables the computer to execute the method in any one of the first to second aspects or their respective implementations.

[0021] In a twelfth aspect, a code stream is provided. The code stream is generated based on the method of the second aspect. Optionally, the code stream includes at least one of a first parameter and a second parameter.

[0022] Based on the above technical solution, when encoding and decoding the planar structure information of the current node, the N domain nodes of the current node are determined, and the planar structure information of the current node is encoded and decoded based on the occupancy information of these N domain nodes. In other words, when predicting and decoding the planar structure information of the current node, the correlation between the planar structure information of adjacent nodes is taken into account, which can effectively improve the efficiency of encoding and decoding the geometric information of the point cloud, thereby improving the predictive encoding and decoding performance of the planar structure information, and improving the encoding and decoding efficiency and performance of the point cloud. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1A is a schematic diagram of a point cloud;

[0024] Figure 1B is a partial enlarged view of the point cloud;

[0025] FIG2 is a schematic diagram of six viewing angles of a point cloud image;

[0026] FIG3 is a schematic block diagram of a point cloud encoding and decoding system according to an embodiment of the present application;

[0027] FIG4A is a schematic block diagram of a point cloud encoder provided in an embodiment of the present application;

[0028] FIG4B is a schematic block diagram of a point cloud decoder provided in an embodiment of the present application;

[0029] FIG5A is a schematic plan view;

[0030] FIG5B is a schematic diagram of node coding sequence;

[0031] FIG5C is a schematic diagram of a plane mark;

[0032] Figure 5D is a schematic diagram of sibling nodes;

[0033] Figure 5E is a schematic diagram of the intersection of the laser radar and the node;

[0034] FIG5F is a schematic diagram of neighborhood nodes at the same partition depth and the same coordinates;

[0035] FIG5G is a schematic diagram of neighboring nodes when the node is located at a lower plane position of the parent node;

[0036] FIG5H is a schematic diagram of neighboring nodes when the node is located at a high plane position of the parent node;

[0037] FIG5I is a schematic diagram of predictive coding of planar position information of a laser radar point cloud;

[0038] Figure 6 is a schematic diagram of IDCM encoding;

[0039] 7A to 7C are schematic diagrams of geometric information encoding based on triangular facets;

[0040] FIG8 is a schematic diagram of a point cloud decoding method according to an embodiment of the present application;

[0041] FIG9 is a schematic diagram of an octree partition;

[0042] FIG10 is a schematic diagram of a domain node;

[0043] FIG11 is another schematic diagram of a domain node;

[0044] Figure 12 is a schematic diagram of primary information and secondary information;

[0045] FIG13 is a schematic diagram of a secondary information partitioning tree;

[0046] FIG14 is a schematic diagram of a partitioning tree for secondary information;

[0047] FIG15 is a schematic diagram showing another division of the secondary information partition tree;

[0048] FIG16 is a schematic diagram showing another division of the secondary information division tree;

[0049] FIG17 is a schematic diagram of a point cloud encoding method according to an embodiment of the present application;

[0050] FIG18 is a schematic block diagram of a point cloud decoding device provided in an embodiment of the present application;

[0051] FIG19 is a schematic block diagram of a point cloud encoding device provided in an embodiment of the present application;

[0052] FIG20 is a schematic block diagram of an electronic device provided in an embodiment of the present application;

[0053] Figure 21 is a schematic block diagram of the point cloud encoding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] The present application can be applied to the field of point cloud upsampling technology, for example, it can be applied to the field of point cloud compression technology.

[0055] To facilitate understanding of the embodiments of the present application, the following briefly introduces the relevant concepts involved in the embodiments of the present application:

[0056] A point cloud is a set of irregularly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or scene. Figure 1A is a schematic diagram of a 3D point cloud image, and Figure 1B is a zoomed-in view of Figure 1A. As can be seen from Figures 1A and 1B, the point cloud surface is composed of densely distributed points.

[0057] 2D images contain information at every pixel, and their distribution is regular, so there's no need to record their location. However, the distribution of points in a point cloud in 3D space is random and irregular, so recording the location of every point in space is necessary to fully represent a point cloud. Similar to 2D images, each location in the data collection process has corresponding attribute information.

[0058] Point cloud data is a specific record format for point clouds. Points in a point cloud can include both their location information and attribute information. For example, the location information of a point can be its 3D coordinate information. This information can also be referred to as its geometric information. For example, the attribute information of a point can include color information, reflectance information, normal vector information, and so on. Color information reflects the color of an object, while reflectance information reflects the surface material of the object. The color information can be information in any color space. For example, the color information can be in RGB. Another example is luminance and chrominance (YCbCr, YUV) information. For example, Y represents luminance (Luma), Cb (U) represents blue color difference, Cr (V) represents red, and U and V represent chroma (Chroma) to describe color difference information. For example, a point cloud obtained using laser measurement principles can include both its 3D coordinate information and its laser reflection intensity (reflectance). Another example is a point cloud obtained using photogrammetry principles, which can include both its 3D coordinate information and its color information. For example, a point cloud is obtained by combining the principles of laser measurement and photogrammetry. The points in the point cloud may include the three-dimensional coordinate information of the point, the laser reflection intensity (reflectance) of the point, and the color information of the point. Figure 2 shows a point cloud image, where Figure 2 shows six viewing angles of the point cloud image. Table 1 shows the point cloud data storage format consisting of a file header information part and a data part:

[0059] Table 1

[0060]

[0061] In Table 1, the header information includes the data format, data representation type, the total number of point cloud points, and the content represented by the point cloud. For example, the point cloud in this example is in the ".ply" format, represented by ASCII code, with a total number of 207242 points. Each point has three-dimensional position information XYZ and three-dimensional color information RGB.

[0062] Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes. Moreover, since point clouds are obtained by directly sampling real objects, they can provide a strong sense of reality while ensuring accuracy. Therefore, they are widely used, including virtual reality games, computer-aided design, geographic information systems, automatic navigation systems, digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive remote presentation, and three-dimensional reconstruction of biological tissues and organs.

[0063] The ways to obtain point cloud data may include but are not limited to at least one of the following: (1) generation by computer equipment. Computer equipment can generate point cloud data based on virtual three-dimensional objects and virtual three-dimensional scenes. (2) 3D (3-Dimension) laser scanning acquisition. 3D laser scanning can obtain point cloud data of static real-world three-dimensional objects or three-dimensional scenes, and millions of point cloud data can be obtained per second; (3) 3D photogrammetry acquisition. 3D photography equipment (i.e., a group of cameras or camera equipment with multiple lenses and sensors) is used to collect real-world visual scenes to obtain point cloud data of real-world visual scenes. 3D photography can obtain point cloud data of dynamic real-world three-dimensional objects or three-dimensional scenes. (4) Point cloud data of biological tissues and organs can be obtained through medical equipment. In the medical field, point cloud data of biological tissues and organs can be obtained through medical equipment such as magnetic resonance imaging (MRI), computed tomography (CT), and electromagnetic positioning information.

[0064] Point clouds can be divided into dense point clouds and sparse point clouds according to the acquisition method.

[0065] Point clouds are divided into the following types according to the time series of the data:

[0066] The first type of static point cloud: the object is stationary and the device used to obtain the point cloud is also stationary;

[0067] The second type of dynamic point cloud: the object is moving, but the device that obtains the point cloud is stationary;

[0068] The third type of dynamic point cloud acquisition: the device that acquires the point cloud is moving.

[0069] Point clouds are divided into two categories according to their uses:

[0070] Category 1: Machine perception point cloud, which can be used in scenarios such as autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, and disaster relief robots;

[0071] Category 2: Human eye perception point cloud, which can be used in point cloud application scenarios such as digital cultural heritage, free viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.

[0072] The aforementioned point cloud acquisition technologies reduce the cost and time required to acquire point cloud data, while improving data accuracy. This evolution in point cloud data acquisition has made it possible to acquire large amounts of point cloud data. However, as application demands grow, the processing of massive amounts of 3D point cloud data is facing bottlenecks due to storage space and transmission bandwidth limitations.

[0073] Taking a point cloud video with a frame rate of 30 fps (frames per second) as an example, each frame contains 700,000 points, each with coordinate information (xyz, float) and color information (RGB, uchar). Therefore, the data volume of a 10-second point cloud video is approximately 0.7 million points x (4 bytes x 3 + 1 byte x 3) x 30 fps x 10 seconds = 3.15 GB. For a 1280 x 720 two-dimensional video with a YUV sampling format of 4:2:0 and a frame rate of 24 fps, the data volume for 10 seconds is approximately 1280 x 720 x 12 bits x 24 frames x 10 seconds, which is approximately 0.33 GB. A 10-second two-view 3D video has a data volume of approximately 0.33 x 2 = 0.66 GB. Therefore, the data volume of a point cloud video far exceeds that of a 2D or 3D video of the same length. Therefore, point cloud compression has become a key issue in promoting the development of the point cloud industry to better manage data, save server storage space, and reduce the transmission traffic and time between the server and client.

[0074] The following introduces the relevant knowledge of point cloud encoding and decoding.

[0075] Figure 3 is a schematic block diagram of a point cloud encoding and decoding system involved in an embodiment of the present application. It should be noted that Figure 3 is only an example, and the point cloud encoding and decoding system of the embodiment of the present application includes but is not limited to that shown in Figure 3. As shown in Figure 3, the point cloud encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compression) the point cloud data to generate a code stream, and transmit the code stream to the decoding device. The decoding device decodes the code stream generated by the encoding device to obtain decoded point cloud data.

[0076] The encoding device 110 of the embodiment of the present application can be understood as a device with a point cloud encoding function, and the decoding device 120 can be understood as a device with a point cloud decoding function, that is, the embodiment of the present application includes a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, point cloud game consoles, vehicle-mounted computers, etc.

[0077] In some embodiments, the encoding device 110 may transmit the encoded point cloud data (such as a code stream) to the decoding device 120 via the channel 130. The channel 130 may include one or more media and / or devices capable of transmitting the encoded point cloud data from the encoding device 110 to the decoding device 120.

[0078] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded point cloud data directly to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded point cloud data according to a communication standard and transmit the modulated point cloud data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media can also include wired communication media, such as one or more physical transmission lines.

[0079] In another example, channel 130 includes a storage medium that can store the point cloud data encoded by encoding device 110. The storage medium includes various locally accessible data storage media, such as optical disks, DVDs, and flash memory. In this example, decoding device 120 can retrieve the encoded point cloud data from the storage medium.

[0080] In another example, the channel 130 may include a storage server that can store the point cloud data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded point cloud data from the storage server. Alternatively, the storage server can store the encoded point cloud data and transmit the encoded point cloud data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.

[0081] In some embodiments, the encoding device 110 includes a point cloud encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.

[0082] In some embodiments, the encoding device 110 may further include a point cloud source 111 in addition to the point cloud encoder 112 and the input interface 113 .

[0083] The point cloud source 111 may include at least one of a point cloud acquisition device (e.g., a scanner), a point cloud archive, a point cloud input interface, and a computer graphics system, wherein the point cloud input interface is used to receive point cloud data from a point cloud content provider, and the computer graphics system is used to generate point cloud data.

[0084] The point cloud encoder 112 encodes the point cloud data from the point cloud source 111 to generate a code stream. The point cloud encoder 112 transmits the encoded point cloud data directly to the decoding device 120 via the output interface 113. The encoded point cloud data can also be stored on a storage medium or storage server for subsequent reading by the decoding device 120.

[0085] In some embodiments, the decoding device 120 includes an input interface 121 and a point cloud decoder 122 .

[0086] In some embodiments, the decoding device 120 may further include a display device 123 in addition to the input interface 121 and the point cloud decoder 122 .

[0087] The input interface 121 includes a receiver and / or a modem and can receive the encoded point cloud data via the channel 130 .

[0088] The point cloud decoder 122 is used to decode the encoded point cloud data to obtain decoded point cloud data, and transmit the decoded point cloud data to the display device 123.

[0089] The decoded point cloud data is displayed on the display device 123. The display device 123 may be integrated with the decoding device 120 or external to the decoding device 120. The display device 123 may include various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0090] In addition, Figure 3 is only an example, and the technical solution of the embodiment of the present application is not limited to Figure 3. For example, the technology of the present application can also be applied to unilateral point cloud encoding or unilateral point cloud decoding.

[0091] Current point cloud encoders can use two point cloud compression coding technology routes proposed by the Moving Picture Experts Group (MPEG) of the International Organization for Standardization: Video-based Point Cloud Compression (VPCC) and Geometry-based Point Cloud Compression (GPCC). VPCC projects a 3D point cloud onto a 2D image and uses existing 2D coding tools to encode the projected 2D image. GPCC uses a hierarchical structure to divide the point cloud into multiple units, encoding the entire point cloud by recording the division process.

[0092] The following uses the GPCC encoding and decoding framework as an example to illustrate the point cloud encoder and point cloud decoder applicable to the embodiments of the present application.

[0093] Figure 4A is a schematic block diagram of the point cloud encoder provided in an embodiment of the present application.

[0094] As can be seen from the above, points in a point cloud can include both their location information and their attribute information. Therefore, the encoding of points in a point cloud mainly includes location encoding and attribute encoding. In some examples, the location information of points in a point cloud is also called geometric information, and the corresponding location encoding of points in the point cloud can also be called geometric encoding.

[0095] In the GPCC coding framework, the geometric information of the point cloud and the corresponding attribute information are encoded separately.

[0096] As shown in Figure 4A below, the current geometric coding and decoding of G-PCC can be divided into octree-based geometric coding and decoding and prediction tree-based geometric coding and decoding.

[0097] The position encoding process involves preprocessing the points in the point cloud, such as coordinate transformation, quantization, and duplicate point removal. Next, geometric encoding is performed on the preprocessed point cloud, such as constructing an octree or prediction tree. Based on the constructed octree or prediction tree, geometric encoding is performed to form a geometric bitstream. Simultaneously, the position information of each point in the point cloud data is reconstructed based on the position information output by the constructed octree or prediction tree, resulting in a reconstructed value for each point's position information.

[0098] The attribute encoding process includes: given the reconstruction information of the input point cloud position information and the original value of the attribute information, selecting one of the three prediction modes for point cloud prediction, quantizing the predicted result, and performing arithmetic coding to form an attribute code stream.

[0099] As shown in Figure 4A, position encoding can be achieved through the following units:

[0100] Coordinate conversion (Tanmsform coordinates) unit 201, voxel (Voxelize) unit 202, octree partition (Analyze octree) unit 203, geometry reconstruction (Reconstruct geometry) unit 204, arithmetic encoding (Arithmetic enconde) unit 205, surface fitting unit (Analyze surface approximation) 206 and prediction tree construction unit 207.

[0101] The coordinate conversion unit 201 can be used to convert the world coordinates of a point in the point cloud into relative coordinates. For example, the geometric coordinates of the point are subtracted from the minimum value of the x, y, and z coordinate axes, which is equivalent to a DC removal operation, to convert the coordinates of the point in the point cloud from world coordinates to relative coordinates.

[0102] Voxelize unit 202, also known as the quantize and remove points unit, reduces the number of coordinates through quantization. After quantization, previously different points may be assigned the same coordinates. Based on this, duplicate points can be removed through deduplication. For example, multiple clouds with the same quantized position but different attribute information can be merged into a single cloud through attribute conversion. In some embodiments of the present application, voxel unit 202 is an optional unit module.

[0103] The octree partitioning unit 203 may encode the quantized point position information using an octree encoding scheme. For example, the point cloud may be partitioned using an octree, so that point positions correspond one-to-one with octree positions. Geometric encoding is performed by counting the point positions in the octree and setting their flags to 1.

[0104] In some embodiments, in the geometric information encoding process based on a triangle soup (trisoup), the point cloud is also octree-partitioned by the octree partitioning unit 203. However, unlike the geometric information encoding based on the octree, the trisoup does not need to divide the point cloud into unit cubes with a side length of 1X1X1 step by step. Instead, the division is stopped when the block (sub-block) has a side length of W. Based on the surface formed by the distribution of the point cloud in each block, at most twelve vertices (intersections) generated by the surface and the twelve edges of the block are obtained. The intersections are surface fitted by the surface fitting unit 206, and the fitted intersections are geometrically encoded.

[0105] The prediction tree construction unit 207 can encode the quantized point position information using a prediction tree encoding method. For example, the point cloud is divided into a prediction tree, so that the point positions correspond one-to-one with the positions of the nodes in the prediction tree. By counting the positions of the points in the prediction tree, different prediction modes are selected to predict the geometric position information of the nodes to obtain prediction residuals, and the geometric prediction residuals are quantized using quantization parameters. Finally, through continuous iteration, the prediction residuals of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary bitstream.

[0106] The geometric reconstruction unit 204 can perform position reconstruction based on the position information output by the octree partitioning unit 203 or the intersection points fitted by the surface fitting unit 206 to obtain a reconstructed value of the position information of each point in the point cloud data. Alternatively, the geometric reconstruction unit 204 can perform position reconstruction based on the position information output by the prediction tree construction unit 207 to obtain a reconstructed value of the position information of each point in the point cloud data.

[0107] The arithmetic coding unit 205 may perform entropy coding on the position information output by the octree analysis unit 203 or the intersection points fitted by the surface fitting unit 206, or the geometric prediction residual values ​​output by the prediction tree construction unit 207 to generate a geometric code stream; the geometric code stream may also be referred to as a geometry bitstream.

[0108] Attribute encoding can be achieved through the following units:

[0109] A color conversion unit 210 , a transfer attributes unit 211 , a region adaptive hierarchical transform (RAHT) unit 212 , a generate LOD unit 213 , a lifting transform unit 214 , a quantize coefficients unit 215 , and an arithmetic coding unit 216 .

[0110] It should be noted that the point cloud encoder 200 may include more, fewer, or different functional components than those shown in FIG. 4A .

[0111] The color conversion unit 210 may be configured to convert the RGB color space of a point in the point cloud into a YCbCr format or other formats.

[0112] The recoloring unit 211 recolors the color information using the reconstructed geometric information so that the uncoded attribute information corresponds to the reconstructed geometric information.

[0113] After the original value of the point attribute information is converted by the recoloring unit 211, any transformation unit can be selected to transform the points in the point cloud. The transformation units may include: RAHT transformation 212 and lifting transformation unit 214. The lifting transformation relies on generating the level of detail (LOD).

[0114] Either the RAHT transform or the lifting transform can be understood as being used to predict the attribute information of a point in a point cloud to obtain a predicted value of the attribute information of the point, and then to obtain a residual value of the attribute information of the point based on the predicted value of the attribute information of the point. For example, the residual value of the attribute information of the point can be the original value of the attribute information of the point minus the predicted value of the attribute information of the point.

[0115] In one embodiment of the present application, the process of generating LOD by the LOD generation unit includes: obtaining the Euclidean distance between points based on the position information of the points in the point cloud; and dividing the points into different detail expression layers based on the Euclidean distance. In one embodiment, the Euclidean distances can be sorted and then Euclidean distances in different ranges can be divided into different detail expression layers. For example, a point can be randomly selected as the first detail expression layer. The Euclidean distances between the remaining points and the point are then calculated, and the points whose Euclidean distances meet the first threshold requirement are classified as the second detail expression layer. The centroid of the points in the second detail expression layer is obtained, and the Euclidean distances between the points other than the first and second detail expression layers and the centroid are calculated, and the points whose Euclidean distances meet the second threshold requirement are classified as the third detail expression layer. And so on, all points are classified into the detail expression layer. By adjusting the threshold of the Euclidean distance, the number of points in each LOD layer can be increased. It should be understood that the LOD division method can also be adopted in other ways, and this application is not limited to this.

[0116] It should be noted that the point cloud can be directly divided into one or more detail expression layers, or the point cloud can be first divided into multiple point cloud slices, and then each point cloud slice can be divided into one or more LOD layers.

[0117] For example, a point cloud can be divided into multiple point cloud tiles, each containing between 550,000 and 1.1 million points. Each point cloud tile can be considered a separate point cloud. Each point cloud tile can be further divided into multiple detail expression layers, each containing multiple points. In one embodiment, the detail expression layers can be divided based on the Euclidean distance between points.

[0118] The quantization unit 215 may be used to quantize the residual value of the attribute information of the point. For example, if the quantization unit 215 is connected to the RAHT transformation unit 212, the quantization unit 215 may be used to quantize the residual value of the attribute information of the point output by the RAHT transformation unit 212.

[0119] The arithmetic coding unit 216 may perform entropy coding on the residual value of the attribute information of the point using zero run length coding to obtain an attribute code stream. The attribute code stream may be bit stream information.

[0120] Figure 4B is a schematic block diagram of the point cloud decoder provided in an embodiment of the present application.

[0121] As shown in Figure 4B, the decoder 300 can obtain the point cloud code stream from the encoding device and obtain the position information and attribute information of the points in the point cloud by parsing the code. The decoding of the point cloud includes position decoding and attribute decoding.

[0122] The position decoding process includes: performing arithmetic decoding on the geometric code stream; constructing an octree and then merging it to reconstruct the point position information to obtain the reconstructed position information of the point; and performing coordinate transformation on the reconstructed position information of the point to obtain the point position information. The point position information can also be called the point's geometric information.

[0123] The attribute decoding process includes: obtaining the residual value of the attribute information of the point in the point cloud by parsing the attribute code stream; obtaining the residual value of the attribute information of the point after dequantization by dequantizing the residual value of the attribute information of the point; based on the reconstruction information of the point position information obtained in the position decoding process, selecting one of the following RAHT inverse transform and lifting inverse transform to perform point cloud prediction to obtain the predicted value, and adding the predicted value to the residual value to obtain the reconstructed value of the attribute information of the point; performing inverse color space conversion on the reconstructed value of the attribute information of the point to obtain the decoded point cloud.

[0124] As shown in Figure 4B, position decoding can be achieved by the following units:

[0125] Arithmetic decoding unit 301, octree reconstruction unit 302, surface reconstruction unit 303, geometry reconstruction unit 304, inverse transform coordinates unit 305 and prediction tree reconstruction unit 306.

[0126] Attribute encoding can be achieved through the following units:

[0127] an arithmetic decoding unit 310 , an inverse quantization unit 311 , an inverse RAHT transform unit 312 , a LOD generation unit 313 , an inverse lifting transform unit 314 , and an inverse color transform unit 315 .

[0128] It should be noted that decompression is the inverse process of compression. Similarly, the functions of each unit in the decoder 300 can refer to the functions of the corresponding units in the encoder 200. In addition, the point cloud decoder 300 may include more, fewer, or different functional components than those in Figure 4B.

[0129] For example, the decoder 300 can divide the point cloud into multiple LODs based on the Euclidean distance between points in the point cloud. The decoder 300 then decodes the attribute information of the points in the LODs in sequence. For example, the number of zeros (zero_cnt) in the zero-run encoding technique is calculated to decode the residual based on zero_cnt. The decoding framework 200 then dequantizes the decoded residual value and adds the dequantized residual value to the predicted value of the current point to obtain the reconstructed value of the point cloud until all point clouds are decoded. The current point will be used as the nearest neighbor of the subsequent LOD point, and the reconstructed value of the current point will be used to predict the attribute information of the subsequent point.

[0130] The above is the basic process of the point cloud codec based on the GPCC codec framework. With the development of technology, some modules or steps of the framework or process may be optimized. This application is applicable to the basic process of the point cloud codec based on the GPCC codec framework, but is not limited to this framework and process.

[0131] The following introduces octree-based geometric coding and prediction tree-based geometric coding.

[0132] The geometric encoding based on octree includes: first, coordinate transformation of the geometric information so that all point clouds are contained in a bounding box. Then quantization is performed. This step of quantization mainly plays a role of scaling. Due to the quantization rounding, the geometric information of some points is the same. The parameters are used to decide whether to remove duplicate points. The process of quantization and removal of duplicate points is also called voxelization. Next, the bounding box is continuously divided into trees (octree / quadtree / binary tree) in the order of breadth-first traversal, and the placeholder code of each node is encoded. In an implicit geometric division method, the bounding box of the point cloud is first calculated. Assume that the d x >d y >d z The bounding box corresponds to a cuboid. When geometrically partitioning, the binary tree partitioning is first performed based on the X coordinate axis to obtain two child nodes; until d is satisfied x =d y >d z When the conditions are met, the quadtree partitioning will be performed based on the x and y coordinate axes to obtain four child nodes; when d is finally satisfied x =d y =d zWhen the conditions are met, the octree partitioning will continue until the leaf node obtained by the partitioning is a 1x1x1 unit cube. The partitioning will stop and the points in the leaf node will be encoded to generate a binary code stream. In the process of binary tree / quadtree / octree partitioning, two parameters are introduced: K and M. Parameter K indicates the maximum number of binary tree / quadtree partitions before octree partitioning; parameter M is used to indicate that the minimum block side length corresponding to binary tree / quadtree partitioning is 2 M . At the same time, K and M must meet the following conditions: Assume d max =max(d x ,d y ,d z ),d min =min(d x ,d y ,d z ), parameter K satisfies: K>=d max -d min ; Parameter M satisfies: M>=d min The parameters K and M meet the above conditions because the priority of the partitioning method in the current G-PCC implicit geometric partitioning process is binary tree, quadtree and octree. When the node block size does not meet the conditions of binary tree / quadtree, the node will be partitioned into octree until the minimum unit of leaf node 1X1X1 is reached.

[0133] The octree-based geometric information coding mode can effectively encode the geometric information of the point cloud by utilizing the correlation between adjacent points in space. However, for some relatively flat nodes or nodes with planar characteristics, the coding efficiency of the point cloud geometric information can be further improved by using plane coding.

[0134] For example, as shown in Figure 5A, the (a) series belongs to the low plane position in the Z coordinate axis direction, and the (b) series belongs to the high plane position in the Z coordinate axis direction. Taking (a) as an example, it can be seen that the four occupied child nodes of the current node are all located in the low plane position of the current node in the Z coordinate axis direction. Therefore, it can be considered that the current node belongs to a Z plane and is a low plane in the Z coordinate axis direction. Similarly, (b) shows that the occupied child nodes of the current node are located in the high plane position of the current node in the Z coordinate axis direction.

[0135] Taking (a) as an example, the efficiency of octree coding and plane coding is compared. As shown in Figure 5B, if the octree coding method is used for (a) in Figure 1, the placeholder information of the current node is represented as: 11001100. However, if the plane coding method is used, first, an identifier needs to be encoded to indicate that the current node is a plane in the Z coordinate axis direction. Secondly, if the current node is a plane in the Z coordinate axis direction, the plane position of the current node needs to be represented. Secondly, only the placeholder information of the low plane node in the Z coordinate axis direction needs to be encoded (that is, the placeholder information of the four child nodes 0246). Therefore, encoding the current node based on the plane coding method only requires encoding 6 bits, which can reduce the representation of 2 bits compared to the original octree coding. Based on this analysis, plane coding has more obvious coding efficiency than octree coding. Therefore, for an occupied node, if the plane coding method is used in a certain dimension, as shown in Figure 5C, first, the plane identification (planarMode) and plane position (PlanePos) information of the current node in the dimension need to be represented, and then the occupancy information of the current node is encoded based on the plane information of the current node. It should be noted that: PlaneMode i (i=0,1,2): 0 means the current node is not a plane in the direction of i axis. When the node is a plane in the direction of i axis, PlanePosition i :0 means the current node is a plane in the i-axis direction and the plane position is a low plane, 1 means the current node is a high plane in the i-axis direction. For example, i=0 represents the X-axis, i=1 represents the Y-axis, and i=2 represents the Z-axis.

[0136] The following details how to determine whether a node meets the plane coding conditions in the current G-PCC standard and predictively encode the node plane identifier and plane position information when the node meets the plane coding conditions.

[0137] Currently, there are three types of conditions in G-PCC to determine whether a node meets the conditions for plane coding. The following describes them one by one:

[0138] The first method is to judge based on the plane probability of the node in each dimension.

[0139] First, determine the local area density (local_node_density) of the current node and the probability Prob(i) of the current node in each dimension.

[0140] When the local area density of a node is less than the threshold Th (Th = 3), the plane probability Prob(i) of the current node in three dimensions is compared with the thresholds Th0, Th1, and Th2, where Th0 < Th1 < Th2 (Th0 = 0.6, Th1 = 0.77, Th2 = 0.88). Next, Eligible i (i = 0, , 2) represents whether plane coding is started in each dimension, where Eligible i The judgment process is shown in formula (1). For example, if Eligible i >= threshold, it means that plane coding is started in the i-th dimension:

[0141] Eligible i = Prob(i) >= threshold (1)

[0142] It should be noted that threshold changes adaptively. For example, when Prob(0) > Prob(1) > Prob(2), the value of threshold is shown in formula (2):

[0143] Eligible0 = Prob(0) >= Th0

[0144] Eligible1 = Prob(1) >= Th1

[0145] Eligible2 = Prob(2) >= Th2 (2)

[0146] Next, the update process of local_node_density and the update of Prob(i) are introduced.

[0147] In one example, Prob(i) is updated by the following formula (3):

[0148] Prob(i) new = (Lx Prob(i) + δ(coded node)) / L + 1 (3)

[0149] where L = 255. When the coded node is a plane, it is 1; otherwise, it is 0.

[0150] In one example, local_node_density is updated by the following formula (4):

[0151] local_node_density new = local_node_density + 4 * numSiblings (4)

[0152] Among them, local_node_density is initialized to 4, numSiblings is the number of sibling nodes of the node, as shown in Figure 5D, the current node is the left node, the right node is the sibling node of the current node, and the number of sibling nodes of the current node is 5 (including itself).

[0153] The second method is to determine whether the nodes in the current layer meet the requirements of plane coding based on the point cloud density of the current layer.

[0154] The density of the points in the current layer is used to determine whether to perform plane coding on the nodes in the current layer. Assuming that the number of points in the current point cloud to be coded is pointCount, the number of points reconstructed after IDCM coding is numPointCountRecon, and because the octree is coded in the order of breadth-first traversal, the number of nodes to be coded in the current layer can be obtained as nodeCount. It is assumed that planarEligibleKOctreeDepth is used to indicate whether the current layer starts plane coding. The judgment process of planarEligibleKOctreeDepth is shown in formula (5):

[0155] planarEligibleKOctreeDepth=(pointCount-numPointCountRecon) <nodeCount*1.3 (5)

[0156] When planarEligibleKOctreeDepth is true, all nodes in the current layer are plane coded; otherwise, no plane coding is performed and only octree coding is used.

[0157] The third method is to determine whether the current node meets the requirements of plane coding based on the acquisition parameters of the lidar point cloud.

[0158] As shown in Figure 5E, the large cube node at the top is simultaneously traversed by two lasers, so the current node is not a plane in the vertical direction of the Z coordinate axis. The small cube node at the bottom is small enough that it cannot be traversed by both nodes simultaneously, so it is likely a plane. Therefore, based on the number of lasers corresponding to the current node, we can determine whether the current node meets the requirements of plane coding.

[0159] The following describes the predictive coding of plane identification information and plane position information for nodes that currently meet the plane coding conditions.

[0160] 1. Predictive Coding of Plane Marking Information

[0161] Currently, three contexts are used to encode plane identification information, that is, the plane representation in each dimension is designed separately.

[0162] The following introduces the encoding of planar position information of non-lidar point clouds and lidar point clouds respectively.

[0163] 1) Encoding of non-lidar point cloud planar position information

[0164] 1. Predictive coding of planar position information.

[0165] The plane position information is predictively coded based on the following information:

[0166] (1) Using the occupancy information of neighboring nodes, the plane position information of the current node is predicted to be three elements: predicted as low plane, predicted as high plane, and unpredictable;

[0167] (2) The spatial distance between the nodes at the same partition depth and the same coordinates as the current node and the current node is “close” or “far”;

[0168] (3) The plane position of the node at the same partition depth and the same coordinates as the current node;

[0169] (4) Coordinate dimension (i=0, 1, 2).

[0170] As shown in Figure 5F, the current node to be encoded is the left node, then the neighboring node is searched for as the right node at the same octree partition depth level and the same vertical coordinate, and the distance between the two nodes is judged as "near" and "far", and the plane position of the reference node is used.

[0171] In one example, as shown in FIG5G , the black node is the current node. If the current node is located on the lower plane of the parent node, the plane position of the current node is determined as follows:

[0172] a) If any of the child nodes 4 to 7 of the dashed node is occupied, and all the dot nodes are unoccupied, it is very likely that there is a plane in the current node, and the plane is at a lower position.

[0173] b) If the child nodes 4 to 7 of the dashed node are not occupied, and any dotted node is occupied, it is very likely that there is a plane in the current node, and the plane is at a higher position.

[0174] c) If the child nodes 4 to 7 of the dashed node are all empty nodes and the dotted nodes are all empty nodes, the plane position cannot be inferred and is therefore marked as unknown.

[0175] If any of the child nodes 4 to 7 of the dashed node are occupied and any of the dotted nodes are occupied, the plane position cannot be inferred and is therefore marked as unknown.

[0176] In another example, as shown in FIG5H , the black node is the current node. If the node is at a high plane position of the parent node, the plane position of the current node is determined as follows:

[0177] a) If any of the dot node's child nodes 4 to 7 is occupied, and the dashed node is not occupied, it is very likely that there is a plane in the current node, and the plane is at a lower position.

[0178] b) If the child nodes 4 to 7 of the dot node are not occupied, but the node with the dashed line is occupied, it is very likely that a plane exists in the current node, and the plane is located at a higher position.

[0179] c) If the child nodes 4 to 7 of the dot node are all unoccupied, and the dashed node is unoccupied, the plane position cannot be inferred and is therefore marked as unknown.

[0180] d) If one of the child nodes 4-7 of the dotted node is occupied and the dashed node is occupied, the plane position cannot be inferred and is therefore marked as unknown.

[0181] 2) Coding of LiDAR point cloud plane position information

[0182] Figure 5I shows the predictive coding of the plane position information of the laser radar point cloud. The plane position of the current node is predicted by using the laser radar acquisition parameters. The position is quantized into four intervals by using the intersection position of the current node and the laser ray, and finally used as the context of the plane position of the current node. The specific calculation process is as follows: Assume that the coordinates of the laser radar are (x Lidar ,y Lidar ,z Lidar ), the geometric coordinates of the current point are (x, y, z), then first calculate the vertical tangent value tanθ of the current point relative to the lidar. The calculation process is shown in formula (6):

[0183]

[0184] Because each laser has a certain offset angle relative to the laser radar, the relative tangent value tanθ of the current node relative to the laser is calculated. corr,L , the specific calculation process is shown in formula (7):

[0185]

[0186] Finally, the corrected tangent value of the current node is used to predict the plane position of the current node. Specifically, assuming that the tangent value of the lower boundary of the current node is tan(θ bottom), and the tangent value of the upper boundary is tan(θ top), according to tanθ corr,L The plane position is quantized into 4 quantization intervals, which are the contexts of the plane position.

[0187] However, the octree-based geometric information coding mode only has an efficient compression rate for points with correlation in space. For points that are isolated in the geometric space, the use of the Direct Coding Model (DCM) can greatly reduce the complexity. For all nodes in the octree, the use of DCM is not indicated by flag information, but is inferred from the parent node and neighbor information of the current node. There are three ways to determine whether the current node is eligible for DCM coding, as shown in Figure 6:

[0188] (1) The current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.

[0189] (2) The parent node of the current node has only one child node, the current node, and the six neighbor nodes that share a face with the current node are all empty nodes.

[0190] (3) The number of sibling nodes of the current node is greater than 1.

[0191] If the current node does not meet the DCM coding qualifications, it will be divided into octrees. If it meets the DCM coding qualifications, the number of points contained in the node will be further determined. When the number of points is less than the threshold 2, the node will be DCM-encoded, otherwise the octree division will continue. When the DCM coding mode is applied, it is first necessary to encode whether the current node is a true isolated point, that is, IDCM_flag. When IDCM_flag is true, the current node uses DCM coding, otherwise octree coding is still used. When the current node meets the DCM coding requirements, the DCM coding mode of the current node needs to be encoded. There are currently two DCM modes: 1: There is only one point (or multiple points, but they are duplicate points); 2: Contains two points. Finally, the geometric information of each point needs to be encoded. Assume that the side length of the node is 2 d When encoding each component of the node's geometric coordinates, d bits are required, and these bits are directly encoded into the bitstream. It is important to note that when encoding LiDAR point clouds, the efficiency of geometric information coding can be further improved by predictively encoding the three-dimensional coordinate information using LiDAR acquisition parameters.

[0192] It is important to note that when partitioning nodes into leaf nodes, the number of duplicate points in the leaf nodes must be encoded in the case of lossless geometric coding. Ultimately, the placeholder information for all nodes is encoded to generate a binary bitstream. Furthermore, G-PCC currently introduces a plane coding mode. During the geometric partitioning process, it determines whether the child nodes of the current node are in the same plane. If the child nodes of the current node meet the conditions for being in the same plane, the child nodes of the current node are represented using that plane.

[0193] In octree-based geometric decoding, the decoder follows a breadth-first traversal. Before decoding each node's occupancy information, it first uses the reconstructed geometric information to determine whether the current node is for plane decoding or IDCM decoding. If the current node meets the requirements for plane decoding, it first decodes the plane identifier and plane position information of the current node. Then, based on the plane information, it decodes the current node's occupancy information. If the current node meets the requirements for IDCM decoding, it first decodes whether the current node is a true IDCM node. If so, it continues to parse the DCM decoding mode of the current node, then obtains the number of points in the current DCM node, and finally decodes the geometric information of each point. For nodes that do not meet either plane decoding or DCM decoding requirements, the current node's occupancy information is decoded. By continuously parsing in this way, the placeholder code of each node is obtained, and the node is continuously divided until a 1x1x1 unit cube is obtained. The number of points contained in each leaf node is parsed, and the geometrically reconstructed point cloud information is finally recovered.

[0194] In octree-based geometric decoding, the decoder follows a breadth-first traversal. Before decoding each node's occupancy information, it first uses the reconstructed geometric information to determine whether the current node is for plane decoding or IDCM decoding. If the current node meets the requirements for plane decoding, it first decodes the plane identifier and plane position information of the current node. Then, based on the plane information, it decodes the current node's occupancy information. If the current node meets the requirements for IDCM decoding, it first decodes whether the current node is a true IDCM node. If so, it continues to parse the DCM decoding mode of the current node, then obtains the number of points in the current DCM node, and finally decodes the geometric information of each point. For nodes that do not meet either plane decoding or DCM decoding requirements, the current node's occupancy information is decoded. By continuously parsing in this way, the placeholder code of each node is obtained, and the node is continuously partitioned until a 1x1x1 unit cube is obtained. The number of points contained in each leaf node is parsed, and the geometrically reconstructed point cloud information is finally recovered.

[0195] In the trisoup (triangle soup)-based geometric information coding framework, geometric partitioning is also performed first. However, unlike geometric information coding based on binary trees, quad trees, and octrees, this method does not need to gradually partition the point cloud into unit cubes with side lengths of 1x1x1. Instead, the partitioning stops when the block (sub-block) has a side length of W. Based on the surface formed by the distribution of the point cloud in each block, the surface and the twelve edges of the block are obtained. The vertex coordinates of each block are encoded in sequence to generate a binary code stream.

[0196] When reconstructing point cloud geometry based on trisoup, the decoding end first decodes vertex coordinates to complete triangle reconstruction. This process is shown in Figures 7A to 7C. The block shown in Figure 7A contains three vertices (v1, v2, v3). The set of triangles formed by these three vertices in a certain order is called triangle soup, or trisoup, as shown in Figure 7B. Afterwards, sampling is performed on this set of triangles, and the resulting sampling points are used as the reconstructed point cloud within the block, as shown in Figure 7C.

[0197] The geometric coding based on the prediction tree includes: first, sorting the input point cloud. The currently used sorting methods include unordered, Morton order, azimuth order, and radial distance order. At the encoding end, the prediction tree structure is established by using two different methods, including: KD-Tree (high-latency slow mode) and using the lidar calibration information to divide each point into different Lasers and establish a prediction structure according to different Lasers (low-latency fast mode). Next, based on the structure of the prediction tree, each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter. Finally, through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream.

[0198] Based on the geometric decoding of the prediction tree, the decoding end reconstructs the prediction tree structure by continuously parsing the bit stream. Secondly, the geometric position prediction residual information and quantization parameters of each prediction node are obtained through parsing, and the prediction residual is dequantized to restore the reconstructed geometric position information of each node, finally completing the geometric reconstruction at the decoding end.

[0199] After the geometric encoding is completed, the geometric information is reconstructed. At present, attribute encoding is mainly performed on color information. First, the color information is converted from the RGB color space to the YUV color space. Then, the point cloud is recolored using the reconstructed geometric information so that the unencoded attribute information corresponds to the reconstructed geometric information. In color information encoding, there are two main transformation methods. One is the distance-based lifting transformation that relies on LOD (Level of Detail) division, and the other is to directly perform RAHT (Region Adaptive Hierarchal Transform) transformation. Both methods will convert the color information from the spatial domain to the frequency domain, and obtain high-frequency coefficients and low-frequency coefficients through transformation. Finally, the coefficients are quantized and encoded to generate a binary code stream.

[0200] When using geometric information to predict attribute information, Morton codes can be used to perform nearest neighbor search. The Morton code corresponding to each point in the point cloud can be obtained from the geometric coordinates of the point. The specific method for calculating the Morton code is described as follows. For a three-dimensional coordinate represented by a d-bit binary number for each component, its three components can be expressed as formula (8):

[0201]

[0202] Among them, x l ,y l ,z l∈{0,1} are the binary values ​​corresponding to the highest bit (l=1) to the lowest bit (l=d) of x, y, and z respectively. The Morton code M is to cross-arrange x, y, and z starting from the highest bit. l ,y l ,z l To the lowest bit, the calculation formula of M is shown in the following formula (9):

[0203]

[0204] Among them, m l′ ∈{0,1} are the values ​​of the highest bit (l′=1) to the lowest bit (l′=3d) of M. After obtaining the Morton code M of each point in the point cloud, the points in the point cloud are arranged in ascending order of the Morton code, and the weight w of each point is set to 1.

[0205] There are 4 general test conditions for GPCC:

[0206] Condition 1: The geometric position is limited and the attributes are lost;

[0207] Condition 2: Geometric position lossless, attribute lossy;

[0208] Condition 3: Geometric position lossless, attribute loss limited;

[0209] Condition 4: Geometric position and attributes are lossless.

[0210] The general test sequences include Cat1A, Cat1B, Cat3-fused, and Cat3-frame, a total of four categories. Among them, Cat2-frame point cloud only contains reflectance attribute information, Cat1A and Cat1B point clouds only contain color attribute information, and Cat3-fused point cloud contains both color and reflectance attribute information.

[0211] There are two technical routes of GPCC, which are distinguished by the algorithm used for geometric compression, and are divided into octree coding branch and prediction tree coding branch.

[0212] Among them, in the octree coding branch, at the encoding end, the bounding box is divided into sub-cubes in sequence, and the non-empty sub-cubes (containing points in the point cloud) are divided until the leaf node obtained by division is a 1X1X1 unit cube. The division stops when the division is completed. In the case of geometric lossless coding, the number of points contained in the leaf node needs to be encoded, and finally the geometric octree encoding is completed to generate a binary code stream. At the decoding end, the decoding end obtains the placeholder code of each node by continuous parsing in the order of breadth-first traversal, and continuously divides the nodes in sequence until the division is a 1x1x1 unit cube. In the case of geometric lossless decoding, the number of points contained in each leaf node needs to be parsed to finally recover the geometric reconstructed point cloud information.

[0213] In the prediction tree coding branch, the encoder establishes the prediction tree structure using two different approaches: a KD-Tree (high-latency, slow mode) and a low-latency, fast mode, where each point is assigned to a different laser using lidar calibration information and the prediction structure is established accordingly. Next, based on the prediction tree structure, each node in the tree is traversed, and the geometric position information of the node is predicted using different prediction modes to obtain a prediction residual. This geometric prediction residual is then quantized using a quantization parameter. Finally, through continuous iteration, the prediction residuals of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary bitstream. On the decoder side, the decoder continuously parses the bitstream to reconstruct the prediction tree structure. The geometric position prediction residual information and quantization parameters for each prediction node are then parsed and dequantized to recover the reconstructed geometric position information for each node, completing the geometric reconstruction at the decoder.

[0214] During point cloud coding, for relatively flat nodes or nodes with planar characteristics, plane coding can further improve the coding efficiency of point cloud geometric information. However, currently, the plane structure information of the current node is only predictively coded based on some prior reference information, resulting in poor predictive coding performance for plane structure information.

[0215] In order to solve the above technical problems, the embodiment of the present application determines the N domain nodes of the current node when encoding and decoding nodes, and predicts and decodes the plane structure of the current node based on the occupancy information of the N domain nodes, thereby improving the predictive encoding and decoding performance of the plane structure information and improving the encoding and decoding efficiency and performance of the point cloud.

[0216] The following describes the point cloud encoding and decoding method involved in the embodiments of the present application in conjunction with specific embodiments.

[0217] First, taking the decoding end as an example, the point cloud decoding method provided in the embodiment of the present application is introduced.

[0218] FIG8 is a flow chart of a point cloud decoding method according to an embodiment of the present application. The point cloud decoding method according to an embodiment of the present application can be implemented by the point cloud decoding device or point cloud decoder shown in FIG3 or FIG4B above.

[0219] As shown in FIG8 , the point cloud decoding method of the embodiment of the present application includes:

[0220] S101. Determine N domain nodes of the current node.

[0221] As can be seen from the above, a point cloud includes geometric information and attribute information, and decoding of a point cloud includes geometric decoding and attribute decoding. The embodiments of the present application relate to geometric decoding of a point cloud.

[0222] In some embodiments, the geometric information of the point cloud is also referred to as the position information of the point cloud. Therefore, the geometric decoding of the point cloud is also referred to as the position decoding of the point cloud.

[0223] In the octree-based encoding method, the encoding end constructs an octree structure of the point cloud based on the geometric information of the point cloud. As shown in Figure 9, the point cloud is enclosed by a minimum rectangular block. The bounding box is first divided into 8 nodes by the octree to obtain 8 nodes. The occupied nodes among these 8 nodes, that is, the nodes including the points, are further divided into octrees, and so on, until the division is to the voxel level, for example, to a 1X1X1 cube. The point cloud octree structure obtained by such division includes multiple layers of nodes, for example, N layers. During encoding, the occupancy information of each layer is encoded layer by layer until the voxel-level leaf nodes of the last layer are encoded. That is to say, in octree encoding, the point cloud is divided into octrees, and finally the points in the point cloud are divided into the voxel-level leaf nodes of the octree. The encoding of the point cloud is achieved by encoding the entire octree.

[0224] Correspondingly, the decoder first decodes the point cloud geometry stream to obtain the occupancy information of the root node of the point cloud's octree. Based on this occupancy information, it determines the child nodes of the root node, that is, the nodes in the second layer of the octree. Next, it decodes the geometry stream to obtain the occupancy information of each node in the second layer. Based on this occupancy information, it determines the nodes in the third layer of the octree, and so on.

[0225] However, for some relatively flat nodes or nodes with planar characteristics, the coding efficiency of the point cloud geometric information can be further improved by utilizing plane coding. For example, as shown in Figure 5A, the four occupied child nodes in the current node are all located in the low plane position of the current node in the Z coordinate axis direction. At this time, the occupancy information of the current node is represented as: 11001100. When encoding the current node using plane coding, it is first necessary to encode an identifier to indicate that the current node is a plane in the Z coordinate axis direction. Secondly, if the current node is a plane in the Z coordinate axis direction, the plane position of the current node needs to be represented. Secondly, it is only necessary to encode the low plane node placeholder information in the Z coordinate axis direction (i.e., the placeholder information of the four child nodes 0246). Therefore, encoding the current node based on the plane coding method only requires encoding 6 bits, which can reduce the representation of 2 bits compared to the original octree coding, thereby improving the coding performance of the point cloud.

[0226] As can be seen from the above, when plane coding is used for the current node, the encoder needs to predictively encode the plane structure information of the current node. Correspondingly, the decoder predictively decodes the plane structure information of the current node and, based on the decoded plane structure information, obtains the geometric information of the current node.

[0227] Currently, the plane structure information of the current node is predictively encoded based on some prior reference information, such as the spatial distance between the nodes at the same division depth and the same coordinates as the current node, and / or the plane position of the node at the same division depth and the same coordinates as the current node, resulting in poor prediction coding performance of the plane structure information.

[0228] In order to solve the above problems, in an embodiment of the present application, the decoding end predicts and decodes the plane structure of the current node based on the occupancy information of the N domain nodes of the current node, thereby improving the prediction encoding and decoding performance of the plane structure information and improving the encoding and decoding efficiency and performance of the point cloud.

[0229] The following describes the specific process of the decoding end determining the N domain nodes of the current node.

[0230] It should be noted that in the embodiment of the present application, there is no restriction on the specific method for the decoding end to determine the N domain nodes of the current node.

[0231] In one example, the N domain nodes of the current node include at least one domain node that is coplanar, colinear, and co-pointed with the current node. As shown in Figure 10, the current node includes 6 coplanar nodes, 12 colinear nodes, and 8 co-pointed nodes.

[0232] In another example, the N domain nodes of the current node may include, in addition to at least one domain node that is coplanar, colinear, and co-point with the current node, other nodes within a preset reference neighborhood range. This embodiment of the present application does not impose any restrictions on this.

[0233] In a specific embodiment, as shown in Figure 11, the thick dashed line node is the current node to be encoded, the solid line node is the three neighboring nodes coplanar with the current node, the dot-dashed line node is the three neighboring nodes colinear with the current node, and the long dashed line node is the neighboring node co-point with the current node. Because according to the order of point cloud decoding, when decoding the occupancy information of the current node, seven neighboring nodes that are coplanar, co-linear, and co-point with the current node (left front and lower direction) can be obtained. The occupancy information of at least one of these seven domain nodes is used to predict and decode the planar structure information of the current node.

[0234] In another specific embodiment, the N domain nodes of the current node include the 7 domain nodes in Figure 11, namely, 3 domain nodes coplanar with the current node, 3 domain nodes colinear with the current node, and 1 domain node copoint with the current node.

[0235] In another specific embodiment, the N domain nodes of the current node include domain nodes that are coplanar and co-point with the current node, for example, 6 domain nodes that are coplanar with the current node and 1 domain node that is co-point with the current node.

[0236] In another specific embodiment, the N domain nodes of the current node include domain nodes that are collinear and co-pointed with the current node, for example, 6 domain nodes that are collinear with the current node and 1 domain node that is co-pointed with the current node.

[0237] In another specific embodiment, the N domain nodes of the current node include domain nodes that are colinear and coplanar with the current node, for example, 6 domain nodes that are coplanar with the current node and 6 domain nodes that are colinear with the current node.

[0238] In another specific embodiment, the N domain nodes of the current node include only domain nodes that are coplanar with the current node, or include only domain nodes that are colinear with the current node, or include only domain nodes that have a common point with the current node.

[0239] The embodiment of the present application does not limit the specific method by which the decoding end determines the N domain nodes of the current node.

[0240] S102: Based on the occupancy information of N domain nodes, predict and decode the plane structure information of the current node.

[0241] In an embodiment of the present application, the plane structure information of the current node includes the plane identification information of the current node and / or the plane position information of the current node.

[0242] From the above, we can see that the plane identification of the current node is PlaneMode i (i=0,1,2) indicates that i=0 represents the X coordinate axis, i=1 represents the Y coordinate axis, and i=2 represents the Z coordinate axis. PlaneMode i =0 means the current node is not a plane in the direction of the i-th coordinate axis. PlaneMode i =1 means that the current node is a plane in the direction of the i-th coordinate axis.

[0243] If the current node is a plane in the direction of the i-th coordinate axis, that is, PlaneMode i = 1, the decoding end continues to decode the plane position information of the current node on the i-th coordinate axis. For example, using PlanePosition i Indicates the plane position information of the current node in the direction of the i-th coordinate axis, such as PlanePosition i =0 means that the current node is a plane in the direction of the i-th coordinate axis, and the plane position is the low plane, PlanePosition i =1 means that the current node is a high plane in the direction of the i-th coordinate axis.

[0244] In the embodiment of the present application, the plane structure information of the current node is predicted and decoded based on the occupancy information of N domain nodes, that is, the plane identification and / or plane position information of the current node is predicted and decoded.

[0245] For example, based on the occupancy information of the N domain nodes of the current node, the plane identifier of the current node on the i-th coordinate axis is predicted and decoded.

[0246] For another example, based on the occupancy information of the N area nodes of the current node, the plane position information of the current node on the i-th coordinate axis is predicted and decoded.

[0247] In an embodiment of the present application, the plane structure information of the current node is predicted and decoded based on the placeholder information of the N domain nodes of the current node. This can be understood as using the placeholder information of the N domain nodes of the current node as the context information of the plane structure information of the current node to predict and decode the plane structure information of the current node. For example, based on the N domain nodes of the current node, a context model index is determined, based on the context model index, a context model is determined, and based on the context model, the plane structure information of the current node is predicted and decoded, for example, the plane identifier of the current node is predicted and decoded based on the context model, or the plane position information of the current node is predicted and decoded based on the context model.

[0248] In some embodiments, if the plane structure information of the current node includes the plane position information of the current node, the decoding end predicts and decodes the plane position information of the current node based on the placeholder information of the N areas of the current node. In this case, the above S102 includes steps S102-A and S102-B:

[0249] S102-A, based on the occupancy information of the N domain nodes, determine the plane structure information of the N domain nodes;

[0250] S102-B. Based on the plane structure information of the N domain nodes, predict and decode the plane position information of the current node.

[0251] In this embodiment, when predicting and decoding the plane position information of the current node using the placeholder information of the N domain nodes of the current node, the decoding end first determines the plane structure information of the N domain nodes, and then predicts and decodes the plane position information of the current node based on the plane structure information of the N domain nodes. For example, based on the plane structure information of the N domain nodes, a context model index is determined, and based on the context model index, a context model is determined, and the plane position information of the current node is predictively decoded based on the context model.

[0252] In an embodiment of the present application, for each of the N domain nodes, the specific process of determining the plane structure information of the domain node based on the occupancy information of the domain node is consistent. For the sake of convenience of description, any domain node among the N domain nodes is used as an example for illustration.

[0253] In some embodiments, the above S102-A includes the following step S102-A1:

[0254] S102-A1. For any domain node among the N domain nodes, determine at least one of the plane identification information and the plane location information of the domain node based on the occupancy information of the domain node.

[0255] In the embodiment of the present application, the decoding end may determine the plane identification information and / or plane position information of the domain node based on the occupancy information of the domain node.

[0256] The following describes the specific process of determining the plane identification information of the domain node based on the placeholder information of the domain node.

[0257] Specifically, the decoding end determines plane0 and plane1 corresponding to the i-th coordinate axis based on the occupancy information of the domain node, and then determines the plane identification information corresponding to the domain node on the i-th coordinate axis based on plane0 and plane1.

[0258] For example, the decoding end determines the plane 0 corresponding to the domain node on the X, Y, and Z coordinate axes respectively based on the following code:

[0259] uint8_t plane0 = 0;

[0260] plane0|=! ! (occupancy&0x0f)<<0;

[0261] plane0|=! ! (occupancy&0x33)<<1;

[0262] plane0|=! ! (occupancy&0x55)<<2;

[0263] Among them, occupancy represents the occupancy information of the domain node. plane0|=! !(occupancy&0x0f)<<0 represents the plane0 corresponding to the domain node on the X-coordinate axis. plane0|=! !(occupancy&0x33)<<1 represents the plane0 corresponding to the domain node on the Y-coordinate axis. plane0|=! !(occupancy&0x33)<<1 represents the plane0 corresponding to the domain node on the Y-coordinate axis. (occupancy&0x55)<<2 indicates the plane 0 corresponding to the domain node on the Z coordinate axis. 0x0f indicates 00001111. The occupancy information of the domain node is ANDed with 0x0f, and the value of the lower plane of the domain node on the X coordinate axis is 0. 0x33 indicates 00110011. The occupancy information of the domain node is ANDed with 0x33, and the value of the lower plane of the domain node on the Y coordinate axis is 0. 0x55 indicates 01010101. The occupancy information of the domain node is ANDed with 0x55, and the value of the lower plane of the domain node on the Z coordinate axis is 0.

[0264] For example, the decoding end determines the plane 1 corresponding to the domain node on the X, Y, and Z coordinate axes respectively based on the following code:

[0265] uint8_t plane1 = 0;

[0266] plane1|=! ! (occupancy&0xf0)<<0;

[0267] plane1|=! ! (occupancy&0xcc)<<1;

[0268] plane1|=! ! (occupancy&0xaa)<<2;

[0269] Where occupancy represents the occupancy information of the domain node, & represents an AND operation, plane1|=! !(occupancy&0xf0)<<0 represents the plane1 corresponding to the domain node on the X-coordinate axis, plane1|=! !(occupancy&0xcc)<<1 represents the plane1 corresponding to the domain node on the Y-coordinate axis, plane1|=! ! (occupancy&0xaa)<<2 indicates plane 1 corresponding to the domain node on the Z coordinate axis. 0xf0 indicates 11110000. The occupancy information of the domain node is ANDed with 0xf0, and the value of the high plane of the domain node on the X coordinate axis is 0. 0xcc indicates 11001100. The occupancy information of the domain node is ANDed with 0xcc, and the value of the high plane of the domain node on the Y coordinate axis is 0. 0xaa indicates 10101010. The occupancy information of the domain node is ANDed with 0xaa, and the value of the high plane of the domain node on the Z coordinate axis is 0.

[0270] Based on the above method, the decoding end can determine plane0 and plane1 corresponding to the domain node on the i-th coordinate axis, and then determine the plane identification information corresponding to the domain node on the i-th coordinate axis based on plane0 and plane1.

[0271] For example, for the i-th coordinate axis, the plane 0 and plane 1 of the i-th coordinate axis determined above are XORed to determine the plane identification information of the domain node on the i-th axis. Specifically, only when a single plane perpendicular to the axis is occupied is it considered a plane.

[0272] Exemplarily, the decoding end determines the plane identification information of the domain node on the i-th axis based on the following formula (10):

[0273] planarMode=plane0^plane1 (10)

[0274] Among them, planarMode represents the plane identification information of the domain node on the i-th axis, and ^ represents the exclusive OR operation.

[0275] As shown in the above formula (10), the decoding end performs an XOR operation on plane0 and plane1 corresponding to the domain node on the X coordinate axis to obtain the plane identification information of the domain node on the X coordinate axis. For another example, the decoding end performs an XOR operation on plane0 and plane1 corresponding to the domain node on the Y coordinate axis to obtain the plane identification information of the domain node on the Y coordinate axis. For another example, the decoding end performs an XOR operation on plane0 and plane1 corresponding to the domain node on the Z coordinate axis to obtain the plane identification information of the domain node on the Z coordinate axis.

[0276] The specific process of determining the plane location information of the domain node is introduced below.

[0277] In an embodiment of the present application, the decoding end can determine the plane identification information planarMode of the domain node based on the above method, and then determine the plane position information of the domain node based on the plane identification information planarMode.

[0278] Exemplarily, the decoding end determines the plane position information of the domain node on the i-th axis based on the following formula (11):

[0279] PlanePos=planarMode&plane1 (11)

[0280] Among them, PlanePos represents the plane position information of the domain node on the i-th axis, and & represents the AND operation.

[0281] As shown in the above formula (11), the decoding end performs an AND operation on the plane identification information of the domain node on the X coordinate axis and the corresponding plane1 on the X coordinate axis to obtain the plane position information of the domain node on the X coordinate axis. For another example, the decoding end performs an AND operation on the plane identification information on the Y coordinate axis and the corresponding plane1 on the Y coordinate axis to obtain the plane position information of the domain node on the Y coordinate axis. For another example, the decoding end performs an AND operation on the plane identification information on the Z coordinate axis and the corresponding plane1 on the Z coordinate axis to obtain the plane position information of the domain node on the Z coordinate axis.

[0282] The following example illustrates how the decoder determines the plane identification and plane position information of a domain node on the X-axis. Assuming the domain node's placeholder information is 10110000, substitute the placeholder information 10110000 into plane0|=! !(10110000 & 00001111)<<0, resulting in the corresponding plane0 for this domain node on the X-axis being 000000000. Substitute the placeholder information 10110000 into plane1|=! !(10110000 & 11110000)<<0, resulting in the corresponding plane1 for this domain node on the X-axis being 10110000. Next, an XOR operation is performed on plane0 and plane1 to obtain the domain node's plane identification information on the X-axis: planarMode = 00000000^10110000 = 10110000. From planarMode = 10110000, we can see that there are no occupied nodes on the high plane of the X-axis, but there are occupied nodes on the low plane of the X-axis. Therefore, it can be determined that the domain node is a plane in the direction of the X-axis. Next, the decoding end performs an AND operation on the domain node's plane identification information on the X-axis, planarMode, and plane1 to obtain the domain node's plane position information on the X-axis: PlanePos = 10110000 & 10110000 = 10110000. From PlanePos = 10110000, we can see that the domain node is a plane in the direction of the X-axis, and its plane position is the low plane.

[0283] The above describes the specific process of determining the plane identification information and plane position information of the domain node on the X-coordinate axis. The specific process of determining the plane identification information and plane position information of the domain node on the Y-coordinate axis and the Z-coordinate axis can refer to the above-mentioned process of determining the plane identification information and plane position information of the X-coordinate axis, which will not be repeated here.

[0284] Based on the above steps, the decoding end determines the plane identification information and / or plane position information of each domain node in the N domain nodes, and then predicts and decodes the plane position information of the current node based on the plane identification information and / or plane position information of each domain node in the N domain nodes.

[0285] The embodiment of the present application does not limit the specific method of predicting and decoding the plane position information of the current node based on the plane structure information of N domain nodes in the above S102-B.

[0286] In some embodiments, the decoding end determines a context model index based on the plane structure information of N domain nodes, and then selects a context model from multiple preset context models based on the context model index, and then predicts and decodes the plane position information of the current node based on the context model.

[0287] In some embodiments, the above S102-B includes the following steps:

[0288] S102-B1. Determine, based on the plane structure information of the N domain nodes, the first context information and / or the second context information corresponding to the i-th coordinate axis, where the i-th coordinate axis is the X-coordinate axis, the Y-coordinate axis, or the Z-coordinate axis;

[0289] S102-B2: Based on the first context information and / or the second context information corresponding to the i-th coordinate axis, predict and decode the plane position information of the current node on the i-th coordinate axis.

[0290] In this embodiment, the decoding end determines at least one of the first context information and the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes, and then predictively decodes the plane position information of the current node on the i-th coordinate axis based on the determined first context information and / or second context information. For example, the decoding end determines at least one of the first context information and the second context information corresponding to the X-coordinate axis based on the plane structure information of the N domain nodes, and then predictively decodes the plane position information of the current node on the X-coordinate axis based on the first context information and / or the second context information corresponding to the X-coordinate axis. For another example, the decoding end determines at least one of the first context information and the second context information corresponding to the Y-coordinate axis based on the plane structure information of the N domain nodes, and then predictively decodes the plane position information of the current node on the Y-coordinate axis based on the first context information and / or the second context information corresponding to the Y-coordinate axis. For another example, the decoding end determines at least one of the first context information and the second context information corresponding to the Z-coordinate axis based on the plane structure information of the N domain nodes, and then predictively decodes the plane position information of the current node on the Z-coordinate axis based on the first context information and / or the second context information corresponding to the Z-coordinate axis.

[0291] The following describes a specific process of determining the first context information corresponding to the i-th coordinate axis at the decoding end based on the plane structure information of N domain nodes.

[0292] It should be noted that, in the embodiment of the present application, the specific manner in which the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes includes but is not limited to the following:

[0293] Method 1: The decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane structure information of some domain nodes among the N domain nodes.

[0294] For example, the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane structure information of P domain nodes among the N domain nodes that are coplanar with the current node, where P is a positive integer.

[0295] The plane structure information of the P domain nodes includes the plane identification information and / or plane position information of the P domain nodes. That is, the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane identification information of the P coplanar domain nodes. Alternatively, the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane position information of the P coplanar domain nodes. Alternatively, the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane identification information and plane position information of the P coplanar domain nodes.

[0296] The embodiment of the present application does not limit the specific manner in which the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane structure information of P domain nodes among N domain nodes that are coplanar with the current node.

[0297] In one possible implementation, the decoding end performs an AND operation on the planar structure information of any of the P domain nodes and the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. The decoding end then weights the first values ​​corresponding to the P domain nodes to obtain the first context information corresponding to the i-th coordinate axis. It should be noted that different coordinate axes correspond to different first preset values, and the embodiments of this application do not limit the specific values ​​of the first preset values ​​corresponding to each coordinate axis.

[0298] Exemplarily, the first preset value corresponding to the X coordinate axis is 0, the first preset value corresponding to the Y coordinate axis is 1, and the first preset value corresponding to the Z coordinate axis is 2.

[0299] As can be seen from the above, the plane structure information of a domain node includes plane identification information and / or plane position information. Therefore, in some embodiments, the decoding end performs an AND operation on the plane structure information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node, including: the decoding end performs an AND operation on the plane identification information and / or plane position information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. In other words, the decoding end performs an AND operation on the plane identification information of P coplanar domain nodes with the first preset value corresponding to the i-th coordinate axis, and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis. Alternatively, the decoding end performs an AND operation on the plane position information of P coplanar domain nodes with the first preset value corresponding to the i-th coordinate axis, and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis. Alternatively, the decoding end performs an AND operation on the plane identification information and plane position information of P coplanar domain nodes with the first preset value corresponding to the i-th coordinate axis, and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis.

[0300] In the embodiment of the present application, the specific method of weighting the first values ​​corresponding to P domain nodes to obtain the first context information corresponding to the i-th coordinate axis is not limited.

[0301] In some embodiments, the weights of the first values ​​corresponding to the P domain nodes are preset values, so that the first values ​​corresponding to the P domain nodes can be weighted based on the weights of the first values ​​corresponding to each of the P domain nodes to obtain the first context information corresponding to the i-th coordinate axis.

[0302] In some embodiments, the weighting of the first values ​​corresponding to the P domain nodes to obtain the first context information corresponding to the i-th coordinate axis includes the following steps A1 and A2:

[0303] Step A1: determining the number of left shifts corresponding to the first value, and determining the weight corresponding to the first value based on the number of left shifts;

[0304] Step A2: Based on the weighted weight of the first value, weight the first value corresponding to the P domain node to obtain the first context information corresponding to the i-th coordinate axis.

[0305] For example, assume that the P domain nodes coplanar with the current node among the N domain nodes are the three coplanar domain nodes in Figure 11. The plane identification information of these three coplanar domain nodes is recorded as: coPlanarLeftPlaneMode, coPlanarFrontPlaneMode, and coPlanarBelowPlaneMode. The plane position information of these three coplanar domain nodes is recorded as: coPlanarLeftPlanePos, coPlanarFrontPlanePos, and coPlanarBelowPlanePos.

[0306] Performing an AND operation on coPlanarLeftPlaneMode and the first preset value yields a first value of 1, performing an AND operation on coPlanarFrontPlaneMode and the first preset value yields a first value of 2, performing an AND operation on coPlanarBelowPlaneMode and the first preset value yields a first value of 3, performing an AND operation on coPlanarLeftPlanePos and the first preset value yields a first value of 4, performing an AND operation on coPlanarFrontPlanePos and the first preset value yields a first value of 5, and performing an AND operation on coPlanarBelowPlanePos and the first preset value yields a first value of 6. At this point, the six first values ​​occupy a total of six bits, so the number of left shifts corresponding to the six first values ​​can be determined. Assume that the number of left shifts corresponding to the first value 1 is 5, the number of left shifts corresponding to the first value 2 is 4, the number of left shifts corresponding to the first value 3 is 3, the number of left shifts corresponding to the first value 4 is 2, the number of left shifts corresponding to the first value 5 is 1, and the number of left shifts corresponding to the first value 6 is 0.

[0307] In this way, the weighted weight corresponding to each first value can be determined based on the number of left shifts corresponding to each first value. For example, if the number of left shifts corresponding to the first value is m, then 2 m Determine the weighted weight corresponding to the first value. In this way, it can be determined that the weighted weight corresponding to the first value 1 is 2 5 , the first value 2 corresponds to a weighted weight of 2 4 , the first value 3 corresponds to a weighted weight of 2 3 , the first value 4 corresponds to a weighted weight of 2 2 , the first value 5 corresponds to a weighted weight of 2 1 , the first value 6 corresponds to a weighted weight of 2 0 .

[0308] Then, based on the weighted weights of the first values, the first values ​​are weighted to obtain the first context information corresponding to the i-th coordinate axis. It is understood that the weighting of the first values ​​can be understood as concatenating the first values, that is, placing the first values ​​in corresponding bits to obtain the first context information corresponding to the i-th coordinate axis.

[0309] In one example, if the decoding end performs an AND operation on the plane identification information of P coplanar area nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the first context information corresponding to the i-th coordinate axis, the decoding end calculates the first context information Ctx1 using the method shown in the following code:

[0310] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0311] Ctx1=! ! (coPlanarLeftPlaneMode&mask)<<2|

[0312] ! ! (coPlanarFrontPlaneMode&mask)<<1|

[0313] ! ! (coPlanarBelowPlaneMode&mask)

[0314] In one example, if the decoding end performs an AND operation on the planar position information of P coplanar area nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the first context information corresponding to the i-th coordinate axis, the decoding end calculates the first context information Ctx1 using the method shown in the following code:

[0315] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0316] Ctx1=! ! (coPlanarLeftPlanePos&mask)<<2

[0317] ! ! (coPlanarFrontPlanePos&mask)<<1|

[0318] ! ! (coPlanarBelowPlanePos&mask)|

[0319] In one example, if the decoding end performs an AND operation on the plane identification information and plane position information of P coplanar area nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the first context information corresponding to the i-th coordinate axis, the decoding end calculates the first context information Ctx1 using the method shown in the following code:

[0320] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0321] Ctx1=! ! (coPlanarLeftPlanePos&mask)<<5

[0322] ! ! (coPlanarFrontPlanePos&mask)<<4|

[0323] ! ! (coPlanarBelowPlanePos&mask)<<3|

[0324] ! ! (coPlanarLeftPlaneMode&mask)<<2|

[0325] ! ! (coPlanarFrontPlaneMode&mask)<<1|

[0326] ! ! (coPlanarBelowPlaneMode&mask)

[0327] The above describes the specific process of determining the first context information corresponding to the i-th coordinate axis at the decoding end based on the plane structure information of P domain nodes coplanar with the current node among the N domain nodes.

[0328] In some embodiments, the decoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of the domain nodes among the N domain nodes that are collinear with the current node.

[0329] In some embodiments, the decoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of the domain nodes among the N domain nodes that have a common point with the current node.

[0330] In some embodiments, the decoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar and colinear with the current node.

[0331] In some embodiments, the decoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar and co-point with the current node.

[0332] In some embodiments, the decoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is collinear and co-point with the current node.

[0333] The above describes the specific process of determining the first context information corresponding to the i-th coordinate axis at the decoding end based on the plane structure information of some domain nodes among the N domain nodes.

[0334] Method 2: The decoding end determines the first context information corresponding to the i-th coordinate axis based on the first plane position information of the N domain nodes.

[0335] The first plane structure information includes plane identification information and / or plane position information of the domain node. That is, the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane identification information of the N domain nodes. Alternatively, the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane position information of the N domain nodes. Alternatively, the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane position information and plane identification information of the N domain nodes.

[0336] The embodiment of the present application does not limit the specific method in which the decoding end determines the first context information corresponding to the i-th coordinate axis based on the first plane position information of N domain nodes.

[0337] In one possible implementation, the decoding end performs an AND operation on the first plane position information of any of the N domain nodes with the first preset value corresponding to the i-th coordinate axis to obtain the second value corresponding to the domain node. The decoding end then weights the second values ​​corresponding to the N domain nodes to obtain the first context information corresponding to the i-th coordinate axis. It should be noted that different coordinate axes correspond to different first preset values, and the embodiments of this application do not limit the specific values ​​of the first preset values ​​corresponding to each coordinate axis.

[0338] Exemplarily, the first preset value corresponding to the X coordinate axis is 0, the first preset value corresponding to the Y coordinate axis is 1, and the first preset value corresponding to the Z coordinate axis is 2.

[0339] As can be seen from the above, the plane structure information of the domain node includes plane identification information and / or plane position information. Therefore, in some embodiments, the decoding end performs an AND operation on the first plane structure information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the second value corresponding to the domain node, including: the decoding end performs an AND operation on the plane identification information and / or plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. In other words, the decoding end performs an AND operation on the plane identification information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis. Alternatively, the decoding end performs an AND operation on the plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis. Alternatively, the decoding end performs an AND operation on the plane identification information and plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis.

[0340] In the embodiment of the present application, the specific method of weighting the second values ​​corresponding to N domain nodes to obtain the first context information corresponding to the i-th coordinate axis is not limited.

[0341] In some embodiments, the weights of the second values ​​corresponding to the N domain nodes are preset values, so that the second values ​​corresponding to the N domain nodes can be weighted based on the weights of the second values ​​corresponding to each of the N domain nodes to obtain the first context information corresponding to the i-th coordinate axis.

[0342] In some embodiments, the weighting of the second values ​​corresponding to the N domain nodes to obtain the first context information corresponding to the i-th coordinate axis includes the following steps B1 and B2:

[0343] Step B1: determining the number of left shifts corresponding to the second value, and determining a weighted weight corresponding to the second value based on the number of left shifts;

[0344] Step B2: Based on the weighted weights of the second values, weight the second values ​​corresponding to the N domain nodes to obtain the first context information corresponding to the i-th coordinate axis.

[0345] For example, assume that the N domain nodes include the three coplanar domain nodes coPlanarLeft, coPlanarFrontPlane, and coPlanarBelow, as shown in Figure 11, as well as the three collinear domain nodes coEdgerLeft, coEdgerFront, and coEdgerBelow, and the one co-vertex domain node coVertex. Assume that the plane identification information of these seven domain nodes is recorded as: coPlanarLeftPlaneMode, coPlanarFrontPlaneMode, coPlanarBelowPlaneMode, coEdgerLeftPlanarMode, coEdgerFrontPlanarMode, coEdgerBelowPlanarMode, and coVertexPlanarMode, respectively. The plane position information of these seven domain nodes are recorded as: coPlanarLeftPlanePos, coPlanarFrontPlanePos, coPlanarBelowPlanePos, coEdgerLeftPlanePos, coEdgerFrontPlanePos, coEdgerBelowPlanePos and coVertexPlanePos.

[0346] For example, coPlanarLeftPlaneMode is ANDed with the first preset value to obtain a second value of 1, coPlanarFrontPlaneMode is ANDed with the first preset value to obtain a second value of 2, coPlanarBelowPlaneMode is ANDed with the first preset value to obtain a second value of 3, coEdgerLeftPlanarMode is ANDed with the first preset value to obtain a second value of 4, coEdgerFrontPlanarMode is ANDed with the first preset value to obtain a second value of 5, coEdgerBelowPlanarMode is ANDed with the first preset value to obtain a second value of 6, and coVertexPlanarMode is ANDed with the first preset value to obtain a second value of 7. At this time, the seven second values ​​occupy a total of seven bits, so the number of left shift bits corresponding to the seven second values ​​can be determined. Assume that the number of left shifts corresponding to the second value 1 is 6, the number of left shifts corresponding to the second value 2 is 5, the number of left shifts corresponding to the second value 3 is 4, the number of left shifts corresponding to the second value 4 is 3, the number of left shifts corresponding to the second value 5 is 2, the number of left shifts corresponding to the second value 6 is 1, and the number of left shifts corresponding to the second value 7 is 0.

[0347] In this way, the weighted weight corresponding to each second value can be determined based on the number of left shifts corresponding to each second value. For example, if the number of left shifts corresponding to the second value is m, then 2 m Determine the weighted weight corresponding to the second value. In this way, it can be determined that the weighted weight corresponding to the second value 1 is 2 6 , the second value 2 corresponds to a weighted weight of 2 5 , the second value 3 corresponds to a weighted weight of 2 4 , the second value 4 corresponds to a weighted weight of 2 3 , the second value 5 corresponds to a weighted weight of 2 2 , the second value 6 corresponds to a weighted weight of 2 1 , the second value 7 corresponds to a weighted weight of 2 0 .

[0348] Then, based on the weighted weights of the second values, the second values ​​are weighted to obtain the first context information corresponding to the i-th coordinate axis. It is understood that the weighting of the second values ​​can be understood as concatenating the second values, that is, placing the second values ​​in corresponding bits to obtain the first context information corresponding to the i-th coordinate axis.

[0349] In one example, if the decoding end performs an AND operation on the plane identification information of the N-domain node and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the first context information corresponding to the i-th coordinate axis, the decoding end calculates the first context information Ctx1 using the method shown in the following code:

[0350] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0351] Ctx1=! ! (coPlanarLeftPlanarMode&mask)<<6|

[0352] ! ! (coPlanarFrontPlanarMode&mask)<<5|

[0353] ! ! (coPlanarBelowPlanarMode&mask)<<4|

[0354] ! ! (coEdgerLeftPlanarMode&mask)<<3|

[0355] ! ! (coEdgerFrontPlanarMode&mask)<<2|

[0356] ! ! (coEdgerBelowPlanarMode&mask)<<1|

[0357] ! ! (coVertexPlanarMode&mask)

[0358] In one example, if the decoding end performs an AND operation on the plane position information of the N-domain node and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the first context information corresponding to the i-th coordinate axis, the decoding end calculates the first context information Ctx1 using the method shown in the following code:

[0359] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0360] Ctx1=! ! (coPlanarLeftPlanePos&mask)<<6|

[0361] ! ! (coPlanarFrontlanePos&mask)<<5|

[0362] ! ! (coPlanarBelowPlanePos&mask)<<4|

[0363] ! ! (coEdgerLeftPlanePos&mask)<<3|

[0364] ! ! (coEdgerFrontPlanePos&mask)<<2|

[0365] ! ! (coEdgerBelowPlanePos&mask)<<1|

[0366] ! ! (coVertexPlanePos&mask)

[0367] In one example, if the decoding end performs an AND operation on the plane identification information and plane position information of the N domain nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the first context information corresponding to the i-th coordinate axis, the decoding end calculates the first context information Ctx1 using the method shown in the following code:

[0368] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0369] Ctx1=! ! (coPlanarLeftPlanarMode&mask)<<13|

[0370] ! ! (coPlanarFrontPlanarMode&mask)<<12|

[0371] ! ! (coPlanarBelowPlanarMode&mask)<<11|

[0372] ! ! (coEdgerLeftPlanarMode&mask)<<10|

[0373] ! ! (coEdgerFrontPlanarMode&mask)<<9|

[0374] ! ! (coEdgerBelowPlanarMode&mask)<<8|

[0375] ! ! (coVertexPlanarMode&mask)<<7|

[0376] ! ! (coPlanarLeftPlanePos&mask)<<6|

[0377] ! ! (coPlanarFrontlanePos&mask)<<5|

[0378] ! ! (coPlanarBelowPlanePos&mask)<<4|

[0379] ! ! (coEdgerLeftPlanePos&mask)<<3|

[0380] ! ! (coEdgerFrontPlanePos&mask)<<2|

[0381] ! ! (coEdgerBelowPlanePos&mask)<<1|

[0382] ! ! (coVertexPlanePos&mask)

[0383] The above describes the specific process by which the decoder determines the first context information corresponding to the i-th coordinate axis based on the planar structure information of N domain nodes. It should be noted that in addition to determining the first context information corresponding to the i-th coordinate axis using the various methods described above, the decoder may also use other methods to determine the first context information corresponding to the i-th coordinate axis.

[0384] The following describes a specific process of determining the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes in S102-B1.

[0385] It should be noted that, in the embodiment of the present application, the specific manner in which the decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes includes but is not limited to the following:

[0386] Method 1: The decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane structure information of some domain nodes among the N domain nodes.

[0387] For example, the decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane structure information of Q domain nodes among the N domain nodes that are collinear and / or co-point with the current node, where Q is a positive integer.

[0388] The plane structure information of the Q domain nodes includes the plane identification information and / or plane position information of the Q domain nodes. That is, the decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane identification information of the P coplanar domain nodes. Alternatively, the decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane position information of the Q domain nodes. Alternatively, the decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane identification information and plane position information of the Q domain nodes.

[0389] The embodiment of the present application does not limit the specific method for the decoding end to determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of Q domain nodes in N domain nodes that are collinear and / or co-point with the current node.

[0390] In one possible implementation, the decoding end performs an AND operation on the planar structure information of any of the Q domain nodes and the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. The decoding end then weights the first values ​​corresponding to the P domain nodes to obtain the second context information corresponding to the i-th coordinate axis. It should be noted that different coordinate axes correspond to different first preset values, and the embodiments of this application do not limit the specific values ​​of the first preset values ​​corresponding to each coordinate axis.

[0391] Exemplarily, the first preset value corresponding to the X coordinate axis is 0, the first preset value corresponding to the Y coordinate axis is 1, and the first preset value corresponding to the Z coordinate axis is 2.

[0392] As can be seen from the above, the plane structure information of the domain node includes plane identification information and / or plane position information. Therefore, in some embodiments, the decoding end performs an AND operation on the plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node, including: the decoding end performs an AND operation on the plane identification information and / or plane position information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. In other words, the decoding end performs an AND operation on the plane identification information of Q domain nodes and the first preset value corresponding to the i-th coordinate axis, and then weights them to obtain the second context information corresponding to the i-th coordinate axis. Alternatively, the decoding end performs an AND operation on the plane position information of Q domain nodes and the first preset value corresponding to the i-th coordinate axis, and then weights them to obtain the second context information corresponding to the i-th coordinate axis. Alternatively, the decoding end performs an AND operation on the plane identification information and plane position information of Q domain nodes and the first preset value corresponding to the i-th coordinate axis, and then weights them to obtain the second context information corresponding to the i-th coordinate axis.

[0393] In the embodiment of the present application, the specific method of weighting the first values ​​corresponding to the Q domain nodes to obtain the second context information corresponding to the i-th coordinate axis is not limited.

[0394] In some embodiments, the weights of the first values ​​corresponding to the Q domain nodes are preset values, so that the first values ​​corresponding to the Q domain nodes can be weighted based on the weights of the first values ​​corresponding to each of the Q domain nodes to obtain the second context information corresponding to the i-th coordinate axis.

[0395] In some embodiments, the weighting of the first values ​​corresponding to the Q domain nodes to obtain the second context information corresponding to the i-th coordinate axis includes the following steps C1 and C2:

[0396] Step C1: determining the number of left shifts corresponding to the first value, and determining a weighted weight corresponding to the first value based on the number of left shifts;

[0397] Step C2: Based on the weighted weight of the first value, weight the first value corresponding to the Q domain node to obtain the second context information corresponding to the i-th coordinate axis.

[0398] For example, assume that among N domain nodes, the Q domain nodes that are collinear or co-point with the current node are the three collinear domain nodes coEdgerLeft, coEdgerFront, and coEdgerBelow in Figure 11, and the one domain node coVertex that is co-point with the current node. The plane identification information of these four domain nodes is recorded as: coEdgerLeftPlaneMode, coEdgerFrontPlaneMode, coEdgerBelowPlaneMode, and coVertexPlaneMode, respectively. The plane position information of these four domain nodes is recorded as: coEdgerLeftPlanePos, coEdgerFrontPlanePos, coEdgerBelowPlanePos, and coVertexPlanePos, respectively.

[0399] For example, coEdgerLeftPlaneMode is ANDed with the first preset value to obtain a first value of 1, coEdgerFrontPlaneMode is ANDed with the first preset value to obtain a first value of 2, coEdgerBelowPlaneMode is ANDed with the first preset value to obtain a first value of 3, coVertexPlaneMode is ANDed with the first preset value to obtain a first value of 4, coEdgerLeftPlanePos is ANDed with the first preset value to obtain a first value of 5, coEdgerFrontPlanePos is ANDed with the first preset value to obtain a first value of 6, coEdgerBelowPlanePos is ANDed with the first preset value to obtain a first value of 7, and coVertexPlanePos is ANDed with the first preset value to obtain a first value of 8. At this time, the eight first values ​​occupy a total of 8 bits, so the number of left shift bits corresponding to the eight first values ​​can be determined. Assume that the number of left shifts corresponding to the first value 1 is 7, the number of left shifts corresponding to the first value 2 is 6, the number of left shifts corresponding to the first value 3 is 5, the number of left shifts corresponding to the first value 4 is 4, the number of left shifts corresponding to the first value 5 is 3, the number of left shifts corresponding to the first value 6 is 2, the number of left shifts corresponding to the first value 7 is 1, and the number of left shifts corresponding to the first value 8 is 0.

[0400] In this way, the weighted weight corresponding to each first value can be determined based on the number of left shifts corresponding to each first value. For example, if the number of left shifts corresponding to the first value is m, then 2 m Determine the weighted weight corresponding to the first value. In this way, it can be determined that the weighted weight corresponding to the first value 1 is 2 7 , the first value 2 corresponds to a weighted weight of 2 6 , the first value 3 corresponds to a weighted weight of 2 5, the first value 4 corresponds to a weighted weight of 2 4 , the first value 5 corresponds to a weighted weight of 2 3 , the first value 6 corresponds to a weighted weight of 2 2 , the first value 7 corresponds to a weighted weight of 2 1 , the first value 8 corresponds to a weighted weight of 2 0 .

[0401] Then, based on the weighted weights of the first values, the first values ​​are weighted to obtain the second context information corresponding to the i-th coordinate axis. It is understood that the weighting of the first values ​​can be understood as concatenating the first values, that is, placing the first values ​​in corresponding bits to obtain the second context information corresponding to the i-th coordinate axis.

[0402] In one example, if the decoding end performs an AND operation on the plane identification information of the Q domain nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information corresponding to the i-th coordinate axis, the decoding end calculates the second context information Ctx2 using the method shown in the following code:

[0403] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0404] Ctx2=! ! (coEdgerFrontPlaneMode&mask)<<3|

[0405] ! ! (coEdgerFrontPlaneMode&mask)<<2|

[0406] ! ! (coEdgerBelowPlaneMode&mask)<<1|

[0407] ! ! (coVertexPlaneMode&mask)

[0408] In one example, if the decoding end performs an AND operation on the plane position information of Q domain nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information corresponding to the i-th coordinate axis, the decoding end calculates the second context information Ctx2 using the method shown in the following code:

[0409] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0410] Ctx2=! ! (coEdgerFrontPlanePos&mask)<<3|

[0411] ! ! (coEdgerFrontPlanePos&mask)<<2|

[0412] ! ! (coEdgerBelowPlanePos&mask)<<1|

[0413] ! ! (coVertexPlanePos&mask)

[0414] In one example, if the decoding end performs an AND operation on the plane identification information and plane position information of the Q domain nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information corresponding to the i-th coordinate axis, the decoding end calculates the second context information Ctx2 using the method shown in the following code:

[0415] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0416] Ctx2=! ! (coEdgerLeftPlanePos&mask)<<7|

[0417] ! ! (coEdgerFrontPlanePos&mask)<<6|

[0418] ! ! (coEdgerBelowPlanePos&mask)<<5|

[0419] ! ! (coVertexPlanePos&mask)<<4|

[0420] ! ! (coEdgerFrontPlaneMode&mask)<<3|

[0421] ! ! (coEdgerFrontPlaneMode&mask)<<2|

[0422] ! ! (coEdgerBelowPlaneMode&mask)<<1|

[0423] ! ! (coVertexPlaneMode&mask)

[0424] The above describes a specific process of determining the second context information corresponding to the i-th coordinate axis at the decoding end based on the plane structure information of Q domain nodes among the N domain nodes that are collinear and / or co-point with the current node.

[0425] In some embodiments, the decoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is collinear with the current node.

[0426] In some embodiments, the decoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that has a common point with the current node.

[0427] In some embodiments, the decoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar with the current node.

[0428] In some embodiments, the decoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar and colinear with the current node.

[0429] In some embodiments, the decoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar and co-point with the current node.

[0430] The above describes the specific process of determining the second context information corresponding to the i-th coordinate axis at the decoding end based on the plane structure information of some domain nodes among the N domain nodes.

[0431] Method 2: The decoding end determines the second context information corresponding to the i-th coordinate axis based on the second plane position information of the N domain nodes.

[0432] The second plane structure information includes plane identification information and / or plane position information of the domain node. That is, the decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane identification information of the N domain nodes. Alternatively, the decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane position information of the N domain nodes. Alternatively, the decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane position information and plane identification information of the N domain nodes.

[0433] The embodiment of the present application does not limit the specific method in which the decoding end determines the second context information corresponding to the i-th coordinate axis based on the first plane position information of N domain nodes.

[0434] In one possible implementation, the decoding end performs an AND operation on the second plane position information of any one of the N domain nodes with the first preset value corresponding to the i-th coordinate axis to obtain a third value corresponding to the domain node. The decoding end then weights the third values ​​corresponding to the N domain nodes to obtain the second context information corresponding to the i-th coordinate axis. It should be noted that different coordinate axes correspond to different first preset values, and the embodiments of this application do not limit the specific values ​​of the first preset values ​​corresponding to each coordinate axis.

[0435] Exemplarily, the first preset value corresponding to the X coordinate axis is 0, the first preset value corresponding to the Y coordinate axis is 1, and the first preset value corresponding to the Z coordinate axis is 2.

[0436] As can be seen from the above, the plane structure information of the domain node includes plane identification information and / or plane position information. Therefore, in some embodiments, the decoding end performs an AND operation on the second plane structure information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the third value corresponding to the domain node, including: the decoding end performs an AND operation on the plane identification information and / or plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis to obtain the third value corresponding to the domain node. In other words, the decoding end performs an AND operation on the plane identification information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the second context information corresponding to the i-th coordinate axis. Alternatively, the decoding end performs an AND operation on the plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the second context information corresponding to the i-th coordinate axis. Alternatively, the decoding end performs an AND operation on the plane identification information and plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the second context information corresponding to the i-th coordinate axis.

[0437] In the embodiment of the present application, the specific method of weighting the third values ​​corresponding to N domain nodes to obtain the second context information corresponding to the i-th coordinate axis is not limited.

[0438] In some embodiments, the weights of the third values ​​corresponding to the N domain nodes are preset values, so that the third values ​​corresponding to the N domain nodes can be weighted based on the weights of the third values ​​corresponding to each of the N domain nodes to obtain the second context information corresponding to the i-th coordinate axis.

[0439] In some embodiments, the weighting of the third values ​​corresponding to the N domain nodes to obtain the second context information corresponding to the i-th coordinate axis includes the following steps D1 and D2:

[0440] Step D1, determining the number of left shifts corresponding to the third value, and determining a weighted weight corresponding to the third value based on the number of left shifts;

[0441] Step D2: Based on the weighted weight of the third value, weight the third values ​​corresponding to the N domain nodes to obtain the second context information corresponding to the i-th coordinate axis.

[0442] For example, assume that the N domain nodes include the three coplanar domain nodes coPlanarLeft, coPlanarFrontPlane, and coPlanarBelow, as shown in Figure 11, as well as the three collinear domain nodes coEdgerLeft, coEdgerFront, and coEdgerBelow, and the one co-vertex domain node coVertex. Assume that the plane identification information of these seven domain nodes is recorded as: coPlanarLeftPlaneMode, coPlanarFrontPlaneMode, coPlanarBelowPlaneMode, coEdgerLeftPlanarMode, coEdgerFrontPlanarMode, coEdgerBelowPlanarMode, and coVertexPlanarMode, respectively. The plane position information of these seven domain nodes are recorded as: coPlanarLeftPlanePos, coPlanarFrontPlanePos, coPlanarBelowPlanePos, coEdgerLeftPlanePos, coEdgerFrontPlanePos, coEdgerBelowPlanePos and coVertexPlanePos.

[0443] For example, coPlanarLeftPlaneMode is ANDed with the first preset value to obtain a third value of 1, coPlanarFrontPlaneMode is ANDed with the first preset value to obtain a third value of 2, coPlanarBelowPlaneMode is ANDed with the first preset value to obtain a third value of 3, coEdgerLeftPlanarMode is ANDed with the first preset value to obtain a third value of 4, coEdgerFrontPlanarMode is ANDed with the first preset value to obtain a third value of 5, coEdgerBelowPlanarMode is ANDed with the first preset value to obtain a third value of 6, and coVertexPlanarMode is ANDed with the first preset value to obtain a third value of 7. At this time, the seven third values ​​occupy a total of seven bits, so the number of left shift bits corresponding to the seven third values ​​can be determined. Assume that the number of left shifts corresponding to the third value 1 is 6, the number of left shifts corresponding to the third value 2 is 5, the number of left shifts corresponding to the third value 3 is 4, the number of left shifts corresponding to the third value 4 is 3, the number of left shifts corresponding to the third value 5 is 2, the number of left shifts corresponding to the third value 6 is 1, and the number of left shifts corresponding to the third value 7 is 0.

[0444] In this way, the weighted weight corresponding to each third value can be determined based on the number of left shifts corresponding to each third value. For example, if the number of left shifts corresponding to the third value is m, then 2 m The weighted weight corresponding to the third value is determined. In this way, it can be determined that the weighted weight corresponding to the third value 1 is 2 6 , the weighted weight corresponding to the third value 2 is 2 5 , the third value 3 corresponds to a weighted weight of 2 4 , the third value 4 corresponds to a weighted weight of 2 3 , the third value 5 corresponds to a weighted weight of 2 2 , the third value 6 corresponds to a weighted weight of 2 1 , the third value 7 corresponds to a weighted weight of 2 0 .

[0445] Then, based on the weighted weights of the third values, the third values ​​are weighted to obtain the second context information corresponding to the i-th coordinate axis. It is understood that the weighting of the third values ​​can be understood as concatenating the third values, that is, placing the third values ​​in corresponding bits to obtain the second context information corresponding to the i-th coordinate axis.

[0446] In one example, if the decoding end performs an AND operation on the plane identification information of the N-domain node and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information corresponding to the i-th coordinate axis, the decoding end calculates the second context information Ctx2 using the method shown in the following code:

[0447] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0448] Ctx2=! ! (coPlanarLeftPlanarMode&mask)<<6|

[0449] ! ! (coPlanarFrontPlanarMode&mask)<<5|

[0450] ! ! (coPlanarBelowPlanarMode&mask)<<4|

[0451] ! ! (coEdgerLeftPlanarMode&mask)<<3|

[0452] ! ! (coEdgerFrontPlanarMode&mask)<<2|

[0453] ! ! (coEdgerBelowPlanarMode&mask)<<1|

[0454] ! ! (coVertexPlanarMode&mask)

[0455] In one example, if the decoding end performs an AND operation on the plane position information of the N-domain node and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information corresponding to the i-th coordinate axis, the decoding end calculates the second context information Ctx2 using the method shown in the following code:

[0456] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0457] Ctx2=! ! (coPlanarLeftPlanePos&mask)<<6|

[0458] ! ! (coPlanarFrontlanePos&mask)<<5|

[0459] ! ! (coPlanarBelowPlanePos&mask)<<4|

[0460] ! ! (coEdgerLeftPlanePos&mask)<<3|

[0461] ! ! (coEdgerFrontPlanePos&mask)<<2|

[0462] ! ! (coEdgerBelowPlanePos&mask)<<1|

[0463] ! ! (coVertexPlanePos&mask)

[0464] In one example, if the decoding end performs an AND operation on the plane identification information and plane position information of the N domain nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information corresponding to the i-th coordinate axis, the decoding end calculates the second context information Ctx2 using the method shown in the following code:

[0465] Const int mask=1< <axisIdx(axisIdx=0(x),1(y),2(z))

[0466] Ctx2=! ! (coPlanarLeftPlanarMode&mask)<<13|

[0467] ! ! (coPlanarFrontPlanarMode&mask)<<12|

[0468] ! ! (coPlanarBelowPlanarMode&mask)<<11|

[0469] ! ! (coEdgerLeftPlanarMode&mask)<<10|

[0470] ! ! (coEdgerFrontPlanarMode&mask)<<9|

[0471] ! ! (coEdgerBelowPlanarMode&mask)<<8|

[0472] ! ! (coVertexPlanarMode&mask)<<7|

[0473] ! ! (coPlanarLeftPlanePos&mask)<<6|

[0474] ! ! (coPlanarFrontlanePos&mask)<<5|

[0475] ! ! (coPlanarBelowPlanePos&mask)<<4|

[0476] ! ! (coEdgerLeftPlanePos&mask)<<3|

[0477] ! ! (coEdgerFrontPlanePos&mask)<<2|

[0478] ! ! (coEdgerBelowPlanePos&mask)<<1|

[0479] ! ! (coVertexPlanePos&mask)

[0480] The above describes the specific process by which the decoder determines the second context information corresponding to the i-th coordinate axis based on the planar structure information of N domain nodes. It should be noted that in addition to determining the second context information corresponding to the i-th coordinate axis using the various methods described above, the decoder may also use other methods to determine the second context information corresponding to the i-th coordinate axis.

[0481] It should be noted that the first context information and the second context information corresponding to the i-th coordinate axis obtained by the decoding end are different. That is, the method used by the decoding end to determine the first context information corresponding to the i-th coordinate axis is different from the method used to determine the second context information corresponding to the i-th coordinate axis, and thus the first context information and the second context information obtained are also different.

[0482] After the decoding end determines the first context information and / or second context information corresponding to the i-th coordinate axis based on the plane structure information of N domain nodes in the above manner, it executes the above S102-B2 steps, and predicts and decodes the plane position information of the current node on the i-th coordinate axis based on the first context information and / or second context information corresponding to the i-th coordinate axis.

[0483] The embodiment of the present application does not limit the specific method of predicting and decoding the planar position information of the current node on the i-th coordinate axis based on the first context information and / or second context information corresponding to the i-th coordinate axis in S102-B2.

[0484] In some embodiments, the decoding end predictively decodes the planar position information of the current node on the i-th coordinate axis based only on the first context information and / or the second context information corresponding to the i-th coordinate axis. For example, the decoding end determines a context model index based on the first context information and / or the second context information corresponding to the i-th coordinate axis, then selects a context model from a plurality of preset context models based on the context model index, and then uses the context model to predictively decode the planar position information of the current node on the i-th coordinate axis.

[0485] In some embodiments, the above S102-B2 includes the following steps S102-B21:

[0486] S102-B21. Based on the first context information and / or the second context information corresponding to the i-th coordinate axis and the preset context information, predict and decode the plane position information of the current node on the i-th coordinate axis.

[0487] In this embodiment, when the decoding end predictively decodes the planar position information of the current node on the i-th coordinate axis, the referenced context information includes other preset context information in addition to the first context information and / or second context information corresponding to the i-th coordinate axis.

[0488] The embodiment of the present application does not limit the specific content of the preset context information, which can be determined according to actual needs.

[0489] In one possible implementation, the preset context information includes at least one of the following four context information:

[0490] 1. Using the occupancy information of neighboring nodes to predict the plane position information of the current node, the plane position information is divided into three elements: predicted as low plane, predicted as high plane, and unpredictable;

[0491] 2. The spatial distance between the nodes at the same partition depth and the same coordinates as the current node and the current node is "close" or "far";

[0492] 3. The plane position of the node at the same partition depth and the same coordinates as the current node, if it is a plane;

[0493] 4. Coordinate dimension (i=0, 1, 2).

[0494] In this embodiment, when the decoding end predicts and decodes the plane position information of the current node on the i-th coordinate axis, it determines the first context information and / or the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes of the current node, and then predicts and decodes the plane position information of the current node on the i-th coordinate axis based on the first context information and / or the second context information corresponding to the i-th coordinate axis and the preset context information. It can be seen that in the embodiment of the present application, when the decoding end predicts and decodes the plane position information of the current node, it not only considers the preset prior information (i.e., the preset context information), but also considers the plane structure information of the domain node (i.e., the first context information and / or the second context information), thereby improving the prediction decoding effect of the plane position information of the current node and improving the decoding efficiency of the point cloud.

[0495] The embodiment of the present application does not limit the specific process of predicting and decoding the planar position information of the current node on the i-th coordinate axis based on the first context information and / or second context information corresponding to the i-th coordinate axis, and the preset context information at the decoding end.

[0496] In some embodiments, the above S102-B21 includes the following steps S102-B211 and S102-B212:

[0497] S102-B211, determining a target context model based on the first context information and / or the second context information corresponding to the i-th coordinate axis, and preset context information;

[0498] S102-B212: Based on the target context model, predict and decode the plane position information of the current node on the i-th coordinate axis.

[0499] In this embodiment, the decoding end determines a context model based on the first context information and / or second context information corresponding to the i-th coordinate axis, as well as the preset context information. For ease of description, this context model is referred to as the target context model. The target context model is then used to predictively decode the planar position information of the current node on the i-th coordinate axis.

[0500] The following describes a specific process in which the decoding end determines the target context model based on the first context information and / or the second context information corresponding to the i-th coordinate axis and the preset context information.

[0501] In some embodiments, the decoding end determines the index of the target context model based on the first context information and / or second context information corresponding to the i-th coordinate axis, as well as the preset context information, and then selects the target context model from the preset multiple context models based on the index of the target context model, and then uses the target context model to predict and decode the planar position information of the current node on the i-th coordinate axis.

[0502] In this embodiment, multiple context models are set for the plane position information. The embodiment of the present application does not limit the specific number of context models corresponding to the plane position information, as long as it is greater than 1. In other words, in the embodiment of the present application, an optimal context model is selected from at least two context models to predict and decode the plane position information of the current node on the i-th coordinate axis.

[0503] For example, the plane position information corresponds to multiple context models as shown in Table 2:

[0504] Table 2

[0505] Index context model 0 context model A1 context model B…………

[0506] In this way, the decoding end determines the index of the target context model based on the first context information and / or second context information corresponding to the i-th coordinate axis, as well as the preset context information. Then, based on the index of the target context model, the target context model is selected from the corresponding context models in Table 2 to perform predictive decoding on the planar position information of the current node on the i-th coordinate axis.

[0507] In some embodiments, the above S102-B211 includes the following steps S102-B2111 and S102-B2112:

[0508] S102-B2111. Divide the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information;

[0509] S102-B2112. Determine a target context model based on the primary information of the current node and part or all of the secondary information of the current node.

[0510] As can be seen from the above, assuming that the context information of the plane position information includes the first context information and the second context information corresponding to the i-th coordinate axis, as well as the above-mentioned four preset context information, the final context of the plane position is as follows:

[0511] 1. Using the occupancy information of neighboring nodes, the plane position information of the current node is predicted to be three elements: predicted to be ground plane, predicted to be high plane, and unpredictable;

[0512] 2. The spatial distance between the nodes at the same partition depth and the same coordinates as the current node and the current node is "close" or "far";

[0513] 3. The plane position of the node at the same partition depth and the same coordinates as the current node, if it is a plane;

[0514] 4. Coordinate dimension (i = 0, 1, 2);

[0515] 5. Ctx1: Planar structure information of three coplanar neighboring nodes;

[0516] 6. Ctx2: Planar structure information of three collinear neighboring nodes and one common neighboring node.

[0517] Assuming that the decoding end determines the first context information corresponding to the i-th coordinate axis based on the plane identification information and plane position information of the three domain nodes coplanar with the current node among the N domain nodes, it can be obtained that Ctx1 includes 2 6= 64 contexts. Assuming that the decoding end determines the second context information corresponding to the i-th coordinate axis based on the plane identification information and plane position information of the three domain nodes that are collinear with the current node and the one domain node that is co-pointed with the current node among the N domain nodes, it can be obtained that Ctx2 includes 2 8 = 256 contexts. In this way, the decoding end can obtain 3×2×2×3×64×256=589824 contexts based on the first context information and the second context information corresponding to the i-th coordinate axis, as well as the above-mentioned 4 preset context information. The memory space occupied by so many contexts is very large. Based on this, when the embodiment of the present application predictively decodes the plane position information of the node, the most advanced coding technology Dynamic-OUBF of G-PCC is added to this algorithm to reduce the number of contexts used for decoding the plane position information, for example, reducing the number of plane position contexts to 3x16=48.

[0518] Specifically, in an embodiment of the present application, as shown in FIG12 , the decoding end divides the first context information and / or second context information corresponding to the above-determined i-th coordinate axis, and the preset context information into primary information and secondary information, and then determines the target context model based on the primary information of the current node and part or all of the secondary information of the current node. It should be noted that in an embodiment of the present application, the target context model is mainly determined based on the primary information and part of the secondary information of the current node, thereby reducing the number of contexts. This not only reduces the memory usage of the context, but also improves the prediction decoding efficiency of the node's planar position information.

[0519] The embodiment of the present application does not limit the specific manner of dividing the first context information and / or second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information.

[0520] In one example, the decoding end divides the first context information corresponding to the i-th coordinate axis, the spatial distance "near" and "far" between the node and the current node at the same division depth and the same coordinates as the current node, and the plane position of the node at the same division depth and the same coordinates as the current node, if it is a plane, into primary information. The second context information corresponding to the i-th coordinate axis and the plane position information of the current node predicted by using the occupancy information of the neighboring nodes are divided into three elements: predicted as the ground plane, predicted as the high plane, and unpredictable as secondary information, and the coordinate dimension (i=0, 1, 2) is used as the index itself and is not divided into primary information and secondary information.

[0521] In another example, the decoding end may divide the first context information and the second context information corresponding to the i-th coordinate axis into the primary information of the current node, and divide at least one of the above-mentioned four preset context information into the secondary information of the current node.

[0522] In another example, the decoding end may assign the second context information corresponding to the i-th coordinate axis to the primary information of the current node, and assign the first context information corresponding to the i-th coordinate axis to the secondary information of the current node. Alternatively, on this basis, at least one of the four preset context information may be assigned to the primary information of the current node, and the remaining preset context information may be assigned to the secondary information of the current node.

[0523] In another example, the decoding end can also divide the first context information and the second context information corresponding to the i-th coordinate axis into the secondary information of the current node, and divide at least one context information of the above-mentioned four preset context information into the secondary information of the current node.

[0524] It should be noted that the way in which the decoding end divides the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information includes but is not limited to the way shown above. The decoding end can adopt other ways to divide the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information.

[0525] Based on the above steps, the decoding end divides the first context information and / or second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information, and then executes the above steps S102-B2112 to determine the target context model based on the primary information of the current node and part or all of the secondary information of the current node.

[0526] The embodiment of the present application does not limit the specific manner in which the decoding end determines the target context model based on the primary information of the current node and part or all of the secondary information of the current node.

[0527] In some embodiments, the decoding end determines an index based on the main information of the current node and part of the secondary information of the current node, determines the index of the target context model based on the index, and then determines the target context model from multiple preset context models based on the index of the target context model.

[0528] In some embodiments, the above S102-B2112 includes the following steps S102-B21121 to S102-B21124:

[0529] S102-B21121, converting the primary information of the current node and the secondary information of the current node into binary representation;

[0530] S102-B21122, determine the number of right-shifted bits of the secondary information corresponding to the current node, and select the first secondary information from the secondary information of the current node after binary representation based on the number of right-shifted bits of the secondary information corresponding to the current node, where the initial number of right-shifted bits of the secondary information is the total number of bits of the binary secondary information;

[0531] S102-B21123, determining a first index based on the primary information and the first secondary information after the binary representation of the current node, and obtaining an index of the target context model from a preset context model index cache based on the first index;

[0532] S102-B21124. Obtain a target context model based on the index of the target context model.

[0533] In this embodiment, the decoding end divides the first context information and / or second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information based on the above steps. Then, the decoding end converts the primary information and secondary information of the current node obtained by the division into binary representation.

[0534] For example, referring to the above example, it is assumed that the decoding end divides the first context information corresponding to the i-th coordinate axis, the spatial distance between the node and the current node at the same division depth and the same coordinate as the current node, and the plane position of the node at the same division depth and the same coordinate as the current node into the main information if it is a plane. Assume that the first context information Ctx1 corresponding to the i-th coordinate axis includes 2 6 =64 contexts, which require 6 bits to represent when converted to binary representation. The spatial distances "near" and "far" between the nodes at the same division depth and the same coordinates as the current node and the current node include 2 contexts, which require 1 bit to represent when converted to binary representation. The plane position of the node at the same division depth and the same coordinates as the current node includes 2 contexts, which require 1 bit to represent when converted to binary representation. Therefore, in this example, when the main information of the current node is converted to binary representation, 6+1+1=8 bits are required to represent.

[0535] Similarly, if the decoder uses the second context information corresponding to the i-th coordinate axis and the neighboring node occupancy information to predict the plane position information of the current node into three elements: predicted as ground plane, predicted as high plane and unpredictable as secondary information. Assume that the second context information Ctx2 corresponding to the i-th coordinate axis includes 2 8 = 256 contexts, which require 8 bits to represent when converted to binary. Context information: Using neighboring node occupancy information, the plane position information of the current node is predicted to consist of three elements: predicted ground plane, predicted high plane, and unpredictable. These three contexts require 2 bits to represent when converted to binary. Therefore, in this example, the secondary information of the current node requires 8 + 2 = 10 bits to represent when converted to binary.

[0536] The above examples illustrate the specific process of converting the primary and secondary information of the current node into binary representation. It should be noted that the division of the primary and secondary information of the current node includes, but is not limited to, the above examples. If the primary and secondary information of the current node also include other contextual information, the method shown in the above examples can be used to convert the primary and secondary information of the current node into binary representation.

[0537] After the decoding end converts the primary information and secondary information of the current node into binary representation, it determines the number of right-shifted bits of the secondary information corresponding to the current node, and then selects the first secondary information from the secondary information after the binary representation of the current node based on the number of right-shifted bits of the secondary information. In the embodiment of the present application, the number of right-shifted bits of the secondary information corresponding to the current node can be understood as the number of secondary information selected from the secondary information of the current node to predict and decode the planar position information of the current node.

[0538] The following describes how to determine the number of right-shifted bits of the secondary information corresponding to the current node.

[0539] The embodiment of the present application does not limit the specific method of determining the number of right-shift bits of the secondary information corresponding to the current node.

[0540] In some embodiments, the number of right-shifted bits of secondary information corresponding to the current node is a preset value. For example, for the nodes in the point cloud octree, a preset number of nodes corresponds to one right-shifted bit of secondary information, so that the number of right-shifted bits of secondary information corresponding to the current node can be determined. Exemplarily, the closer the node is to the root node of the octree, the larger the number of right-shifted bits of secondary information corresponding to the node. Optionally, the initial value of the right-shifted bit of secondary information is the total number of bits of binary secondary information. For example, if the current node is the root node of the octree, the number of right-shifted bits of secondary information corresponding to the current node is the above-mentioned 10 bits.

[0541] In some embodiments, determining the number of right-shifted bits of the secondary information corresponding to the current node in S102-B21122 includes the following steps S102-B211221 and S102-B211222:

[0542] S102-B211221, determining the number of right-shifted bits of the secondary information corresponding to the last layer of the current secondary information partition tree, where the secondary information partition tree is obtained by performing binary tree partitioning on the secondary information starting from the highest bit of the secondary information;

[0543] S102-B211222. Determine the number of right shifts of the secondary information corresponding to the last layer as the number of right shifts of the secondary information corresponding to the current node.

[0544] The following is an introduction to the division process of secondary information.

[0545] Specifically, when the decoding end decodes the current point cloud, during the entire Dynamic-OUBF initialization process, assuming that the integer representation of the main information is ct1 and the integer representation of the secondary information is ct2, a context model index cache ContextBuffer is initialized, and the size of ContextBuffer is ct1×ct2. Exemplarily, referring to the above example, assuming that the main information includes 8 bits and the secondary information includes 10 bits, an 8×10 context model index cache ContextBuffer can be determined, and the context model index cache ContextBuffer stores 8×10 context model indexes. In addition, the initial context probability of each state is set to 127 (i.e., 0.5).

[0546] In some embodiments, the process of recovering the secondary information accuracy is shown in FIG13 :

[0547] First, the entire secondary information context is represented in binary format. Next, the secondary information is partitioned into a binary tree, starting with the highest bit. As shown in Figure 13, above a certain level, a non-full binary tree is created, meaning the partitioning is based on the secondary information itself. However, once the MinDepth (currently set to 3) is reached, the accuracy of the secondary information is fully restored. The partitioning of secondary information is described in detail below.

[0548] Exemplarily, a countBuffer counter is initialized with a size of ct1×(ct2>>MinDepth) and is initialized to 0.

[0549] In addition, a KDown is initialized to represent the precision (i.e., the number of right shifts) of the secondary information corresponding to each first index (state). The initial value of the number of right shifts of the secondary information is the total number of bits in the binary secondary information, i.e., the maximum precision of the secondary information. For example, if the secondary information is 10 bits, the initial value of the number of right shifts of the secondary information is 10 bits.

[0550] Furthermore, a table named CountTimeTh is initialized to control the maximum number of occurrences of the first index (state) at each level of the secondary information partitioning tree. When the number of occurrences of a first index (state) exceeds the limit for that level, the low-bit precision of the secondary information is restored, and the number of occurrences of the current first index (state) is reset to zero. The context probability of the new first index (state) is restored and inherits the probability of its parent node.

[0551] Specifically, as shown in FIG13 , for the first node 1 in the point cloud, when predicting and decoding the planar position information of node 1, first, based on the above steps, the first context information and / or second context information corresponding to node 1 is determined, and the first context information and / or second context information corresponding to node 1, as well as the preset context information, are divided into primary information and secondary information, for example, 8 bits of primary information and 10 bits of secondary information. Next, the number of right-shifted bits of the secondary information corresponding to node 1 is obtained in KDown. Since node 1 is the first point in the point cloud, the number of right-shifted bits of the secondary information corresponding to node 1 is the initial value of the right-shifted bits of the secondary information, for example, 10 bits. In this way, when the decoding end determines that the number of right-shifted bits of the secondary information corresponding to node 1 is 10 bits, the secondary information of node 1 is right-shifted by 10 bits. Since the secondary information of node 1 is a total of 10 bits, after the right shift, the first secondary information of node 1 is 0 bits. Next, the decoding end determines the first index 1 based on the main information and the first secondary information after the binary representation of node 1, and then obtains the index of the target context model corresponding to node 1 from the context model index cache ContextBuffer based on the first index 1, and then obtains the target context model corresponding to node 1 based on the index of the target context model corresponding to node 1, and then uses the target context model corresponding to the node to predict and decode the plane position information of node 1 on the i-th coordinate axis. At the same time, the number of times the first index 1 in countBuffer is added by 1, and at the same time, the number of times the first index 1 in countBuffer is compared with the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh. If the number of times the first index 1 in countBuffer is less than the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh, the secondary information partition tree is not divided.

[0552] Next, for node 2 in the point cloud, when predicting and decoding the planar position information of node 2, first, based on the above steps, the first context information and / or second context information corresponding to node 2 is determined, and the first context information and / or second context information corresponding to node 2, as well as the preset context information, are divided into primary information and secondary information, for example, 8 bits of primary information and 10 bits of secondary information. Next, the number of right-shifted bits of the secondary information corresponding to node 2 is obtained in KDown. Since the secondary information partitioning tree is not divided, the number of right-shifted bits of the secondary information corresponding to node 2 is the same as the number of right-shifted bits of the secondary information corresponding to node 1, which is the initial value of the right-shifted bits of the secondary information, for example, 10 bits. In this way, when the decoding end determines that the number of right-shifted bits of the secondary information corresponding to node 2 is 10 bits, the secondary information of node 2 is right-shifted by 10 bits. Since the secondary information of node 2 is a total of 10 bits, after the right shift, the first secondary information of node 2 is 0 bits. Next, the decoding end determines the first index 2 based on the main information and the first secondary information after the binary representation of node 2, and then obtains the index of the target context model corresponding to node 2 from the context model index cache ContextBuffer based on the first index 2, and then obtains the target context model corresponding to node 2 based on the index of the target context model corresponding to node 2, and then uses the target context model corresponding to the node to predict and decode the plane position information of node 2 on the i-th coordinate axis. At the same time, the number of first index 2 in countBuffer is increased by 1, and at the same time, the number of first index 2 in countBuffer is compared with the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh. If the number of first index 2 in countBuffer is less than the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh, the secondary information partition tree is not divided.

[0553] Assuming that the above-mentioned first index 1 is the same as the above-mentioned first index 2, and the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh is 2, it can be determined that the number of times the first index 1 in countBuffer is equal to the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh. At this time, the secondary information partition tree is divided, specifically, the first layer of the secondary information partition tree is divided into a non-full binary tree to obtain a new secondary information partition tree. At this time, the new secondary information partition tree includes 2 layers, the first layer includes 1 node, and the second layer includes 2 nodes.

[0554] At the same time, the number of right-shifted bits of the secondary information in KDown is updated to obtain the number of right-shifted bits of the secondary information corresponding to the second level of the secondary information partitioning tree. For example, the number of right-shifted bits of the secondary information corresponding to the second level is the number of right-shifted bits of the secondary information corresponding to the first level minus 1, that is, 10 bits - 1 bit = 9 bits.

[0555] Furthermore, countBuffer is set to 0.

[0556] With reference to the above steps, the accuracy of the secondary information is gradually restored, and the secondary information partitioning tree shown in FIG13 can be obtained.

[0557] In this way, when predicting and decoding the plane position information of the current node on the i-th coordinate axis in the point cloud, based on the above steps, the first context information and / or second context information corresponding to the current node is determined, and the first context information and / or second context information corresponding to the current node, as well as the preset context information, are divided into primary information and secondary information, for example, 8 bits of primary information and 10 bits of secondary information. Next, the number of secondary information right-shift bits corresponding to the last layer of the current secondary information partitioning tree is determined. As can be seen from the above, the number of secondary information right-shift bits corresponding to the last layer of the current secondary information partitioning tree (i.e., the current layer obtained by the most recent partitioning) is stored in KDown. Therefore, the decoding end can obtain the number of secondary information right-shift bits corresponding to the last layer of the current secondary information partitioning tree from KDown, and then determine the number of secondary information right-shift bits corresponding to the last layer as the number of secondary information right-shift bits corresponding to the current node.

[0558] Next, the decoding end selects the first secondary information from the secondary information represented in binary format at the current node based on the number of right-shifted bits of the secondary information corresponding to the current node.

[0559] For example, the number of right shift bits of the secondary information corresponding to the current node is n bits, so the decoding end can right shift the secondary information of the current node after binary representation by n+1 bits or n-1 bits to obtain the first secondary information.

[0560] For another example, the secondary information of the current node after binary representation is right-shifted by the number of bits of the secondary information corresponding to the current node to obtain the first secondary information. Assuming that the number of bits of the secondary information corresponding to the current node after right shift is n bits, the secondary information of the current node after binary representation is right-shifted by n bits to obtain the first secondary information.

[0561] Next, the decoding end determines a first index based on the primary information and the first secondary information after the binary representation of the current node.

[0562] The embodiment of the present application does not limit the specific method for the decoding end to determine the first index based on the main information and the first secondary information after the binary representation of the current node.

[0563] In one example, the decoding end obtains the first index corresponding to the current node based on the following formula (12):

[0564] state=ct1×(ct2>>shift) (12)

[0565] Among them, state is the first index corresponding to the current node, ct1 is the main information of the current node after binary representation, ct2 is the secondary information of the current node after binary representation, shift is the number of bits of right shift of the secondary information corresponding to the current node, and ct2>>shift is the first secondary information corresponding to the current node.

[0566] After obtaining the first index corresponding to the current node based on formula (12), the decoder obtains the context model index corresponding to the first index from the preset context model index cache, and then records the context model index as the index of the target context model. In this way, the decoder selects the target context model from the preset multiple context models based on the target context model index, and then uses the target context model to predict and decode the planar position information of the current node on the i-th coordinate axis.

[0567] In some embodiments, after the decoding end determines the index of the target context model based on the above steps, it updates the index of the target upper and lower models in the context model index cache to increase the probability of the index of the target context model.

[0568] In an embodiment of the present application, the decoding end determines the target context model based on the above steps, and also includes the steps of data updating and partitioning the secondary information partition tree.

[0569] The embodiment of the present application does not limit the specific division method of the secondary information partition tree.

[0570] In one example, each layer in the secondary information partition tree is partitioned into a non-full binary tree.

[0571] In another example, each layer in the secondary information partition tree is partitioned into a full binary tree.

[0572] In another example, some layers in the secondary information partition tree are partitioned into non-full binary trees, and some layers are partitioned into full binary trees.

[0573] The following is an introduction to the division process of the secondary information division tree.

[0574] In some embodiments, if the secondary information partition tree of the embodiment of the present application includes an incomplete binary tree layer, the method of the embodiment of the present application further includes the following step 1:

[0575] Step 1: If the last layer of the current secondary information partitioning tree is a non-full binary tree layer, and the number of occurrences of the first index in the last layer is greater than or equal to the first preset threshold corresponding to the last layer, then the last layer is binary-divided to obtain a new secondary information partitioning tree.

[0576] Based on the above steps, the decoding end determines the first index corresponding to the current node and the index of the target context model corresponding to the current node, and also determines whether to continue to divide the last layer of the current secondary information partitioning tree. Specifically, if the last layer of the current secondary information partitioning tree is a non-full binary tree, the decoding end determines whether the number of times the first index corresponding to the current node appears in the last layer (i.e., the latest layer) of the current secondary information partitioning tree is greater than or equal to the first preset threshold corresponding to the last layer. If the decoding end determines that the number of times the first index corresponding to the current node appears in the last layer (i.e., the latest layer) of the current secondary information partitioning tree is greater than or equal to the first preset threshold corresponding to the last layer, the last layer of the current secondary information partitioning tree is divided into a binary tree to obtain a new secondary information partitioning tree.

[0577] Exemplarily, the decoding end determines based on the following formula (13) to further divide the secondary information:

[0578] countBuffer[state]>=CountTimeTh[shift] (13)

[0579] Among them, countBuffer[state] represents the number of times the first index state corresponding to the current node appears in the last layer (i.e., the latest layer) of the current secondary information partitioning tree, and CountTimeTh[shift] is the first preset threshold corresponding to the last layer of the current secondary information partitioning tree.

[0580] In the embodiment of the present application, the decoding end performs binary tree partitioning on the last layer of the current secondary information partitioning tree to obtain a new secondary information partitioning tree, which includes at least the following two situations:

[0581] Case 1: If the last layer of the current secondary information partition tree is not the last non-full binary tree layer of the secondary information partition tree, then the last layer is partitioned into a non-full binary tree to obtain a new secondary information partition tree.

[0582] For example, as shown in Figure 13, assume that the secondary information partition tree includes 4 non-full binary tree layers and 2 full binary tree layers. As shown in Figure 14, assume that the last layer of the current secondary information partition tree is layer 2, that is, the secondary information is currently partitioned to layer 2. At this time, layer 2 is not the last non-full binary tree layer, because the last non-full binary tree layer is layer 4. Therefore, when layer 2 is partitioned, a non-full binary tree partition is performed on layer 2 to obtain a new secondary information partition tree. This new secondary information partition tree includes 3 layers, and these 3 layers are all non-full binary tree layers.

[0583] Case 2: If the last layer of the current secondary information partition tree is the last non-full binary tree layer of the secondary information partition tree, then the last layer is partitioned into a full binary tree to obtain a new secondary information partition tree.

[0584] For example, as shown in Figure 13, it is assumed that the secondary information partition tree includes 4 non-full binary tree layers and 2 full binary tree layers. As shown in Figure 15, it is assumed that the last layer of the current secondary information partition tree is the 4th layer, that is, the secondary information is currently divided to the 4th layer. At this time, the 4th layer is the last non-full binary tree layer. Therefore, when the 4th layer is divided, the 4th layer is divided into a full binary tree, and a new secondary information partition tree is obtained. The new secondary information partition tree includes 5 layers, and the first 4 layers of these 5 layers are non-full binary tree layers, and the last layer is a full binary tree layer.

[0585] In an embodiment of the present application, in addition to performing full binary tree division on the last layer of the current secondary information partition tree to obtain a new secondary information partition tree, it also includes a step of updating the number of secondary information right shift bits, that is, reducing the number of secondary information right shift bits corresponding to the current node by one to obtain a new number of secondary information right shift bits.

[0586] Exemplarily, the decoding end obtains the number of right-shifted bits of the new secondary information based on the following formula (14):

[0587] newShift=shift-1 (14)

[0588] Among them, shift is the number of times the secondary information corresponding to the current node is shifted right, and newShift is the number of times the new secondary information is shifted right.

[0589] Correspondingly, the calculation formula of stateUpdate after the update is shown in formula (15):

[0590] stateUpdate=ct1×(ct2>>newShift) (15)

[0591] Correspondingly, the context probability corresponding to the updated stateUpdate inherits the context of its parent node, as shown in formula (16):

[0592] ContextBuffer[stateUpdate]=ContextBuffer[state] (16)

[0593] Correspondingly, the accuracy of the secondary information corresponding to the current state is reduced, that is, KDown[state]--.

[0594] Finally, the decoder resets the occurrence counter of the current state to 0, that is, countBuffer[state]=0.

[0595] In some embodiments, if the secondary information partition tree of the embodiment of the present application includes a full binary tree layer, the method of the embodiment of the present application further includes the following steps 21 to 24:

[0596] Step 21: If the last layer of the current secondary information partition tree is a full binary tree layer, determine the number of secondary information right shift bits corresponding to the last non-full binary tree layer of the current secondary information partition tree and a first preset threshold;

[0597] Step 22: Based on the number of right shifts of the secondary information corresponding to the last non-full binary tree layer, select the second secondary information from the secondary information after binary representation of the current node;

[0598] Step 23: Determine a second index based on the primary information and the second secondary information after the binary representation of the current node;

[0599] Step 24: If the number of occurrences of the second index in the last layer is greater than or equal to the first preset threshold corresponding to the last non-full binary tree layer, perform full binary tree partitioning on the last layer to obtain a new secondary information partitioning tree.

[0600] In this embodiment, if the secondary information partitioning tree includes a non-full binary tree layer and a full binary tree layer, it is determined whether the full binary tree layer is to be further divided based on the number of right-shifted bits of the secondary information corresponding to the last non-full binary tree layer of the secondary information partitioning tree and the first preset threshold. Specifically, if the last layer of the current secondary information partitioning tree is a full binary tree layer, the number of right-shifted bits of the secondary information corresponding to the last non-full binary tree layer of the current secondary information partitioning tree and the first preset threshold are determined, and based on the number of right-shifted bits of the secondary information corresponding to the last non-full binary tree layer, the second secondary information is selected from the secondary information after the binary representation of the current node. For example, the secondary information after the binary representation of the current node is right-shifted by the number of right-shifted bits of the secondary information corresponding to the last non-full binary tree layer to obtain the second secondary information corresponding to the current node. Then, based on the main information after the binary representation of the current node and the second secondary information, the second index is determined. For example, the main information after the binary representation of the current node and the second secondary information are multiplied to determine the second index.

[0601] Then, based on the following formula (17), it is determined whether the number of occurrences of the second index in the last layer of the current secondary information partition tree is greater than or equal to the first preset threshold corresponding to the last non-full binary tree layer:

[0602] countBuffer[state]1>=CountTimeTh[shift]1 (17)

[0603] Among them, countBuffer[state]1 is the number of occurrences of the second index in the last layer of the current secondary information partitioning tree, and CountTimeTh[shift]1 is the first preset threshold corresponding to the last non-full binary tree layer.

[0604] If the number of occurrences of the second index in the last layer of the current secondary information partition tree is greater than or equal to the first preset threshold corresponding to the last non-full binary tree layer, the last layer is partitioned into a full binary tree to obtain a new secondary information partition tree.

[0605] At the same time, the decoding end updates the number of right shift bits of the secondary information, that is, subtracts one from the number of right shift bits of the secondary information corresponding to the current node to obtain a new number of right shift bits of the secondary information.

[0606] Exemplarily, the decoding end obtains the new number of right-shifted bits of the secondary information based on the following formula (14).

[0607] Correspondingly, the calculation formula of stateUpdate after the update is shown in the above formula (15).

[0608] Correspondingly, the context probability corresponding to the updated stateUpdate inherits the context from the last non-full binary tree layer.

[0609] Correspondingly, the accuracy of the secondary information corresponding to the current state is reduced, that is, KDown[state]--.

[0610] Finally, the decoder resets the occurrence counter of the current state to 0, that is, countBuffer[state]=0.

[0611] For example, as shown in Figure 13, it is assumed that the secondary information partitioning tree includes 4 non-full binary tree layers and 2 full binary tree layers. As shown in Figure 16, it is assumed that the last layer of the current secondary information partitioning tree is the 5th layer, that is, the secondary information at the current moment is divided into the 5th layer. At this time, the 5th layer is a full binary tree layer. Therefore, when judging whether to divide the 5th layer, the decoding end first determines the last non-full binary tree layer of the current secondary information partitioning tree, that is, the secondary information right shift bit a corresponding to the 4th layer and the first preset threshold b. Then, the decoding end shifts the secondary information after the binary representation of the current node right by the secondary information right shift bit a corresponding to the last non-full binary tree layer to obtain the second secondary information. Then, the decoding end multiplies the main information after the binary representation of the current node and the second secondary information to obtain the second index corresponding to the current node. Then determine whether the number of occurrences of the second index in the current last layer (i.e., the 5th layer) is greater than or equal to the first preset threshold b corresponding to the last non-full binary tree layer. If the number of occurrences of the second index in the 5th layer is greater than or equal to the first preset threshold b corresponding to the last non-full binary tree layer, perform full binary tree partitioning on the 5th layer to obtain a new secondary information partitioning tree.

[0612] In summary, the entire processing flow of Dynamic-OUBF can be as follows: Dynamic-OUBF is used as a processor, with the primary information and secondary information of the current node as input, and finally an index context of the target context model between 0 and 255 is output.

[0613] In some embodiments, in order to further reduce the amount of context information, obtaining the target context model based on the index of the target context model in the above S102-B21124 includes the following steps S102-B211241 and S102-B211242:

[0614] S102-B211241. Quantize the index of the target context model to obtain a quantized model index;

[0615] S102-B211242. Based on the quantized model index, obtain the target context model.

[0616] In this embodiment, in order to further reduce the amount of context information, the index of the target context model determined above is quantized to obtain a quantized model index, and then based on the quantized model index, the target context model is obtained from multiple preset context models.

[0617] In the embodiment of the present application, the index of the target context model is quantized, and the specific method of obtaining the quantized model index is not limited.

[0618] In one possible implementation, the index of the target context model is right-shifted by n bits to obtain a quantized model index, where n is a positive integer.

[0619] The embodiment of the present application does not limit the specific value of n.

[0620] In one example, the above n=2 bits. At this time, if quantization is not performed, the number of contexts is 256, and the index of the context model is shifted right by 2 bits, which can reduce the total number of contexts to 256 / 4=64. This can greatly reduce the number of contexts and improve the decoding efficiency of the point cloud.

[0621] In one example, if n = 4 bits, and if quantization is not performed, the number of contexts is 256, then right-shifting the context model index by 4 bits can reduce the total number of contexts to 256 / 16 = 16. This yields 3 × 16 = 48 context models for the three coordinate axes, significantly reducing the number of contexts and improving point cloud decoding efficiency.

[0622] The embodiment of the present application predictively encodes the planar position information of the current node by considering the planar structure information of the neighboring nodes, which can improve the geometric coding efficiency of the point cloud.

[0623] The following takes the geometric lossless attribute lossless test environment as an example, where the number of bits per pixel (BPP) is used as a performance indicator to measure compression efficiency. When BPP is less than 100%, it means that the encoding and decoding efficiency is improved compared to existing encoding and decoding solutions.

[0624] Table 3

[0625] Test sequence geometry information_bppegyptian_mask_vox1297.124% facade_00009_vox1296.340% facade_00015_vox1496.740%

[0626] frog_00067_vox1295.015%house_without_roof_00057_vox1296.777%shiva_00035_vox1297.763%ulb_unicor n_vox1399.722%arco_valentino_dense_vox1299.866%arco_valentino_dense_vox2099.941%egyptian_mask_v ox2099.071%facade_00009_vox2099.154%facade_00015_vox2098.989%facade_00064_vox1496.055%facade_00 064_vox2099.070% frog_00067_vox2098.862% head_00039_vox2099.230% house_without_roof_00057_vox2098. 800%landscape_00014_vox2098.710%palazzo_carignano_dense_vox1499.879%palazzo_carignano_dense_vo x2099.947%shiva_00035_vox2099.345%stanford_area_2_vox1699.590%stanford_area_2_vox2099.610%staue _klimt_vox1295.984%staue_klimt_vox2098.876%ulb_unicorn_hires_vox1598.008%ulb_unicorn_hires_vox2 099.327%ulb_unicorn_vox2099.907%citytunnel_q1mm99.066%overpass_q1mm98.886%tollbooth_q1mm98.665%

[0627] As shown in Table 3, after experimental testing, it can be seen that the point cloud decoding method provided in the embodiment of the present application can improve the compression performance of a single sequence by up to 5% on the selected test sequence set (frog_00067_vox12).

[0628] Under the condition of lossless geometry and lossless attributes, the performance of the point cloud decoding method provided by the embodiment of the present application is shown in Table 4:

[0629] Table 4

[0630]

[0631] As shown in Table 3, when the point cloud decoding method of the embodiment of the present application is applied to the Cat1-A test set, the decoding performance of the geometric information is improved by 2.1%.

[0632] Under lossy geometry and lossy attributes, the performance of the point cloud decoding method provided by the embodiment of the present application is shown in Table 5:

[0633] Table 5:

[0634]

[0635] As shown in Table 4, when the point cloud decoding method of the embodiment of the present application is applied to the Cat3-frame test set, the decoding performance of the geometric information is improved by 1.3%.

[0636] With all plane decoding enabled for all test sequences, the test performance of this application and the existing TMC13-v19 is shown in Table 6:

[0637] Table 6

[0638]

[0639] As shown in Table 6, when the point cloud decoding method of the embodiment of the present application is combined with TMC13-v19 and applied to the Cat1-A test set, the decoding performance of the geometric information is improved by 5%.

[0640] As can be seen from the above, in the embodiment of the present application, when decoding the plane position information of the node, the plane position information of the current node is predicted and encoded based on the plane structure information of the neighboring nodes, taking into account the correlation between the plane structure information of the adjacent nodes, thereby effectively improving the geometric information encoding efficiency of the point cloud. Furthermore, the embodiment of the present application uses the plane structure information of the coplanar, colinear and co-point neighboring nodes of the current node to predict the plane position information of the current node, and finally uses the Dynamic-OUBF technology to map the plane position context to a preset number (for example, 48) contexts. Under the premise of improving the decoding effect of the node's plane position information, the number of contexts is reduced, thereby saving memory space for storing context information, and further improving the decoding efficiency of the point cloud.

[0641] The point cloud decoding method provided in an embodiment of the present application determines N domain nodes of the current node when decoding the plane structure information of the current node in the current decoding frame, and predictively decodes the plane structure information of the current node based on the occupancy information of the N domain nodes. In other words, when predicting and decoding the plane structure information of the current node, the embodiment of the present application takes into account the correlation between the plane structure information of adjacent nodes, thereby effectively improving the efficiency of decoding the geometric information of the point cloud, thereby improving the predictive decoding performance of the plane structure information, and improving the decoding efficiency and performance of the point cloud.

[0642] The above takes the decoding end as an example to introduce in detail the point cloud decoding method provided in the embodiment of the present application. The following takes the encoding end as an example to introduce the point cloud encoding method provided in the embodiment of the present application.

[0643] Figure 17 is a schematic diagram of a point cloud encoding method according to an embodiment of the present application. The point cloud encoding method according to the embodiment of the present application can be implemented by the point cloud encoding device shown in Figure 3 or Figure 4A above.

[0644] As shown in FIG17 , the point cloud encoding method of the embodiment of the present application includes:

[0645] S201. Determine N domain nodes of the current node.

[0646] As can be seen from the above, a point cloud includes geometric information and attribute information, and encoding of a point cloud includes geometric encoding and attribute encoding. The embodiments of the present application relate to geometric encoding of a point cloud.

[0647] In some embodiments, the geometric information of the point cloud is also referred to as the position information of the point cloud. Therefore, the geometric encoding of the point cloud is also referred to as the position encoding of the point cloud.

[0648] In the octree-based encoding method, the encoding end constructs an octree structure of the point cloud based on the geometric information of the point cloud. As shown in Figure 9, the point cloud is enclosed by a minimum rectangular block. The bounding box is first divided into 8 nodes by the octree to obtain 8 nodes. The occupied nodes among these 8 nodes, that is, the nodes including the points, are further divided into octrees, and so on, until the division is to the voxel level, for example, to a 1X1X1 cube. The point cloud octree structure obtained by such division includes multiple layers of nodes, for example, N layers. During encoding, the occupancy information of each layer is encoded layer by layer until the voxel-level leaf nodes of the last layer are encoded. That is to say, in octree encoding, the point cloud is divided into octrees, and finally the points in the point cloud are divided into the voxel-level leaf nodes of the octree. The encoding of the point cloud is achieved by encoding the entire octree.

[0649] However, for some relatively flat nodes or nodes with planar characteristics, the coding efficiency of the point cloud geometric information can be further improved by utilizing plane coding. For example, as shown in Figure 5A, the four occupied child nodes in the current node are all located in the low plane position of the current node in the Z coordinate axis direction. At this time, the occupancy information of the current node is represented as: 11001100. When encoding the current node using plane coding, it is first necessary to encode an identifier to indicate that the current node is a plane in the Z coordinate axis direction. Secondly, if the current node is a plane in the Z coordinate axis direction, the plane position of the current node needs to be represented. Secondly, it is only necessary to encode the low plane node placeholder information in the Z coordinate axis direction (i.e., the placeholder information of the four child nodes 0246). Therefore, encoding the current node based on the plane coding method only requires encoding 6 bits, which can reduce the representation of 2 bits compared to the original octree coding, thereby improving the coding performance of the point cloud.

[0650] As can be seen from the above, when the current node is encoded using the plane coding method, the encoding end needs to perform predictive coding on the plane structure information of the current node.

[0651] Currently, the plane structure information of the current node is predictively encoded based on some prior reference information, such as the spatial distance between the nodes at the same division depth and the same coordinates as the current node, and / or the plane position of the node at the same division depth and the same coordinates as the current node, resulting in poor prediction coding performance of the plane structure information.

[0652] In order to solve the above problems, in an embodiment of the present application, the encoding end predictively encodes the plane structure of the current node based on the occupancy information of the N domain nodes of the current node, thereby improving the predictive encoding performance of the plane structure information and improving the encoding efficiency and performance of the point cloud.

[0653] The following describes the specific process of the encoding end determining the N domain nodes of the current node.

[0654] It should be noted that in the embodiment of the present application, there is no restriction on the specific method for the encoding end to determine the N domain nodes of the current node.

[0655] In one example, the N domain nodes of the current node include at least one domain node that is coplanar, colinear, and co-pointed with the current node. As shown in Figure 10, the current node includes 6 coplanar nodes, 12 colinear nodes, and 8 co-pointed nodes.

[0656] In another example, the N domain nodes of the current node may include, in addition to at least one domain node that is coplanar, colinear, and co-point with the current node, other nodes within a preset reference neighborhood range. This embodiment of the present application does not impose any restrictions on this.

[0657] In a specific embodiment, as shown in Figure 11, the thick dashed line node is the current node to be encoded, the solid line node is the three neighboring nodes coplanar with the current node, the dot-dashed line node is the three neighboring nodes colinear with the current node, and the long dashed line node is the neighboring node co-pointed with the current node. Because according to the order of point cloud encoding, when encoding the occupancy information of the current node, seven neighboring nodes that are coplanar, co-linear, and co-pointed with the current node (left front and lower direction) can be obtained. The occupancy information of at least one of these seven domain nodes is used to predict the planar structure information of the current node.

[0658] In another specific embodiment, the N domain nodes of the current node include the 7 domain nodes in Figure 11, namely, 3 domain nodes coplanar with the current node, 3 domain nodes colinear with the current node, and 1 domain node copoint with the current node.

[0659] In another specific embodiment, the N domain nodes of the current node include domain nodes that are coplanar and co-point with the current node, for example, 6 domain nodes that are coplanar with the current node and 1 domain node that is co-point with the current node.

[0660] In another specific embodiment, the N domain nodes of the current node include domain nodes that are collinear and co-pointed with the current node, for example, 6 domain nodes that are collinear with the current node and 1 domain node that is co-pointed with the current node.

[0661] In another specific embodiment, the N domain nodes of the current node include domain nodes that are colinear and coplanar with the current node, for example, 6 domain nodes that are coplanar with the current node and 6 domain nodes that are colinear with the current node.

[0662] In another specific embodiment, the N domain nodes of the current node include only domain nodes that are coplanar with the current node, or include only domain nodes that are colinear with the current node, or include only domain nodes that have a common point with the current node.

[0663] The embodiment of the present application does not limit the specific method by which the encoding end determines the N domain nodes of the current node.

[0664] S202: Based on the placeholder information of N domain nodes, predictive coding is performed on the plane structure information of the current node.

[0665] In an embodiment of the present application, the plane structure information of the current node includes the plane identification information of the current node and / or the plane position information of the current node.

[0666] From the above, we can see that the plane identification of the current node is PlaneMode i (i=0,1,2) indicates that i=0 represents the X coordinate axis, i=1 represents the Y coordinate axis, and i=2 represents the Z coordinate axis. PlaneMode i =0 means the current node is not a plane in the direction of the i-th coordinate axis. PlaneMode i =1 means that the current node is a plane in the direction of the i-th coordinate axis.

[0667] If the current node is a plane in the direction of the i-th coordinate axis, that is, PlaneMode i = 1, the encoder continues to encode the plane position information of the current node on the i-th coordinate axis. For example, using PlanePosition i Indicates the plane position information of the current node in the direction of the i-th coordinate axis, such as PlanePosition i =0 means that the current node is a plane in the direction of the i-th coordinate axis, and the plane position is the low plane, PlanePosition i =1 means that the current node is a high plane in the direction of the i-th coordinate axis.

[0668] In the embodiment of the present application, the plane structure information of the current node is predictively encoded based on the placeholder information of N domain nodes, that is, the plane identification and / or plane position information of the current node is predictively encoded.

[0669] For example, based on the occupancy information of the N domain nodes of the current node, the plane identifier of the current node on the i-th coordinate axis is predictively encoded.

[0670] For another example, based on the occupancy information of the N domain nodes of the current node, the plane position information of the current node on the i-th coordinate axis is predictively encoded.

[0671] In an embodiment of the present application, the plane structure information of the current node is predictively encoded based on the placeholder information of the N domain nodes of the current node. This can be understood as using the placeholder information of the N domain nodes of the current node as the context information of the plane structure information of the current node to predictively encode the plane structure information of the current node. For example, based on the N domain nodes of the current node, a context model index is determined, based on the context model index, a context model is determined, and based on the context model, the plane structure information of the current node is predictively encoded, for example, the plane identifier of the current node is predictively encoded based on the context model, or the plane position information of the current node is predictively encoded based on the context model.

[0672] In some embodiments, if the plane structure information of the current node includes the plane position information of the current node, the encoding end performs predictive encoding on the plane position information of the current node based on the placeholder information of the N fields of the current node. In this case, the above S202 includes steps S202-A and S202-B:

[0673] S202-A: Determine the plane structure information of the N domain nodes based on the placeholder information of the N domain nodes;

[0674] S202-B. Based on the plane structure information of the N domain nodes, predictive coding is performed on the plane position information of the current node.

[0675] In this embodiment, when predictively encoding the plane position information of the current node using the placeholder information of the N domain nodes of the current node, the encoder first determines the plane structure information of the N domain nodes, and then predictively encodes the plane position information of the current node based on the plane structure information of the N domain nodes. For example, a context model index is determined based on the plane structure information of the N domain nodes, a context model is determined based on the context model index, and predictively encoding is performed on the plane position information of the current node based on the context model.

[0676] In an embodiment of the present application, for each of the N domain nodes, the specific process of determining the plane structure information of the domain node based on the occupancy information of the domain node is consistent. For the sake of convenience of description, any domain node among the N domain nodes is used as an example for illustration.

[0677] In some embodiments, the above S202-A includes the following step S202-A1:

[0678] S202-A1. For any domain node among the N domain nodes, determine at least one of the plane identification information and the plane location information of the domain node based on the occupancy information of the domain node.

[0679] In an embodiment of the present application, the encoding end may determine the plane identification information and / or plane position information of the domain node based on the occupancy information of the domain node.

[0680] The following describes the specific process of determining the plane identification information of the domain node based on the placeholder information of the domain node.

[0681] Specifically, the encoding end determines plane0 and plane1 corresponding to the i-th coordinate axis based on the occupancy information of the domain node, and then determines the plane identification information corresponding to the domain node on the i-th coordinate axis based on plane0 and plane1.

[0682] For example, the encoder determines the plane 0 corresponding to the domain node on the X, Y, and Z coordinate axes respectively based on the following code:

[0683] uint8_t plane0 = 0;

[0684] plane0|=! ! (occupancy&0x0f)<<0;

[0685] plane0|=! ! (occupancy&0x33)<<1;

[0686] plane0|=! ! (occupancy&0x55)<<2;

[0687] Among them, occupancy represents the occupancy information of the domain node. plane0|=! !(occupancy&0x0f)<<0 represents the plane0 corresponding to the domain node on the X-coordinate axis. plane0|=! !(occupancy&0x33)<<1 represents the plane0 corresponding to the domain node on the Y-coordinate axis. plane0|=! !(occupancy&0x33)<<1 represents the plane0 corresponding to the domain node on the Y-coordinate axis. (occupancy&0x55)<<2 indicates the plane 0 corresponding to the domain node on the Z coordinate axis. 0x0f indicates 00001111. The occupancy information of the domain node is ANDed with 0x0f, and the value of the lower plane of the domain node on the X coordinate axis is 0. 0x33 indicates 00110011. The occupancy information of the domain node is ANDed with 0x33, and the value of the lower plane of the domain node on the Y coordinate axis is 0. 0x55 indicates 01010101. The occupancy information of the domain node is ANDed with 0x55, and the value of the lower plane of the domain node on the Z coordinate axis is 0.

[0688] For example, the encoder determines the plane 1 corresponding to the domain node on the X, Y, and Z coordinate axes respectively based on the following code:

[0689] uint8_t plane1 = 0;

[0690] plane1|=! ! (occupancy&0xf0)<<0;

[0691] plane1|=! ! (occupancy&0xcc)<<1;

[0692] plane1|=! ! (occupancy&0xaa)<<2;

[0693] Where occupancy represents the occupancy information of the domain node, & represents an AND operation, plane1|=! !(occupancy&0xf0)<<0 represents the plane1 corresponding to the domain node on the X-coordinate axis, plane1|=! !(occupancy&0xcc)<<1 represents the plane1 corresponding to the domain node on the Y-coordinate axis, plane1|=! ! (occupancy&0xaa)<<2 indicates plane 1 corresponding to the domain node on the Z coordinate axis. 0xf0 indicates 11110000. The occupancy information of the domain node is ANDed with 0xf0, and the value of the high plane of the domain node on the X coordinate axis is 0. 0xcc indicates 11001100. The occupancy information of the domain node is ANDed with 0xcc, and the value of the high plane of the domain node on the Y coordinate axis is 0. 0xaa indicates 10101010. The occupancy information of the domain node is ANDed with 0xaa, and the value of the high plane of the domain node on the Z coordinate axis is 0.

[0694] Based on the above method, the encoding end can determine plane0 and plane1 corresponding to the domain node on the i-th coordinate axis, and then determine the plane identification information corresponding to the domain node on the i-th coordinate axis based on plane0 and plane1.

[0695] For example, for the i-th coordinate axis, the plane 0 and plane 1 of the i-th coordinate axis determined above are XORed to determine the plane identification information of the domain node on the i-th axis. Specifically, only when a single plane perpendicular to the axis is occupied is it considered a plane.

[0696] Exemplarily, the encoding end determines the plane identification information of the domain node on the i-th axis based on the following formula (10).

[0697] The specific process of determining the plane location information of the domain node is introduced below.

[0698] In an embodiment of the present application, the encoding end can determine the plane identification information planarMode of the domain node based on the above method, and then determine the plane position information of the domain node based on the plane identification information planarMode.

[0699] Exemplarily, the encoding end determines the plane position information of the domain node on the i-th axis based on the following formula (11).

[0700] The above describes the specific process of determining the plane identification information and plane position information of the domain node on the X-coordinate axis. The specific process of determining the plane identification information and plane position information of the domain node on the Y-coordinate axis and the Z-coordinate axis can refer to the above-mentioned process of determining the plane identification information and plane position information of the X-coordinate axis, which will not be repeated here.

[0701] Based on the above steps, the encoding end determines the plane identification information and / or plane position information of each domain node in the N domain nodes, and then predictively encodes the plane position information of the current node based on the plane identification information and / or plane position information of each domain node in the N domain nodes.

[0702] The embodiment of the present application does not limit the specific method of predictively encoding the plane position information of the current node based on the plane structure information of N domain nodes in the above S202-B.

[0703] In some embodiments, the encoding end determines a context model index based on the plane structure information of N domain nodes, and then selects a context model from multiple preset context models based on the context model index, and then predictively encodes the plane position information of the current node based on the context model.

[0704] In some embodiments, the above S202-B includes the following steps:

[0705] S202-B1. Determine, based on the plane structure information of the N domain nodes, the first context information and / or the second context information corresponding to the i-th coordinate axis, where the i-th coordinate axis is the X-coordinate axis, the Y-coordinate axis, or the Z-coordinate axis;

[0706] S202-B2. Based on the first context information and / or the second context information corresponding to the i-th coordinate axis, predictively encode the plane position information of the current node on the i-th coordinate axis.

[0707] In this embodiment, the encoder determines at least one of the first context information and the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes, and then predictively encodes the plane position information of the current node on the i-th coordinate axis based on the determined first context information and / or second context information. For example, the encoder determines at least one of the first context information and the second context information corresponding to the X-coordinate axis based on the plane structure information of the N domain nodes, and then predictively encodes the plane position information of the current node on the X-coordinate axis based on the first context information and / or the second context information corresponding to the X-coordinate axis. For another example, the encoder determines at least one of the first context information and the second context information corresponding to the Y-coordinate axis based on the plane structure information of the N domain nodes, and then predictively encodes the plane position information of the current node on the Y-coordinate axis based on the first context information and / or the second context information corresponding to the Y-coordinate axis. For another example, the encoder determines at least one of the first context information and the second context information corresponding to the Z-coordinate axis based on the plane structure information of the N domain nodes, and then predictively encodes the plane position information of the current node on the Z-coordinate axis based on the first context information and / or the second context information corresponding to the Z-coordinate axis.

[0708] The following describes a specific process of determining the first context information corresponding to the i-th coordinate axis based on the planar structure information of N domain nodes at the encoding end.

[0709] It should be noted that, in the embodiment of the present application, the encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes in the following specific ways, but not limited to:

[0710] Method 1: The encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane structure information of some domain nodes among the N domain nodes.

[0711] For example, the encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane structure information of P domain nodes among the N domain nodes that are coplanar with the current node, where P is a positive integer.

[0712] The plane structure information of the P domain nodes includes the plane identification information and / or plane position information of the P domain nodes. That is, the encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane identification information of the P coplanar domain nodes. Alternatively, the encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane position information of the P coplanar domain nodes. Alternatively, the encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane identification information and plane position information of the P coplanar domain nodes.

[0713] The embodiment of the present application does not limit the specific method by which the encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane structure information of P domain nodes among N domain nodes that are coplanar with the current node.

[0714] In one possible implementation, the encoding end performs an AND operation on the planar structure information of any of the P domain nodes and the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. The encoding end then weights the first values ​​corresponding to the P domain nodes to obtain the first context information corresponding to the i-th coordinate axis. It should be noted that different coordinate axes correspond to different first preset values, and the embodiments of this application do not limit the specific values ​​of the first preset values ​​corresponding to each coordinate axis.

[0715] Exemplarily, the first preset value corresponding to the X coordinate axis is 0, the first preset value corresponding to the Y coordinate axis is 1, and the first preset value corresponding to the Z coordinate axis is 2.

[0716] As can be seen from the above, the plane structure information of a domain node includes plane identification information and / or plane position information. Therefore, in some embodiments, the encoding end performs an AND operation on the plane structure information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node, including: the encoding end performs an AND operation on the plane identification information and / or plane position information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. In other words, the encoding end performs an AND operation on the plane identification information of P coplanar domain nodes with the first preset value corresponding to the i-th coordinate axis, and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis. Alternatively, the encoding end performs an AND operation on the plane position information of P coplanar domain nodes with the first preset value corresponding to the i-th coordinate axis, and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis. Alternatively, the encoding end performs an AND operation on the plane identification information and plane position information of P coplanar domain nodes with the first preset value corresponding to the i-th coordinate axis, and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis.

[0717] For example, assume that the P domain nodes coplanar with the current node among the N domain nodes are the three coplanar domain nodes in Figure 11. The plane identification information of these three coplanar domain nodes is recorded as: coPlanarLeftPlaneMode, coPlanarFrontPlaneMode, and coPlanarBelowPlaneMode. The plane position information of these three coplanar domain nodes is recorded as: coPlanarLeftPlanePos, coPlanarFrontPlanePos, and coPlanarBelowPlanePos.

[0718] In one example, the encoder performs an AND operation on the plane identification information of the P coplanar area nodes and a first preset value corresponding to the i-th coordinate axis, and then weights the resultant values ​​to obtain the first context information Ctx1 corresponding to the i-th coordinate axis.

[0719] In one example, the encoder performs an AND operation on the planar position information of the P coplanar area nodes and a first preset value corresponding to the i-th coordinate axis, and then weights the resultant value to obtain the first context information Ctx1 corresponding to the i-th coordinate axis.

[0720] In one example, the encoder performs an AND operation on the plane identification information and plane position information of the P coplanar area nodes and a first preset value corresponding to the i-th coordinate axis, and then weights the resultant values ​​to obtain the first context information Ctx1 corresponding to the i-th coordinate axis.

[0721] The above describes a specific process in which the encoder determines the first context information corresponding to the i-th coordinate axis based on the planar structure information of P domain nodes among the N domain nodes that are coplanar with the current node.

[0722] In some embodiments, the encoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of the domain nodes in the N domain nodes that are collinear with the current node.

[0723] In some embodiments, the encoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of the domain nodes among the N domain nodes that have a common point with the current node.

[0724] In some embodiments, the encoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar and colinear with the current node.

[0725] In some embodiments, the encoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar and co-point with the current node.

[0726] In some embodiments, the encoding end may also determine the first context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is collinear and co-point with the current node.

[0727] The above describes the specific process of determining the first context information corresponding to the i-th coordinate axis by the encoder based on the plane structure information of some domain nodes among the N domain nodes.

[0728] Method 2: The encoding end determines the first context information corresponding to the i-th coordinate axis based on the first plane position information of the N domain nodes.

[0729] The first plane structure information includes plane identification information and / or plane location information of the domain node. That is, the encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane identification information of the N domain nodes. Alternatively, the encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane location information of the N domain nodes. Alternatively, the encoding end determines the first context information corresponding to the i-th coordinate axis based on the plane location information and plane identification information of the N domain nodes.

[0730] The embodiment of the present application does not limit the specific method by which the encoding end determines the first context information corresponding to the i-th coordinate axis based on the first plane position information of N domain nodes.

[0731] In one possible implementation, the encoding end performs an AND operation on the first plane position information of any of the N domain nodes with the first preset value corresponding to the i-th coordinate axis to obtain a second value corresponding to the domain node. The encoding end then weights the second values ​​corresponding to the N domain nodes to obtain the first context information corresponding to the i-th coordinate axis. It should be noted that different coordinate axes correspond to different first preset values, and the embodiments of this application do not limit the specific values ​​of the first preset values ​​corresponding to each coordinate axis.

[0732] Exemplarily, the first preset value corresponding to the X coordinate axis is 0, the first preset value corresponding to the Y coordinate axis is 1, and the first preset value corresponding to the Z coordinate axis is 2.

[0733] As can be seen from the above, the plane structure information of the domain node includes plane identification information and / or plane position information. Therefore, in some embodiments, the encoding end performs an AND operation on the first plane structure information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the second value corresponding to the domain node, including: the encoding end performs an AND operation on the plane identification information and / or plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. In other words, the encoding end performs an AND operation on the plane identification information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis. Alternatively, the encoding end performs an AND operation on the plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis. Alternatively, the encoding end performs an AND operation on the plane identification information and plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the first context information corresponding to the i-th coordinate axis.

[0734] For example, assume that the N domain nodes include the three coplanar domain nodes coPlanarLeft, coPlanarFrontPlane, and coPlanarBelow, as shown in Figure 11, as well as the three collinear domain nodes coEdgerLeft, coEdgerFront, and coEdgerBelow, and the one co-vertex domain node coVertex. Assume that the plane identification information of these seven domain nodes is recorded as: coPlanarLeftPlaneMode, coPlanarFrontPlaneMode, coPlanarBelowPlaneMode, coEdgerLeftPlanarMode, coEdgerFrontPlanarMode, coEdgerBelowPlanarMode, and coVertexPlanarMode, respectively. The plane position information of these seven domain nodes are recorded as: coPlanarLeftPlanePos, coPlanarFrontPlanePos, coPlanarBelowPlanePos, coEdgerLeftPlanePos, coEdgerFrontPlanePos, coEdgerBelowPlanePos and coVertexPlanePos.

[0735] In one example, the encoding end performs an AND operation on the plane identification information of the N domain nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the first context information Ctx1 corresponding to the i-th coordinate axis.

[0736] In one example, the encoding end performs an AND operation on the plane position information of the N domain nodes and a first preset value corresponding to the i-th coordinate axis and then weights the resultant values ​​to obtain the first context information Ctx1 corresponding to the i-th coordinate axis.

[0737] In one example, the encoding end performs an AND operation on the plane identification information and plane position information of the N domain nodes and a first preset value corresponding to the i-th coordinate axis, and then weights them to obtain the first context information Ctx1 corresponding to the i-th coordinate axis.

[0738] The above describes the specific process by which the encoder determines the first context information corresponding to the i-th coordinate axis based on the planar structure information of N domain nodes. It should be noted that in addition to determining the first context information corresponding to the i-th coordinate axis using the various methods described above, the encoder may also use other methods to determine the first context information corresponding to the i-th coordinate axis.

[0739] The following describes a specific process of determining the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes in S202-B1.

[0740] It should be noted that, in the embodiment of the present application, the encoding end determines the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes in the following specific ways, including but not limited to:

[0741] Method 1: The encoding end determines the second context information corresponding to the i-th coordinate axis based on the plane structure information of some domain nodes among the N domain nodes.

[0742] For example, the encoding end determines the second context information corresponding to the i-th coordinate axis based on the plane structure information of Q domain nodes in the N domain nodes that are collinear and / or co-point with the current node, where Q is a positive integer.

[0743] The plane structure information of the Q domain nodes includes the plane identification information and / or plane position information of the Q domain nodes. That is, the encoder determines the second context information corresponding to the i-th coordinate axis based on the plane identification information of the P coplanar domain nodes. Alternatively, the encoder determines the second context information corresponding to the i-th coordinate axis based on the plane position information of the Q domain nodes. Alternatively, the encoder determines the second context information corresponding to the i-th coordinate axis based on the plane identification information and plane position information of the Q domain nodes.

[0744] The embodiment of the present application does not limit the specific method for the encoding end to determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of Q domain nodes in N domain nodes that are collinear and / or co-point with the current node.

[0745] In one possible implementation, the encoding end performs an AND operation on the planar structure information of any of the Q domain nodes and the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. The encoding end then weights the first values ​​corresponding to the P domain nodes to obtain the second context information corresponding to the i-th coordinate axis. It should be noted that different coordinate axes correspond to different first preset values, and the embodiments of this application do not limit the specific values ​​of the first preset values ​​corresponding to each coordinate axis.

[0746] Exemplarily, the first preset value corresponding to the X coordinate axis is 0, the first preset value corresponding to the Y coordinate axis is 1, and the first preset value corresponding to the Z coordinate axis is 2.

[0747] As can be seen from the above, the plane structure information of the domain node includes plane identification information and / or plane position information. Therefore, in some embodiments, the above-mentioned encoding end performs an AND operation on the plane structure information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node, including: the encoding end performs an AND operation on the plane identification information and / or plane position information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node. In other words, the encoding end performs an AND operation on the plane identification information of Q domain nodes with the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information corresponding to the i-th coordinate axis. Alternatively, the encoding end performs an AND operation on the plane position information of Q domain nodes with the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information corresponding to the i-th coordinate axis. Alternatively, the encoding end performs an AND operation on the plane identification information and plane position information of Q domain nodes with the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information corresponding to the i-th coordinate axis.

[0748] For example, assume that among N domain nodes, the Q domain nodes that are collinear or co-point with the current node are the three collinear domain nodes coEdgerLeft, coEdgerFront, and coEdgerBelow in Figure 11, and the one domain node coVertex that is co-point with the current node. The plane identification information of these four domain nodes is recorded as: coEdgerLeftPlaneMode, coEdgerFrontPlaneMode, coEdgerBelowPlaneMode, and coVertexPlaneMode, respectively. The plane position information of these four domain nodes is recorded as: coEdgerLeftPlanePos, coEdgerFrontPlanePos, coEdgerBelowPlanePos, and coVertexPlanePos, respectively.

[0749] In one example, the encoding end performs an AND operation on the plane identification information of the Q domain nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information Ctx2 corresponding to the i-th coordinate axis.

[0750] In one example, the encoding end performs an AND operation on the plane position information of the Q domain nodes and a first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information Ctx2 corresponding to the i-th coordinate axis.

[0751] In one example, the encoding end performs an AND operation on the plane identification information and plane position information of the Q domain nodes and a first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information Ctx2 corresponding to the i-th coordinate axis.

[0752] The above describes a specific process of determining the second context information corresponding to the i-th coordinate axis on the encoder side based on the plane structure information of Q domain nodes among the N domain nodes that are collinear and / or co-point with the current node.

[0753] In some embodiments, the encoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is collinear with the current node.

[0754] In some embodiments, the encoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that has a common point with the current node.

[0755] In some embodiments, the encoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar with the current node.

[0756] In some embodiments, the encoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar and colinear with the current node.

[0757] In some embodiments, the encoding end may also determine the second context information corresponding to the i-th coordinate axis based on the plane structure information of at least one domain node among the N domain nodes that is coplanar and co-point with the current node.

[0758] The above describes the specific process of determining the second context information corresponding to the i-th coordinate axis by the encoding end based on the plane structure information of some domain nodes among the N domain nodes.

[0759] Method 2: The encoding end determines the second context information corresponding to the i-th coordinate axis based on the second plane position information of the N domain nodes.

[0760] The second plane structure information includes plane identification information and / or plane position information of the domain node. That is, the encoding end determines the second context information corresponding to the i-th coordinate axis based on the plane identification information of the N domain nodes. Alternatively, the encoding end determines the second context information corresponding to the i-th coordinate axis based on the plane position information of the N domain nodes. Alternatively, the encoding end determines the second context information corresponding to the i-th coordinate axis based on the plane position information and plane identification information of the N domain nodes.

[0761] The embodiment of the present application does not limit the specific method in which the encoding end determines the second context information corresponding to the i-th coordinate axis based on the first plane position information of N domain nodes.

[0762] In one possible implementation, the encoding end performs an AND operation on the second plane position information of any one of the N domain nodes with the first preset value corresponding to the i-th coordinate axis to obtain a third value corresponding to the domain node. The encoding end then weights the third values ​​corresponding to the N domain nodes to obtain the second context information corresponding to the i-th coordinate axis. It should be noted that different coordinate axes correspond to different first preset values, and the embodiments of this application do not limit the specific values ​​of the first preset values ​​corresponding to each coordinate axis.

[0763] Exemplarily, the first preset value corresponding to the X coordinate axis is 0, the first preset value corresponding to the Y coordinate axis is 1, and the first preset value corresponding to the Z coordinate axis is 2.

[0764] As can be seen from the above, the plane structure information of the domain node includes plane identification information and / or plane position information. Therefore, in some embodiments, the encoding end performs an AND operation on the second plane structure information of the domain node with the first preset value corresponding to the i-th coordinate axis to obtain the third value corresponding to the domain node, including: the encoding end performs an AND operation on the plane identification information and / or plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis to obtain the third value corresponding to the domain node. In other words, the encoding end performs an AND operation on the plane identification information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the second context information corresponding to the i-th coordinate axis. Alternatively, the encoding end performs an AND operation on the plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the second context information corresponding to the i-th coordinate axis. Alternatively, the encoding end performs an AND operation on the plane identification information and plane position information of N domain nodes with the first preset value corresponding to the i-th coordinate axis and then performs a weighted addition to obtain the second context information corresponding to the i-th coordinate axis.

[0765] For example, assume that the N domain nodes include the three coplanar domain nodes coPlanarLeft, coPlanarFrontPlane, and coPlanarBelow, as shown in Figure 11, as well as the three collinear domain nodes coEdgerLeft, coEdgerFront, and coEdgerBelow, and the one co-vertex domain node coVertex. Assume that the plane identification information of these seven domain nodes is recorded as: coPlanarLeftPlaneMode, coPlanarFrontPlaneMode, coPlanarBelowPlaneMode, coEdgerLeftPlanarMode, coEdgerFrontPlanarMode, coEdgerBelowPlanarMode, and coVertexPlanarMode, respectively. The plane position information of these seven domain nodes are recorded as: coPlanarLeftPlanePos, coPlanarFrontPlanePos, coPlanarBelowPlanePos, coEdgerLeftPlanePos, coEdgerFrontPlanePos, coEdgerBelowPlanePos and coVertexPlanePos.

[0766] In one example, the encoding end performs an AND operation on the plane identification information of the N domain nodes and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information Ctx2 corresponding to the i-th coordinate axis.

[0767] In one example, the encoding end performs an AND operation on the plane position information of the N-domain node and the first preset value corresponding to the i-th coordinate axis and then weights them to obtain the second context information Ctx2 corresponding to the i-th coordinate axis.

[0768] In one example, the encoding end performs an AND operation on the plane identification information and plane position information of the N domain nodes and a first preset value corresponding to the i-th coordinate axis, and then weights them to obtain the second context information Ctx2 corresponding to the i-th coordinate axis.

[0769] The above describes the specific process by which the encoder determines the second context information corresponding to the i-th coordinate axis based on the planar structure information of N domain nodes. It should be noted that in addition to determining the second context information corresponding to the i-th coordinate axis using the various methods described above, the encoder may also use other methods to determine the second context information corresponding to the i-th coordinate axis.

[0770] It should be noted that the first context information and the second context information corresponding to the i-th coordinate axis obtained by the encoding end are different. That is, the method used by the encoding end to determine the first context information corresponding to the i-th coordinate axis is different from the method used to determine the second context information corresponding to the i-th coordinate axis, and thus the first context information and the second context information obtained are also different.

[0771] In the embodiment of the present application, the specific processes of weighting the first values ​​corresponding to P domain nodes, weighting the first values ​​corresponding to Q domain nodes, weighting the second values ​​corresponding to N domain nodes, and weighting the third values ​​corresponding to N domain nodes are basically the same.

[0772] The weighting process is introduced below.

[0773] It should be noted that, for the sake of convenience of description, the target value below can be understood as the first value corresponding to the above-mentioned domain node, the second value corresponding to the domain node, or the third value corresponding to the domain node. The at least one domain node below can be understood as the above-mentioned P domain nodes, Q domain nodes, or N domain nodes. The target context information below can be understood as the above-mentioned first context information or second context information.

[0774] In some embodiments, the weight of the target value corresponding to the domain node is a preset value, so that the target value corresponding to at least one domain node can be weighted based on the weight of the target value corresponding to each domain node in at least one domain node to obtain the target context information corresponding to the i-th coordinate axis.

[0775] In some embodiments, the weighting of the target value corresponding to at least one domain node to obtain the target context information corresponding to the i-th coordinate axis includes the following steps E1 and E2:

[0776] Step E1: Determine the number of left-shifted bits corresponding to the target value, and determine the weighted weight corresponding to the target value based on the number of left-shifted bits;

[0777] Step E2: Based on the weighted weight of the target value, weight the target value corresponding to the at least one domain node to obtain the target context information corresponding to the i-th coordinate axis.

[0778] For example, determine the number of left shift bits corresponding to the above-mentioned first value, and based on the number of left shift bits, determine the weighted weight corresponding to the first value; based on the weighted weight of the first value, weight the first values ​​corresponding to the P domain nodes to obtain the first context information corresponding to the i-th coordinate axis.

[0779] For another example, determine the number of left shift bits corresponding to the above-mentioned second value, and based on the number of left shift bits, determine the weighted weight corresponding to the second value; based on the weighted weight of the second value, weight the second values ​​corresponding to the N domain nodes to obtain the first context information corresponding to the i-th coordinate axis.

[0780] For another example, determine the number of left shift bits corresponding to the above-mentioned first value, and based on the number of left shift bits, determine the weighted weight corresponding to the first value; based on the weighted weight of the first value, weight the first values ​​corresponding to the Q domain nodes to obtain the second context information corresponding to the i-th coordinate axis.

[0781] For another example, determine the number of left shift bits corresponding to the above-mentioned third value, and based on the number of left shift bits, determine the weighted weight corresponding to the third value; based on the weighted weight of the third value, weight the third values ​​corresponding to the N domain nodes to obtain the second context information corresponding to the i-th coordinate axis.

[0782] After the encoding end determines the first context information and / or second context information corresponding to the i-th coordinate axis based on the plane structure information of N domain nodes in the above manner, it executes the above S202-B2 step, and predictively encodes the plane position information of the current node on the i-th coordinate axis based on the first context information and / or second context information corresponding to the i-th coordinate axis.

[0783] The embodiment of the present application does not limit the specific method of predictively encoding the planar position information of the current node on the i-th coordinate axis based on the first context information and / or second context information corresponding to the i-th coordinate axis in S202-B2.

[0784] In some embodiments, the encoder performs predictive encoding on the plane position information of the current node on the i-th coordinate axis based only on the first context information and / or the second context information corresponding to the i-th coordinate axis. For example, the encoder determines a context model index based on the first context information and / or the second context information corresponding to the i-th coordinate axis, then selects a context model from a plurality of preset context models based on the context model index, and then uses the context model to perform predictive encoding on the plane position information of the current node on the i-th coordinate axis.

[0785] In some embodiments, the above S202-B2 includes the following steps S202-B21:

[0786] S202-B21. Based on the first context information and / or the second context information corresponding to the i-th coordinate axis and the preset context information, predictively encode the plane position information of the current node on the i-th coordinate axis.

[0787] In this embodiment, when the encoding end predictively encodes the planar position information of the current node on the i-th coordinate axis, the referenced context information includes other preset context information in addition to the first context information and / or second context information corresponding to the i-th coordinate axis.

[0788] The embodiment of the present application does not limit the specific content of the preset context information, which can be determined according to actual needs.

[0789] In one possible implementation, the preset context information includes at least one of the following four context information:

[0790] 1. Using the occupancy information of neighboring nodes to predict the plane position information of the current node, the plane position information is divided into three elements: predicted as low plane, predicted as high plane, and unpredictable;

[0791] 2. The spatial distance between the nodes at the same partition depth and the same coordinates as the current node and the current node is "close" or "far";

[0792] 3. The plane position of the node at the same partition depth and the same coordinates as the current node, if it is a plane;

[0793] 4. Coordinate dimension (i=0, 1, 2).

[0794] In this embodiment, when the encoding end predictively encodes the plane position information of the current node on the i-th coordinate axis, it determines the first context information and / or the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes of the current node, and then predictively encodes the plane position information of the current node on the i-th coordinate axis based on the first context information and / or the second context information corresponding to the i-th coordinate axis and the preset context information. It can be seen that in the embodiment of the present application, when the encoding end predictively encodes the plane position information of the current node, it not only considers the preset prior information (i.e., the preset context information), but also considers the plane structure information of the domain node (i.e., the first context information and / or the second context information), thereby improving the predictive encoding effect of the plane position information of the current node and improving the encoding efficiency of the point cloud.

[0795] The embodiment of the present application does not limit the specific process of predictive encoding of the planar position information of the current node on the i-th coordinate axis based on the first context information and / or second context information corresponding to the i-th coordinate axis, and the preset context information.

[0796] In some embodiments, the above S202-B21 includes the following steps S202-B211 and S202-B212:

[0797] S202-B211, determining a target context model based on the first context information and / or the second context information corresponding to the i-th coordinate axis, and preset context information;

[0798] S202-B212: Based on the target context model, predictively encode the plane position information of the current node on the i-th coordinate axis.

[0799] In this embodiment, the encoder determines a context model based on the first context information and / or second context information corresponding to the i-th coordinate axis, as well as the preset context information. For ease of description, this context model is referred to as the target context model. The target context model is then used to predictively encode the planar position information of the current node on the i-th coordinate axis.

[0800] The following describes a specific process in which the encoder determines the target context model based on the first context information and / or the second context information corresponding to the i-th coordinate axis, as well as the preset context information.

[0801] In some embodiments, the encoding end determines the index of the target context model based on the first context information and / or second context information corresponding to the i-th coordinate axis, and the preset context information, and then selects the target context model from the preset multiple context models based on the index of the target context model, and then uses the target context model to predict and encode the planar position information of the current node on the i-th coordinate axis.

[0802] In this embodiment, multiple context models are set for the plane position information. The embodiment of the present application does not limit the specific number of context models corresponding to the plane position information, as long as it is greater than 1. In other words, in the embodiment of the present application, an optimal context model is selected from at least two context models to perform predictive coding on the plane position information of the current node on the i-th coordinate axis.

[0803] Exemplarily, the plane position information corresponds to multiple context models as shown in Table 2. Thus, the encoder determines the index of the target context model based on the first context information and / or second context information corresponding to the i-th coordinate axis, as well as the preset context information. Then, based on the index of the target context model, the encoder selects the target context model from the corresponding context models in Table 2 to perform predictive coding on the plane position information of the current node on the i-th coordinate axis.

[0804] In some embodiments, the above S202-B211 includes the following steps S202-B2111 and S202-B2112:

[0805] S202-B2111, dividing the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information;

[0806] S202-B2112: Determine a target context model based on the primary information of the current node and part or all of the secondary information of the current node.

[0807] As can be seen from the above, assuming that the context information of the plane position information includes the first context information and the second context information corresponding to the i-th coordinate axis, as well as the above-mentioned four preset context information, the final context of the plane position is as follows:

[0808] 1. Using the occupancy information of neighboring nodes, the plane position information of the current node is predicted to be three elements: predicted to be ground plane, predicted to be high plane, and unpredictable;

[0809] 2. The spatial distance between the nodes at the same partition depth and the same coordinates as the current node and the current node is "close" or "far";

[0810] 3. The plane position of the node at the same partition depth and the same coordinates as the current node, if it is a plane;

[0811] 4. Coordinate dimension (i = 0, 1, 2);

[0812] 5. Ctx1: Planar structure information of three coplanar neighboring nodes;

[0813] 6. Ctx2: Planar structure information of three collinear neighboring nodes and one common neighboring node.

[0814] Assuming that the encoder determines the first context information corresponding to the i-th coordinate axis based on the plane identification information and plane position information of the three domain nodes coplanar with the current node among the N domain nodes, it can be obtained that Ctx1 includes 2 6 = 64 contexts. Assuming that the encoder determines the second context information corresponding to the i-th coordinate axis based on the plane identification information and plane position information of the three domain nodes that are collinear with the current node and the one domain node that is co-pointed among the N domain nodes, it can be obtained that Ctx2 includes 2 8= 256 contexts. In this way, the encoding end can obtain 3×2×2×3×64×256=589824 contexts based on the first context information and the second context information corresponding to the i-th coordinate axis, as well as the above-mentioned 4 preset context information. The memory space occupied by so many contexts is very large. Based on this, when the embodiment of the present application predictively encodes the plane position information of the node, the most advanced encoding technology Dynamic-OUBF of G-PCC is added to this algorithm to reduce the number of contexts used for encoding the plane position information, for example, reducing the number of plane position contexts to 3x16=48.

[0815] Specifically, in an embodiment of the present application, as shown in FIG12 , the encoding end divides the first context information and / or second context information corresponding to the above-determined i-th coordinate axis, and the preset context information into primary information and secondary information, and then determines the target context model based on the primary information of the current node and part or all of the secondary information of the current node. It should be noted that in an embodiment of the present application, the target context model is mainly determined based on the primary information and part of the secondary information of the current node, thereby reducing the number of contexts. This not only reduces the memory usage of the context, but also improves the prediction coding efficiency of the plane position information of the node.

[0816] The embodiment of the present application does not limit the specific manner of dividing the first context information and / or second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information.

[0817] In one example, the encoding end divides the first context information corresponding to the i-th coordinate axis, the spatial distance "near" and "far" between the node and the current node at the same division depth and the same coordinates as the current node, and the plane position of the node at the same division depth and the same coordinates as the current node, if it is a plane, into primary information. The second context information corresponding to the i-th coordinate axis and the plane position information of the current node predicted by using the occupancy information of the neighboring nodes are divided into three elements: predicted as the ground plane, predicted as the high plane, and unpredictable as secondary information, and the coordinate dimension (i=0, 1, 2) is used as the index itself and is not divided into primary information and secondary information.

[0818] In another example, the encoding end may divide the first context information and the second context information corresponding to the i-th coordinate axis into the primary information of the current node, and divide at least one of the above-mentioned four preset context information into the secondary information of the current node.

[0819] In another example, the encoder may assign the second context information corresponding to the i-th coordinate axis to the primary information of the current node, and assign the first context information corresponding to the i-th coordinate axis to the secondary information of the current node. Alternatively, based on this, at least one of the four preset context information may be assigned to the primary information of the current node, and the remaining preset context information may be assigned to the secondary information of the current node.

[0820] In another example, the encoding end can also divide the first context information and the second context information corresponding to the i-th coordinate axis into the secondary information of the current node, and divide at least one context information of the above-mentioned four preset context information into the secondary information of the current node.

[0821] It should be noted that the way in which the encoding end divides the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information includes but is not limited to the way shown above. The encoding end can adopt other ways to divide the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information.

[0822] Based on the above steps, the encoding end divides the first context information and / or second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information, and then executes the above steps S202-B2112 to determine the target context model based on the primary information of the current node and part or all of the secondary information of the current node.

[0823] The embodiment of the present application does not limit the specific manner in which the encoding end determines the target context model based on the primary information of the current node and part or all of the secondary information of the current node.

[0824] In some embodiments, the encoding end determines an index based on the main information of the current node and part of the secondary information of the current node, determines the index of the target context model based on the index, and then determines the target context model from multiple preset context models based on the index of the target context model.

[0825] In some embodiments, the above S202-B2112 includes the following steps S202-B21121 to S202-B21124:

[0826] S202-B21121, converting the primary information of the current node and the secondary information of the current node into binary representation;

[0827] S202-B21122, determine the number of right-shifted bits of the secondary information corresponding to the current node, and select the first secondary information from the secondary information of the current node after binary representation based on the number of right-shifted bits of the secondary information corresponding to the current node, where the initial number of right-shifted bits of the secondary information is the total number of bits of the binary secondary information;

[0828] S202-B21123, determining a first index based on the primary information and the first secondary information after the binary representation of the current node, and obtaining an index of the target context model from a preset context model index cache based on the first index;

[0829] S202-B21124. Obtain a target context model based on the index of the target context model.

[0830] In this embodiment, based on the above steps, the encoding end divides the first context information and / or second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information. Then, the encoding end converts the primary information and secondary information of the current node obtained by the division into binary representation.

[0831] For example, referring to the above example, it is assumed that the encoder divides the first context information corresponding to the i-th coordinate axis, the spatial distance between the node at the same division depth and the same coordinate as the current node and the current node "near" and "far", and the plane position of the node at the same division depth and the same coordinate as the current node, if it is a plane division as the main information. Assume that the first context information Ctx1 corresponding to the i-th coordinate axis includes 2 6 =64 contexts, which require 6 bits to represent when converted to binary representation. The spatial distances "near" and "far" between the nodes at the same division depth and the same coordinates as the current node and the current node include 2 contexts, which require 1 bit to represent when converted to binary representation. The plane position of the node at the same division depth and the same coordinates as the current node includes 2 contexts, which require 1 bit to represent when converted to binary representation. Therefore, in this example, when the main information of the current node is converted to binary representation, 6+1+1=8 bits are required to represent.

[0832] Similarly, if the encoder uses the second context information corresponding to the i-th coordinate axis and the neighboring node occupancy information to predict the plane position information of the current node into three elements: predicted as ground plane, predicted as high plane and unpredictable, it is divided into secondary information. Assume that the second context information Ctx2 corresponding to the i-th coordinate axis includes 2 8= 256 contexts, which require 8 bits to represent when converted to binary. Context information: Using neighboring node occupancy information, the plane position information of the current node is predicted to consist of three elements: predicted ground plane, predicted high plane, and unpredictable. These three contexts require 2 bits to represent when converted to binary. Therefore, in this example, the secondary information of the current node requires 8 + 2 = 10 bits to represent when converted to binary.

[0833] The above examples illustrate the specific process of converting the primary and secondary information of the current node into binary representation. It should be noted that the division of the primary and secondary information of the current node includes, but is not limited to, the above examples. If the primary and secondary information of the current node also include other contextual information, the method shown in the above examples can be used to convert the primary and secondary information of the current node into binary representation.

[0834] After the encoding end converts the primary information and secondary information of the current node into binary representation, it determines the number of right-shifted bits of the secondary information corresponding to the current node, and then selects the first secondary information from the secondary information after the binary representation of the current node based on the number of right-shifted bits of the secondary information. In the embodiment of the present application, the number of right-shifted bits of the secondary information corresponding to the current node can be understood as the number of secondary information selected from the secondary information of the current node to perform predictive coding on the planar position information of the current node.

[0835] The following describes how to determine the number of right-shifted bits of the secondary information corresponding to the current node.

[0836] The embodiment of the present application does not limit the specific method of determining the number of right-shift bits of the secondary information corresponding to the current node.

[0837] In some embodiments, the number of right-shifted bits of secondary information corresponding to the current node is a preset value. For example, for the nodes in the point cloud octree, a preset number of nodes corresponds to one right-shifted bit of secondary information, so that the number of right-shifted bits of secondary information corresponding to the current node can be determined. Exemplarily, the closer the node is to the root node of the octree, the larger the number of right-shifted bits of secondary information corresponding to the node. Optionally, the initial value of the right-shifted bit of secondary information is the total number of bits of binary secondary information. For example, if the current node is the root node of the octree, the number of right-shifted bits of secondary information corresponding to the current node is the above-mentioned 10 bits.

[0838] In some embodiments, determining the number of right-shifted bits of the secondary information corresponding to the current node in S202-B21122 includes the following steps S202-B211221 and S202-B211222:

[0839] S202-B211221, determining the number of right-shifted bits of the secondary information corresponding to the last layer of the current secondary information partition tree, where the secondary information partition tree is obtained by performing binary tree partitioning on the secondary information starting from the highest bit of the secondary information;

[0840] S202-B211222. Determine the number of right-shifted bits of the secondary information corresponding to the last layer as the number of right-shifted bits of the secondary information corresponding to the current node.

[0841] The following is an introduction to the division process of secondary information.

[0842] Specifically, when the encoder encodes the current point cloud, during the entire Dynamic-OUBF initialization process, assuming that the integer representation of the main information is ct1 and the integer representation of the secondary information is ct2, a context model index cache ContextBuffer is initialized, and the size of ContextBuffer is ct1×ct2. Exemplarily, referring to the above example, assuming that the main information includes 8 bits and the secondary information includes 10 bits, an 8×10 context model index cache ContextBuffer can be determined, and the context model index cache ContextBuffer stores 8×10 context model indexes. In addition, the initial probability of each context is set to 127 (i.e., 0.5).

[0843] In some embodiments, the process of recovering the secondary information accuracy is shown in FIG13 :

[0844] First, the entire secondary information context is represented in binary format. Next, the secondary information is partitioned into a binary tree, starting with the highest bit. As shown in Figure 13, above a certain level, a non-full binary tree is created, meaning the partitioning is based on the secondary information itself. However, once the MinDepth (currently set to 3) is reached, the accuracy of the secondary information is fully restored. The partitioning of secondary information is described in detail below.

[0845] Exemplarily, a countBuffer counter is initialized with a size of ct1×(ct2>>MinDepth) and is initialized to 0.

[0846] In addition, a KDown is initialized to represent the precision (i.e., the number of right shifts) of the secondary information corresponding to each first index (state). The initial value of the number of right shifts of the secondary information is the total number of bits in the binary secondary information, i.e., the maximum precision of the secondary information. For example, if the secondary information is 10 bits, the initial value of the number of right shifts of the secondary information is 10 bits.

[0847] Furthermore, a table named CountTimeTh is initialized to control the maximum number of occurrences of the first index (state) at each level of the secondary information partitioning tree. When the number of occurrences of a first index (state) exceeds the limit for that level, the low-bit precision of the secondary information is restored, and the number of occurrences of the current first index (state) is reset to zero. The context probability of the new first index (state) is restored and inherits the probability of its parent node.

[0848] Specifically, as shown in FIG13 , for the first node 1 in the point cloud, when predictively encoding the planar position information of node 1, first, based on the above steps, the first context information and / or second context information corresponding to node 1 is determined, and the first context information and / or second context information corresponding to node 1, as well as the preset context information, are divided into primary information and secondary information, for example, 8 bits of primary information and 10 bits of secondary information. Next, the number of right-shifted bits of the secondary information corresponding to node 1 is obtained in KDown. Since node 1 is the first point in the point cloud, the number of right-shifted bits of the secondary information corresponding to node 1 is the initial value of the right-shifted bits of the secondary information, for example, 10 bits. In this way, when the encoding end determines that the number of right-shifted bits of the secondary information corresponding to node 1 is 10 bits, the secondary information of node 1 is right-shifted by 10 bits. Since the secondary information of node 1 is a total of 10 bits, after the right shift, the first secondary information of node 1 is 0 bits. Next, the encoding end determines the first index 1 based on the main information and the first secondary information after the binary representation of node 1, and then obtains the index of the target context model corresponding to node 1 from the context model index cache ContextBuffer based on the first index 1, and then obtains the target context model corresponding to node 1 based on the index of the target context model corresponding to node 1, and then uses the target context model corresponding to the node to predict the planar position information of node 1 on the i-th coordinate axis. At the same time, the number of times the first index 1 in countBuffer is added by 1, and at the same time, the number of times the first index 1 in countBuffer is compared with the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh. If the number of times the first index 1 in countBuffer is less than the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh, the secondary information partition tree is not divided.

[0849] Next, for node 2 in the point cloud, when predicting the planar position information of node 2, first, based on the above steps, the first context information and / or second context information corresponding to node 2 is determined, and the first context information and / or second context information corresponding to node 2, as well as the preset context information, are divided into primary information and secondary information, for example, 8 bits of primary information and 10 bits of secondary information. Next, the number of right-shifted bits of the secondary information corresponding to node 2 is obtained in KDown. Since the secondary information partitioning tree is not divided, the number of right-shifted bits of the secondary information corresponding to node 2 is the same as the number of right-shifted bits of the secondary information corresponding to node 1, which is the initial value of the right-shifted bits of the secondary information, for example, 10 bits. In this way, when the encoding end determines that the number of right-shifted bits of the secondary information corresponding to node 2 is 10 bits, the secondary information of node 2 is right-shifted by 10 bits. Since the secondary information of node 2 is a total of 10 bits, after the right shift, the first secondary information of node 2 is 0 bits. Next, the encoding end determines the first index 2 based on the main information and the first secondary information after the binary representation of node 2, and then obtains the index of the target context model corresponding to node 2 from the context model index cache ContextBuffer based on the first index 2, and then obtains the target context model corresponding to node 2 based on the index of the target context model corresponding to node 2, and then uses the target context model corresponding to the node to predict the plane position information of node 2 on the i-th coordinate axis. At the same time, the number of first index 2 in countBuffer is increased by 1, and at the same time, the number of first index 2 in countBuffer is compared with the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh. If the number of first index 2 in countBuffer is less than the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh, the secondary information partition tree is not divided.

[0850] Assuming that the above-mentioned first index 1 is the same as the above-mentioned first index 2, and the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh is 2, it can be determined that the number of times the first index 1 in countBuffer is equal to the first preset threshold corresponding to the first layer of the secondary information partition tree stored in CountTimeTh. At this time, the secondary information partition tree is divided, specifically, the first layer of the secondary information partition tree is divided into a non-full binary tree to obtain a new secondary information partition tree. At this time, the new secondary information partition tree includes 2 layers, the first layer includes 1 node, and the second layer includes 2 nodes.

[0851] At the same time, the number of right-shifted bits of the secondary information in KDown is updated to obtain the number of right-shifted bits of the secondary information corresponding to the second level of the secondary information partitioning tree. For example, the number of right-shifted bits of the secondary information corresponding to the second level is the number of right-shifted bits of the secondary information corresponding to the first level minus 1, that is, 10 bits - 1 bit = 9 bits.

[0852] Furthermore, countBuffer is set to 0.

[0853] With reference to the above steps, the accuracy of the secondary information is gradually restored, and the secondary information partitioning tree shown in FIG13 can be obtained.

[0854] In this way, when predictive coding is performed on the plane position information of the current node on the i-th coordinate axis in the point cloud, based on the above steps, the first context information and / or second context information corresponding to the current node is determined, and the first context information and / or second context information corresponding to the current node, as well as the preset context information, are divided into primary information and secondary information, for example, into 8-bit primary information and 10-bit secondary information. Next, the number of secondary information right-shift bits corresponding to the last layer of the current secondary information partitioning tree is determined. As can be seen from the above, the number of secondary information right-shift bits corresponding to the last layer of the current secondary information partitioning tree (i.e., the current layer obtained by the most recent partitioning) is stored in KDown. Therefore, the encoding end can obtain the number of secondary information right-shift bits corresponding to the last layer of the current secondary information partitioning tree from KDown, and then determine the number of secondary information right-shift bits corresponding to the last layer as the number of secondary information right-shift bits corresponding to the current node.

[0855] Next, the encoder selects the first secondary information from the secondary information represented in binary format at the current node based on the number of right shifts of the secondary information corresponding to the current node.

[0856] For example, the number of right shift bits of the secondary information corresponding to the current node is n bits, so the encoding end can right shift the secondary information of the current node after binary representation by n+1 bits or n-1 bits to obtain the first secondary information.

[0857] For another example, the secondary information of the current node after binary representation is right-shifted by the number of bits of the secondary information corresponding to the current node to obtain the first secondary information. Assuming that the number of bits of the secondary information corresponding to the current node after right shift is n bits, the secondary information of the current node after binary representation is right-shifted by n bits to obtain the first secondary information.

[0858] Next, the encoder determines a first index based on the primary information and the first secondary information after the binary representation of the current node.

[0859] The embodiment of the present application does not limit the specific method for the encoding end to determine the first index based on the main information and the first secondary information after the binary representation of the current node.

[0860] In one example, the encoding end obtains the first index corresponding to the current node based on the above formula (12).

[0861] After obtaining the first index corresponding to the current node based on formula (12), the encoder obtains the context model index corresponding to the first index from the preset context model index cache, and then records the context model index as the index of the target context model. In this way, the encoder selects the target context model from the preset multiple context models based on the target context model index, and then uses the target context model to predictively encode the planar position information of the current node on the i-th coordinate axis.

[0862] In some embodiments, after the encoding end determines the index of the target context model based on the above steps, it updates the index of the target upper and lower models in the context model index cache to increase the probability of the index of the target context model.

[0863] In an embodiment of the present application, the encoding end determines the target context model based on the above steps, and also includes the steps of data updating and dividing the secondary information partition tree.

[0864] The embodiment of the present application does not limit the specific division method of the secondary information partition tree.

[0865] In one example, each layer in the secondary information partition tree is partitioned into a non-full binary tree.

[0866] In another example, each layer in the secondary information partition tree is partitioned into a full binary tree.

[0867] In another example, some layers in the secondary information partition tree are partitioned into non-full binary trees, and some layers are partitioned into full binary trees.

[0868] The following is an introduction to the division process of the secondary information division tree.

[0869] In some embodiments, if the secondary information partition tree of the embodiment of the present application includes an incomplete binary tree layer, the method of the embodiment of the present application further includes the following step 1:

[0870] Step 1: If the last layer of the current secondary information partitioning tree is a non-full binary tree layer, and the number of occurrences of the first index in the last layer is greater than or equal to the first preset threshold corresponding to the last layer, then the last layer is binary-divided to obtain a new secondary information partitioning tree.

[0871] Based on the above steps, the encoding end determines the first index corresponding to the current node and the index of the target context model corresponding to the current node, and also determines whether to continue to divide the last layer of the current secondary information partitioning tree. Specifically, if the last layer of the current secondary information partitioning tree is a non-full binary tree, the encoding end determines whether the number of times the first index corresponding to the current node appears in the last layer (i.e., the latest layer) of the current secondary information partitioning tree is greater than or equal to the first preset threshold corresponding to the last layer. If the encoding end determines that the number of times the first index corresponding to the current node appears in the last layer (i.e., the latest layer) of the current secondary information partitioning tree is greater than or equal to the first preset threshold corresponding to the last layer, the last layer of the current secondary information partitioning tree is divided into a binary tree to obtain a new secondary information partitioning tree.

[0872] Exemplarily, the encoding end determines based on the above formula (13) to further divide the secondary information.

[0873] In the embodiment of the present application, the encoder performs binary tree partitioning on the last layer of the current secondary information partition tree to obtain a new secondary information partition tree, which includes at least the following two situations:

[0874] Case 1: If the last layer of the current secondary information partition tree is not the last non-full binary tree layer of the secondary information partition tree, then the last layer is partitioned into a non-full binary tree to obtain a new secondary information partition tree.

[0875] Case 2: If the last layer of the current secondary information partition tree is the last non-full binary tree layer of the secondary information partition tree, then the last layer is partitioned into a full binary tree to obtain a new secondary information partition tree.

[0876] In an embodiment of the present application, in addition to performing full binary tree division on the last layer of the current secondary information partition tree to obtain a new secondary information partition tree, it also includes a step of updating the number of secondary information right shift bits, that is, reducing the number of secondary information right shift bits corresponding to the current node by one to obtain a new number of secondary information right shift bits.

[0877] Exemplarily, the encoding end obtains the new number of right-shifted bits of the secondary information based on the above formula (14).

[0878] Correspondingly, the calculation formula of stateUpdate after the update is shown in formula (15).

[0879] Correspondingly, the context probability corresponding to the updated stateUpdate inherits the context of its parent node, as shown in formula (16).

[0880] Correspondingly, the accuracy of the secondary information corresponding to the current state is reduced, that is, KDown[state]--.

[0881] Finally, the encoder resets the occurrence counter of the current state to 0, that is, countBuffer[state] = 0.

[0882] In some embodiments, if the secondary information partition tree of the embodiment of the present application includes a full binary tree layer, the method of the embodiment of the present application further includes the following steps 21 to 24:

[0883] Step 21: If the last layer of the current secondary information partition tree is a full binary tree layer, determine the number of secondary information right shift bits corresponding to the last non-full binary tree layer of the current secondary information partition tree and a first preset threshold;

[0884] Step 22: Based on the number of right shifts of the secondary information corresponding to the last non-full binary tree layer, select the second secondary information from the secondary information after binary representation of the current node;

[0885] Step 23: Determine a second index based on the primary information and the second secondary information after the binary representation of the current node;

[0886] Step 24: If the number of occurrences of the second index in the last layer is greater than or equal to the first preset threshold corresponding to the last non-full binary tree layer, perform full binary tree partitioning on the last layer to obtain a new secondary information partitioning tree.

[0887] In this embodiment, if the secondary information partitioning tree includes a non-full binary tree layer and a full binary tree layer, it is determined whether the full binary tree layer is to be further divided based on the number of right-shifted bits of the secondary information corresponding to the last non-full binary tree layer of the secondary information partitioning tree and the first preset threshold. Specifically, if the last layer of the current secondary information partitioning tree is a full binary tree layer, the number of right-shifted bits of the secondary information corresponding to the last non-full binary tree layer of the current secondary information partitioning tree and the first preset threshold are determined, and based on the number of right-shifted bits of the secondary information corresponding to the last non-full binary tree layer, the secon...

Claims

1. A point cloud decoding method, characterized in that: include: Determine N domain nodes of the current node, where N is a positive integer; Based on the placeholder information of the N domain nodes, the plane structure information of the current node is predicted and decoded.

2. The method according to claim 1, characterized in that: The plane structure information of the current node includes the plane position information of the current node, and the predictive decoding of the plane structure information of the current node based on the placeholder information of the N domain nodes includes: Based on the placeholder information of the N domain nodes, determine the plane structure information of the N domain nodes; Based on the plane structure information of the N domain nodes, the plane position information of the current node is predicted and decoded.

3. The method according to claim 2, characterized in that The determining of the plane structure information of the N domain nodes based on the placeholder information of the N domain nodes includes: For any domain node among the N domain nodes, at least one of the plane identification information and the plane location information of the domain node is determined based on the placeholder information of the domain node.

4. The method according to claim 3, characterized in that The predicting and decoding the plane position information of the current node based on the plane structure information of the N domain nodes includes: Based on the plane structure information of the N domain nodes, determine the first context information and / or the second context information corresponding to the i-th coordinate axis, where the i-th coordinate axis is the X-coordinate axis, the Y-coordinate axis or the Z-coordinate axis; Based on the first context information and / or the second context information corresponding to the i-th coordinate axis, the plane position information of the current node on the i-th coordinate axis is predicted and decoded.

5. The method according to claim 4, characterized in that The determining, based on the plane structure information of the N domain nodes, the first context information corresponding to the i-th coordinate axis includes: Based on the plane structure information of P domain nodes among the N domain nodes that are coplanar with the current node, the first context information corresponding to the i-th coordinate axis is determined, where P is a positive integer.

6. The method according to claim 5, characterized in that The determining, based on the plane structure information of P domain nodes in the N domain nodes that are coplanar with the current node, the first context information corresponding to the i-th coordinate axis includes: For any domain node among the P domain nodes, perform an AND operation on the plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain a first value corresponding to the domain node; The first values ​​corresponding to the P domain nodes are weighted to obtain the first context information corresponding to the i-th coordinate axis.

7. The method according to claim 4, characterized in that Determining the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes includes: Based on the plane structure information of Q domain nodes among the N domain nodes that are co-linear and / or co-point with the current node, the second context information corresponding to the i-th coordinate axis is determined, where Q is a positive integer.

8. The method according to claim 7, characterized in that The determining the second context information corresponding to the i-th coordinate axis based on the plane structure information of Q domain nodes in the N domain nodes that are co-linear and / or co-point with the current node includes: For any domain node among the Q domain nodes, perform an AND operation on the plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain a first value corresponding to the domain node; The first values ​​corresponding to the Q domain nodes are weighted to obtain the second context information corresponding to the i-th coordinate axis.

9. The method according to claim 6 or 8, characterized in that: The step of performing an AND operation on the plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node includes: An AND operation is performed on the plane identification information and / or the plane position information of the domain node and the first preset value to obtain a first value corresponding to the domain node.

10. The method according to claim 4, characterized in that The determining, based on the plane structure information of the N domain nodes, the first context information corresponding to the i-th coordinate axis includes: Based on the first plane structure information of the N domain nodes, the first context information corresponding to the i-th coordinate axis is determined, where the first plane structure information includes the plane identification information or the plane position information of the domain node.

11. The method according to claim 10, characterized in that The determining, based on the first plane structure information of the N domain nodes, the first context information corresponding to the i-th coordinate axis includes: For any domain node among the N domain nodes, performing an AND operation on the first plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain a second value corresponding to the domain node; The second values ​​corresponding to the N domain nodes are weighted to obtain the first context information corresponding to the i-th coordinate axis.

12. The method according to claim 4, characterized in that The determining the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes includes: Based on the second plane structure information of the N domain nodes, the second context information corresponding to the i-th coordinate axis is determined, where the second plane structure information is the plane identification information or the plane position information of the domain node.

13. The method according to claim 12, characterized in that The determining, based on the second plane structure information of the N domain nodes, the second context information corresponding to the i-th coordinate axis includes: For any domain node among the N domain nodes, performing an AND operation on the second plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain a third value corresponding to the domain node; The third values ​​corresponding to the N domain nodes are weighted to obtain the second context information corresponding to the i-th coordinate axis.

14. The method according to claim 6, 8, 11 or 13, characterized in that The method further comprises: Determine the number of left shifts corresponding to a target value, and based on the number of left shifts, determine a weighted weight corresponding to the target value, the target value being a first value corresponding to the domain node, a second value corresponding to the domain node, or a third value corresponding to the domain node; Based on the weighted weight of the target value, the target value corresponding to at least one domain node is weighted to obtain the target context information corresponding to the i-th coordinate axis, the at least one domain node is P domain nodes, Q domain nodes or N domain nodes, and the target context information is the first context information or the second context information.

15. The method according to claim 4, characterized in that The predicting and decoding the plane position information of the current node on the i-th coordinate axis based on the first context information and / or the second context information corresponding to the i-th coordinate axis includes: Based on the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information, the plane position information of the current node on the i-th coordinate axis is predicted and decoded.

16. The method according to claim 15, characterized in that The preset context information includes at least one of the following: The plane position information of the current node is predicted by using the neighboring node occupancy information to obtain three elements: predicted as a low plane, predicted as a high plane, and unpredictable; The spatial distances between the nodes at the same partition depth and the same coordinates as the current node and the current node are "near" and "far"; The plane position of the node at the same partition depth and the same coordinates as the current node, if it is a plane; Coordinate dimension (i=0, 1, 2).

17. The method according to claim 15, characterized in that The predicting and decoding of the plane position information of the current node on the i-th coordinate axis based on the first context information and / or the second context information corresponding to the i-th coordinate axis and the preset context information includes: Determine a target context model based on the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information; Based on the target context model, the plane position information of the current node on the i-th coordinate axis is predicted and decoded.

18. The method according to claim 17, characterized in that The determining the target context model based on the first context information and / or the second context information corresponding to the i-th coordinate axis and the preset context information includes: Dividing the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information; The target context model is determined based on the primary information of the current node and part or all of the secondary information of the current node.

19. The method according to claim 18, characterized in that The determining the target context model based on the primary information of the current node and part or all of the secondary information of the current node includes: Converting the primary information of the current node and the secondary information of the current node into binary representation; Determine the number of right-shifted bits of the secondary information corresponding to the current node, and select the first secondary information from the secondary information represented by the binary representation of the current node based on the number of right-shifted bits of the secondary information corresponding to the current node, wherein the initial value of the number of right-shifted bits of the secondary information is the total number of bits of the binary secondary information; Determine a first index based on the main information after the binary representation of the current node and the first secondary information, and obtain an index of the target context model from a preset context model index cache based on the first index; Based on the index of the target context model, the target context model is obtained.

20. The method according to claim 19, characterized in that The selecting the first secondary information from the secondary information represented by the binary representation of the current node based on the number of right shift bits of the secondary information corresponding to the current node comprises: The secondary information of the current node represented in binary form is right-shifted by the number of right-shifted bits of the secondary information corresponding to the current node to obtain the first secondary information.

21. The method according to claim 19, characterized in that The determining the right shift position of the secondary information corresponding to the current node includes: Determine the number of right shifts of the secondary information corresponding to the last layer of the current secondary information partition tree, wherein the secondary information partition tree is obtained by performing binary tree partitioning on the secondary information starting from the highest bit of the secondary information; The number of right-shifted bits of the secondary information corresponding to the last layer is determined as the number of right-shifted bits of the secondary information corresponding to the current node.

22. The method according to claim 21, characterized in that The method further comprises: If the last layer of the current secondary information partitioning tree is a non-full binary tree layer, and the number of occurrences of the first index in the last layer is greater than or equal to the first preset threshold corresponding to the last layer, the last layer is binary tree partitioned to obtain a new secondary information partitioning tree.

23. The method according to claim 22, characterized in that The binary tree partitioning of the last layer to obtain a new secondary information partitioning tree includes: If the last layer is not the last non-full binary tree layer of the secondary information partition tree, a non-full binary tree partition is performed on the last layer to obtain the new secondary information partition tree.

24. The method according to claim 22, characterized in that The binary tree partitioning of the last layer to obtain a new secondary information partitioning tree includes: If the last layer is the last non-full binary tree layer of the secondary information partition tree, full binary tree partitioning is performed on the last layer to obtain the new secondary information partition tree.

25. The method according to claim 20, characterized in that The method further comprises: If the last layer of the current secondary information partition tree is a full binary tree layer, determining the number of right shift bits of secondary information corresponding to the last non-full binary tree layer of the current secondary information partition tree and a first preset threshold; Selecting second secondary information from the secondary information represented by the binary representation of the current node based on the number of right shifts of the secondary information corresponding to the last non-full binary tree layer; Determine a second index based on the primary information after the binary representation of the current node and the second secondary information; If the number of occurrences of the second index in the last layer is greater than or equal to the second preset threshold corresponding to the last non-full binary tree layer, the last layer is partitioned into a full binary tree to obtain a new secondary information partition tree.

26. The method according to claim 22 or 25, characterized in that The method further comprises: The number of right-shifted bits of the secondary information corresponding to the current node is reduced by one to obtain a new number of right-shifted bits of the secondary information.

27. The method according to claim 19, characterized in that The obtaining the target context model based on the index of the target context model comprises: quantizing the index of the target context model to obtain a quantized model index; Based on the quantized model index, the target context model is obtained.

28. The method according to claim 27, characterized in that The step of quantizing the index of the target context model to obtain a quantized model index includes: The index of the target context model is right-shifted by n bits to obtain the quantized model index, where n is a positive integer.

29. The method according to claim 19, characterized in that The method further comprises: The index of the target context model in the context model index cache is updated.

30. A point cloud encoding method, characterized in that: include: Determine N domain nodes of the current node, where N is a positive integer; Based on the placeholder information of the N domain nodes, predictive encoding is performed on the plane structure information of the current node.

31. The method according to claim 30, characterized in that The plane structure information of the current node includes the plane position information of the current node, and the predictive coding of the plane structure information of the current node based on the placeholder information of the N domain nodes includes: Based on the placeholder information of the N domain nodes, determine the plane structure information of the N domain nodes; Based on the plane structure information of the N domain nodes, the plane position information of the current node is predictively encoded.

32. The method according to claim 31, characterized in that The determining of the plane structure information of the N domain nodes based on the placeholder information of the N domain nodes includes: For any domain node among the N domain nodes, at least one of the plane identification information and the plane location information of the domain node is determined based on the placeholder information of the domain node.

33. The method according to claim 32, characterized in that The predictive coding of the plane position information of the current node based on the plane structure information of the N domain nodes includes: Based on the plane structure information of the N domain nodes, determine the first context information and / or the second context information corresponding to the i-th coordinate axis, where the i-th coordinate axis is the X-coordinate axis, the Y-coordinate axis or the Z-coordinate axis; Based on the first context information and / or the second context information corresponding to the i-th coordinate axis, the plane position information of the current node on the i-th coordinate axis is predictively encoded.

34. The method according to claim 33, characterized in that The determining, based on the plane structure information of the N domain nodes, the first context information corresponding to the i-th coordinate axis includes: Based on the plane structure information of P domain nodes among the N domain nodes that are coplanar with the current node, the first context information corresponding to the i-th coordinate axis is determined, where P is a positive integer.

35. The method according to claim 34, characterized in that The determining, based on the plane structure information of P domain nodes in the N domain nodes that are coplanar with the current node, the first context information corresponding to the i-th coordinate axis includes: For any domain node among the P domain nodes, perform an AND operation on the plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain a first value corresponding to the domain node; The first values ​​corresponding to the P domain nodes are weighted to obtain the first context information corresponding to the i-th coordinate axis.

36. The method according to claim 33, characterized in that Determining the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes includes: Based on the plane structure information of Q domain nodes among the N domain nodes that are co-linear and / or co-point with the current node, the second context information corresponding to the i-th coordinate axis is determined, where Q is a positive integer.

37. The method according to claim 36, characterized in that The determining the second context information corresponding to the i-th coordinate axis based on the plane structure information of Q domain nodes in the N domain nodes that are co-linear and / or co-point with the current node includes: For any domain node among the Q domain nodes, perform an AND operation on the plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain a first value corresponding to the domain node; The first values ​​corresponding to the Q domain nodes are weighted to obtain the second context information corresponding to the i-th coordinate axis.

38. The method according to claim 35 or 37, characterized in that The step of performing an AND operation on the plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain the first value corresponding to the domain node includes: An AND operation is performed on the plane identification information and / or the plane position information of the domain node and the first preset value to obtain a first value corresponding to the domain node.

39. The method according to claim 33, characterized in that The determining, based on the plane structure information of the N domain nodes, the first context information corresponding to the i-th coordinate axis includes: Based on the first plane structure information of the N domain nodes, the first context information corresponding to the i-th coordinate axis is determined, where the first plane structure information includes the plane identification information or the plane position information of the domain node.

40. The method according to claim 39, characterized in that The determining, based on the first plane structure information of the N domain nodes, the first context information corresponding to the i-th coordinate axis includes: For any domain node among the N domain nodes, performing an AND operation on the first plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain a second value corresponding to the domain node; The second values ​​corresponding to the N domain nodes are weighted to obtain the first context information corresponding to the i-th coordinate axis.

41. The method according to claim 33, characterized in that The determining the second context information corresponding to the i-th coordinate axis based on the plane structure information of the N domain nodes includes: Based on the second plane structure information of the N domain nodes, the second context information corresponding to the i-th coordinate axis is determined, where the second plane structure information is the plane identification information or the plane position information of the domain node.

42. The method according to claim 41, characterized in that The determining, based on the second plane structure information of the N domain nodes, the second context information corresponding to the i-th coordinate axis includes: For any domain node among the N domain nodes, performing an AND operation on the second plane structure information of the domain node and the first preset value corresponding to the i-th coordinate axis to obtain a third value corresponding to the domain node; The third values ​​corresponding to the N domain nodes are weighted to obtain the second context information corresponding to the i-th coordinate axis.

43. The method of claim 35, 37, 40 or 42, wherein: The method further comprises: Determine the number of left shifts corresponding to a target value, and based on the number of left shifts, determine a weighted weight corresponding to the target value, the target value being a first value corresponding to the domain node, a second value corresponding to the domain node, or a third value corresponding to the domain node; Based on the weighted weight of the target value, the target value corresponding to at least one domain node is weighted to obtain the target context information corresponding to the i-th coordinate axis, the at least one domain node is P domain nodes, Q domain nodes or N domain nodes, and the target context information is the first context information or the second context information.

44. The method according to claim 33, characterized in that The predictive encoding of the plane position information of the current node on the i-th coordinate axis based on the first context information and / or the second context information corresponding to the i-th coordinate axis includes: Based on the first context information and / or the second context information corresponding to the i-th coordinate axis, and preset context information, predictive encoding is performed on the plane position information of the current node on the i-th coordinate axis.

45. The method according to claim 44, characterized in that The preset context information includes at least one of the following: The plane position information of the current node is predicted by using the neighboring node occupancy information to obtain three elements: predicted as a low plane, predicted as a high plane, and unpredictable; The spatial distances between the nodes at the same partition depth and the same coordinates as the current node and the current node are "near" and "far"; The plane position of the node at the same partition depth and the same coordinates as the current node, if it is a plane; Coordinate dimension (i=0, 1, 2).

46. ​​The method according to claim 44, characterized in that The predictive encoding of the plane position information of the current node on the i-th coordinate axis based on the first context information and / or the second context information corresponding to the i-th coordinate axis and the preset context information includes: Determine a target context model based on the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information; Based on the target context model, predictive encoding is performed on the plane position information of the current node on the i-th coordinate axis.

47. The method according to claim 46, characterized in that The determining the target context model based on the first context information and / or the second context information corresponding to the i-th coordinate axis and the preset context information includes: Dividing the first context information and / or the second context information corresponding to the i-th coordinate axis, and the preset context information into primary information and secondary information; The target context model is determined based on the primary information of the current node and part or all of the secondary information of the current node.

48. The method according to claim 47, characterized in that The determining the target context model based on the primary information of the current node and part or all of the secondary information of the current node includes: Converting the primary information of the current node and the secondary information of the current node into binary representation; Determine the number of right-shifted bits of the secondary information corresponding to the current node, and select the first secondary information from the secondary information represented by the binary representation of the current node based on the number of right-shifted bits of the secondary information corresponding to the current node, wherein the initial value of the number of right-shifted bits of the secondary information is the total number of bits of the binary secondary information; Determine a first index based on the main information after the binary representation of the current node and the first secondary information, and obtain an index of the target context model from a preset context model index cache based on the first index; Based on the index of the target context model, the target context model is obtained.

49. The method according to claim 48, characterized in that The selecting the first secondary information from the secondary information represented by the binary representation of the current node based on the number of right shift bits of the secondary information corresponding to the current node comprises: The secondary information of the current node represented in binary form is right-shifted by the number of right-shifted bits of the secondary information corresponding to the current node to obtain the first secondary information.

50. The method according to claim 48, characterized in that The determining the right shift position of the secondary information corresponding to the current node includes: Determine the number of right shifts of the secondary information corresponding to the last layer of the current secondary information partition tree, wherein the secondary information partition tree is obtained by performing binary tree partitioning on the secondary information starting from the highest bit of the secondary information; The number of right-shifted bits of the secondary information corresponding to the last layer is determined as the number of right-shifted bits of the secondary information corresponding to the current node.

51. The method according to claim 50, characterized in that The method further comprises: If the last layer of the current secondary information partitioning tree is a non-full binary tree layer, and the number of occurrences of the first index in the last layer is greater than or equal to the first preset threshold corresponding to the last layer, the last layer is binary tree partitioned to obtain a new secondary information partitioning tree.

52. The method according to claim 51, characterized in that The binary tree partitioning of the last layer to obtain a new secondary information partitioning tree includes: If the last layer is not the last non-full binary tree layer of the secondary information partition tree, a non-full binary tree partition is performed on the last layer to obtain the new secondary information partition tree.

53. The method according to claim 51, characterized in that The binary tree partitioning of the last layer to obtain a new secondary information partitioning tree includes: If the last layer is the last non-full binary tree layer of the secondary information partition tree, full binary tree partitioning is performed on the last layer to obtain the new secondary information partition tree.

54. The method according to claim 50, characterized in that The method further comprises: If the last layer of the current secondary information partition tree is a full binary tree layer, determining the number of right shift bits of secondary information corresponding to the last non-full binary tree layer of the current secondary information partition tree and a first preset threshold; Selecting second secondary information from the secondary information represented by the binary representation of the current node based on the number of right shifts of the secondary information corresponding to the last non-full binary tree layer; Determine a second index based on the primary information after the binary representation of the current node and the second secondary information; If the number of occurrences of the second index in the last layer is greater than or equal to the second preset threshold corresponding to the last non-full binary tree layer, the last layer is partitioned into a full binary tree to obtain a new secondary information partition tree.

55. The method according to claim 51 or 54, characterized in that The method further comprises: The number of right-shifted bits of the secondary information corresponding to the current node is reduced by one to obtain a new number of right-shifted bits of the secondary information.

56. The method of claim 48, wherein: The obtaining the target context model based on the index of the target context model comprises: quantizing the index of the target context model to obtain a quantized model index; Based on the quantized model index, the target context model is obtained.

57. The method according to claim 56, characterized in that The step of quantizing the index of the target context model to obtain a quantized model index includes: The index of the target context model is right-shifted by n bits to obtain the quantized model index, where n is a positive integer.

58. The method of claim 48, wherein: The method further comprises: The index of the target context model in the context model index cache is updated.

59. A point cloud decoding device, characterized in that: include: A determination unit, used to determine N domain nodes of the current node, where N is a positive integer; A decoding unit is used to predict and decode the plane structure information of the current node based on the occupancy information of the N domain nodes.

60. A point cloud encoding device, characterized in that: include: A determination unit, used to determine N domain nodes of the current node, where N is a positive integer; The encoding unit is used to predict and encode the plane structure information of the current node based on the placeholder information of the N domain nodes.

61. An electronic device, characterized in that: include: Processor and memory; The memory is used to store computer programs; The processor is used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 29 or 30 to 58.

62. A computer-readable storage medium, characterized in that Used to store a computer program, the computer program causing a computer to execute the method according to any one of claims 1 to 29 or 30 to 58.