Point cloud encoding method, point cloud decoding method, codec, code stream, and storage medium
By determining whether to parse the second identification information based on the first identification information in the point cloud encoding and decoding framework, the identification information conflict problem in the bidirectional inter-frame prediction of attribute information is solved, and the encoding and decoding efficiency is improved and the encoding bits are saved.
Patent Information
- Application Number
- PCT/CN2024/070875
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-10
AI Technical Summary
In the point cloud codec framework, during the bidirectional inter-frame prediction process of attribute information, the identification information at different levels may indicate conflicts, resulting in decoding errors, and improper use of encoding bits leads to waste.
By analyzing the first identification information, determine whether to enable bidirectional inter prediction for at least one frame of image, and without turning on bidirectional inter prediction, the second identification information is not parsed or encoded, avoid identification information conflicts and save encoding bits.
It effectively avoids decoding errors caused by identification information conflicts, saves encoding bits, and improves encoding and decoding efficiency.
Smart Images

Figure CN2024070875_10072025_PF_FP_ABST
Abstract
Description
Point cloud encoding and decoding methods, codecs, bitstreams, and storage media Technical Field
[0001] The present application relates to the field of point cloud encoding and decoding technology, and in particular to a point cloud encoding and decoding method, codec, bit stream and storage medium. Background Art
[0002] In the point cloud codec framework, when predicting and decoding point cloud attribute information, a decision is made as to whether to enable bidirectional inter-frame prediction. In related technologies, when bidirectional inter-frame prediction is used for attribute information, conflicting indications at different levels of identification information (e.g., sequence-level identification information and slice-level identification information) can result in decoding errors.
[0003] Summary of the Invention
[0004] The present invention provides a point cloud encoding and decoding method, codec, code stream, and storage medium. The following describes various aspects of the present invention.
[0005] In a first aspect, a point cloud decoding method is provided, which is applied to a decoder, including: parsing first identification information, the first identification information is used to indicate whether bidirectional inter-frame prediction is enabled for at least one frame of image; if the first identification information indicates that bidirectional inter-frame prediction is enabled for at least one frame of image, parsing second identification information, the second identification information is used to indicate whether bidirectional inter-frame prediction is enabled for attribute information of a first image area, where the first image area is an image area in a frame of image.
[0006] In a second aspect, a point cloud encoding method is provided, which is applied to an encoder, including: determining whether bidirectional inter-frame prediction is enabled for attribute information of at least one frame of image; if bidirectional inter-frame prediction is enabled for attribute information of at least one frame of image, encoding second identification information, the second identification information being used to indicate whether bidirectional inter-frame prediction is enabled for attribute information of a first image area, where the first image area is an image area in a frame of image.
[0007] According to a third aspect, a decoder is provided, comprising: a first decoding unit configured to parse first identification information, the first identification information being used to indicate whether bidirectional inter-frame prediction is enabled for at least one frame of image; and a second decoding unit configured to parse second identification information if the first identification information indicates that bidirectional inter-frame prediction is enabled for at least one frame of image, the second identification information being used to indicate whether bidirectional inter-frame prediction is enabled for attribute information of a first image area, the first image area being an image area in a frame of image.
[0008] In a fourth aspect, a decoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method of the first aspect when running the computer program.
[0009] In a fifth aspect, an encoder is provided, comprising: a first encoding unit, configured to determine whether bidirectional inter-frame prediction is enabled for attribute information of at least one frame of image; a second encoding unit, configured to encode second identification information if bidirectional inter-frame prediction is enabled for attribute information of at least one frame of image, the second identification information being used to indicate whether bidirectional inter-frame prediction is enabled for attribute information of a first image area, the first image area being an image area in a frame of image.
[0010] In a sixth aspect, an encoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method of the second aspect when running the computer program.
[0011] In a seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the method of the first aspect or the second aspect is implemented.
[0012] In an eighth aspect, a non-volatile computer-readable storage medium for storing a bit stream is provided, wherein the bit stream is generated by an encoding method using an encoder, or the bit stream is decoded by a decoding method using a decoder, wherein the decoding method is the method of the first aspect and the encoding method is the method of the second aspect.
[0013] According to a ninth aspect, a code stream is provided, comprising a code stream generated according to the method of the second aspect.
[0014] In the related art, when the decoding end decodes the point cloud attribute information, it determines whether to enable bidirectional inter-frame prediction for the attribute information of the first image area based on the second identification information, but does not determine whether to enable bidirectional inter-frame prediction for the attribute information of the first image area based on the first identification information. Compared with the second identification information, the first identification information is a higher-level identification information. If the first identification information indicates that bidirectional inter-frame prediction is not enabled, the decoding end will not cache the bidirectional reference frame used for bidirectional inter-frame prediction. However, in this case, if the second identification information indicates that bidirectional inter-frame prediction is enabled for the attribute information of the first image area, the decoding end will encounter the problem of being unable to decode. In response to the above problem, the embodiment of the present application determines whether to parse the second identification information based on the first identification information, thereby avoiding the problem of being unable to decode caused by the conflict between the two identification information. In addition, the decoding end determines whether to parse the second identification information based on the first identification information, which means that when the first identification information indicates that bidirectional inter-frame prediction is not enabled, the encoding end can not write the second identification information into the code stream, thereby saving coding bits. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIG1A is a schematic diagram of a three-dimensional point cloud image.
[0016] FIG1B is a partially enlarged view of a three-dimensional point cloud image.
[0017] FIG2A is a schematic diagram of six viewing angles of a point cloud image.
[0018] FIG2B is a schematic diagram of a data storage format corresponding to a point cloud image.
[0019] FIG3 is a schematic diagram of a network architecture for point cloud encoding and decoding.
[0020] FIG4A is a schematic diagram showing the composition framework of a geometry-based point cloud compression (G-PCC) encoder.
[0021] FIG4B is a schematic diagram showing a composition framework of a G-PCC decoder.
[0022] FIG5A is a schematic diagram of a low plane position in the Z-axis direction.
[0023] FIG5B is a schematic diagram of a high plane position in the Z-axis direction.
[0024] FIG6 is a schematic diagram of a node encoding sequence.
[0025] FIG. 7A is a schematic diagram of plane identification information.
[0026] FIG. 7B is a schematic diagram of another type of planar identification information.
[0027] FIG8 is a schematic diagram of sibling nodes of a current node.
[0028] FIG9 is a schematic diagram of the intersection of a laser radar and a node.
[0029] FIG10 is a schematic diagram of neighborhood nodes at the same partition depth and the same coordinates.
[0030] FIG11 is a schematic diagram showing a current node being located at a low plane position of a parent node.
[0031] FIG12 is a schematic diagram showing a current node being located at a high plane position of a parent node.
[0032] FIG13 is a schematic diagram of predictive coding of planar position information of a laser radar point cloud.
[0033] FIG14 is a schematic diagram of IDCM encoding.
[0034] FIG15 is a schematic diagram of coordinate transformation for obtaining a point cloud using a rotating laser radar.
[0035] FIG16 is a schematic diagram of predictive coding in the X-axis or Y-axis direction.
[0036] FIG. 17A is a schematic diagram showing an angle of the X-plane predicted by the horizontal azimuth angle.
[0037] FIG17B is a schematic diagram showing an angle of the Y plane predicted by the horizontal azimuth angle.
[0038] FIG18 is another schematic diagram of predictive coding in the X-axis or Y-axis direction.
[0039] FIG. 19A is a schematic diagram showing three intersection points included in a sub-block.
[0040] FIG19B is a schematic diagram of a triangular facet set fitted using three intersection points.
[0041] FIG19C is a schematic diagram of upsampling of a triangle face set.
[0042] FIG20 is a schematic diagram of a distance-based level of detail (LOD) construction.
[0043] FIG21 is a schematic diagram of a distance-based LOD point cloud generation process.
[0044] FIG22 is a schematic diagram of the encoding process of attribute information of an LOD point cloud.
[0045] FIG23 is a schematic diagram of the structure of a refinement layer based on LOD division.
[0046] FIG24 is a schematic diagram of an inter-layer nearest neighbor search based on LOD.
[0047] FIG25 is a schematic diagram showing the spatial relationship between a child block and a parent block.
[0048] FIG26 is a schematic diagram of a neighbor block that is coplanar, colinear, and co-point with the current parent block.
[0049] FIG27 is a schematic diagram of a method for performing nearest neighbor search for a current point.
[0050] FIG28 is a schematic diagram of a method for searching the nearest neighbor within an attribute information layer.
[0051] FIG29 is a schematic diagram of a method for performing a quick search within an LOD layer.
[0052] FIG30 is a schematic diagram of a neighborhood search prediction structure based on Morton code.
[0053] FIG31 is a schematic diagram of an encoding process of a lifting transformation.
[0054] FIG32 is an example diagram of a region adaptive hierarchical transform (RAHT) transformation process.
[0055] FIG33 is another example diagram of the RAHT transformation process.
[0056] FIG34 is a schematic diagram of RAHT transformation and inverse RAHT transformation.
[0057] FIG35 is a schematic diagram of the coding block structure of attribute information.
[0058] FIG36 is a schematic diagram of the overall process of RAHT intra-frame prediction transform coding of attribute information.
[0059] FIG37 is an example diagram of a linear fitting method for the neighborhood attribute information of the current block.
[0060] FIG38 is a schematic diagram of RAHT intra-frame prediction transform coding of attribute information.
[0061] Figure 39 is a flow chart of the decoding method provided in an embodiment of the present application.
[0062] Figure 40 is a flow chart of the encoding method provided in an embodiment of the present application.
[0063] Figure 41 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application.
[0064] Figure 42 is a schematic diagram of the structure of a decoder provided in another embodiment of the present application.
[0065] Figure 43 is a schematic diagram of the structure of the encoder provided in one embodiment of the present application.
[0066] Figure 44 is a schematic structural diagram of an encoder provided in another embodiment of the present application. DETAILED DESCRIPTION
[0067] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.
[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0069] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0070] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0071] Point cloud is a three-dimensional representation of the surface of an object. Point cloud (data) of the surface of an object can be collected through acquisition equipment such as photoelectric radar, lidar, laser scanner, and multi-view camera.
[0072] A point cloud is a set of irregularly distributed discrete points in space that express the spatial structure and surface properties of a three-dimensional object or scene. Figure 1A shows a three-dimensional point cloud image and Figure 1B shows a partially enlarged view of the three-dimensional point cloud image. It can be seen that the point cloud surface is composed of densely distributed points.
[0073] In a two-dimensional image, each pixel contains information and is distributed regularly, so there's no need to record its location. However, the distribution of points in a point cloud in three-dimensional space is random and irregular, so recording the location of each point in space is necessary to fully represent the point cloud. Similar to a two-dimensional image, each location in the acquisition process has corresponding attribute information, typically an RGB color value, which reflects the object's color. For a point cloud, in addition to color information, each point's corresponding attribute information often includes reflectance values, which reflect the surface texture of the object. Therefore, point cloud data typically includes both point location information and point attribute information. Point location information can also be referred to as point geometric information. For example, point geometric information can be the point's three-dimensional coordinates (x, y, z). Point attribute information can include color information and / or reflectance. For example, reflectance can be one-dimensional reflectance information (r). Color information can be information in any color space, or it can be three-dimensional color information, such as RGB. Here, R represents red (red), G represents green (green), and B represents blue (blue). For another example, the color information may be luminance and chrominance (YCbCr, YUV) information, where Y represents brightness (luma), Cb (U) represents blue color difference, and Cr (V) represents red color difference.
[0074] For example, a point cloud generated using laser measurement principles can include both its 3D coordinate information and its reflectivity. For another example, a point cloud generated using photogrammetry principles can include both its 3D coordinate information and its 3D color information. For another example, a point cloud generated using a combination of laser measurement and photogrammetry principles can include both its 3D coordinate information, its reflectivity value, and its 3D color information.
[0075] Figures 2A and 2B show a point cloud image and its corresponding data storage format. Figure 2A provides six viewing angles of the point cloud image, while Figure 2B consists of a file header and data. The header includes the data format, data representation type, the total number of points in the point cloud, and the content represented by the point cloud. For example, the point cloud is in ".ply" format, represented by ASCII code, with a total of 207,242 points. Each point has 3D coordinate information (x, y, z) and 3D color information (r, g, b).
[0076] Point clouds can be divided into the following categories according to the acquisition method:
[0077] Static point cloud: the object is stationary and the device that acquires the point cloud is also stationary;
[0078] Dynamic point cloud: The object is moving, but the device that obtains the point cloud is stationary;
[0079] Dynamic point cloud acquisition: The device used to acquire the point cloud is in motion.
[0080] For example, point clouds can be divided into two categories according to their usage:
[0081] Category 1: Machine perception point cloud, which can be used in scenarios such as autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, and disaster relief robots;
[0082] Category 2: Human eye perception point cloud, which can be used in point cloud application scenarios such as digital cultural heritage, free viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.
[0083] Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes. Moreover, since point clouds are obtained by directly sampling real objects, they can provide a strong sense of reality while ensuring accuracy. Therefore, they are widely used, including virtual reality games, computer-aided design, geographic information systems, automatic navigation systems, digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive remote presentation, and three-dimensional reconstruction of biological tissues and organs.
[0084] Point clouds are primarily collected through computer generation, 3D laser scanning, and 3D photogrammetry. Computers can generate point clouds of virtual 3D objects and scenes; 3D laser scanning can obtain point clouds of static real-world 3D objects or scenes, generating millions of point clouds per second; and 3D photogrammetry can obtain point clouds of dynamic real-world 3D objects or scenes, generating tens of millions of point clouds per second. These technologies reduce the cost and time required to acquire point cloud data while improving data accuracy. While changes in point cloud data acquisition methods have made it possible to acquire large amounts of point cloud data, the processing of this massive amount of 3D point cloud data is facing bottlenecks due to storage space and transmission bandwidth constraints, as application demands grow.
[0085] For example, taking a point cloud video with a frame rate of 30 frames per second (fps), each frame contains 700,000 points, and each point has coordinate information (xyz, float) and color information (RGB, uchar). Therefore, the data volume of a 10-second point cloud video is approximately 0.7 million × (4 bytes × 3 + 1 byte × 3) × 30 fps × 10 seconds = 3.15 GB, where 1 byte is 10 bits. For a 1280 × 720 2D video with a YUV sampling format of 4:2:0 and a frame rate of 24 fps, the data volume for 10 seconds is approximately 1280 × 720 × 12 bits × 24 fps × 10 seconds ≈ 0.33 GB. The data volume of a 10-second two-view 3D video is approximately 0.33 × 2 = 0.66 GB. This shows that the data volume of a point cloud video far exceeds that of 2D and 3D videos of the same length. Therefore, in order to better realize data management, save server storage space, and reduce the transmission traffic and transmission time between the server and the client, point cloud compression has become a key issue in promoting the development of the point cloud industry.
[0086] That is to say, since the point cloud is a collection of massive points, storing the point cloud not only consumes a lot of memory, but is also not conducive to transmission. There is also not enough bandwidth to support direct transmission of the point cloud at the network layer without compression. Therefore, the point cloud needs to be compressed.
[0087] Currently, point cloud coding frameworks that can compress point clouds can be the geometry-based point cloud compression (G-PCC) codec framework or the video-based point cloud compression (V-PCC) codec framework provided by the Moving Picture Experts Group (MPEG), or the AVS-PCC codec framework provided by AVS. The G-PCC codec framework can be used to compress the first type of static point clouds and the third type of dynamically acquired point clouds, and it can be based on the point cloud compression test platform (test model compression 13, TMC13). The V-PCC codec framework can be used to compress the second type of dynamic point clouds, and it can be based on the point cloud compression test platform (test model compression 2, TMC2). Therefore, the G-PCC codec framework is also called the point cloud codec TMC13, and the V-PCC codec framework is also called the point cloud codec TMC2.
[0088] An embodiment of the present application provides a network architecture of a point cloud encoding and decoding system including a decoding method and an encoding method. FIG3 is a schematic diagram of a network architecture of a point cloud encoding and decoding system provided by an embodiment of the present application. As shown in FIG3 , the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01. During the implementation process, the electronic device can be various types of devices with point cloud encoding and decoding functions. For example, the electronic device can include a mobile phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensor device, a server, etc., which is not limited by the embodiment of the present application. Among them, the decoder or encoder in the embodiment of the present application can be the above-mentioned electronic device.
[0089] Among them, the electronic device in the embodiment of the present application has a point cloud encoding and decoding function, generally including a point cloud encoder (ie, encoder) and a point cloud decoder (ie, decoder).
[0090] The following describes the related technologies using the G-PCC codec framework and the AVS codec framework as examples.
[0091] As you can understand, in the point cloud G-PCC codec framework, the point cloud data to be encoded is first divided into multiple slices through slice partitioning. In each slice, the geometric information of the point cloud and the attribute information corresponding to each point are encoded separately.
[0092] Figure 4A shows a schematic diagram of the G-PCC encoder's architecture. As shown in Figure 4A, during the geometry encoding process, the geometric information is transformed so that the entire point cloud is contained within a bounding box. Quantization is then performed. This quantization step primarily serves a scaling purpose. Due to quantization rounding, the geometric information of some point clouds becomes identical. Parameters are then used to determine whether to remove duplicate points. This process of quantization and removing duplicate points is also known as voxelization. The bounding box is then partitioned into an octree or constructed as a prediction tree. During this process, arithmetic coding is performed on the points within the leaf nodes of the partition to generate a binary geometry bitstream. Alternatively, arithmetic coding is performed on the intersections (vertex) generated by the partition (surface fitting is performed based on the intersections) to generate a binary geometry bitstream. During the attribute encoding process, after the geometry encoding is completed and the geometric information is reconstructed, color conversion is performed to convert the color information (i.e., attribute information) from the RGB color space to the YUV color space. The reconstructed geometry information is then used to recolor the point cloud so that the unencoded attribute information corresponds to the reconstructed geometry information. Attribute coding is mainly performed on color information. In the process of color information coding, there are two main transformation methods. One is the distance-based lifting transformation that relies on the level of detail (LOD) division, and the other is direct RAHT. Both methods convert color information from the spatial domain to the frequency domain, and obtain high-frequency coefficients and low-frequency coefficients through transformation. Finally, the coefficients are quantized and then arithmetic coding is performed on the quantized coefficients to generate a binary attribute bit stream.
[0093] Figure 4B shows a schematic diagram of the composition framework of a G-PCC decoder. As shown in Figure 4B, for the acquired binary bit stream, the geometric bit stream and attribute bit stream in the binary bit stream are first decoded independently. When decoding the geometric bit stream, the geometric information of the point cloud is obtained through arithmetic decoding-reconstruction of the octree / reconstruction of the prediction tree-reconstruction of the geometry-coordinate inverse conversion; when decoding the attribute bit stream, the attribute information of the point cloud is obtained through arithmetic decoding-inverse quantization-LOD partitioning / RAHT-color inverse conversion, and the point cloud data to be encoded (i.e., the output point cloud) is restored based on the geometric information and attribute information.
[0094] It should be noted that, as shown in FIG4A or FIG4B , the current geometric coding and decoding of G-PCC can be divided into octree-based geometric coding and decoding (marked by a dotted box) and prediction tree-based geometric coding and decoding (marked by a dotted box).
[0095] For octree geometry encoding (OctGeomEnc), octree geometry encoding includes: first, coordinate conversion of the geometric information so that all point clouds are contained in a bounding box. Then quantization is performed. This step of quantization mainly plays a role of scaling. Due to the quantization rounding, the geometric information of some points is the same. The parameters are used to decide whether to remove duplicate points. The process of quantization and removal of duplicate points is also called voxelization. Next, the bounding box is continuously divided into trees (such as octrees, quadtrees, binary trees, etc.) in the order of breadth-first traversal, and the placeholder code of each node is encoded. In related technologies, a company proposed an implicit geometric division method. First, the bounding box of the point cloud is calculated. Assume d x >d y >d z , the bounding box corresponds to a cuboid. When geometrically partitioning, the binary tree partitioning will be performed based on the x-axis to obtain two child nodes; until d is satisfied x =d y >d z When the conditions are met, the quadtree partitioning will be performed based on the x and y axes to obtain four child nodes; when d is finally satisfied x =d y =d z When the conditions are met, the octree partitioning will continue until the leaf node obtained by the partitioning is a 1×1×1 unit cube. The partitioning will stop and the points in the leaf node will be encoded to generate a binary code stream. In the process of binary tree / quadtree / octree partitioning, two parameters are introduced: K and M. Parameter K indicates the maximum number of binary tree / quadtree partitions before octree partitioning; parameter M is used to indicate that the minimum block side length corresponding to binary tree / quadtree partitioning is 2 M . At the same time, K and M must meet the following conditions: Assume d max =max(d x ,d y ,d z ), d min =min(d x ,d y ,d z ), parameter K satisfies: K ≥ d max -d min ; Parameter M satisfies: M≥d minThe reason why the parameters K and M meet the above conditions is that in the current G-PCC geometric implicit partitioning process, the priority of the partitioning method is binary tree, quadtree and octree. When the node block size does not meet the conditions of binary tree / quadtree, the node will be divided into octree until the minimum unit of leaf node is 1×1×1. The octree-based geometric information coding mode can effectively encode the geometric information of the point cloud by utilizing the correlation between adjacent points in space. However, for some relatively flat nodes or nodes with planar characteristics, the coding efficiency of the point cloud geometric information can be further improved by using plane coding.
[0096] Exemplarily, Figure 5A and Figure 5B provide a kind of plane position schematic diagram.Wherein, Figure 5A shows a kind of low plane position schematic diagram in the Z-axis direction, and Figure 5B shows a kind of high plane position schematic diagram in the Z-axis direction.As shown in Figure 5A, here (a), (a0), (a1), (a2), (a3) all belong to the low plane position in the Z-axis direction. Taking (a) as an example, it can be seen that the four child nodes occupied in the current node are all located at the low plane position of the current node in the Z-axis direction, so it can be considered that the current node belongs to a Z plane and is a low plane in the Z-axis direction. Similarly, as shown in Figure 5B, here (b), (b0), (b1), (b2), (b3) all belong to the high plane position in the Z-axis direction. Taking (b) as an example, it can be seen that the four child nodes occupied in the current node are located at the high plane position of the current node in the Z-axis direction, so it can be considered that the current node belongs to a Z plane and is a high plane in the Z-axis direction.
[0097] Taking (a) in Figure 5A as an example, the efficiency of octree encoding and plane encoding is compared. Figure 6 provides a schematic diagram of the node encoding sequence, that is, node encoding is performed in the order of 0, 1, 2, 3, 4, 5, 6, and 7 as shown in Figure 6. Here, if octree encoding is used for (a) in Figure 5A, the placeholder information of the current node is represented as: 10101010. However, if plane encoding is used, first, an identifier must be encoded to indicate that the current node is a plane in the Z-axis direction. Secondly, if the current node is a plane in the Z-axis direction, the plane position of the current node must also be represented. Finally, only the placeholder information of the lower plane node in the Z-axis direction needs to be encoded (i.e., the placeholder information of the four child nodes 0, 2, 4, and 6). Therefore, encoding the current node based on plane encoding only requires 6 bits, which can reduce the representation by 2 bits compared to the related art octree encoding. Based on this analysis, plane encoding has significantly higher coding efficiency than octree encoding. Therefore, for an occupied node, if a plane encoding method is used in a certain dimension, it is first necessary to represent the plane identification (planarMode) and plane position (PlanePos) information of the current node in that dimension, and then encode the occupancy information of the current node based on the plane information of the current node. For example, FIG7A shows a schematic diagram of plane identification information. As shown in FIG7A, there is a low plane in the Z-axis direction; correspondingly, the value of the plane identification information is true (true) or 1, that is, planarMode_ Z = true; the plane position information is the low plane (low), that is, PlanePosition_ Z =low. FIG7B shows another schematic diagram of plane identification information. As shown in FIG7B, here it is not a plane in the Z-axis direction; correspondingly, the value of the plane identification information is false or 0, that is, planarMode_ Z =false.
[0098] It should be noted that for PlaneMode_ i :0 means the current node is not a plane in the i-axis direction, 1 means the current node is a plane in the i-axis direction. If the current node is a plane in the i-axis direction, then for PlanePosition_ i : 0 means the current node is a plane in the i-axis direction and the plane position is low, 1 means the current node is a high plane in the i-axis direction. Here, i represents the coordinate dimension, which can be the X-axis direction, Y-axis direction, or Z-axis direction, so i = 0, 1, 2.
[0099] In the G-PCC standard, when determining whether a node meets the conditions for planar coding and when the node meets the conditions for planar coding, predictive coding needs to be performed on the planar identifier and planar position information of the node.
[0100] There are currently three judgment conditions in the G-PCC standard for determining whether a node meets the conditions for planar coding, and they will be explained in detail one by one below.
[0101] I. Judge according to the planar probability of the node in each dimension.
[0102] (1) Determine the local area density (local_node_density) of the current node;
[0103] (2) Determine the probability Prob(i) of the current node in each dimension.
[0104] When the local area density of the node is less than the threshold Th (for example, Th = 3), compare the planar probability Prob(i) of the current node in the three coordinate dimensions with the thresholds Th0, Th1, and Th2, where Th0 < Th1 < Th2 (for example, Th0 = 0.6, Th1 = 0.77, Th2 = 0.88). Here, Eligible i (i = 0, 1, 2) represents whether planar coding is enabled in each dimension: Eligible i = Prob(i) >= threshold.
[0105] It should be noted that the threshold is adaptively changed. For example, when Prob(0) > Prob(1) > Prob(2), then Eligible i is set as follows: Eligible0 = Prob(0) >= Th0; Eligible1 = Prob(1) >= Th1; Eligible2 = Prob(2) >= Th2 (1)
[0106] When Prob(1) > Prob(0) > Prob(2), then Eligible i is set as follows: Eligible0 = Prob(0) >= Th1; Eligible1 = Prob(1) >= Th0; Eligible2 = Prob(2) >= Th2 (2)
[0107] Here, the update of Prob(i) is specifically as follows: Prob(i) new=(L×Prob(i)+δ(coded node)) / L+1 (3)
[0108] Where L = 255; in addition, if the coded node is a plane, δ(coded node) is 1; otherwise, δ(coded node) is 0.
[0109] Here, the update of local_node_density is as follows: local_node_density new =local_node_density+4*numSiblings (4)
[0110] Where local_node_density is initialized to 4, and numSiblings is the number of sibling nodes of the node. For example, FIG8 shows a schematic diagram of the sibling nodes of the current node. As shown in FIG8 , the current node is a node filled with slashes, and the nodes filled with grids are the sibling nodes of the current node. Then, the number of sibling nodes of the current node is 5 (including the current node itself).
[0111] Second, determine whether the current layer nodes meet the plane coding requirements based on the point cloud density of the current layer.
[0112] The point density of the current layer is used to determine whether to perform planar coding on the nodes of the current layer. Assuming that the number of points in the current point cloud to be coded is pointCount, the number of points reconstructed by the infer direct coding model (IDCM) coding is numPointCountRecon, and because the octree is coded in the order of breadth-first traversal, the number of nodes to be coded in the current layer can be obtained as nodeCount. Then, the assumption to determine whether to start planar coding in the current layer is planarEligibleKOctreeDepth, specifically: planarEligibleK OctreeDepth = (pointCount-numPointCountRecon) <nodeCount×1.3。
[0113] Among them, if (pointCount-numPointCountRecon) is less than nodeCount×1.3, then planarEligibleK OctreeDepth is true; if (pointCount-numPointCountRecon) is not less than nodeCount×1.3, then planarEligibleKOctreeDepth is false. In this way, when planarEligibleKOctreeDepth is true, all nodes in the current layer are planar coded; otherwise, all nodes in the current layer are not planar coded and only octree coding is used.
[0114] 3. Determine whether the current node meets the plane coding requirements based on the acquisition parameters of the lidar point cloud.
[0115] Figure 9 shows a schematic diagram of the intersection of a laser radar and a node. As shown in Figure 9, a node filled with a grid is simultaneously traversed by two laser beams, so the current node is not a plane in the direction perpendicular to the Z axis. A node filled with a diagonal line is small enough to not be traversed by two laser beams simultaneously, so it is possible that the node filled with a diagonal line is a plane in the direction perpendicular to the Z axis.
[0116] Furthermore, for nodes that meet the plane coding conditions, predictive coding may be performed on the plane identification information and the plane position information.
[0117] First, predictive coding of plane identification information.
[0118] Here, only three context information are used for encoding, that is, the plane identification in each coordinate dimension is designed separately for context.
[0119] Secondly, predictive coding of plane position information.
[0120] It should be understood that for the encoding of non-lidar point cloud planar position information, the predictive encoding of the planar position information may include:
[0121] (a) Using the occupancy information of neighboring nodes, the plane position information of the current node is predicted to be three elements: predicted as low plane, predicted as high plane, and unpredictable;
[0122] (b) Spatial distance between nodes at the same partition depth and coordinates as the current node and the current node: “near” and “far”;
[0123] (c) If the node at the same partition depth and the same coordinates as the current node is a plane, determine the plane position of the node;
[0124] (d) Coordinate dimension (i=0, 1, 2).
[0125] It should be noted that in an embodiment of the present application, after determining the spatial distance between the node at the same division depth and the same coordinates as the current node and the current node, if the spatial distance is less than the preset distance threshold, then the spatial distance can be determined to be "near"; or, if the spatial distance is greater than the preset distance threshold, then the spatial distance can be determined to be "far".
[0126] For example, Figure 10 shows a schematic diagram of neighboring nodes at the same partition depth and coordinates. As shown in Figure 10, the bold large cube represents the parent node, the small cube filled with a grid inside it represents the current node, and the vertex position of the current node is shown. The small cube filled with white represents neighboring nodes at the same partition depth and coordinates. The distance between the current node and the neighboring node is the spatial distance, which can be judged as "near" or "far." In addition, if the neighboring node is a plane, the planar position of the neighboring node is also required.
[0127] In this way, as shown in Figure 10, the current node is a small cube filled with a grid, and the neighboring node is a small cube filled with white at the same octree partition depth level and the same vertical coordinate, and the distance between the two nodes is judged as "near" and "far", and the plane position of the reference node is used.
[0128] Furthermore, in an embodiment of the present application, FIG11 shows a schematic diagram of a current node being located at a low plane position of a parent node. As shown in FIG11 , (a), (b), and (c) show three examples of the current node being located at a low plane position of a parent node. Specific descriptions are as follows:
[0129] ① If any of the child nodes 4 to 7 of the point fill node is occupied, and all the grid fill nodes are not occupied, it is very likely that there is a plane in the current node (filled with a slash), and the plane is located lower.
[0130] ② If the child nodes 4 to 7 of the point fill node are not occupied, and any grid fill node is occupied, it is very likely that there is a plane in the current node (filled with diagonal lines), and the plane is located higher.
[0131] ③ If the child nodes 4 to 7 of the point fill node are all empty nodes and the grid fill nodes are all empty nodes, the plane position cannot be inferred and is therefore marked as unknown.
[0132] ④ If any of the child nodes 4 to 7 of the point fill node is occupied and any of the grid fill nodes is occupied, the plane position cannot be inferred at this time, so it is marked as unknown.
[0133] In an embodiment of the present application, FIG12 shows a schematic diagram of a current node being located at a high plane position of a parent node. As shown in FIG12, (a), (b), and (c) show three examples of the current node being located at a high plane position of a parent node. The specific description is as follows:
[0134] ① If any of the child nodes 4 to 7 of the grid fill node is occupied, and the point fill node is not occupied, it is very likely that there is a plane in the current node (filled with diagonal lines), and the plane position is low.
[0135] ② If the child nodes 4 to 7 of the grid fill node are not occupied, and the point fill node is occupied, it is very likely that there is a plane in the current node (filled with a slash), and the plane position is higher.
[0136] ③If the child nodes 4 to 7 of the grid fill node are all unoccupied, and the point fill node is unoccupied, the plane position cannot be inferred at this time, so it is marked as unknown.
[0137] ④ If one of the child nodes 4 to 7 of the grid fill node is occupied and the point fill node is occupied, the plane position cannot be inferred at this time and is therefore marked as unknown.
[0138] It should also be understood that, with respect to the coding of the laser radar point cloud plane position information, FIG13 shows a schematic diagram of the predictive coding of the laser radar point cloud plane position information. As shown in FIG13, when the laser radar emission angle is θ bottom When , it can be mapped to the bottom virtual plane; when the laser radar emission angle is θ top At this time, it can be mapped to the top virtual plane.
[0139] That is, by using the laser radar acquisition parameters to predict the plane position of the current node, and by using the position where the current node intersects with the laser ray to quantize the position into multiple intervals, the final result is the context information of the plane position of the current node. The specific calculation process is as follows: Assume that the coordinates of the laser radar are (x Lidar ,y Lidar ,z Lidar ), the geometric coordinates of the current node are (x, y, z), then first calculate the vertical tangent value tanθ of the current node relative to the lidar, the calculation formula is as follows:
[0140] Furthermore, because each laser has a certain offset angle relative to the laser radar, it is also necessary to calculate the relative tangent value tanθ of the current node relative to the laser corr,L , the specific calculation is as follows:
[0141] Finally, the relative tangent value tanθ of the current node will be used corr,L To predict the plane position of the current node, as follows, assuming that the tangent value of the lower boundary of the current node is tan(θ bottom ), the tangent value of the upper boundary is tan(θ top ), according to tanθ corr,L The plane position is quantized into four quantization intervals, that is, the context information of the plane position is determined.
[0142] However, the octree-based geometric information coding mode only has an efficient compression rate for points that are correlated in space. For points that are isolated in the geometric space, the use of the direct coding model (DCM) can greatly reduce the complexity. For all nodes in the octree, the use of DCM is not indicated by flag information, but is inferred from the parent node and neighbor information of the current node. There are three ways to determine whether the current node is eligible for DCM coding, as follows:
[0143] (1) The current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.
[0144] (2) The parent node of the current node has only one child node, the current node, and the six neighbor nodes that share a face with the current node are all empty nodes.
[0145] (3) The number of sibling nodes of the current node is greater than 1.
[0146] Exemplarily, FIG14 provides a schematic diagram of IDCM coding. If the current node does not have the DCM coding qualification, it will be divided into octrees. If it has the DCM coding qualification, the number of points contained in the node will be further determined. When the number of points is less than a threshold (for example, 2), the node will be DCM-encoded, otherwise the octree division will continue. When the DCM coding mode is applied, it is first necessary to encode whether the current node is a true isolated point, that is, IDCM_flag. When IDCM_flag is true, the current node adopts DCM coding, otherwise octree coding is still adopted. When the current node meets the DCM coding requirements, the DCM coding mode of the current node needs to be encoded. There are currently two DCM modes: (a) there is only one point (or multiple points, but they are duplicate points); (b) there are two points. Finally, the geometric information of each point needs to be encoded. Assuming that the side length of the node is 2 d When encoding each component of the node's geometric coordinates, d bits are required, and these bits are directly encoded into the bitstream. It is important to note that when encoding LiDAR point clouds, the efficiency of geometric information coding can be further improved by predictively encoding the three-dimensional coordinate information using LiDAR acquisition parameters.
[0147] Furthermore, the IDCM encoding process is described in detail below.
[0148] When the current node meets the DCM encoding mode, the number of points of the current node, numPoints, is encoded first; the number of points of the current node is encoded according to different DirectModes:
[0149] (1) If the current node does not meet the requirements of the DCM node, exit directly (that is, the number of points is greater than 2 points and is not a duplicate point).
[0150] (2) If the number of points numPonts in the current node is less than or equal to 2, the encoding process is as follows:
[0151] i) First encode whether the numPonts of the current node is greater than 1;
[0152] ii) If the current node has only one point and the geometry coding environment is geometry lossless coding, it is necessary to encode that the second point of the current node is not a duplicate point.
[0153] (3) If the number of points numPonts in the current node is greater than 2, the encoding process is as follows:
[0154] i) First encode the numPonts of the current node to be less than or equal to 1;
[0155] ii) Secondly, it is encoded that the second point of the current node is a repeated point, and then it is encoded whether the number of repeated points of the current node is greater than 1. When the number of repeated points is greater than 1, it is necessary to perform exponential Golomb decoding on the remaining number of repeated points.
[0156] After encoding the number of points in the current node, the coordinate information of the points contained in the current node is encoded. The following will introduce the lidar point cloud and the human eye point cloud in detail.
[0157] (1) Point cloud facing the human eye.
[0158] (1) If the current node contains only one point, the geometric information of the point in three dimensions will be directly encoded (bypass coding);
[0159] (2) If the current node contains two points, the first coordinate axis dirextAxis will be obtained by using the geometric coordinates of the points. It should be noted that the coordinate axes currently compared only include the x-axis and y-axis, not the z-axis. Assuming that the geometric coordinates of the current node are nodePos, the judgment method is as follows: dirextAxis = ! (nodePos[0] <nodePos[1])
[0160] That is, the axis with the smallest node coordinate geometry position will be used as the priority encoding axis dirextAxis, and then the geometry information of the priority encoding axis dirextAxis will be encoded as follows. Assume that the encoding geometry bit depth corresponding to the priority encoding axis is nodeSizeLog2, and assume that the coordinates of the two points are pointPos[0] and pointPos[1]. The specific encoding process is as follows:
[0161] After encoding the priority axis dirextAxis, the geometric coordinates of the current node are directly encoded. Assuming that the remaining encoding bit depth of each point is nodeSizeLog2, the specific encoding process is as follows:
[0162] for(int axisIdx=0; axisIdx<3; ++axisIdx)
[0163] for(int mask=(1<<nodeSizeLog2[axisIdx])> >1;mask;mask>>1)
[0164] encodePosBit(!!(pointPos[axisIdx]&mask)).
[0165] (2) LiDAR point cloud.
[0166] If the current node contains two points, the priority coded coordinate axis dirextAxis will be obtained first by using the geometric coordinates of the points. Assuming that the geometric coordinates of the current node are nodePos, the judgment method is as follows:
[0167] dirextAxis=!(nodePos[0] <nodePos[1])
[0168] That is, the axis with the smaller node coordinate geometry position will be used as the priority encoding axis dirextAxis. It should be noted that the currently compared coordinate axes only include the x-axis and y-axis, not the z-axis. Then, the geometric information of the priority encoded coordinate axis dirextAxis is first encoded as follows, assuming that the encoding geometry bit depth corresponding to the priority encoded axis is nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1]. The specific encoding process is as follows:
[0169] After encoding the priority-encoded coordinate axis dirextAxis, the geometric coordinates of the current node are encoded.
[0170] Since the LiDAR point cloud can obtain the acquisition parameters of the LiDAR point cloud, the geometric coordinate information of the current node can be predicted, thereby further improving the efficiency of the geometric information encoding of the point cloud. Similarly, the geometric information nodePos of the current node is first used to obtain a directly encoded main axis direction, and then the geometric information of the already encoded direction is used to predict the geometric information of another dimension. Assuming that the directly encoded axis direction is directAxis and the bit depth of the direct encoding is nodeSizeLog2, the encoding method is as follows:
[0171] for(int mask=(1<<nodeSizeLog2)> >1;mask;mask>>1)
[0172] encodePosBit(!!(pointPos[directAxis]&mask)).
[0173] It should be noted here that all geometric accuracy information in the directAxis direction will be encoded here.
[0174] For example, FIG15 provides a schematic diagram of coordinate transformation for obtaining point clouds using a rotating laser radar. In the Cartesian coordinate system, the (x, y, z) coordinates of each node can be converted to (R, φ, i). In addition, the laser scanner can perform laser scanning at a preset angle, and different θ(i) can be obtained under different values of i. For example, when i is equal to 1, θ(1) can be obtained, and the corresponding scanning angle is -15°; when i is equal to 2, θ(2) can be obtained, and the corresponding scanning angle is -13°; when i is equal to 10, θ(10) can be obtained, and the corresponding scanning angle is +13°; when i is equal to 9, θ(19) can be obtained, and the corresponding scanning angle is +15°.
[0175] In this way, after encoding all the precision of the directAxis coordinate direction, the LaserIdx corresponding to the current node will be calculated first, that is, the pointLaserIdx number in Figure 15, and the LaserIdx of the current node, that is, nodeLaserIdx; secondly, the LaserIdx of the node, that is, nodeLaserIdx, will be used to predict the LaserIdx of the point, that is, pointLaserIdx. The calculation method of the LaserIdx of the node or point is as follows. Assuming that the geometric coordinates of the point are pointPos, the starting coordinates of the laser ray are LidarOrigin, and assuming that the number of Lasers is LaserNum, the tangent value of each Laser is tanθ i , the vertical offset position of each Laser is Z i ,but:
[0176] After calculating the current node's LaserIdx, the pointLaserIdx of the point is predictively encoded using the current node's LaserIdx. After encoding the current node's LaserIdx, the three-dimensional geometric information of the current node is predictively encoded using the LiDAR acquisition parameters.
[0177] For example, FIG16 shows a schematic diagram of predictive coding in the X-axis or Y-axis direction. As shown in FIG16 , the box filled with a grid represents the current node, and the box filled with a slash represents the already coded node. Here, the LaserIdx corresponding to the current node is first used to obtain the corresponding predicted value of the horizontal azimuth, that is, Secondly, the node geometry information corresponding to the current node is used to obtain the horizontal azimuth angle corresponding to the node Assuming that the geometric coordinates of the node are nodePos, the horizontal azimuth angle The calculation method between the node geometry information is as follows:
[0178] By using the laser radar acquisition parameters, we can get the number of rotation points of each laser, numPoints, which represents the number of points obtained by each laser ray rotating one circle. Then, we can use the number of rotation points of each laser to calculate the rotation angular velocity deltaPhi of each laser. The calculation method is as follows:
[0179] Furthermore, using the horizontal azimuth angle of the node And the horizontal azimuth of the previous Laser code point corresponding to the current node Calculate the horizontal azimuth angle prediction value corresponding to the current node That is, the predicted value of the horizontal azimuth angle as shown in Figure 17A and Figure 17B. Figure 17A shows a schematic diagram of predicting the angle of the Y plane by the horizontal azimuth angle, and Figure 17B shows a schematic diagram of predicting the angle of the X plane by the horizontal azimuth angle. Here, for the predicted value of the horizontal azimuth angle corresponding to the current node The calculation is as follows:
[0180] For example, FIG18 shows another schematic diagram of predictive coding in the X-axis or Y-axis direction. As shown in FIG18 , the portion filled with a grid (left side) represents a low plane, and the portion filled with dots (right side) represents a high plane. Indicates the low plane horizontal azimuth of the current node, Indicates the horizontal azimuth of the high plane of the current node, Indicates the predicted horizontal azimuth angle corresponding to the current node.
[0181] Thus, by using the predicted value of the horizontal azimuth and the low plane horizontal azimuth of the current node and the high plane horizontal azimuth To predict the geometric information of the current node. The details are as follows:
[0182] int context=(angLel≥0&&angLeR≥0)||(angLel<0&&angLeR<0)? 0:2;
[0183] int minAngle=std∷min(abs(angLel),abs(angLeR));
[0184] int maxAngle=std∷max(abs(angLel),abs(angLeR));
[0185] context+=maxAngle>minAngle? 0:1;
[0186] context+=maxAngle>minAngle? 0:4.
[0187] After encoding the LaserIdx of the completed point, the Z-axis direction of the current node will be predicted and encoded using the LaserIdx corresponding to the current node. That is, the radius of the radar coordinate system is calculated using the x and y information of the current node. Then, the tangent value of the current node and the vertical offset are obtained using the LaserIdx of the current node. The predicted value of the Z-axis direction of the current node, namely Z_pred, can be obtained. The details are as follows:
[0188] int tanTheta=tanθ laserIdx ;
[0189] int zOffset = Z laserIdx ;
[0190] Z_pred=radius×tanTheta-zOffset.
[0191] Furthermore, Z_pred is used to perform predictive coding on the geometric information of the current node in the Z-axis direction to obtain the prediction residual Z_res, and finally Z_res is encoded.
[0192] It is important to note that when partitioning nodes into leaf nodes, the number of duplicate points in the leaf nodes must be encoded in the case of lossless geometric coding. Ultimately, the placeholder information for all nodes is encoded to generate a binary bitstream. Furthermore, G-PCC currently introduces a plane coding mode. During the geometric partitioning process, it determines whether the child nodes of the current node are in the same plane. If the child nodes of the current node meet the condition of being in the same plane, the child nodes of the current node are represented by that plane.
[0193] For octree-based geometric decoding, the decoder follows a breadth-first traversal. Before decoding each node's occupancy information, it first uses the reconstructed geometric information to determine whether the current node is for plane decoding or IDCM decoding. If the current node meets the requirements for plane decoding, it first decodes the plane identifier and plane position information of the current node. Then, based on the plane information, it decodes the current node's occupancy information. If the current node meets the requirements for IDCM decoding, it first decodes whether the current node is a true IDCM node. If so, it continues to parse the DCM decoding mode of the current node, then obtains the number of points in the current DCM node, and finally decodes the geometric information of each point. For nodes that do not meet either plane decoding or DCM decoding requirements, the current node's occupancy information is decoded. By continuously parsing in this way, the placeholder code of each node is obtained, and the node is continuously partitioned until a 1×1×1 unit cube is obtained. The number of points contained in each leaf node is parsed, and the geometrically reconstructed point cloud information is finally recovered.
[0194] The following is a detailed introduction to the IDCM decoding process.
[0195] Similar to the processing at the encoding end, we first use prior information to determine whether the node should start IDCM. The starting conditions of IDCM are as follows:
[0196] (1) The current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.
[0197] (2) The parent node of the current node has only one child node, the current node, and the six neighbor nodes that share a face with the current node are all empty nodes.
[0198] (3) The number of sibling nodes of the current node is greater than 1.
[0199] Furthermore, when a node meets the conditions for DCM coding, it is first decoded to determine whether the current node is a true DCM node, that is, IDCM_flag; when IDCM_flag is true, the current node adopts DCM coding, otherwise it still adopts octree coding.
[0200] Next, decode the number of points numPoints of the current node. The specific decoding method is as follows:
[0201] i) First decode whether numPonts of the current node is greater than 1;
[0202] ii) If the numPonts of the current node is greater than 1, the second point is decoded to determine whether it is a duplicate point. If the second point is not a duplicate point, it can be implicitly inferred that the second type of DCM pattern contains only two points.
[0203] iii) If the numPonts of the current node obtained by decoding is less than or equal to 1, then decode whether the second point is a repeated point; if the second point is not a repeated point, then it can be implicitly inferred that the second type of DCM pattern is satisfied, which contains only one point; if the second point obtained by decoding is a repeated point, then it can be inferred that the third type of DCM pattern is satisfied, which contains multiple points, but all are repeated points, then decode whether the number of repeated points is greater than 1 (entropy decoding), and if it is greater than 1, decode the number of remaining repeated points (using exponential Columbus decoding).
[0204] If the current node does not meet the requirements of the DCM node, it will exit directly (that is, the number of points is greater than 2 points and it is not a duplicate point).
[0205] After decoding the number of points in the current node, the coordinate information of the points contained in the current node is decoded. The following will introduce the lidar point cloud and the human eye point cloud in detail.
[0206] (1) Point cloud facing the human eye.
[0207] (1) If the current node contains only one point, the geometric information of the point in three dimensions will be directly decoded;
[0208] (2) If the current node contains two points, the first decoded coordinate axis dirextAxis will be obtained by using the geometric coordinates of the points. It should be noted that the coordinate axes currently compared only include the x and y axes, not the z axis. Assuming that the geometric coordinates of the current node are nodePos, the judgment method is as follows:
[0209] dirextAxis=!(nodePos[0] <nodePos[1])(10)
[0210] That is, the axis with the smallest node coordinate geometry position will be used as the priority decoding axis dirextAxis, and then the geometry information of the priority decoding axis dirextAxis will be decoded first in the following way. Assume that the geometry bit depth to be decoded corresponding to the priority decoding axis is nodeSizeLog2, and assume that the coordinates of the two points are pointPos[0] and pointPos[1]. The specific encoding process is as follows:
[0211] After decoding the prioritized axis dirextAxis, the geometric coordinates of the current node are directly decoded. Assuming that the remaining encoding bit depth of each point is nodeSizeLog2 and the coordinate information of the point is pointPos, the specific decoding process is as follows:
[0212] (2) LiDAR point cloud.
[0213] If the current node contains two points, the priority decoding coordinate axis dirextAxis will be obtained first by using the geometric coordinates of the points. Assuming that the geometric coordinates of the current node are nodePos, the judgment method is as follows: dirextAxis = ! (nodePos[0] <nodePos[1]) (11)
[0214] That is, the axis with the smaller node coordinate geometry position will be used as the priority decoding axis dirextAxis. It should be noted that the currently compared coordinate axes only include the x-axis and y-axis, not the z-axis. Secondly, the priority encoded coordinate axis dirextAxis geometry information is first decoded as follows, assuming that the encoding geometry bit depth corresponding to the priority decoded axis is nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1]. The specific encoding process is as follows:
[0215] After decoding the priority coordinate axis dirextAxis, the geometric coordinates of the current node are decoded.
[0216] Similarly, we first use the current node's geometry information nodePos to get a direct decoding main axis direction, and then use the geometry information of the decoded direction to decode the geometry information of the other dimension. Assuming that the axis direction of direct decoding is directAxis and the bit depth to be decoded in direct decoding is nodeSizeLog2, the decoding method is as follows:
[0217] It should be noted here that all geometric accuracy information in the directAxis direction will be decoded here.
[0218] After decoding all the precision of the directAxis coordinate direction, the LaserIdx of the current node, i.e., nodeLaserIdx, is calculated first. Then, the LaserIdx of the node, i.e., nodeLaserIdx, is used to predict and decode the LaserIdx of the point, i.e., pointLaserIdx. The calculation method of the LaserIdx of the node or point is the same as that of the encoder. Finally, the LaserIdx of the current node and the predicted residual information of the LaserIdx of the node are decoded to obtain ResLaserIdx. The decoding method is as follows: PointLaserIdx=nodeLaserIdx+ResLaserIdx (12)
[0219] After decoding the LaserIdx of the current node, the three-dimensional geometric information of the current node is predicted and decoded using the acquisition parameters of the laser radar. The specific algorithm is as follows:
[0220] As shown in Figure 11, first use the LaserIdx corresponding to the current node to obtain the corresponding horizontal azimuth prediction value, that is, Secondly, the node geometry information corresponding to the current node is used to obtain the horizontal azimuth angle corresponding to the node Assuming that the geometric coordinates of the node are nodePos, the horizontal azimuth angle The calculation method between the node geometry information is as follows:
[0221] By using the acquisition parameters of the laser radar, we can get the number of rotation points of each laser, numPoints, which represents the number of points obtained by each laser ray rotating one circle. Then, we can use the number of rotation points of each laser to calculate the rotation angular velocity deltaPhi of each laser. The calculation method is as follows:
[0222] Furthermore, using the horizontal azimuth angle of the node And the horizontal azimuth of the previous Laser code point corresponding to the current node Calculate the horizontal azimuth angle prediction value corresponding to the current node That is, the predicted value of the horizontal azimuth angle as shown in Figures 17A and 17B. The calculation method is as follows:
[0223] Thus, by using the predicted value of the horizontal azimuth angle φ predPoint and the low plane horizontal azimuth angle φ of the current node left and the horizontal azimuth of the high plane To predict and decode the geometric information of the current node. The details are as follows:
[0224] int context=(angLel≥0&&angLeR≥0)||(angLel<0&&angLeR<0)? 0:2;
[0225] int absAngleL=abs(angLel);
[0226] int absAngleR=abs(angLeR);
[0227] context+=absAngleL>absAngleR? 0:1;
[0228] context+=maxAngle>minAngle<<1?4:0.
[0229] After decoding the LaserIdx of the node, the Z-axis direction of the current node is predicted and decoded using the LaserIdx corresponding to the current node. That is, the radius of the radar coordinate system is calculated using the x and y information of the current node. Then, the tangent value of the current node and the vertical offset are obtained using the LaserIdx of the current node. The predicted value of the Z-axis direction of the current node, namely Z_pred, can be obtained. The details are as follows:
[0230] int tanTheta=tanθ laserIdx ;
[0231] int zOffset = Z laserIdx ;
[0232] Z_pred=radius×tanTheta-zOffset.
[0233] Furthermore, the decoded Z_res and Z_pred are used to reconstruct and restore the geometric information of the current node in the Z-axis direction.
[0234] For triangle soup (trisoup)-based geometric information coding, geometric partitioning must also be performed first in the trisoup-based geometric information coding framework. However, unlike geometric information coding based on binary trees, quadtrees, and octrees, this method does not require step-by-step partitioning of the point cloud into unit cubes with side lengths of 1×1×1. Instead, the partitioning stops when the sub-blocks (blocks) have a side length of W. Based on the surface formed by the distribution of the point cloud in each block, the surface and the twelve edges of the block are obtained. The vertex coordinates of each block are encoded in sequence to generate a binary code stream.
[0235] For trisoup-based point cloud geometry reconstruction, when performing point cloud geometry reconstruction at the decoding end, the vertex coordinates are first decoded to complete the triangle face reconstruction. This process is shown in Figures 19A, 19B, and 19C. Among them, there are three intersection points (v1, v2, v3) in the block shown in Figure 19A. The set of triangle faces formed by these three intersection points in a certain order is called triangle soup, or trisoup, as shown in Figure 19B. Afterwards, sampling is performed on the triangle face set, and the obtained sampling points are used as the reconstructed point cloud within the block, as shown in Figure 19C.
[0236] Predictive geometry coding (PredGeomTree) involves first sorting the input point cloud. Currently used sorting methods include unordered, Morton order, azimuth order, and radial distance order. At the encoding end, the prediction tree structure is established using two different methods: KD-Tree (high-latency slow mode) and low-latency fast mode (using lidar calibration information). When using lidar calibration information, each point is assigned to a different laser, and a prediction tree structure is established based on the different lasers. Next, based on the prediction tree structure, each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain a prediction residual. The geometric prediction residual is then quantized using a quantization parameter. Finally, through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary bitstream.
[0237] For geometric decoding based on the prediction tree, the decoding end reconstructs the prediction tree structure by continuously parsing the bit stream. Secondly, the geometric position prediction residual information and quantization parameters of each prediction node are obtained through parsing, and the prediction residual is dequantized to restore the reconstructed geometric position information of each node, finally completing the geometric reconstruction at the decoding end.
[0238] After the geometric encoding is completed, the geometric information needs to be reconstructed. At present, attribute encoding is mainly performed on color information. First, the color information is converted from the RGB color space to the YUV color space. Then, the point cloud is recolored using the reconstructed geometric information so that the unencoded attribute information corresponds to the reconstructed geometric information. In color information encoding, there are two main transformation methods. One is the distance-based lifting transformation that relies on LOD partitioning, and the other is to directly perform RAHT transformation. Both methods will convert the color information from the spatial domain to the frequency domain, and obtain high-frequency coefficients and low-frequency coefficients through transformation. Finally, the coefficients are quantized and encoded to generate a binary code stream, as shown in Figures 4A and 4B.
[0239] Furthermore, when using geometric information to predict attribute information, Morton codes can be used to perform nearest neighbor search. The Morton code corresponding to each point in the point cloud can be obtained from the geometric coordinates of the point. The specific method for calculating the Morton code is described below. For a three-dimensional coordinate represented by a d-bit binary number for each component, its three components can be expressed as:
[0240] in, The highest bits of x, y, and z respectively To the lowest position The corresponding binary value. The Morton code M is x, y, z starting from the highest bit, arranged in sequence To the lowest bit, the calculation formula of M is as follows:
[0241] in, The highest bit of M To the lowest position After obtaining the Morton code M of each point in the point cloud, the points in the point cloud are arranged in ascending order of the Morton code, and the weight value w of each point is set to 1.
[0242] It can also be understood that for the G-PCC codec framework, the general test conditions are as follows:
[0243] (1) There are 4 test conditions:
[0244] Condition 1: The geometric position is limited and the attributes are lost;
[0245] Condition 2: Geometric position lossless, attribute lossy;
[0246] Condition 3: Geometric position lossless, attribute loss limited;
[0247] Condition 4: Geometric position and attributes are lossless.
[0248] (2) The general test sequence includes four categories: Cat1A, Cat1B, Cat3-fused, and Cat3-frame. Among them, the Cat3-frame point cloud only contains reflectance attribute information, the Cat1A and Cat1B point clouds only contain color attribute information, and the Cat3-fused point cloud contains both color and reflectance attribute information.
[0249] (3) Technical routes: There are two types in total, which are distinguished by the algorithm used for geometric compression.
[0250] Technical route 1: Octree encoding branch.
[0251] At the encoding end, the bounding box is divided into sub-cubes in sequence, and the non-empty sub-cubes (containing points in the point cloud) are divided again until the leaf node obtained is a 1×1×1 unit cube. In the case of geometric lossless coding, the number of points contained in the leaf node needs to be encoded, and finally the geometric octree encoding is completed to generate a binary code stream.
[0252] At the decoding end, the decoding end obtains the placeholder code of each node by continuously parsing in the order of breadth-first traversal, and continuously divides the nodes in turn until a 1×1×1 unit cube is obtained. In the case of geometric lossless decoding, it is necessary to parse the number of points contained in each leaf node and finally recover the geometrically reconstructed point cloud information.
[0253] Technical route 2: prediction tree encoding branch.
[0254] On the encoding side, the prediction tree structure is established using two different methods: based on KD-Tree (high-latency slow mode) and using lidar calibration information (low-latency fast mode). Using lidar calibration information, each point can be divided into different lasers, and the prediction tree structure is established according to different lasers. Next, based on the structure of the prediction tree, each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter. Finally, through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream.
[0255] At the decoding end, the decoding end reconstructs the prediction tree structure by continuously parsing the bit stream. Secondly, the geometric position prediction residual information and quantization parameters of each prediction node are obtained through parsing, and the prediction residual is dequantized to restore the reconstructed geometric position information of each node, finally completing the geometric reconstruction at the decoding end.
[0256] There are mainly three transformation methods for encoding attribute information, namely lifting transformation, prediction transformation, and RAHT transformation. Lifting transformation and prediction transformation perform point cloud prediction and transformation based on the generation order of LOD, while RAHT transformation adaptively transforms attribute information from bottom to top according to the construction levels of the octree. Next, these three transformation encoding methods will be introduced. First, the prediction transformation encoding method will be introduced in detail below.
[0257] In the current attribute prediction module of G-PCC, a nearest neighbor attribute prediction coding scheme based on the LOD structure is adopted. The construction methods of LOD include the LOD construction scheme based on distance, the LOD construction scheme based on a fixed sampling rate, and the LOD construction scheme based on the octree, etc. In the LOD construction scheme based on distance, before constructing LOD, the point cloud is first sorted by Morton to ensure strong attribute correlation between adjacent points. Fig. 20 shows a schematic diagram of the LOD construction based on distance. As shown in Fig. 20, the point cloud is divided into L different point cloud detail levels (Rl)l = 0, 1, … L - 1 according to the pre-set L Manhattan distances (dl)l = 0, 1, … L - 1, where (dl)l = 0, 1, … L - 1 satisfies dl < dl-1.
[0258] The construction process of LOD is as follows: (1) First, mark all points in the point cloud as unvisited, and establish a set V to store the set of visited points; (2) In each iteration, traverse the points in the point cloud. If the current point has been visited, ignore this point, otherwise calculate the minimum distance D from the current point to the point set V. If D < dl, ignore this point; if D ≥ dl, mark the current point as visited and add the current point to the refinement level Rl and the point set V; (3) The points in the detail level LODl are composed of the points in the refinement levels R0, R1, R2 … Rl; (4) Continuously repeat the above steps until all points are marked as visited.
[0259] Based on the LOD structure, the attribute information of each point is linearly weighted predicted by using the reconstructed attribute information of points in the same level or a higher level of LOD. Among them, the maximum number of reference prediction neighbor points is determined by the high-level syntax elements of the encoder. For the attribute of each point, at the encoding end, the rate-distortion optimization algorithm is used to select the attribute information of the N nearest neighbor points found for weighted prediction; alternatively, the attribute information of a single nearest neighbor point is selected for prediction, and finally the selected prediction mode and prediction residual are encoded. In the LOD structure, the formula for predicting the attribute information of the current point can be shown as follows:
[0260] Among them, N represents the number of predicted points in the nearest neighbor point set of the current point i, Pi represents the sum of the N nearest neighbor points of the current point i, Dm represents the spatial geometric distance from the nearest neighbor point m to the current point i, Attrm represents the attribute information of the nearest neighbor point m after reconstruction, Attr i ′ represents the attribute prediction information of the current point i, and the number of points N is a preset value.
[0261] To balance attribute encoding performance and parallel processing between different LOD layers, a switch can be introduced in the encoder's high-level syntax elements. This switch controls whether intra-LOD prediction is used. For example, if the switch is turned on, intra-LOD prediction is enabled, allowing predictions to be made using points within the same LOD layer. It should be noted that if the number of LOD layers is 1, intra-LOD prediction is always used.
[0262] Figure 21 shows a schematic diagram of the distance-based LOD point cloud generation process. As shown in Figure 21, the first image on the left is the original point cloud, and the second image on the left represents the point cloud's outer contours. In subsequent images, as the number of detail layers increases, the point cloud's details become increasingly clear. The following section details the process of predicting attribute information for LOD-structured point clouds.
[0263] Figure 22 shows a schematic diagram of the encoding process of the attribute information of the LOD point cloud. After the LOD is constructed, according to the generation order of the LOD, the three nearest neighbor points of the current point to be encoded are first found from the encoded data points. The attribute reconstruction values of these three nearest neighbor points are used as candidate prediction values of the current point to be encoded; then, the optimal prediction value is selected from them according to the rate-distortion optimization algorithm. For example, as shown in Table 1, when encoding the attribute value of point P2 in Figure 10, the prediction variable index of the attribute value of the nearest neighbor point P4 can be set to 1; the attribute prediction variable index of the second nearest neighbor point P5 and the third nearest neighbor point P0 can be set to 2 and 3 respectively; the prediction variable index of the weighted average of points P0, P5 and P4 is set to 0; finally, the rate-distortion optimization algorithm is used to select the best prediction variable. Among them, the formula for weighted average is as follows:
[0264] In the formula Represents the spatial geometric weight of the neighboring point j to the current point i:
[0265] Represents the attribute prediction value of the current point i, j represents the index of the three neighboring points, Represents the attribute value after reconstruction of the neighboring points, x i ,y i ,z i is the geometric position coordinate of the current point i, x ij,y ij ,z ij is the geometric coordinate of the neighboring point j.
[0266] Table 1
[0267] The attribute prediction value of the current point i is obtained through the above prediction (k is the total number of points in the point cloud). Let (a i ) i∈0…k-1 is the original attribute value of the current point, then the attribute residual (r i ) i∈0…k-1 Denoted as:
[0268] Further quantify the prediction residuals:
[0269] In formula (22), Q i represents the quantized attribute residual at the current point i. Qs is the quantization step (Qs), which can be calculated from the quantization parameter (QP). After quantization, the quantization coefficients are arithmetic coded to generate the attribute bit rate.
[0270] During the encoding process, the encoder reconstructs the attribute value of the current point i. The purpose of reconstruction is to predict the subsequent points. Before reconstructing the attribute value, the residual must be dequantized. is the residual after inverse quantization:
[0271] and predicted value Add up to get the reconstruction value of point i
[0272] As described above, based on LOD partitioning, predicting the attribute value of the current point requires a nearest neighbor search. Currently, there are two main types of nearest neighbor search methods: intra-frame nearest neighbor search and inter-frame nearest neighbor search. The following sections describe these two nearest neighbor search methods in detail.
[0273] Intra-frame nearest neighbor search can be divided into two methods: inter-layer nearest neighbor search and intra-layer nearest neighbor search. First, let's introduce inter-layer nearest neighbor search. Figure 23 shows a schematic diagram of the structure of the refinement layer based on LOD division. As shown in Figure 23, after LOD division, different refinement layers R will form a pyramid-like structure. The method of inter-layer nearest neighbor search can be shown in Figure 24. First, based on the method shown in Figure 20, the geometric information is divided into different LOD layers, and LOD0, LOD1 and LOD2 are obtained. In the process of inter-layer nearest neighbor search, the points in LOD0 are used to predict the attributes of the points in the next layer LOD. Next, the process of inter-layer nearest neighbor search is introduced in detail.
[0274] During the entire LOD partitioning process, there are three sets: O(k), L(k), and I(k). Among them, k is the index of the LOD layer during LOD partitioning, and I(k) is the input point set during the current LOD layer partitioning. After LOD partitioning, the O(k) set and L(k) set are obtained. The O(k) set stores the sampling point set, and L(k) is the point set in the current LOD layer. The entire LOD partitioning process is as follows:
[0275] (1) Initialization;
[0276] if k=0,L(k)←{}otherwise L(k)←L(k-1)
[0277] O(k)←{}
[0278] (2) Using the LOD partitioning algorithm, the sampling points are stored in O(k), and the remaining points are divided into L(k);
[0279] (3) When the next iteration is performed, I←O(k).
[0280] It's important to note that since the LOD partitioning process is based on Morton codes, O(k), L(k), and I(k) store the Morton code index corresponding to the point. When performing inter-layer nearest neighbor search, the nearest neighbor search for a point in the L(k) set is performed in the O(k) set. The specific search method is described in detail below.
[0281] First, the nearest neighbor search is performed based on the spatial relationship. As shown in Figure 25, when predicting the current point P, a neighbor search is performed by using the parent block (Block B) corresponding to point P. Figure 26 shows a schematic diagram of neighbor blocks that are coplanar, colinear, and co-point with the current parent block. As shown in Figure 26, points in the coplanar and co-linear neighbor blocks with the current parent block are searched in Figure 26 to perform attribute prediction. That is to say, the coordinates of the current point are used to obtain the corresponding spatial block, and then, the nearest neighbor search is performed in the previously encoded LOD layer to find the spatial blocks that are coplanar, colinear, and co-point with the current block to obtain the N nearest neighbors of the current point.
[0282] If the N nearest neighbors of the current point are still not obtained after the nearest neighbor search for coplanar, colinear and co-point points, then the N nearest neighbors of the current point will be obtained based on the fast search algorithm. The specific method can be seen in Figure 27. Figure 27 shows a schematic diagram of the method of performing the nearest neighbor search for the current point. As shown in Figure 27, when performing inter-layer prediction of attributes, the geometric coordinates of the current point can be used to obtain the Morton code corresponding to the current point. Then, based on the Morton code of the current point, the first reference point (j) that is larger than the Morton code of the current point is found in the reference frame, and the nearest neighbor search is performed within the range of [j-searchRange, j+searchRange]. The specific method of updating the nearest neighbor is the same as the method of the inter-frame nearest neighbor search, which will be described when introducing the inter-frame nearest neighbor search and will not be repeated here. The following is a detailed introduction to the intra-layer nearest neighbor search.
[0283] Figure 28 shows a schematic diagram of the method of nearest neighbor search within the attribute information layer. As shown in Figure 28, when the intra-layer prediction method is turned on, the nearest neighbor search will be performed in the encoded point set within the same layer LOD to obtain the N nearest neighbors of the current point (the inter-layer nearest neighbor search is also performed). The method of performing the nearest neighbor search can be based on a quick search. For example, as shown in Figure 29, assuming that the Morton code index of the current point is i, the nearest neighbor search will be performed in [i+1, i+searchRange]. The specific nearest neighbor search method is consistent with the inter-frame block-based quick search method, which will not be repeated here. The inter-frame nearest neighbor search method is introduced in detail below.
[0284] Continuing to refer to Figure 27, when performing attribute inter-frame prediction, the geometric coordinates of the current point are used to obtain the Morton code corresponding to the current point. Based on the Morton code of the current point, the first reference point (j) with a Morton code larger than the current point is found in the reference frame, and then the nearest neighbor search is performed within the range of [j-searchRange, j+searchRange].
[0285] Currently, when performing nearest neighbor searches within and between frames, the neighborhood search is performed based on blocks. FIG30 shows a schematic diagram of a neighborhood search prediction structure based on Morton codes. For example, as shown in FIG30 , when performing neighborhood search for the current point (Morton code index is i), the points in the reference frame are divided into N (N=3) layers according to the Morton code. The division method can be as follows:
[0286] First layer: Assuming that the number of points in the reference frame is numPoints, first divide the points in the reference frame into M (M=2 5 =32) points are divided into one block;
[0287] Second layer: Based on the first layer, the blocks of the first layer are also processed every M (M=2 5 =32) blocks are divided into one block;
[0288] The third layer: Based on the second layer, the blocks of the first layer are also processed every M (M=2 5 =32) blocks are divided into one block.
[0289] Finally, the predicted structure shown in Figure 30 is obtained.
[0290] When performing attribute prediction based on the prediction structure shown in Figure 30, assume that the Morton code index of the current point to be encoded is i, and the first point in the reference frame with a Morton code greater than or equal to the current point has an index of j. The block index of the reference point is calculated based on j, and the specific calculation method is as follows:
[0291] First layer: BucketSize_0 = 2 5 =32;
[0292] Second layer: BucketSize_1=2 5 =32×BucketSize_0=1024;
[0293] Third layer: BucketSize_2=2 5 =32×BucketSize_1=32768.
[0294] Assume that the reference range in the prediction frame for the current point is [j-searchRange, j+searchRange]. Use j-searchRange to calculate the starting index of the third layer, and j+searchRange to calculate the ending index of the third layer. First, determine whether some blocks in the second layer need to be searched for their nearest neighbors within the blocks in the third layer. Then, for each block in the first layer, determine whether a search is required. If some blocks in the first layer need to be searched for their nearest neighbors, a point-by-point search is performed on some of the blocks in the first layer to update the nearest neighbors. The following describes the method for calculating blocks based on indexes.
[0295] Assuming that the Morton code index corresponding to the current point is index, then the index of the corresponding third-layer block is: idx_2=index / BucketSize_2 (25)
[0296] After obtaining the block index idx_2 of the third layer, the start index and end index of the block corresponding to the current block in the second layer can be obtained using idx_2: startIdx1=idx_2×BucketSize_1 (26) endIdx=idx_2×BucketSize_1+BucketSize_1-1 (27)
[0297] Based on the same algorithm, the index of the first layer block is obtained according to the index of the second layer block.
[0298] When performing a block-based nearest neighbor search, it will determine whether the current block needs to be searched for the nearest neighbor, that is, the nearest neighbor search of the filtered block. Each spatial block can obtain minPos and maxPos through two variables. MinPos represents the minimum value of the block, and maxPos represents the maximum value of the block. Assume that the distance to the farthest point among the N nearest neighbors of the current point is Dist, the coordinates of the point to be encoded are (x, y, z), and the current block is represented by (minPos, maxPos), where minPos is the minimum value of the three dimensions of the bounding box and maxPos is the maximum value of the three dimensions of the bounding box. The distance D between the current point and the bounding box is calculated as follows: int dx=int(std::max(std::max(minPos[0]-point[0],0),point[0]-maxPos[0])) (28) int dy=int(std::max(std::max(minPos[1]-point[1],0),point[1]-maxPos[1])) (29) int dz=int(std::max(std::max(minPos[2]-point[2],0),point[2]-maxPos[2])) (30) D=dx+dy+dz (31)
[0299] When D is less than or equal to Dist, the points in the current block will be traversed.
[0300] The above describes the method of predictive transform encoding of point cloud attribute information. Next, the method of lifting transform encoding of point cloud attribute information will be described in detail.
[0301] Figure 31 shows a schematic diagram of the encoding process of the lifting transform. As shown in Figure 31, the lifting transform also predicts and encodes the point cloud attributes based on LOD. The difference from the prediction transform described above is that the lifting transform divides the LOD into high and low layers. Then, the prediction is performed in the reverse order of the LOD generation layer, and an update operator is introduced in the prediction process to update the quantization weights of the low-level LOD midpoints to improve the accuracy of the prediction. The attribute values of the low-level LOD midpoints are frequently used to predict the attribute values of the high-level LOD midpoints, so the points in the low-level LOD should have greater influence. Continuing to refer to Figure 31, the encoding method of the lifting transform can be divided into three steps, namely: segmentation process, prediction process and update process. The following will introduce these three steps in detail.
[0302] Step 1: Segmentation Process
[0303] The segmentation process is to divide the complete LOD layer into a low LOD layer L(N) and a high LOD layer H(N). If a point cloud has three LOD layers, namely (LOD l ) l=0,1,2 , after segmentation, LOD2 is the high LOD layer, denoted as H(N), (LOD l ) l=0,1 It is the low LOD layer, denoted as L(N).
[0304] Step 2: Prediction Process
[0305] The point in the high-level LOD selects the attribute information of the nearest neighbor point from the low-level LOD as the attribute prediction value P(N) of the current point to be coded, and the prediction residual D(N) is recorded as: D(N) = H(N) - P(N) (32)
[0306] Step 3: Update Process
[0307] The attribute prediction residual D(N) in the high-level LOD is updated to obtain U(N), and the attribute value of the midpoint of the low-level LOD is improved using U(N), as shown in the following formula: L′(N)=L(N)+U(N) (33)
[0308] The above process will iterate continuously until the lowest LOD according to the order of LOD from high to low.
[0309] The LOD-based prediction scheme gives points in the lower LOD layers greater influence. The lifting wavelet transform-based transformation introduces quantization weights and updates the prediction residual based on the prediction residual D(N) and the distance between the prediction point and its adjacent points. Finally, the prediction residual is adaptively quantized using the quantization weights from the transformation process. It should be noted that the quantization weight value for each point can be determined by geometric reconstruction at the decoder, so the quantization weights should not be encoded.
[0310] The RAHT transform uses the Haar wavelet transform, which can transform the attribute information of the point cloud from the spatial domain to the frequency domain, thereby further reducing the correlation between the attribute information of the point cloud. Figure 32 is an example diagram of the RAHT transform process. As shown in Figure 32, RAHT performs wavelet transform based on the hierarchical structure of the octree, thereby associating the attribute information with the octree nodes. The attribute information of the occupied nodes in the same parent node is recursively transformed in a bottom-up manner, and the nodes in each layer are transformed from the three dimensions of x, y, and z (see Figure 33) until the root node of the octree is transformed. In the process of hierarchical transformation, the direct current (DC) coefficient (or low-pass coefficient) obtained after the transformation of the nodes in the same layer is passed to the nodes in the upper layer for further transformation, and all alternating current (AC) coefficients (or high-pass coefficients) will be quantized and encoded.
[0311] Figure 34 is a schematic diagram of RAHT transformation and inverse RAHT transformation. Assume that g′ L,2x,y,z And g′L,2x+1,y,z are the DC coefficients of two neighboring points in the L layer. After RAHT transformation, the information of the L-1 layer is the AC coefficient f′ L-1,x,y,z and DC coefficient g′ L-1,x,y,z . f′ L-1,x,y,z No more transformation will be performed, and quantization coding will be performed directly, g′ L-1,x,y,z We will continue to search for neighboring points to transform. If no neighboring points are found, we can transform g′ L-1,x,y,z Directly passed to the L-2 layer. That is to say, RAHT transformation is only valid for nodes with neighboring points, and nodes without neighboring points will be directly passed to the previous layer. In the above transformation process, g′ L,2x,y,z The weights corresponding to g′L, 2x+2, y, and z (the weights can be determined based on the number of non-empty child nodes in the node) are w′ L,2x,y,z and w′L,2x+1,y,z (abbreviated as w′0 and w′1), g′ L-1,x,y,z The weight is w′ L-1,x,y,z , then the general transformation formula of RAHT transformation is:
[0312] In formula (19), T w0,w1 is the transformation matrix, which can be determined based on formula (35):
[0313] The transformation matrix will be updated adaptively as the weights corresponding to each point are transformed. The above transformation process will be iterated and updated continuously according to the partitioning structure of the octree until the root node of the octree is reached.
[0314] Based on the RAHT transform, RAHT intra-frame prediction can be performed on the attribute information, that is, RAHT intra-frame prediction combined with transform coding can be performed on the attribute information. This coding mode is described in detail below.
[0315] As shown in Figure 35, the RAHT transform is based on the order of the octree hierarchy, and continuously transforms from the voxel level until the root node is obtained, thereby completing the hierarchical transform coding of the entire attribute information. In RAHT intra-frame prediction combined with transform coding, the attribute information can also be predicted and transformed based on the hierarchical order of the octree. The difference is that the process of RAHT intra-frame prediction combined with transform coding can continuously transform from the root node to the voxel level. In each RAHT transform process, the attribute information can be predicted and transformed based on a 2×2×2 block.
[0316] The structure of the encoding block of attribute information can be seen in Figure 35. The dark gray block in Figure 35 is the current block to be encoded, and the light gray block is the neighboring block coplanar and colinear with the current block. The attribute information of the current block can be normalized based on equations (36) to (38): node =∑ p∈node attribute(p) (36) w node =∑ p∈node 1=#{p∈node} (37) a node =A node / w node (38)
[0317] Specifically, we can first obtain the attribute information of the current block based on the attribute information of the nodes in the current block, that is, A node For example, a simple sum operation can be performed on the attribute information of the nodes in the current block to determine A node Then, we can use the attribute information of the current block and the number of nodes in the current block (i.e., w node ) is normalized to obtain the mean value a of the attribute information of the current block node Then, the mean value of the attribute information of the current block can be used for transform coding.
[0318] Figure 36 shows the overall process of RAHT intra-frame prediction combined with transform coding of attribute information. (d) in Figure 36 shows the attribute information of the current block, and (e) in Figure 36 shows the attribute information of the predicted block obtained by linear weighted fitting using the neighborhood attribute information of the current block. Then, the attribute information of the current block and the attribute information of the predicted block can be attribute transformed respectively to obtain DC coefficients and AC coefficients. Then, the AC coefficients can be predictively encoded. Among them, the attribute information of the predicted block is obtained by linear fitting based on the method shown in Figure 37.
[0319] FIG37 is an example diagram of a linear fitting method for the neighborhood attribute information of the current block. As shown in FIG37 , first, the 19 neighborhood blocks of the current block can be determined. Secondly, the spatial geometric distance between the neighborhood block and each sub-block in the current block can be used to perform linear weighted prediction on the attribute information of each sub-block to obtain the attribute information of the predicted block. Then, the attribute information of the predicted block can be transformed. Exemplarily, equations (39) to (41) in FIG38 can be used to predict and transform the attribute information (Equation (39) represents the transformation method of the attribute information of the current block, equation (40) represents the transformation method of the attribute information of the predicted block, and equation (41) outputs the predicted residual information):
[0320] When performing inter-frame prediction coding of attribute information, if inter-frame prediction coding is started, the RAHT attribute transform coding structure will first be constructed based on the geometric information of the current node, that is, the nodes will be continuously merged at the voxel level until the root node of the entire RAHT transform tree is obtained, thereby obtaining the transform coding hierarchical structure corresponding to the attribute information. Then, according to the RAHT transform structure, the root node can be divided to obtain N child nodes of each node (N is less than or equal to 8). Unlike the RAHT intra-frame prediction combined with transformation coding mode, the RAHT inter-frame prediction combined with transformation coding mode will utilize the node information of the reference frame. For example, the attribute information of the N child nodes of the current node can be RAHT transformed to obtain DC and AC coefficients. Secondly, the AC coefficients of the N child nodes can be inter-frame predicted in the following way.
[0321] For example, if the inter-frame prediction node of the current node is valid (ie, the co-located node of the current node in the reference frame exists), the attribute information of the prediction node is directly used as the attribute prediction value of the current node.
[0322] For another example, if the current node can find a node with exactly the same position as the current node in the cache of the reference frame (that is, the current node exists in the same node in the reference frame), then the attribute prediction values of the AC coefficients of the N child nodes of the current node can be determined based on the AC coefficients of the M child nodes contained in the same node. For example, if the AC coefficient of the inter-frame prediction node corresponding to a child node is not zero, the AC coefficient of the inter-frame prediction node is directly used as the prediction value of the child node; if the AC coefficient of the inter-frame prediction node corresponding to a child node is zero, the AC coefficient of the intra-frame prediction node corresponding to the child node can be used as the prediction value.
[0323] For another example, if the inter-frame prediction node of the current node is invalid (ie, the co-located node of the current node in the reference frame does not exist), the attribute prediction value of the adjacent node in the frame can be used as the attribute prediction value of the current node.
[0324] In addition, after RAHT inter prediction is enabled, the optimal RAHT prediction mode can be selected for each layer. The RAHT prediction mode can be either RAHT intra prediction mode or RAHT inter prediction mode. If the cost of the RAHT intra prediction mode is less than the cost of the RAHT inter prediction mode, RAHT intra prediction can be performed on the current layer; otherwise, RAHT inter prediction is performed.
[0325] In the point cloud codec framework, when predicting and decoding point cloud attribute information, a decision is made as to whether to enable bidirectional inter-frame prediction. In related technologies, when bidirectional inter-frame prediction is used for attribute information, conflicting identification information at different levels (e.g., sequence-level identification information and slice-level identification information) can result in decoding errors.
[0326] For example, under the G-PCC codec framework, when the related technology performs inter-frame prediction coding on attribute information, first, in the high-level attribute parameter set (APS), a syntax element (such as attrInterPredictionEnabled) is used to determine whether to use attribute inter-frame prediction coding. Then, when predictive coding is performed on the attribute information in each slice, geometric information is used to adaptively determine whether the attribute information of the current slice enables attribute inter-frame prediction coding. Next, the syntax element slice_attr_inter_prediction is passed to the decoding end, and the decoding end determines whether inter-frame prediction coding is enabled for the current slice by parsing the current syntax element. If the attribute information is bidirectionally inter-frame predictive coded, then at the encoding end, first determine whether attrInterPredictionEnabled in APS is turned on, and determine whether the biPredictionEnabledFlag parameter in the geometry parameter set (GPS) is turned on; if both are turned on at the same time, then when encoding the slice in the current frame, add a disableAttrInterPredForRefFrame2 parameter to the high-level syntax element of the attribute information of each slice to determine whether bidirectional inter-frame attribute predictive coding is turned on for the current slice.
[0327] According to the above introduction, in the related art, when inter-frame prediction encoding of attribute information is performed based on G-PCC, at the decoding end, attrInterPredictionEnabled in APS is used to determine whether inter-frame prediction decoding of attribute information is enabled for each frame (or sequence); biPredictionEnabledFlag in GPS is used to determine whether bidirectional inter-frame prediction decoding is enabled for each frame (or sequence). When inter-frame prediction decoding is enabled for the attribute information of each frame (or sequence), for each slice, the syntax element slice_attr_inter_prediction is used to determine whether the current slice enables inter-frame prediction decoding of attribute information. Whether to parse the syntax element disableAttrInterPredForRefFrame2 in the slice depends only on whether attrInterPredictionEnabled in APS is enabled. That is, in the related art, whether the syntax element disableAttrInterPredForRefFrame2 is parsed has nothing to do with whether the syntax element biPredictionEnabledFlag in GPS is enabled. There may be some problems with parsing the syntax element disableAttrInterPredForRefFrame2 in the above manner, which are described in detail below.
[0328] For example, if APS indicates that inter-frame prediction decoding of attribute information is turned on, but bidirectional inter-frame prediction decoding is turned off (that is, biPredictionEnabledFlag in GPS is turned off); at this time, not only is it necessary to encode slice_attr_inter_prediction in the slice to determine whether the current slice has inter-frame prediction turned on, but it is also necessary to encode the syntax element disableAttrInterPredForRefFrame2 in the slice to determine whether the current slice has bidirectional attribute inter-frame prediction turned on. However, since the above-mentioned bidirectional inter-frame prediction decoding has been turned off, the syntax element disableAttrInterPredForRefFrame2 has no effect on the decoding end, that is, bidirectional inter-frame prediction cannot be started. In this case, there is a problem of wasting coding bits.
[0329] In addition, when decoding the point cloud attribute information, the decoding end determines whether to enable bidirectional inter-frame prediction for the attribute information of the first image area based on disableAttrInterPredForRefFrame2, but does not determine whether to enable bidirectional inter-frame prediction for the attribute information of the first image area based on biPredictionEnabledFlag. Compared with disableAttrInterPredForRefFrame2, biPredictionEnabledFlag is a higher-level identification information. If biPredictionEnabledFlag indicates that bidirectional inter-frame prediction is not enabled, the decoding end will not cache the bidirectional reference frame used for bidirectional inter-frame prediction. However, in this case, if disableAttrInterPredForRefFrame2 indicates that bidirectional inter-frame prediction is enabled for the attribute information of the first image area, the decoding end will encounter the problem of being unable to decode.
[0330] In response to the above problems, an embodiment of the present application provides an encoding method, including: determining whether to enable bidirectional inter-frame prediction for the attribute information of at least one frame of image; if bidirectional inter-frame prediction is enabled for the attribute information of the at least one frame of image, encoding second identification information, wherein the second identification information is used to indicate whether bidirectional inter-frame prediction is enabled for the attribute information of a first image area, where the first image area is an image area in a frame of image.
[0331] An embodiment of the present application also provides a decoding method, including: parsing first identification information, the first identification information is used to indicate whether bidirectional inter-frame prediction is enabled for at least one frame of image; if the first identification information indicates that bidirectional inter-frame prediction is enabled for the at least one frame of image, parsing second identification information, the second identification information is used to indicate whether bidirectional inter-frame prediction is enabled for attribute information of a first image area, the first image area being an image area in a frame of image.
[0332] Based on the above, the embodiment of the present application determines whether to parse the second identification information based on the first identification information (the first identification information is identification information of a higher level than the second identification information), thereby avoiding the problem of decoding failure caused by the conflict between the two identification information. In addition, the decoding end determines whether to parse the second identification information based on the first identification information, which means that if the first identification information indicates that bidirectional inter-frame prediction is not enabled, the encoding end can not write the second identification information into the bitstream, thereby saving coding bits.
[0333] The following will describe in detail the point cloud decoding method provided in the embodiment of the present application with reference to the accompanying drawings.
[0334] Figure 39 is a flow chart of the point cloud decoding method provided in an embodiment of the present application. The point cloud decoding method of Figure 39 can be applied to a decoder. The point cloud decoding method of Figure 39 can be used to decode the attribute information of a point cloud. In some implementations, the point cloud decoding method can be applied to G-PCC. Alternatively, in other implementations, the point cloud decoding method can be applied to a geometry-based solid content test model (GES-TM). GES-TM is a coding and decoding framework proposed for dense point clouds, such as point clouds collected in augmented reality (AR) or virtual reality (VR) scenes.
[0335] Referring to FIG. 39 , at step 3910, first identification information is parsed. The first identification information is used to indicate whether bidirectional inter-frame prediction is enabled for at least one image frame (or, the first identification information is used to indicate whether bidirectional inter-frame prediction is enabled for geometric information and / or attribute information of at least one image frame). The at least one image frame may be a single image frame, or the at least one image frame may be an image sequence.
[0336] In some implementations, the first identification information may be sequence-level or frame-level information. For example, the first identification information may belong to a geometry parameter set (GPS). In another example, the first identification information may belong to an attribute parameter set (APS). In another example, the first identification information may belong to a sequence parameter set.
[0337] In step S3920, if the first identification information indicates that bidirectional inter-frame prediction is enabled for at least one frame of image, the second identification information is parsed. The second identification information is used to indicate whether bidirectional inter-frame prediction is enabled for attribute information of a first image region. The first image region is an image region within a frame of image.
[0338] The embodiments of the present application do not specifically limit the first image area. For example, the first image area may include a frame of image. In another example, the first image area may include one or more slices. In another example, the first image area may include one or more tiles. In another example, the first image area may include one or more RAHT layers.
[0339] The second identification information may be carried at any position in the code stream. The second identification information may be slice-level information. For example, the second identification information may be carried in an attribute brick header (ABH).
[0340] In some implementations, if the first identification information indicates that bidirectional inter-frame prediction is not enabled for at least one frame of image, the method of FIG39 further includes: not parsing the second identification information. That is, if the attribute information in at least one frame of image does not enable bidirectional inter-frame prediction, then bidirectional inter-frame prediction is not performed when decoding the image area (such as a strip) in the current frame (belonging to the at least one frame of image). In this way, the second identification information can be encoded at the encoding end (in the related art, even if the first identification information indicates that bidirectional inter-frame prediction is not enabled for at least one frame of image, the decoding end will determine whether to enable bidirectional inter-frame prediction for the attribute information by parsing the second identification information), thereby avoiding redundant information from being encoded into the bitstream, which helps to improve encoding and decoding performance. In some implementations, if the second identification information is not parsed, the method of FIG39 further includes: not enabling bidirectional inter-frame prediction for the attribute information of the first image area.
[0341] Furthermore, in some implementations, the method of Figure 39 also includes: performing inter-frame prediction on the attribute information of the first image area based on a unidirectional reference frame, thereby determining the predicted value of the attribute information of the first image area; and then, determining the reconstructed value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area and the residual value of the attribute information of the first image area.
[0342] In some implementations, if the second identification information indicates that bidirectional inter-frame prediction is enabled for the attribute information of the first image area, then the method of Figure 39 also includes: performing inter-frame prediction on the attribute information of the first image area based on the bidirectional reference frame, thereby obtaining a predicted value of the attribute information of the first image area; and then, determining a reconstructed value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area and the residual value of the attribute information of the first image area.
[0343] The bidirectional reference frames mentioned above can be of various types. For example, a bidirectional reference frame can be two forward reference frames of the current frame in the time domain. In another example, a bidirectional reference frame can be two backward reference frames of the current frame in the time domain. In another example, a bidirectional reference frame can be one forward reference frame and one backward reference frame of the current frame in the time domain.
[0344] In some implementations, if the second identification information indicates that bidirectional inter-frame prediction is not enabled for the attribute information of the first image area, the method of Figure 39 also includes: performing inter-frame prediction on the attribute information of the first image area based on a unidirectional reference frame, thereby determining the predicted value of the attribute information of the first image area; and then determining the reconstructed value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area and the residual value of the attribute information of the first image area.
[0345] In some implementations, the method of FIG. 39 may include parsing third identification information, where the third identification information indicates whether inter-frame prediction is enabled for attribute information of at least one frame of image. The third identification information may be sequence-level information or frame-level information. As an example, the third identification information may belong to an attribute parameter set.
[0346] As mentioned above, whether to resolve the second identification information can be determined based on the indication information of the first identification information. In some cases, whether the second identification information is resolved is also related to the third identification information. In other words, the first identification information can be combined with the third identification information to determine whether to resolve the second identification information.
[0347] For example, if the first identification information indicates that bidirectional inter-frame prediction is enabled for at least one frame of image, and the third identification information indicates that inter-frame prediction is enabled for attribute information of at least one frame of image, then the second identification information is parsed.
[0348] For another example, if the first identification information indicates that bidirectional inter-frame prediction is not enabled for at least one frame of image, and the third identification information indicates that inter-frame prediction is enabled for attribute information of at least one frame of image, the second identification information is not parsed.
[0349] For another example, if the first identification information indicates that bidirectional inter-frame prediction is not enabled for at least one frame of image, and the third identification information indicates that inter-frame prediction is enabled for attribute information of at least one frame of image, the second identification information is not parsed.
[0350] In the related art, if at least one frame of image turns on inter-frame prediction, the relevant syntax elements will be parsed during the decoding process of the current strip to confirm whether the attribute information of the strip turns on bidirectional inter-frame prediction. If at least one frame of image turns off bidirectional inter-frame prediction, even if the parsed syntax elements indicate that the attribute information of the current strip turns on bidirectional inter-frame prediction, the strip will not turn on bidirectional inter-frame prediction. In other words, the above syntax elements will not work at the decoding end. In an embodiment of the present application, the second identification information will be parsed only when at least one frame of image turns on bidirectional inter-frame prediction. In other words, when at least one frame of image turns off bidirectional inter-frame prediction, the second identification information will not be parsed. Accordingly, the second identification information can be not encoded into the bitstream at the encoding end, that is, unnecessary information can be avoided from being encoded into the bitstream, which helps to improve encoding and decoding performance.
[0351] In some implementations, if the third identification information indicates that inter-frame prediction is enabled for the attribute information of at least one image frame, the method of FIG39 further includes parsing fourth identification information. The fourth identification information is used to indicate whether inter-frame prediction is enabled for the attribute information of the first image region. The fourth identification information may be slice-level information. For example, the fourth identification information may be carried in the ABH.
[0352] Alternatively, in some other implementations, if the third identification information indicates that inter-frame prediction is not enabled for the attribute information of at least one frame of image, then the method of FIG39 further includes: not parsing the fourth identification information. If the attribute information in a frame of image does not enable inter-frame prediction, then inter-frame prediction is not performed on the image region (such as a strip) in the image during decoding. In other words, the fourth identification information can be omitted from encoding at the encoding end, thereby avoiding redundant information from being included in the bitstream, which helps improve encoding and decoding performance.
[0353] In some implementations, if the fourth identification information is not parsed, the method of FIG. 39 further includes: not enabling inter-frame prediction for the attribute information of the first image region.
[0354] In some implementations, if the third identification information indicates that inter-frame prediction is not enabled for the attribute information of at least one frame of image, that is, the prediction mode of the attribute information of at least one frame of image is intra-frame prediction, the method of FIG39 further includes: performing intra-frame prediction on the attribute information of the first image region to determine a predicted value of the attribute information of the first image region; and then determining a reconstructed value of the attribute information of the first image region based on the predicted value of the attribute information of the first image region and a residual value of the attribute information of the first image region.
[0355] The aforementioned method of determining the reconstruction value may include, for example, taking the sum of the predicted value of the attribute information of the first image region and the residual value of the attribute information of the first image region as the reconstruction value of the attribute information of the first image region.
[0356] As mentioned above, the point cloud decoding method shown in FIG39 can be executed by using various identification information. The following describes in detail how to set these identification information.
[0357] In some implementations, the first identification information may be represented by biPredictionEnabledFlag (of course, the first identification information may also be represented by any other letters and / or numbers). For example, if biPredictionEnabledFlag is true, it may indicate that bidirectional inter-frame prediction is enabled for at least one image frame; if biPredictionEnabledFlag is false, it may indicate that bidirectional inter-frame prediction is not enabled for at least one image frame.
[0358] In some implementations, the second identification information may be represented by disableAttrInterPredForRefFrame2 (of course, the second identification information may also be represented by any other letters and / or numbers). For example, if disableAttrInterPredForRefFrame2 is true, it may indicate that bidirectional inter-frame prediction is enabled for the attribute information of the first image region; if disableAttrInterPredForRefFrame2 is false, it may indicate that bidirectional inter-frame prediction is not enabled for the attribute information of the first image region.
[0359] In some implementations, the third identification information may be represented by attrInterPredictionEnabled (of course, the third identification information may also be represented by any other letters and / or numbers). For example, if attrInterPredictionEnabled is true, it may indicate that inter-frame prediction is enabled for the attribute information of at least one frame of image; if attrInterPredictionEnabled is false, it may indicate that inter-frame prediction is not enabled for the attribute information of at least one frame of image.
[0360] In some implementations, the fourth identification information may be represented by slice_attr_inter_prediction (of course, the fourth identification information may also be represented by any other letters and / or numbers). For example, if slice_attr_inter_prediction is true, it may indicate that inter-frame prediction is enabled for the attribute information of the first image region; if slice_attr_inter_prediction is false, it may indicate that inter-frame prediction is not enabled for the attribute information of the first image region.
[0361] In some implementations, the method of Figure 39 may also include: parsing the code stream to determine the quantization coefficients of the attribute information of the first image area; then, inverse quantizing the quantization coefficients to determine the transformation coefficients of the attribute information of the first image area; then, inverse transforming the transformation coefficients to determine the residual value of the attribute information of the first image area.
[0362] The above describes in detail the point cloud decoding method provided by the embodiment of the present application in conjunction with Figure 39. The following describes in detail the point cloud encoding method provided by the embodiment of the present application in conjunction with Figure 40.
[0363] Figure 40 is a flow chart of the point cloud coding method provided in an embodiment of the present application. The point cloud coding method of Figure 40 can be applied to an encoder. The point cloud coding method of Figure 40 can be used to encode the attribute information of the point cloud. In some implementations, the point cloud coding method can be applied to G-PCC. Alternatively, in other implementations, the point cloud coding method can be applied to a geometry-based solid content test model (GES-TM). GES-TM is a coding and decoding framework proposed for dense point clouds (such as point clouds collected in augmented reality (AR) or virtual reality (VR) scenes).
[0364] 40 , in step 4010 , it is determined whether bidirectional inter-frame prediction is enabled for attribute information of at least one frame of image. The at least one frame of image may be a single frame of image, or may be an image sequence.
[0365] The bidirectional reference frames mentioned above can be of various types. For example, a bidirectional reference frame can be two forward reference frames of the current frame in the time domain. In another example, a bidirectional reference frame can be two backward reference frames of the current frame in the time domain. In another example, a bidirectional reference frame can be one forward reference frame and one backward reference frame of the current frame in the time domain.
[0366] To facilitate the decoding end in determining the inter-frame prediction mode of at least one image frame, in some implementations, the method of FIG. 40 further includes encoding first identification information. The first identification information is used to indicate whether bidirectional inter-frame prediction is enabled for the at least one image frame (or the first identification information is used to indicate whether bidirectional inter-frame prediction is enabled for geometric information and / or attribute information of the at least one image frame).
[0367] In some implementations, the first identification information may be sequence-level or frame-level information. For example, the first identification information may belong to a geometry parameter set (GPS). In another example, the first identification information may belong to an attribute parameter set (APS). In another example, the first identification information may belong to a sequence parameter set. In step S4020, if bidirectional inter-frame prediction is enabled for the attribute information of at least one frame of image, the second identification information is encoded. The first image region is an image region within a frame of image.
[0368] The embodiments of the present application do not specifically limit the first image area. For example, the first image area may include a frame of image. In another example, the first image area may include one or more slices. In another example, the first image area may include one or more tiles. In another example, the first image area may include one or more RAHT layers.
[0369] The second identification information mentioned above can be used to indicate whether to perform bidirectional inter-frame prediction on the attribute information of the first image region. The second identification information is encoded into the bitstream so that the decoding end can determine the inter-frame prediction mode of the attribute information.
[0370] The second identification information may be carried at any position in the code stream. The second identification information may be slice-level information. For example, the second identification information may be carried in an attribute brick header (ABH).
[0371] In some implementations, if bidirectional inter-frame prediction is enabled for the attribute information of at least one frame of image, the method of FIG40 further includes: performing inter-frame prediction on the attribute information of the first image area based on the bidirectional reference frame, thereby determining the predicted value of the attribute information of the first image area; and then, determining the residual value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area.
[0372] In some implementations, if bidirectional inter-frame prediction is not enabled for the attribute information of at least one frame of image, the method of FIG40 further includes: not encoding the second identification information. That is, if bidirectional inter-frame prediction is not enabled for the attribute information in at least one frame of image, then bidirectional inter-frame prediction is not performed when decoding the image area (such as a strip) in the current frame (belonging to the at least one frame of image). In this way, the second identification information may not be encoded at the encoding end (in the related art, even if the first identification information indicates that bidirectional inter-frame prediction is not enabled for at least one frame of image, the decoding end will determine whether to enable bidirectional inter-frame prediction for the attribute information by parsing the second identification information), thereby avoiding redundant information from being encoded into the bitstream, which helps to improve encoding and decoding performance.
[0373] As mentioned above, step 4010 can determine whether bidirectional inter-frame prediction is enabled for the attribute information of at least one frame of image. If step 4010 determines that bidirectional inter-frame prediction is not enabled for the attribute information of at least one frame of image, the prediction mode of the attribute information can be unidirectional inter-frame prediction or intra-frame prediction. Therefore, the method of Figure 40 can also determine whether inter-frame prediction is enabled for the attribute information of at least one frame of image.
[0374] In some implementations, the method of FIG. 40 further includes encoding third identification information. The third identification information is used to indicate whether inter-frame prediction is enabled for attribute information of at least one frame of image. The third identification information can be sequence-level information or frame-level information. As an example, the third identification information can belong to an attribute parameter set.
[0375] In some implementations, if the attribute information of at least one frame of image is inter-frame predicted, and the inter-frame prediction here is unidirectional inter-frame prediction, then the method of Figure 40 also includes: the attribute information of the first image area can be inter-frame predicted based on the unidirectional reference frame, so as to determine the predicted value of the attribute information of the first image area; then, based on the predicted value of the attribute information of the first image area, the residual value of the attribute information of the first image area is determined.
[0376] In some implementations, if inter-frame prediction is enabled for the attribute information of at least one image frame, it is also possible to determine whether inter-frame prediction is performed on a portion of the image frame, and to encode the result into the bitstream as identification information. Therefore, the method of FIG40 further includes encoding fourth identification information. This fourth identification information is used to indicate whether inter-frame prediction is performed on the attribute information of the first image region. The fourth identification information can be slice-level information. For example, the fourth identification information can be carried in the ABH.
[0377] Alternatively, in some other implementations, if inter-frame prediction is not enabled for the attribute information of at least one frame of image, the method of FIG40 further includes: not encoding the fourth identification information. If inter-frame prediction is not performed for the attribute information of at least one frame of image, that is, inter-frame prediction is not performed when encoding the image region (such as a strip) in the current frame (belonging to the at least one frame of image). In this way, the fourth identification information can be omitted from encoding at the encoding end, thereby avoiding encoding redundant information into the bitstream, which helps improve encoding and decoding performance.
[0378] Furthermore, the method of Figure 40 also includes: performing intra-frame prediction on the attribute information of the first image area to determine the predicted value of the attribute information of the first image area; and then, determining the residual value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area.
[0379] The aforementioned method of determining the residual value may include, for example, taking the difference between the predicted value and the original value of the attribute information of the first image region as the residual value of the attribute information of the first image region.
[0380] As mentioned above, the point cloud encoding method shown in FIG40 can be executed by using various identification information. The following describes in detail how to set these identification information.
[0381] In some implementations, the first identification information may be represented by biPredictionEnabledFlag (of course, the first identification information may also be represented by any other letters and / or numbers). For example, if biPredictionEnabledFlag is true, it may indicate that bidirectional inter-frame prediction is enabled for at least one image frame; if biPredictionEnabledFlag is false, it may indicate that bidirectional inter-frame prediction is not enabled for at least one image frame.
[0382] In some implementations, the second identification information may be represented by disableAttrInterPredForRefFrame2 (of course, the second identification information may also be represented by any other letters and / or numbers). For example, if disableAttrInterPredForRefFrame2 is true, it may indicate that bidirectional inter-frame prediction is enabled for the attribute information of the first image region; if disableAttrInterPredForRefFrame2 is false, it may indicate that bidirectional inter-frame prediction is not enabled for the attribute information of the first image region.
[0383] In some implementations, the third identification information may be represented by attrInterPredictionEnabled (of course, the third identification information may also be represented by any other letters and / or numbers). For example, if attrInterPredictionEnabled is true, it may indicate that inter-frame prediction is enabled for the attribute information of at least one frame of image; if attrInterPredictionEnabled is false, it may indicate that inter-frame prediction is not enabled for the attribute information of at least one frame of image.
[0384] In some implementations, the fourth identification information may be represented by slice_attr_inter_prediction (of course, the fourth identification information may also be represented by any other letters and / or numbers). For example, if slice_attr_inter_prediction is true, it may indicate that inter-frame prediction is enabled for the attribute information of the first image region; if slice_attr_inter_prediction is false, it may indicate that inter-frame prediction is not enabled for the attribute information of the first image region.
[0385] In some implementations, the method of FIG. 40 may further include: transforming the residual value of the attribute information of the first image region to determine a transform coefficient; and then, quantizing the transform coefficient to determine a quantization coefficient.
[0386] The following examples are used to describe the embodiments of the present application in more detail. It should be noted that the examples below are only intended to help those skilled in the art understand the embodiments of the present application, rather than to limit the embodiments of the present application to the specific numerical values or specific scenarios illustrated. It is apparent that those skilled in the art can make various equivalent modifications or changes based on the examples given below, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0387] This example modifies the inter-frame prediction encoding and decoding parameters of the attribute information in ABH based on the G-PCC codec framework, as shown below:
[0388] The syntax elements from the 13th to the 10th line from the bottom of the above syntax elements are the new syntax elements introduced in this example based on the syntax elements provided by the relevant technology. According to the above syntax elements, when this example parses the attribute syntax element disableAttrInterPredForRefFrame2 of the current slice, when the bidirectional inter-frame prediction is turned on (that is, biPredictionEnabledFlag is true), it will parse the bidirectional inter-frame prediction syntax element disableAttrInterPredForRefFrame2 of the current slice. When the bidirectional inter-frame prediction is turned off in GPS or APS, the bidirectional inter-frame prediction related parameters at the slice level do not need to be passed, thereby reducing the redundancy in the high-level syntax elements. For other contents of the above syntax elements, please refer to the relevant technology and will not be described in detail here.
[0389] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 40. The device embodiment of the present application is described in detail below in conjunction with Figures 41 to 44. It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.
[0390] FIG41 is a schematic diagram of the structure of a decoder provided by an embodiment of the present application. As shown in FIG41 , the decoder 4100 includes a first decoding unit 4110 and a second decoding unit 4120 .
[0391] The first decoding unit 4110 is configured to parse first identification information, where the first identification information is used to indicate whether bidirectional inter-frame prediction is enabled for at least one frame of image.
[0392] The second decoding unit 4120 is configured to parse the second identification information if the first identification information indicates that bidirectional inter-frame prediction is enabled for the at least one frame image, wherein the second identification information is used to indicate whether bidirectional inter-frame prediction is enabled for the attribute information of the first image area, where the first image area is an image area in a frame image.
[0393] In some implementations, the decoder 4100 further includes: a third decoding unit 4125 configured to skip parsing the second identification information if the first identification information indicates that bidirectional inter-frame prediction is not enabled for the at least one frame of image.
[0394] In some implementations, the decoder 4100 also includes: a fourth decoding unit 4130, configured to parse third identification information, wherein the third identification information is used to indicate whether inter-frame prediction is enabled for the attribute information of the at least one frame image; if the third identification information indicates that inter-frame prediction is not enabled for the attribute information of the at least one frame image, the parsing of the second identification information is skipped.
[0395] In some implementations, the third decoding unit 4125 or the fourth decoding unit 4130 is configured to not enable bidirectional inter-frame prediction for the attribute information of the first image region if parsing of the second identifier is skipped.
[0396] In some implementations, the decoder 4100 further includes: a fifth decoding unit 4135 configured to determine whether to parse the second identification information based on the first identification information if the third identification information indicates that inter-frame prediction is enabled for the attribute information of the at least one frame of image.
[0397] In some implementations, the decoder 4100 also includes: a sixth decoding unit 4140, configured to parse fourth identification information if the third identification information indicates that inter-frame prediction is enabled for the attribute information of the at least one frame of image, and the fourth identification information is used to indicate whether inter-frame prediction is enabled for the attribute information of the first image area.
[0398] In some implementations, the decoder 4100 further includes: a seventh decoding unit 4145 configured to skip parsing the fourth identification information if the third identification information indicates that inter-frame prediction is not enabled for the attribute information of the at least one frame of image.
[0399] In some implementations, the seventh decoding unit 4145 is configured to not enable inter-frame prediction for the attribute information of the first image region if parsing of the fourth identification information is skipped.
[0400] In some implementations, the third identification information belongs to an attribute parameter set.
[0401] In some implementations, the decoder 4100 also includes: an eighth decoding unit 4150, configured to perform intra-frame prediction on the attribute information of the first image area if the third identification information indicates that inter-frame prediction is not enabled for the attribute information of the at least one frame image, and determine the predicted value of the attribute information of the first image area; and determine the reconstructed value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area and the residual value of the attribute information of the first image area.
[0402] In some implementations, the first identification information belongs to an attribute parameter set; or, the first identification information belongs to a geometry parameter set; or, the first identification information belongs to a sequence parameter set.
[0403] In some implementations, the second identification information belongs to an attribute slice header.
[0404] In some implementations, the decoder 4100 also includes: a ninth decoding unit 4155, configured to perform inter-frame prediction on the attribute information of the first image area based on a bidirectional reference frame to determine the predicted value of the attribute information of the first image area if the second identification information indicates that bidirectional inter-frame prediction is turned on for the attribute information of the first image area; and determine the reconstructed value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area and the residual value of the attribute information of the first image area.
[0405] In some implementations, the decoder 4100 also includes: a tenth decoding unit 4160, configured to perform inter-frame prediction on the attribute information of the first image area based on a unidirectional reference frame to determine the predicted value of the attribute information of the first image area if the second identification information indicates that bidirectional inter-frame prediction is not enabled for the attribute information of the first image area; and determine the reconstructed value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area and the residual value of the attribute information of the first image area.
[0406] In some implementations, the decoder 4100 also includes: an eleventh decoding unit 4165, configured to perform inter-frame prediction on the attribute information of the first image area based on a unidirectional reference frame to determine the predicted value of the attribute information of the first image area if the first identification information indicates that bidirectional inter-frame prediction is not enabled for the at least one frame of image; and determine the reconstructed value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area and the residual value of the attribute information of the first image area.
[0407] In some implementations, the decoder 4100 also includes: a twelfth decoding unit 4170, configured to parse the code stream to determine the quantization coefficients of the attribute information of the first image area; perform inverse quantization on the quantization coefficients to determine the transformation coefficients of the attribute information of the first image area; and perform inverse transformation on the transformation coefficients to determine the residual value of the attribute information of the first image area.
[0408] In some implementations, the at least one frame of image is one frame of image; or, the at least one frame of image is an image sequence.
[0409] In some implementations, the first image region is a frame image; or the first image region is a strip; or the first image region is a block; or the first image region is a RAHT layer.
[0410] It is understood that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular device. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0411] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0412] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the decoder 4100. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the decoding method described in any one of the aforementioned embodiments.
[0413] Based on the composition of the above-mentioned decoder 4100 and the computer-readable storage medium, refer to Figure 42, which shows a specific hardware structure diagram of the encoder 4200 provided in an embodiment of the present application. As shown in Figure 42, the encoder 4200 may include: a communication interface 4210, a memory 4220 and a processor 4230; each component is coupled together through a bus system 4240. It can be understood that the bus system 4240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 4240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 4240 in Figure 42. Among them,
[0414] Communication interface 4210, used for sending and receiving signals during the process of sending and receiving information with other external network elements;
[0415] Memory 4220, for storing computer programs;
[0416] The processor 4230 is configured to, when running the computer program, execute:
[0417] Parse the first identification information, where the first identification information is used to indicate whether bidirectional inter-frame prediction is enabled for at least one frame of image; if the first identification information indicates that bidirectional inter-frame prediction is enabled for the at least one frame of image, parse the second identification information, where the second identification information is used to indicate whether bidirectional inter-frame prediction is enabled for attribute information of a first image area, where the first image area is an image area in a frame of image.
[0418] It is understood that the memory 4220 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM). The memory 4220 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0419] Processor 4230 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 4230. The above-mentioned processor 4230 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 4220, and the processor 4230 reads the information in the memory 4220 and completes the steps of the above method in combination with its hardware.
[0420] It is understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSP devices, DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application or a combination thereof. For software implementation, the technology described in the present application can be implemented by a module (such as a process, a function, etc.) that performs the functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0421] Optionally, as another embodiment, the processor 4230 is further configured to execute the decoding method described in any one of the aforementioned embodiments when running the computer program.
[0422] FIG43 is a schematic diagram of the structure of an encoder provided by an embodiment of the present application. As shown in FIG43 , the encoder 4300 includes a first encoding unit 4310 and a second encoding unit 4320 .
[0423] The first encoding unit 4310 is configured to determine whether to enable bidirectional inter-frame prediction for attribute information of at least one frame of image.
[0424] The second encoding unit 4320 is configured to encode second identification information if bidirectional inter-frame prediction is enabled for the attribute information of the at least one frame image, wherein the second identification information is used to indicate whether bidirectional inter-frame prediction is enabled for the attribute information of the first image area, where the first image area is an image area in a frame image.
[0425] In some implementations, the encoder 4300 further includes: a third encoding unit 4330 configured to skip encoding of the second identification information if bidirectional inter-frame prediction is not enabled for the attribute information of the at least one frame of image.
[0426] In some implementations, the encoder 4300 further includes: a fourth encoding unit 4335 configured to encode first identification information, where the first identification information is used to indicate whether bidirectional inter-frame prediction is enabled for at least one frame of image.
[0427] In some implementations, the first identification information belongs to an attribute parameter set; or, the first identification information belongs to a geometry parameter set; or, the first identification information belongs to a sequence parameter set.
[0428] In some implementations, the encoder 4300 also includes: a fifth encoding unit 4340, configured to determine whether inter-frame prediction is enabled for the attribute information of the at least one frame image; if inter-frame prediction is not enabled for the attribute information of the at least one frame image, the encoding of the second identification information is skipped.
[0429] In some implementations, the encoder 4300 further includes: a sixth encoding unit 4345 configured to encode third identification information, where the third identification information is used to indicate whether inter-frame prediction is enabled for attribute information of the at least one frame of image.
[0430] In some implementations, the third identification information belongs to an attribute parameter set.
[0431] In some implementations, the encoder 4300 also includes: a seventh encoding unit 4350, configured to encode fourth identification information if inter-frame prediction is enabled for the attribute information of the at least one frame of image, wherein the fourth identification information is used to indicate whether inter-frame prediction is enabled for the attribute information of the first image area.
[0432] In some implementations, the encoder 4300 further includes: a seventh encoding unit 4355 configured to skip encoding of the fourth identification information if inter-frame prediction is not enabled for the attribute information of the at least one frame of image.
[0433] In some implementations, the encoder 4300 also includes: an eighth encoding unit 4360, configured to perform intra-frame prediction on the attribute information of the first image area if inter-frame prediction is not enabled for the attribute information of the at least one frame of image, and determine the predicted value of the attribute information of the first image area; and determine the residual value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area.
[0434] In some implementations, the second identification information belongs to an attribute slice header.
[0435] In some implementations, the encoder 4300 also includes: a ninth encoding unit 4365, configured to perform inter-frame prediction on the attribute information of the first image area based on a bidirectional reference frame to determine a predicted value of the attribute information of the first image area if bidirectional inter-frame prediction is turned on for the attribute information of the first image area; and determine a residual value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area.
[0436] In some implementations, the encoder 4300 also includes: a tenth encoding unit 4370, configured to, if bidirectional inter-frame prediction is not enabled for the attribute information of the first image area, perform inter-frame prediction on the attribute information of the first image area based on a unidirectional reference frame to determine the predicted value of the attribute information of the first image area; and determine the residual value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area.
[0437] In some implementations, the encoder 4300 also includes: an eleventh encoding unit 4375, configured to, if bidirectional inter-frame prediction is not enabled for the attribute information of the at least one frame of image, perform inter-frame prediction on the attribute information of the first image area based on a unidirectional reference frame to determine the predicted value of the attribute information of the first image area; and determine the residual value of the attribute information of the first image area based on the predicted value of the attribute information of the first image area.
[0438] In some implementations, the encoder 4300 further includes: a twelfth encoding unit 4380 configured to transform the residual value of the attribute information of the first image area to determine a change coefficient; and quantize the transform coefficient to determine a quantization coefficient.
[0439] In some implementations, the at least one frame of image is one frame of image; or, the at least one frame of image is an image sequence.
[0440] In some implementations, the first image region is a frame image; or, the first image region is a strip; or, the first image region is a block; or, the first image region is a RAHT layer.
[0441] It is understood that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular device. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0442] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, ROM, RAM, a magnetic disk, or an optical disk.
[0443] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 4300. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the decoding method described in any one of the aforementioned embodiments.
[0444] Based on the composition of the above-mentioned encoder 4300 and the computer-readable storage medium, refer to Figure 44, which shows a specific hardware structure diagram of the encoder 4400 provided in an embodiment of the present application. As shown in Figure 44, the encoder 4400 may include: a communication interface 4410, a memory 4420 and a processor 4430; each component is coupled together through a bus system 4440. It can be understood that the bus system 4440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 4440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 4440 in Figure 44. Among them,
[0445] Communication interface 4410, used for sending and receiving signals during the process of sending and receiving information with other external network elements;
[0446] Memory 4420, for storing computer programs;
[0447] The processor 4430 is configured to, when running the computer program, execute:
[0448] Determine whether to enable bidirectional inter-frame prediction for the attribute information of at least one frame of image; if bidirectional inter-frame prediction is enabled for the attribute information of the at least one frame of image, encode second identification information, wherein the second identification information is used to indicate whether bidirectional inter-frame prediction is enabled for the attribute information of a first image area, where the first image area is an image area in a frame of image.
[0449] It will be appreciated that the memory 4420 in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a ROM, PROM, EPROM, EEPROM, or flash memory. The volatile memory may be a RAM, which serves as an external cache. By way of example and not limitation, many forms of RAM are available, such as SRAM, DRAM, SDRAM, DDRSDRAM, ESDRAM, SLDRAM, and DRRAM. The memory 4420 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0450] Processor 4430 may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be performed by hardware integrated logic circuits within processor 4430 or by software instructions. Processor 4430 may be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software modules can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 4420. Processor 4430 reads information from memory 4420 and, in conjunction with its hardware, completes the steps of the above method.
[0451] It is understood that the embodiments described herein can be implemented with hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more ASICs, DSPs, DSPDs, PLDs, FPGAs, general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof. For software implementation, the technology described herein can be implemented by modules (e.g., processes, functions, etc.) that perform the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0452] Optionally, as another embodiment, the processor 4430 is further configured to execute the encoding method described in any one of the aforementioned embodiments when running the computer program.
[0453] An embodiment of the present application also provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium for storing a bit stream. The bit stream can be generated by an encoding method of an encoder, or the bit stream can be decoded by a decoding method of a decoder, wherein the decoding method can be the decoding method described in any of the foregoing embodiments, and the encoding method can be the encoding method described in any of the foregoing embodiments.
[0454] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0455] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0456] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0457] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0458] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0459] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A point cloud decoding method, applied to a decoder, comprising: Parsing first identification information, where the first identification information is used to indicate whether to enable bidirectional inter-frame prediction for at least one frame of image; If the first identification information indicates to enable bidirectional inter-frame prediction for the at least one frame of image, then parsing second identification information, where the second identification information is used to indicate whether to enable bidirectional inter-frame prediction for the attribute information of a first image region, and the first image region is an image region in a frame of image.
2. The method according to claim 1, wherein The method further comprises: If the first identification information indicates not to enable bidirectional inter-frame prediction for the at least one frame of image, then skipping the parsing of the second identification information.
3. The method according to claim 1 or 2, wherein The method comprises: Parsing third identification information, where the third identification information is used to indicate whether to enable inter-frame prediction for the attribute information of the at least one frame of image; If the third identification information indicates not to enable inter-frame prediction for the attribute information of the at least one frame of image, then skipping the parsing of the second identification information.
4. The method according to claim 2 or 3, wherein The method further comprises: If the parsing of the second identification is skipped, then not enabling bidirectional inter-frame prediction for the attribute information of the first image region.
5. The method according to claim 4, wherein, The method further comprises: If the third identification information indicates to enable inter-frame prediction for the attribute information of the at least one frame of image, then determining whether to parse the second identification information according to the first identification information.
6. The method according to claim 5, wherein The method further comprises: If the third identification information indicates to enable inter-frame prediction for the attribute information of the at least one frame of image, then parsing fourth identification information, where the fourth identification information is used to indicate whether to enable inter-frame prediction for the attribute information of the first image region.
7. The method according to claim 6, wherein The method further comprises: If the third identification information indicates not to enable inter-frame prediction for the attribute information of the at least one frame of image, then skipping the parsing of the fourth identification information; If the parsing of the fourth identification information is skipped, then not enabling inter-frame prediction for the attribute information of the first image region.
8. The method according to claim 3, wherein The third identification information belongs to an attribute parameter set.
9. The method according to claim 3, wherein The method further comprises: If the third identification information indicates not to enable inter-frame prediction for the attribute information of the at least one frame of image, then performing intra-frame prediction on the attribute information of the first image region to determine a predicted value of the attribute information of the first image region; Determining a reconstructed value of the attribute information of the first image region according to the predicted value of the attribute information of the first image region and a residual value of the attribute information of the first image region.
10. The method according to any one of claims 1 to 9, wherein: The first identification information belongs to an attribute parameter set; or, The first identification information belongs to a geometric parameter set; or, The first identification information belongs to a sequence parameter set.
11. The method according to any one of claims 1 to 10, wherein: The second identification information belongs to an attribute header.
12. The method according to claim 1, wherein, The method further comprises: If the second identification information indicates to enable bidirectional inter-frame prediction for the attribute information of the first image region, then performing inter-frame prediction on the attribute information of the first image region according to bidirectional reference frames to determine a predicted value of the attribute information of the first image region; Determine the reconstructed value of the attribute information of the first image region based on the predicted value of the attribute information of the first image region and the residual value of the attribute information of the first image region.
13. The method according to claim 1, wherein The method further includes: If the second identification information indicates that bidirectional inter-frame prediction is not enabled for the attribute information of the first image region, perform inter-frame prediction on the attribute information of the first image region according to a unidirectional reference frame to determine the predicted value of the attribute information of the first image region; Determine the reconstructed value of the attribute information of the first image region based on the predicted value of the attribute information of the first image region and the residual value of the attribute information of the first image region.
14. The method according to claim 1, wherein, The method further includes: If the first identification information indicates that bidirectional inter-frame prediction is not enabled for the at least one frame of image, perform inter-frame prediction on the attribute information of the first image region according to a unidirectional reference frame to determine the predicted value of the attribute information of the first image region; Determine the reconstructed value of the attribute information of the first image region based on the predicted value of the attribute information of the first image region and the residual value of the attribute information of the first image region.
15. The method according to claim 1, wherein, The method further includes: Parse the bitstream to determine the quantization coefficient of the attribute information of the first image region; Inverse-quantize the quantization coefficient to determine the transform coefficient of the attribute information of the first image region; Inverse-transform the transform coefficient to determine the residual value of the attribute information of the first image region.
16. The method according to any one of claims 1 to 15, wherein: The at least one frame of image is one frame of image; or, The at least one frame of image is an image sequence.
17. The method according to any one of claims 1 to 16, wherein: The first image region is one frame of image; or, The first image region is a stripe; or, The first image region is a block; or, The first image region is a region adaptive hierarchical transform (RAHT) layer.
18. A point cloud encoding method, applied to an encoder, includes: Determine whether to enable bidirectional inter-frame prediction for the attribute information of at least one frame of image; If bidirectional inter-frame prediction is enabled for the attribute information of the at least one frame of image, encode second identification information, where the second identification information is used to indicate whether to enable bidirectional inter-frame prediction for the attribute information of a first image region, and the first image region is an image region in one frame of image.
19. The method according to claim 18, wherein, The method further includes: If bidirectional inter-frame prediction is not enabled for the attribute information of the at least one frame of image, skip the encoding of the second identification information.
20. The method according to claim 18 or 19, wherein The method further includes: Encode first identification information, where the first identification information is used to indicate whether to enable bidirectional inter-frame prediction for at least one frame of image.
21. The method according to claim 20, wherein: The first identification information belongs to an attribute parameter set; or, The first identification information belongs to a geometric parameter set; or, The first identification information belongs to a sequence parameter set.
22. The method according to any one of claims 18 to 21, wherein The method includes: Determine whether to enable inter-frame prediction for the attribute information of the at least one frame of image; If inter - frame prediction is not enabled for the attribute information of the at least one frame of image, the encoding of the second identification information is skipped.
23. The method according to claim 22, wherein, The method further includes: Encoding third identification information, where the third identification information is used to indicate whether inter - frame prediction is enabled for the attribute information of the at least one frame of image.
24. The method according to claim 23, wherein The third identification information belongs to an attribute parameter set.
25. The method according to any one of claims 22 to 24, wherein, The method further includes: If inter - frame prediction is enabled for the attribute information of the at least one frame of image, encoding fourth identification information, where the fourth identification information is used to indicate whether inter - frame prediction is enabled for the attribute information of the first image region.
26. The method according to claim 25, wherein, The method further includes: If inter - frame prediction is not enabled for the attribute information of the at least one frame of image, the encoding of the fourth identification information is skipped.
27. The method according to claim 22, wherein, The method further includes: If inter - frame prediction is not enabled for the attribute information of the at least one frame of image, performing intra - frame prediction on the attribute information of the first image region to determine a predicted value of the attribute information of the first image region; Determining a residual value of the attribute information of the first image region according to the predicted value of the attribute information of the first image region.
28. The method according to any one of claims 18 to 27, wherein: The second identification information belongs to an attribute header.
29. The method according to claim 18, wherein The method further includes: If bidirectional inter - frame prediction is enabled for the attribute information of the first image region, performing inter - frame prediction on the attribute information of the first image region according to bidirectional reference frames to determine a predicted value of the attribute information of the first image region; Determining a residual value of the attribute information of the first image region according to the predicted value of the attribute information of the first image region.
30. The method according to claim 18, wherein The method further includes: If bidirectional inter - frame prediction is not enabled for the attribute information of the first image region, performing inter - frame prediction on the attribute information of the first image region according to a unidirectional reference frame to determine a predicted value of the attribute information of the first image region; Determining a residual value of the attribute information of the first image region according to the predicted value of the attribute information of the first image region.
31. The method according to claim 18, wherein The method further includes: If bidirectional inter - frame prediction is not enabled for the attribute information of the at least one frame of image, performing inter - frame prediction on the attribute information of the first image region according to a unidirectional reference frame to determine a predicted value of the attribute information of the first image region; Determining a residual value of the attribute information of the first image region according to the predicted value of the attribute information of the first image region.
32. The method according to claim 18, wherein The method further includes: Transforming the residual value of the attribute information of the first image region to determine transformation coefficients; Quantizing the transformation coefficients to determine quantization coefficients.
33. The method according to any one of claim 32, wherein: The at least one frame of image is one frame of image; or, The at least one frame of image is an image sequence.
34. The method according to any one of claims 18 to 33, wherein: The first image region is one frame of image; or, The first image region is a strip; or, The first image region is a block; or, The first image region is a region - adaptive hierarchical transform (RAHT) layer.
35. A decoder, comprising: A first decoding unit, configured to parse first identification information, where the first identification information is used to indicate whether to enable bidirectional inter-frame prediction for at least one frame of image; A second decoding unit, configured to parse second identification information if the first identification information indicates to enable bidirectional inter-frame prediction for the at least one frame of image, where the second identification information is used to indicate whether to enable bidirectional inter-frame prediction for attribute information of a first image region, and the first image region is an image region in a frame of image.
36. A decoder, comprising: A memory, configured to store a computer program; A processor, configured to execute the method according to any one of claims 1 to 17 when running the computer program.
37. An encoder, comprising: A first encoding unit, configured to determine whether to enable bidirectional inter-frame prediction for attribute information of at least one frame of image; A second encoding unit, configured to encode second identification information if bidirectional inter-frame prediction is enabled for the attribute information of the at least one frame of image, where the second identification information is used to indicate whether to enable bidirectional inter-frame prediction for attribute information of a first image region, and the first image region is an image region in a frame of image.
38. An encoder, comprising: A memory, configured to store a computer program; A processor, configured to execute the method according to any one of claims 18 to 34 when running the computer program.
39. A non-volatile computer-readable storage medium storing a bitstream, the bitstream being generated by using an encoding method of an encoder or the bitstream being decoded by using a decoding method of a decoder, wherein, The decoding method is the method according to any one of claims 1 to 17, and the encoding method is the method according to any one of claims 18 to 34.
40. A bitstream, where the bitstream includes the bitstream generated by the method according to any one of claims 18 to 34.
41. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 17, or 18 to 34 is implemented.
Citation Information
Patent Citations
Encoding and decoding methods and encoding and decoding device of range image
CN103067715A
Video image decoding method and device
CN111372086A
Encoder, decoder and corresponding methods
CN114009040A
Video image decoding method and apparatus, and video image encoding method and apparatus
WO2020143589A1
Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data transmission method
WO2023075453A1