Method, encoder and decoder for encoding and decoding 3d point clouds

By adopting the octree structure and TriSoup encoding scheme in point cloud compression technology, the 3D point cloud is encoded into a bitstream, and the geometric structure of the 3D point cloud is implemented in the decoder, the problem of difficulty in effectively compressing dense and dynamic 3D point cloud data in the existing technology is solved, and efficient decoding accuracy and compression efficiency are achieved.

CN119948527APending Publication Date: 2025-05-06BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280100571.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing point cloud compression technologies are difficult to effectively compress dense and dynamic 3D point cloud data, especially under the need to transmit real-time and maintain high decoding quality.

Method used

By using an octree structure and a TriSoup encoding scheme, a 3D point cloud is encoded into a bitstream and a method of decoding the geometric structure of a 3D point cloud from a bitstream is implemented in a decoder. The method includes decoding the octree information and vertex information, reconstructing the triangle by calculation of virtual locations, normal vectors, and centroid locations, and voxelizing by ray tracing to determine the point of the point cloud.

Benefits of technology

It improves the decoding accuracy and compression efficiency of 3D point clouds, and can adapt to different types of point cloud data while transmitting real-time and high decoding quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948527A_ABST
    Figure CN119948527A_ABST
Patent Text Reader

Abstract

A method, preferably implemented in a decoder, for decoding a geometry of a 3D point cloud from a bitstream, the method comprising: receiving and decoding the bitstream, where the bitstream comprises octree information and vertex information, the octree information comprising information about an octree structure of a point cloud volume, the vertex information comprising information about an octree structure of the point cloud volume; the vertex information comprises information about vertex existence and vertex position on the edge of the cube of the leaf node of the octree structure; determining a virtual position by averaging positions of vertices of one cube associated with leaf nodes of the octree structure; constructing a triangle by connecting two continuous vertexes and virtual positions of a cube clockwise; determining a normal vector of the cube based on the constructed triangle; determining a centroid position based on the normal vector; reconstructing a triangle by clockwise connecting two continuous vertexes and centroid positions of the cube; the reconstructed triangle is voxelized to determine points of the point cloud, where a normal vector of the cube is determined based on the determined area of the triangle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for decoding a 3D point cloud from a bitstream. Furthermore, the object of the present invention is to provide a method for encoding a 3D point cloud into a bitstream. Further, the object of the present invention is to provide an encoder and a decoder, a bitstream encoded according to the present invention and software. In particular, the object of the present invention is to provide a method for improving the accuracy of a decoding or reconstruction process of a 3D point cloud. Background Art

[0002] Point clouds have recently gained attention as a format for 3D data representation due to their versatility in representing all types of 3D objects or scenes. As a result, many use cases can be solved with point clouds, including:

[0003] Film post-production,

[0004] Real-time 3D immersive telepresence or virtual reality (VR) / augmented reality (AR) applications,

[0005] Free viewpoint video (e.g. for watching sports),

[0006] Geographic Information Systems (also known as Cartography),

[0007] Cultural heritage (scans of rare objects stored in digital form),

[0008] Autonomous driving, including 3D mapping of the environment and real-time lidar data collection.

[0009] A point cloud is a set of points located in 3D space, optionally with additional values ​​attached to each point. These additional values ​​are often called point attributes. Thus, a point cloud is a combination of geometry (the 3D position of each point) and attributes.

[0010] The attributes may be, for example, a three-component color, a material property such as reflectivity, and / or a two-component normal vector of a surface associated with the point.

[0011] Point clouds can be captured by various types of devices, for example, camera arrays, depth sensors, Lidar, scanners, or can be computer generated (for example, in movie post-production). Depending on the use case, a point cloud may have thousands to billions of points for mapping applications.

[0012] The raw representation of point clouds requires a very high number of bits per point, with at least twelve bits per spatial component X, Y, or Z, and optionally more bits for attributes, e.g., three times 10 bits for color. Practical deployment of point cloud-based applications requires compression techniques that can store and distribute point clouds with reasonable storage and transmission infrastructure.

[0013] For distribution to and visualization by end users, for example, on AR / VR glasses or any other 3D-enabled device, compression can be lossy (like in video compression). Other use cases do require lossless compression, for example, medical applications or autonomous driving, to avoid altering the decision results obtained from the analysis of the compressed and transmitted point cloud.

[0014] Until recently, point cloud compression (aka PCC) had not been addressed by the mass market and no standardized point cloud codec was available. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11, also known as the Moving Picture Experts Group or MPEG, started a work item on point cloud compression. This resulted in two standards, namely

[0015] MPEG-I Part 5 (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC),

[0016] MPEG-I Part 9 (ISO / IEC 23090-9) or Geometry-based Point Cloud Compression (G-PCC).

[0017] Both the V-PCC and G-PCC standards completed their first versions at the end of 2020.

[0018] The V-PCC encoding method compresses point clouds by performing multiple projections of 3D objects to obtain 2D patches that are packed into images (or videos when processing mobile point clouds). The acquired images or videos are then compressed using existing image / video codecs, allowing the use of already deployed image and video solutions. By its nature, V-PCC is only effective on dense and continuous point clouds, because image / video codecs cannot compress non-smooth patches like those obtained from the projection of sparse geometric data collected by, for example, Lidar.

[0019] The G-PCC coding method has two schemes for geometry compression. The first scheme is based on an occupancy tree (octree / quadtree / binary tree) representation of the point cloud geometry. Occupied nodes are split until a certain size is reached, and the occupied leaf nodes provide the locations of the points, usually at the centers of these nodes. A high level of compression for dense point clouds can be obtained by using neighbor-based prediction techniques. Sparse point clouds are also addressed by directly encoding the locations of points within nodes with non-minimum sizes, by stopping tree construction when only isolated points exist in the node; this technique is called Direct Coding Mode (DCM).

[0020] The second scheme is based on prediction trees, where each node represents the 3D position of a point and the relationship between nodes is the spatial prediction from parent to child. This approach can only handle sparse point clouds and has the advantages of lower latency and simpler decoding compared to occupancy trees. However, the compression performance is only slightly better than the first occupancy-based approach, and the encoding is complex, with intensive searching for the best predictor (in a long list of potential predictors) when building the prediction tree.

[0021] In both schemes, the attribute (de)coding can be performed after the geometry (de)coding is completed, resulting in a two-pass encoding. Thus, low latency is obtained by using slices that decompose the 3D space into independently coded sub-volumes without prediction between sub-volumes. When many slices are used, the compression performance may be severely affected.

[0022] An important use case is the transmission of dynamic AR / VR point clouds. Dynamic means that the point cloud evolves over time. Moreover, AR / VR point clouds are usually locally 2D, as they represent the surface of objects most of the time. Therefore, AR / VR point clouds are highly connected (or dense), as points are rarely isolated but have many neighbors.

[0023] A dense (or solid) point cloud represents a continuous surface at a resolution such that the volumes (small cubes called voxels) associated with the points touch each other without showing any visible holes in the surface.

[0024] Such point clouds are typically used in AR / VR environments and viewed by end users through devices such as TVs, smartphones or headsets. They are either transmitted to the device or stored locally. Many AR / VR applications use dynamic point clouds, which change over time, rather than static point clouds. Therefore, the amount of data is huge and must be compressed. Today, lossless compression based on an octree representation of the point cloud geometry can achieve slightly less than one bit per point (1bpp). This may not be enough for real-time transmission, which may involve several million points per frame and frame rates of up to 50 frames per second (fps), resulting in hundreds of megabits of data per second.

[0025] Therefore, lossy compression can be used, maintaining the usual requirements of acceptable visual quality while compressing sufficiently to fit within the bandwidth provided by the transmission channel while maintaining real-time transmission of frames. In many applications, bit rates as low as 0.1bpp (10 times higher than lossless coding compression rates) have made real-time transmission possible.

[0026] Codecs based on MPEG-I Part 5 (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC) such as VPCC can achieve such low bitrates by using lossy compression with a video codec that compresses 2D frames obtained from the projection of the point cloud onto a plane. The geometry is represented in one frame by a sequence of projected patches, each of which is a small local depth map. However, VPCC is not general and is limited to narrow types of point clouds that do not exhibit locally complex geometry (e.g., trees, hair), since the obtained projected depth maps are not smooth enough to be effectively compressed by video codecs.

[0027] Pure 3D compression techniques can handle any type of point cloud. It is still an open question whether 3D compression techniques can compete with VPCC (or any projection + image coding scheme) on dense point clouds. Standardization is still moving towards providing an extension (revision) of GPCC that provides competitive lossy compression that can compress dense point clouds as well as VPCC intra-frame coding, while maintaining the versatility of GPCC to handle any type of point cloud (dense point cloud, Lidar, 3D map). This extension will probably use the so-called TriSoup (triangle soup) encoding scheme, which works with octrees, which will be described in detail in the next section. TriSoup is being explored by the ISO / IEC standardization working group JTC1 / SC29 / WG7.

[0028] However, for all lossy compression schemes, the reconstruction quality of the points in the point cloud is crucial. Summary of the invention

[0029] Therefore, an object of the present invention is to provide a method for decoding the geometry of a 3D point cloud from a bitstream and a method for encoding a 3D point cloud into a bitstream, which method improves efficiency.

[0030] This problem is solved by a decoding method according to claim 1 , an encoding method according to claim 2 , an encoder according to claim 6 , a decoder according to claim 7 , a bit stream according to claim 8 and software according to claim 9 .

[0031] In a first aspect, a method for decoding the geometry of a 3D point cloud from a bitstream is provided, the method preferably being implemented in a decoder. The method comprises:

[0032] Receive and decode a bitstream, wherein the bitstream includes octree information and vertex information, the octree information includes information about an octree structure of a point cloud volume, and the vertex information includes information about the existence and position of vertices on edges of cubes of leaf nodes of the octree structure;

[0033] Determining a virtual position by averaging the positions of vertices of a cube associated with a leaf node of the octree structure;

[0034] constructing a triangle by connecting two consecutive vertices of a cube and the virtual position clockwise;

[0035] Determining a normal vector of the cube based on the constructed triangles;

[0036] determining a center of mass position based on the virtual position and the normal vector;

[0037] Reconstructing a triangle by connecting two consecutive vertices of the cube and the centroid position clockwise;

[0038] voxelizing the reconstructed triangles to determine the points of the point cloud,

[0039] The normal vector of the cube is determined based on the determined area of ​​the triangle.

[0040] Therefore, in the first step, a bitstream is received, and the bitstream contains information about the octree structure of the volume of the decoded point cloud. Preferably, the geometric structure of the point cloud is GPCC encoded. Therefore, by decoding from the bitstream, octree information about the volume of the point cloud can be provided. In addition, the bitstream also includes vertex information, which includes information about the existence and position of vertices on the edges of the cube associated with the leaf nodes in the octree structure. Therefore, vertex information is provided by decoding from the bitstream. Wherein, the bitstream is preferably decoded at the encoder by the TriSoup encoding scheme.

[0041] After decoding the octree information and vertex information from the bitstream described in the previous step, in a further step, triangles are determined for each cube in order to reconstruct the point cloud geometry. In particular, a virtual position is first determined by averaging the positions of the vertices of a cube. Then, a first set of triangles is constructed by connecting two consecutive vertices on the edge of the cube in a clockwise direction and the determined virtual positions. Thus, the surface of the first set of triangles is determined by the positions of the vertices included in the bitstream and the determined virtual positions (i.e. the average positions of the vertices).

[0042] Subsequently, the normal vector of the cube is determined based on the constructed triangles. Among them, the normal vector is a vector perpendicular to a given object. Therefore, the normal vector of each triangle in the first group of triangles can be determined separately. Then, the normal vector of the cube can be determined based on the sum of the normal vectors of the determined first group of triangles. According to the present invention, the normal vectors of the first group of triangles are weighted. Preferably, the weights are determined based on the area of ​​the first group of triangles. Triangles with larger areas are more likely to contain more points of the point cloud and are therefore given larger weights. Then, the center of mass position is determined based on the virtual position of the cube and the weighted normal vector. Through such adjustments, the determined center of mass position can be closer to the original position of the point cloud in the leaf node. Triangles are constructed based on vertices and center of mass in a similar manner to that described above.

[0043] To reconstruct the points of the point cloud from the reconstructed triangles, voxelization is performed by a ray tracing process, in which rays are emitted along three directions parallel to any of the three axes. Their origin is a point with integer coordinates corresponding to the sampling accuracy required for rendering. The intersection of the ray with one of the reconstructed triangles (if any) is then determined and added to the list of rendered points, i.e., to the points of the point cloud. During voxelization, the surface of the reconstructed triangle is sampled with rays to determine the points of the point cloud.

[0044] Preferably, the bitstream further comprises an additional value α, and the centroid position is further determined based on the additional value α, wherein the additional value may be determined by an encoder and may be used to calculate the centroid position.

[0045] Preferably, the additional value α is determined based on the positions of all points of the point cloud belonging to the current leaf node. Therefore, the additional value can be determined by the encoder by considering all points of the point cloud belonging to the current leaf node. Preferably, only points that are within a predetermined distance of the offline (virtual point, normal vector) are considered. With the additional value, the accuracy of the centroid point can be further improved.

[0046] In another aspect of the present invention, a method for encoding a 3D point cloud into a bitstream is provided, preferably implemented in an encoder. The method for encoding a 3D point cloud comprises:

[0047] Obtaining octree information, the octree information including an octree structure of a body, the body including a plurality of cubes;

[0048] Obtaining vertex information from a surface of a point cloud of each cube associated with a leaf node, wherein the vertex information includes information about the existence and position of vertices on an edge of the cube;

[0049] Encoding the octree information and the vertex information into a bitstream;

[0050] The geometric data of the point cloud is reconstructed by using the octree information and vertex information obtained in the aforementioned encoding process, wherein the geometric data of the point cloud is reconstructed including:

[0051] Determining a virtual position by averaging the positions of vertices of a cube associated with a leaf node of the octree structure;

[0052] constructing a triangle by connecting two consecutive vertices of a cube and the virtual position clockwise;

[0053] Determining a normal vector of the cube based on the constructed triangles;

[0054] determining a center of mass position based on the normal vector;

[0055] Reconstructing a triangle by connecting two consecutive vertices of the cube and the centroid position clockwise;

[0056] voxelizing the reconstructed triangles to determine the points of the point cloud,

[0057] The normal vector of the cube is determined based on the determined area of ​​the triangle.

[0058] Thus, by the encoding method, octree information and vertex information are generated. This information is encoded into the bitstream. A reconstruction step is then performed on the encoder side. In this reconstruction step, the point cloud geometry information is reconstructed, wherein the reconstruction steps are the same as in the above-mentioned decoding method. The reconstructed geometry of the point cloud is then used on the encoder side to encode the attributes of the points of the point cloud (color, reflectivity, ...), for example, by RAHT (Region Adaptive Hierarchical Transform), prediction transform or lifting transform.

[0059] Preferably, the geometric structure of the point cloud is encoded into the bitstream via Geometry-based Point Cloud Compression (G-PCC).

[0060] Preferably, the bitstream is an MPEG G-PCC compliant bitstream.

[0061] Preferably, the encoding method is further constructed according to the features described above in conjunction with the decoding method.

[0062] In another aspect of the present invention, an encoder for encoding a 3D point cloud into a bit stream is provided. The encoder comprises a memory and a processor, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the steps of the aforementioned encoding method are performed.

[0063] In another aspect of the present invention, a decoder for decoding a 3D point cloud from a bitstream is provided. The decoder comprises a memory and a processor, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the steps of the above-mentioned decoding method are performed.

[0064] In another aspect of the present invention, a bit stream is provided, wherein the bit stream is encoded based on the steps of the aforementioned encoding method.

[0065] In another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium comprising instructions for executing the steps of the above method for encoding a 3D point cloud into a bitstream.

[0066] In another aspect of the present invention, a computer-readable storage medium is provided, comprising instructions for executing the steps of the above method for decoding a 3D point cloud from a bitstream. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Hereinafter, the present invention is described in more detail with reference to the accompanying drawings.

[0068] These figures show:

[0069] Figure 1 A flowchart showing a method for decoding 3D point cloud geometry according to the present invention;

[0070] Figure 2 An example of generating an octree structure is shown;

[0071] Figure 3 Show according to Figure 2 The octree of

[0072] Figure 4 An example of determining the vertices on the edge of a cube is shown;

[0073] Figure 5 An example of generating a triangle is shown;

[0074] Figure 6 An example showing vertices on the edges of a cube;

[0075] Figure 7 Shows the generation of triangles from vertices;

[0076] Figure 8 Show the basis for determination Figure 7 An example of the order of triangles;

[0077] Fig. 9 A schematic diagram showing the voxelization step;

[0078] Fig.10 An example of reconstructing a triangle using the centroid point C as the pivot point is shown;

[0079] Fig.11 An example of a normal vector is shown;

[0080] Fig.12 An example showing the 1D residual along the normal vector;

[0081] Fig.13 An example of the normal vector determination process is shown;

[0082] Fig.14 An example of calculating the area of ​​a parallelogram is shown. DETAILED DESCRIPTION

[0083] Reference Figure 1 , which shows a schematic diagram of a method for decoding geometric information of a 3D point cloud from a bitstream.

[0084] A method for decoding the geometry of a 3D point cloud from a bitstream, preferably implemented in a decoder, comprises the following steps:

[0085] Receive and decode a bit stream, wherein the bit stream includes octree information and vertex information, the octree information includes information about an octree structure of a point cloud volume, and the vertex information includes information about the existence and position of vertices on edges of cubes of leaf nodes of the octree structure;

[0086] Determining the virtual position by averaging the positions of the vertices of a cube associated with the leaf nodes of the octree structure;

[0087] Construct a triangle by connecting two consecutive vertices of a cube and the virtual position clockwise;

[0088] Determine the normal vector of the cube based on the constructed triangles;

[0089] Determine the center of mass position based on the normal vector;

[0090] Reconstruct the triangle by connecting two consecutive vertices and centroid positions of the cube clockwise;

[0091] voxelize the reconstructed triangles to determine the points of the point cloud,

[0092] The normal vector of the cube is determined based on the area of ​​the determined triangle.

[0093] To determine the octree information, the first step of the geometry encoding process is to construct and encode the octree, such as Figure 2 and Figure 3As shown. The bounding box is the body 100 that contains all the points and is associated with a root node 112 (i.e., a single node at the top of the tree 110). The body 100 is first divided into eight sub-volumes 102, called octants, each of which is represented by a node 114 in the tree 110. Then, octants 106 are recursively split out of the sub-volumes 104 until the target level is reached, where an octant 106 is occupied by at least one point. Figure 2 and Figure 3 Indicated by shading.

[0094] Each octant (or node) is represented by an occupation byte containing one bit for each sub-octant, with the corresponding bit set to one if the sub-octant is occupied by at least one node, and set to zero otherwise. The occupation bytes 118 of all octants are serialized (in breadth-first order) and entropy encoded using a binary arithmetic encoder.

[0095] Figure 4 A block representation of a 3D surface 210 is shown, along with an example of a block 220 in TriSoup. The surface 210 intersects the block 220, so the block 220 is an occupied block, and the block 220 exists between multiple blocks 200 in the 3D space. Within the block 220, the closed portion of the surface 210 intersects the edge of the block at the six illustrated vertices of the polygon 230. If an edge of the block 220 contains a vertex, the edge is said to be selected.

[0096] Figure 5 Block 220 in TriSoup is shown, with surface 210 omitted for clarity, and showing unselected edge 270, selected edge 260, and i-th edge 250. Assume that i-th edge 250 is selected. To specify a vertex vi on edge i, a scalar value is specified that indicates the corresponding fraction of the length of edge 250.

[0097] like Figure 4 and Figure 5 As shown, within each octant 220 of the target level of the octree, trisoup represents the original surface 210 as a set of triangles 245. The surface is encoded and used to obtain the position of the reconstructed (or decoded) points. First, the intersection of the surface represented by the original points with the edges of the octant is estimated by averaging the positions of the points closest to those edges within the octant among the original points used to represent the surface. Secondly, the twelve edges of all octants and their associated intersections (if any) are stored as segments and vertices respectively. Each (unique) segment is then encoded as follows. The first single bit is arithmetically encoded, and the bit is set to 1 if the segment is occupied by a vertex, otherwise it is set to 0. If it is occupied, the relative position of the vertex on the segment is also arithmetically encoded.

[0098] The vertices 310 of the triangles are encoded along the edges 320 of the volume associated with the leaf nodes 300 of the tree, as Figure 6 As shown. These vertices 310 on the edge 320 are shared between multiple leaf nodes 300 with a common edge 320. This means that each edge belonging to at least one leaf node encodes at most one vertex. In this way, the continuity of the model is ensured through the leaf nodes.

[0099] As mentioned above, the encoding of TriSoup vertices requires two pieces of information per edge:

[0100] A vertex flag indicating whether a TriSoup vertex exists on the edge, and

[0101] The positions of vertices along the edges, when present.

[0102] Therefore, the encoded data includes Octree data as well as TriSoup data.

[0103] The vertex flags are encoded by an adaptive binary arithmetic encoder that uses a specific context to encode the vertex flags. The length is N = 2 s The positions of vertices on the edges of can be encoded with unit precision by pushing s bits into the (bypassing / non-entropy encoding) bitstream.

[0104] Within a leaf node, if there are at least three vertices 310 on an edge 320 of the leaf node 300, a triangle is constructed from the TriSoup vertices. Figure 7 The reconstructed triangles 330, 340 are depicted in FIG.

[0105] Obviously, other combinations of triangles 330, 340 are possible. The selection of triangles results from a three-step process:

[0106] 1. Determine the dominant direction along one of the three axes;

[0107] 2. Sort the TriSoup vertices based on the dominant direction;

[0108] 3. Build a triangle based on the ordered list of vertices.

[0109] Knowledge of the exact position of the triangle in the current leaf node is not required and can be inferred from the vertices.

[0110] Figure 8 will be used to explain this process. Each of the three axes is tested and the one that maximizes the total surface of the triangle is chosen as the principal axis. For simplicity of the diagram, Figure 8 Only tests on two axes are described.

[0111] The first test along the vertical axis (top) is performed by vertically projecting the cube and TriSoup vertices 310 on the 2D plane. The vertices 310 are then sorted in clockwise order relative to the center of the projection node (square). Then, based on the ordered vertices, triangles 330, 340 are constructed according to a fixed rule. Here, when 4 vertices are involved, triangle 123 as well as triangle 134 are systematically constructed. When there are 3 vertices, the only possible triangle is 123. When there are 5 vertices, the fixed rule can be to construct triangles 123, 134 and 451. And so on, up to a maximum of 12 vertices.

[0112] The second test (left side) along the horizontal vertical axis is performed by projecting the cube and trisoup vertices horizontally onto the 2D plane.

[0113] The vertical projection shows the maximum total 2D surface of the triangle, therefore, the principal axis is chosen as the vertical axis, and the constructed TriSoup triangles are obtained in the order of vertical projection, such as Figure 8 As shown, it is located inside the node. It should be noted that taking the horizontal axis as the main axis will result in another construction of the triangle.

[0114] The principal axes are appropriately selected by maximizing the projection surface, thus achieving continuous reconstruction of point clouds without holes.

[0115] Rendering of TriSoup triangles into points is performed by ray tracing. The set of all points rendered by ray tracing will form the decoded point cloud.

[0116] for Fig. 9 In the ray tracing shown, rays are emitted in three directions parallel to the axis. Their origin is a point with integer (voxelized) coordinates of a precision corresponding to the sampling precision required for rendering. The intersection with one of the trisoup triangles (dashed point if any) is then voxelized (= rounded to the nearest point of the required sampling precision) and added to the list of rendered points.

[0117] After applying Trisoup to all leaf nodes, i.e., building triangles and obtaining points by ray tracing, discard all copies of the same point in the list of rendered points (i.e., keep only one voxel among all voxels sharing the same position and volume) to obtain a set of decoded (unique) points.

[0118] Based on the above basic concepts, the Trisoup encoding can be improved. For example, by calculating the centroid point whose coordinates are the average coordinates of all (ordered) vertices Vi, see Fig.10 , where the centroid C is depicted using a checkerboard fill. The centroid is used as the pivot point. By rotating around the centroid C, the ordered vertices (V 1 , V2 , …, V M ) Construct the following M triangles:

[0119] ·V 1 V 2 C,

[0120] ·V 2 V 3 C,

[0121] ·…

[0122] ·V M-1 V M C,

[0123] ·V M V 1 C.

[0124] Among them, the vertices are sorted in a clockwise direction, and which vertex is selected as V 1 It doesn't matter.

[0125] This construction preserves the natural symmetry of the model without privileging any arbitrary triangles. In addition, this construction provides an additional degree of freedom to improve the accuracy of the model, namely, the location of the centroid point C.

[0126] Therefore, the position of the centroid point C can be further improved by encoding the residual position in the bitstream, so that the position of the centroid point C is closer to the original point of the point cloud.

[0127] For example, C = C mean +C res ,

[0128] Among them, C mean is the average position obtained by averaging the coordinates of all (ordered) vertices, and C res is the coding residual.

[0129] The coded residual can be a 3D residual. However, it can be observed that a 3D residual is rarely advantageous because it requires many bits to encode and these many bits are not fully compensated by the better accuracy of the model. Therefore, it is preferred to encode the 1D residual C res to encode.

[0130] For example, the normal vector can be constructed like Fig.11 As shown, the residual can be determined by the following equation:

[0131]

[0132] where α is a 1D signed scalar value encoded in the bitstream, see Fig.12 . Normal vector It can be exported in two steps:

[0133] 1.

[0134] 2. Then, normalize

[0135] where × is the cross product between two vectors (also called the vector (cross) product), and the side yes:

[0136]

[0137] It is understood that in order to simplify its determination and calculation of the value α, the vector Can be parallel to an axis. A good approximation for the vector calculated above is that it can be parallel to the principal axis.

[0138] The value α may be determined by an encoder, encoded into a bitstream and obtained by a decoder by decoding the bitstream. The value α may be binarized and each bit may be encoded using a binary entropy encoder such as an arithmetic encoder or a context adaptive binary encoder like CABAC.

[0139] The value α can be binarized as:

[0140] A flag f indicating whether α is equal to 0 0 ,

[0141] a sign indicating α>0 or α<0,

[0142] A flag f indicating whether |α| is equal to 1 1 ,

[0143] The remainder |α|-2 encoded by the expGolomb encoder.

[0144] The value α can be obtained by the encoder by considering all points P in the point cloud belonging to the current leaf node k Determine. For each point P k , which is consistent with the line (C mean , ) k Obtained by the following formula:

[0145]

[0146] And at this distance d k If the value of point P is lower than a predefined threshold value th (for example, th=2), k Used to calculate the value α. Point P k Relative to the average point C meanThe 1D residual r k Obtained by the scalar product (also called inner product or dot product):

[0147]

[0148] Therefore, the value α is obtained by:

[0149]

[0150] Where S is point P k A set of such that their distance d k is below a threshold th, and |S| is the number of points belonging to this set.

[0151] In the vertex and C mean In the triangle formed (assuming the index of the triangle is i), the normal vector of the triangle Through the edge The cross product of For example, Fig.13 As shown, triangle V 1 V 2 C mean The normal vector Can be and The cross product of is obtained, which can be described by the equation:

[0152]

[0153] Triangle V 2 V 3 C mean The normal vector Can be and The cross product of is obtained, which can be described by the equation:

[0154]

[0155] Similarly, triangle V 3 V 4 C mean The normal vector and triangle V 4 V 1 C mean of It can be obtained by the following formula:

[0156]

[0157] Then, the normal vector of the leaf node in this example is This is determined in two steps:

[0158] 1. Calculate each triangle (V 1 V 2 C mean , V 2 V 3 C mean , …, V M-1 V M C mean , V M V 1 C mean ) of and,

[0159]

[0160] 2. Then, calculate the normalization of the normal vector

[0161]

[0162] As mentioned above, the normal vectors of the leaf nodes used for trisoup encoding By using the vertex and C mean The normal vector of each triangle constructed Then, along the average point C mean and the normal vector The constructed line finds the residual α, the purpose is to make the position of the centroid point C closer to the original point of the point cloud in the leaf node, where However, in the sum calculation, each normal vector is not taken into account This may result in obtaining a normal vector for each leaf node (for trisoup encoding) It is not optimal because in many cases the number of points in a leaf node is not the same for each triangle of the leaf node, for example in the case where the surface of the point cloud has many concave and convex areas.

[0163] Therefore, simply for each normal vector of each triangle The sum may not be the normal vector in the leaf node This will result in a non-optimal position of the centroid point C. In addition, the position of the reconstructed point may have a larger error relative to the original point in the point cloud, which will lead to reduced compression efficiency.

[0164] Therefore, it is recommended to use the normal vector of each triangle in the leaf node The weighted average of To achieve better compression efficiency. The weights can be based on the vertices and C mean The area of ​​each triangle constructed is determined. In this way, the normal vector of the triangle with the larger area The normal vector Therefore, the normal vector obtained is will be greater than the original normal vector Closer Because a larger triangle area is more likely to contain more points in the original point cloud. Therefore, by utilizing the area information of each triangle, the normal vector obtained can be The refinement is to point to the direction where more original points may exist in the point cloud. As a result, a better position of the centroid point C is obtained, which will make the reconstruction error of the reconstructed point position smaller, thereby further improving the compression efficiency.

[0165] In some embodiments, the normal vector of a leaf node is obtained by using a weighted average based on the area of ​​each triangle, the vector of the leaf node method This can be determined by following two steps:

[0166] 1. Calculate the normal vector of each triangle using the following equation The weighted average of:

[0167]

[0168] 2. Vector Normalize to get the normal vector

[0169]

[0170] Among them A i represents the area of ​​the i-th triangle in the leaf node, A Sum Defined as the sum of the areas of each triangle, M is the number of triangles built in the leaf nodes for trisoup encoding.

[0171] In some embodiments, the normal vector The weights used in determining the leaf node are based on the square root of the area of ​​each triangle, so the normal vector for each triangle is calculated. The weighted average of can be obtained as follows:

[0172]

[0173] in

[0174] In some embodiments, the weights used in the normal vector determination are based on the square of the area of ​​each triangle in the leaf node, so the normal vector for each triangle is calculated as The weighted average of can be obtained as follows:

[0175]

[0176] Among them A Sum =A 1 2 +A 2 2 +…+A M 2 .

[0177] In a more general way, the normal vector The weights used in determining the area of ​​each triangle in the leaf node are based on the exponent, then the normal vector of each triangle is calculated. The weighted average of can be obtained as follows:

[0178]

[0179] Among them A Sum =A 1 p +A 2 p +…+A M p , the parameter p controls the influence of the normal vector on the triangle area, and can be set according to different types of point cloud data to achieve the best compression efficiency. For example, for point cloud data with many concave and convex areas, p can be set to be greater than 1.

[0180] In some embodiments, each triangle (V 1 V 2 C mean , V 2 V 3 C mean , …, V M-1 V M C mean , V M V 1 C mean ) can be expressed by two edge vectors. For example, the area of ​​the first triangle A is 1 It can be obtained by

[0181]

[0182] because is the area of ​​the parallelogram, such as Fig.14 As shown, triangle V 1 V 2 C mean The area of ​​the parallelogram is half the area of ​​the parallelogram.

[0183] In general, the area A of the i-th triangle i It can be obtained by

[0184]

[0185] Where i=j. Different variables are used here just to distinguish the index of the triangle and the index of the vertex.

[0186] In some embodiments, the area of ​​each triangle can be obtained by:

[0187]

[0188] Since the normal vector is the average result of the area weights, where the normal vector In the weighted average equation of , scalar 2 will be omitted. For example,

[0189] will become

[0190]

[0191] According to the present invention, the normal vector of the proposed leaf node is The determination method can be used in the encoding process and the decoding process of trisoup encoding.

[0192] According to some embodiments of the trisoup encoding process, it follows the following steps:

[0193] First, the vertex information of each leaf node is determined, which includes the vertex existence flag information of each edge of the leaf node and the vertex position of the edge (if the vertex existence flag of the edge is true), and then the trisoup information is encoded into a bit stream through entropy coding.

[0194] Then, in order to encode the position information of the centroid point C of each leaf node into the bit stream, iterate on each leaf node. For each leaf node with a vertex number greater than 2,

[0195] ··Compute C by using the vertices obtained from the leaf nodes mean ,

[0196] Then, by using the proposed normal vector The determination method of (i.e., based on the weighted area of ​​the triangle), based on C mean And the obtained vertex to calculate the normal vector of the leaf node

[0197] • Then, the centroid residual value α is obtained and encoded into a bitstream through entropy coding.

[0198] According to some embodiments of the trisoup decoding process, it follows the following steps:

[0199] First, the vertex information of each leaf node is decoded from the bitstream. The information includes the vertex existence flag information of each edge of the leaf node and the vertex position of the edge (if the vertex existence flag of the edge is true).

[0200] Then, in order to reconstruct the position information of the centroid point C of each leaf node, iterate on each leaf node. For each leaf node with more than 2 vertices,

[0201] · Compute C by using the decoded vertices in the leaf nodes mean ,

[0202] Then, by using the proposed normal vector The determination method of (i.e., based on the weighted area of ​​the triangle), based on C mean and the decoded vertices to calculate the normal vector of the leaf node

[0203] Then, the centroid residual value α is decoded from the bitstream. Then, it can be obtained by Reconstruct the position of the centroid point C of the leaf node,

[0204] Finally, construct triangles in leaf nodes (V 1 V 2 C, V 2 V 3 C, …, V M-1 V M C, V M V 1 C), and then use the ray tracing method to construct the triangle (V 1 V 2 C, V 2 V 3 C, …, V M-1 V M C, V M V 1 C) Reconstruct the points in the leaf nodes.

Claims

1. A method for decoding the geometry of a 3D point cloud from a bitstream, preferably implemented in a decoder, the method comprising: Receive and decode a bitstream, wherein the bitstream includes octree information and vertex information, the octree information includes information about an octree structure of a point cloud volume, and the vertex information includes information about the existence and position of vertices on edges of cubes of leaf nodes of the octree structure; Determining a virtual position by averaging the positions of vertices of a cube associated with a leaf node of the octree structure; constructing a triangle by connecting two consecutive vertices of a cube and the virtual position clockwise; Determining a normal vector of the cube based on the constructed triangles; determining a center of mass position based on the normal vector; Reconstructing a triangle by connecting two consecutive vertices of the cube and the centroid position clockwise; voxelizing the reconstructed triangles to determine the points of the point cloud, The normal vector of the cube is determined based on the determined area of ​​the triangle.

2. A method for encoding a 3D point cloud into a bitstream, preferably implemented in an encoder, the method comprising: Obtaining octree information, the octree information including an octree structure of a body, the body including a plurality of cubes; Obtaining vertex information from a surface of a point cloud of each cube associated with a leaf node, wherein the vertex information includes information about the existence and position of vertices on an edge of the cube; Encoding the octree information and the vertex information into a bitstream; The geometric data of the point cloud is reconstructed by using the octree information and vertex information obtained in the aforementioned encoding process, wherein the geometric data of the point cloud is reconstructed including: Determining a virtual position by averaging the positions of vertices of a cube associated with a leaf node of the octree structure; constructing a triangle by connecting two consecutive vertices of a cube and the virtual position clockwise; Determining a normal vector of the cube based on the constructed triangles; determining a center of mass position based on the virtual position and the normal vector; Reconstructing a triangle by connecting two consecutive vertices of the cube and the centroid position clockwise; voxelizing the reconstructed triangles to determine the points of the point cloud, The normal vector of the cube is determined based on the determined area of ​​the triangle.

3. The method according to claim 1, wherein: The bitstream further comprises an additional value α, and the centroid position is further determined based on the additional value α.

4. The method according to claim 2, further comprising: Determine the added value α; The additional value is encoded into the bitstream, wherein the centroid position is further determined based on the additional value α. 5 . The method according to claim 3 , wherein the additional value α is determined based on the positions of all points of the point cloud.

6. An encoder for encoding a 3D point cloud into a bitstream, the encoder comprising at least one processor and a memory, wherein: The memory stores instructions which, when executed by the processor, perform the steps of the method according to any one of claim 2 and claims 4 and 5 dependent on claim 2.

7. A decoder for decoding a 3D point cloud from a bitstream, the decoder comprising at least one processor and a memory, wherein: The memory stores instructions which, when executed by the processor, perform the steps of the method according to any one of claim 1 and claims 3 and 5 dependent on claim 1.

8. A bit stream encoded by the method according to claim 2 and any one of claims 4 and 5 when dependent on claim 2.

9. A computer-readable storage medium, comprising instructions, which, when executed by a processor, perform the steps of the method according to any one of claims 1 to 5.