Method, encoder and decoder for encoding and decoding 3d point clouds

By including octree information and vertex information in the bitstream and extending triangles with an adaptive halo method for voxelization, the problem of low compression and decoding efficiency of dense 3D point clouds in the prior art is solved, and higher decoding accuracy and compression performance are achieved.

CN119998837APending Publication Date: 2025-05-13BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202280100691.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress and decode intensive 3D point clouds, especially in dynamic AR/VR applications, resulting in low transmission efficiency and insufficient reconstruction accuracy.

Method used

By including octree information and vertex information in the bitstream, the triangle is extended by using an adaptive halo method to voxelize, thereby improving the decoding accuracy and compression performance of the point cloud.

Benefits of technology

It achieves higher 3D point cloud decoding accuracy and compression performance, reduces sampling errors during voxelization, and is suitable for dense and dynamic 3D point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998837A_ABST
    Figure CN119998837A_ABST
Patent Text Reader

Abstract

A method, preferably implemented in a decoder, for decoding a geometry of a 3D point cloud from a bitstream, the method comprising: receiving and decoding the bitstream, where the bitstream comprises octree information and vertex information, the octree information comprising information about an octree structure of a point cloud volume, the vertex information comprising information about an octree structure of the point cloud volume; the vertex information comprises information about vertex existence and vertex position on the edge of the cube of the leaf node of the octree structure; determining a triangle by connecting vertices of a cube associated with leaf nodes of the octree structure; and voxelizing the triangle to determine the points of the point cloud, characterized in that the method further comprises: whether additional information contained in the bit stream satisfies a predefined condition, the additional information being determined based on the density of the point cloud, preferably evaluated by the sampling distance dsampl of the point cloud; when a predefined condition is satisfied, at least one triangle is expanded along at least one edge based on the sampling distance dsampl for voxelization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for decoding a 3D point cloud from a bitstream. Furthermore, the object of the present invention is to provide a method for encoding a 3D point cloud into a bitstream. Further, the object of the present invention is to provide an encoder and a decoder, a bitstream encoded according to the present invention and software. In particular, the object of the present invention is to provide a method for improving the accuracy of the decoding or reconstruction process of a 3D point cloud, so that an overall better compression performance can be achieved. Background Art

[0002] Point clouds have recently gained attention as a format for 3D data representation due to their versatility in representing all types of 3D objects or scenes. As a result, many use cases can be solved with point clouds, including:

[0003] Film post-production,

[0004] Real-time 3D immersive telepresence or virtual reality (VR) / augmented reality (AR) applications,

[0005] Free viewpoint video (e.g. for watching sports),

[0006] Geographic Information Systems (also known as Cartography),

[0007] Cultural heritage (scans of rare objects stored in digital form),

[0008] Autonomous driving, including 3D mapping of the environment and real-time lidar data collection.

[0009] A point cloud is a set of points located in 3D space, optionally with additional values ​​attached to each point. These additional values ​​are often called point attributes. Thus, a point cloud is a combination of geometry (the 3D position of each point) and attributes.

[0010] The attributes may be, for example, a three-component color, a material property such as reflectivity, and / or a two-component normal vector of a surface associated with the point.

[0011] Point clouds can be captured by various types of devices, such as camera arrays, depth sensors, Lidar, scanners, or can be computer generated (e.g., in movie post-production). Depending on the use case, a point cloud may have thousands to billions of points for mapping applications.

[0012] The raw representation of point clouds requires a very high number of bits per point, with at least twelve bits per spatial component X, Y, or Z, and optionally more bits for attributes, e.g., three times 10 bits for color. Practical deployment of point cloud-based applications requires compression techniques that can store and distribute point clouds with reasonable storage and transmission infrastructure.

[0013] For distribution to and visualization by end users, for example, on AR / VR glasses or any other 3D-enabled device, compression can be lossy (like in video compression). Other use cases do require lossless compression, for example, medical applications or autonomous driving, to avoid altering the decision results obtained from the analysis of the compressed and transmitted point cloud.

[0014] Until recently, point cloud compression (aka PCC) had not been addressed by the mass market and no standardized point cloud codec was available. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11, also known as the Moving Picture Experts Group or MPEG, started a work item on point cloud compression. This resulted in two standards, namely

[0015] MPEG-I Part 5 (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC),

[0016] MPEG-I Part 9 (ISO / IEC 23090-9) or Geometry-based Point Cloud Compression (G-PCC).

[0017] Both the V-PCC and G-PCC standards completed their first versions at the end of 2020 and will be launched on the market soon.

[0018] The V-PCC encoding method compresses point clouds by performing multiple projections of 3D objects to obtain 2D patches that are packed into images (or videos when processing mobile point clouds). The acquired images or videos are then compressed using existing image / video codecs, allowing the use of already deployed image and video solutions. By its nature, V-PCC is only effective on dense and continuous point clouds, because image / video codecs cannot compress non-smooth patches like those obtained from the projection of sparse geometric data collected by, for example, Lidar.

[0019] The G-PCC coding method has two schemes for geometry compression.

[0020] The first scheme is based on an occupancy tree (octree / quadtree / binarytree) representation of the point cloud geometry. Occupied nodes are split until a certain size is reached, and the occupied leaf nodes provide the locations of the points, usually in the centers of these nodes. A high level of compression for dense point clouds can be obtained by using neighbor-based prediction techniques. Sparse point clouds are also addressed by directly encoding the locations of points within nodes with non-minimum sizes, by stopping the tree construction when only isolated points exist in the node; this technique is called Direct Coding Mode (DCM).

[0021] The second scheme is based on prediction trees, where each node represents the 3D position of a point and the relationship between nodes is the spatial prediction from parent to child. This approach can only handle sparse point clouds and has the advantages of lower latency and simpler decoding compared to occupancy trees. However, the compression performance is only slightly better than the first occupancy-based approach, and the encoding is complex, with intensive searching for the best predictor (in a long list of potential predictors) when building the prediction tree.

[0022] In both schemes, the attribute (de)coding can be performed after the geometry (de)coding is completed, resulting in a two-pass encoding. Thus, low latency is obtained by using slices that decompose the 3D space into independently coded sub-volumes without prediction between sub-volumes. When many slices are used, the compression performance may be severely affected.

[0023] An important use case is the transmission of dynamic AR / VR point clouds. Dynamic means that the point cloud evolves over time. Moreover, AR / VR point clouds are usually locally 2D, as they represent the surface of objects most of the time. Therefore, AR / VR point clouds are highly connected (or dense), as points are rarely isolated but have many neighbors.

[0024] A dense (or solid) point cloud represents a continuous surface at a resolution such that the volumes (small cubes called voxels) associated with the points touch each other without showing any visible holes in the surface.

[0025] Such point clouds are typically used in AR / VR environments and viewed by end users through devices such as TVs, smartphones or headsets. They are either transmitted to the device or stored locally. Many AR / VR applications use dynamic point clouds, which change over time, rather than static point clouds. Therefore, the amount of data is huge and must be compressed. Today, lossless compression based on an octree representation of the point cloud geometry can achieve slightly less than one bit per point (1bpp). This may not be enough for real-time transmission, which may involve several million points per frame and frame rates of up to 50 frames per second (fps), resulting in hundreds of megabits of data per second.

[0026] Therefore, lossy compression can be used, maintaining the usual requirements of acceptable visual quality while compressing sufficiently to fit within the bandwidth provided by the transmission channel while maintaining real-time transmission of frames. In many applications, bit rates as low as 0.1bpp (10 times higher than lossless coding compression rates) have made real-time transmission possible.

[0027] Codecs based on MPEG-I Part 5 (ISO / IEC 23090-5) or Video-based Point Cloud Compression (V-PCC) such as VPCC can achieve such low bitrates by using lossy compression with a video codec that compresses 2D frames obtained from the projection of the point cloud onto a plane. The geometry is represented in one frame by a sequence of projected patches, each of which is a small local depth map. However, VPCC is not general and is limited to narrow types of point clouds that do not exhibit locally complex geometry (e.g., trees, hair), since the obtained projected depth maps are not smooth enough to be effectively compressed by video codecs.

[0028] Pure 3D compression techniques can handle any type of point cloud. It is still an open question whether 3D compression techniques can compete with VPCC (or any projection + image coding scheme) on dense point clouds. Standardization is still moving towards providing an extension (revision) of GPCC that provides competitive lossy compression that can compress dense point clouds as well as VPCC intra-frame coding, while maintaining the versatility of GPCC to handle any type of point cloud (dense point cloud, Lidar, 3D map). This extension will probably use the so-called TriSoup (triangle soup) encoding scheme, which works with octrees. TriSoup is being explored by the ISO / IEC standardization working group JTC1 / SC29 / WG7. TriSoup coding is also described in "Adaptive multi-level triangle soup forgeometry-based point cloud coding" by A.DRICOT et al., 2019 IEEE 21st International Symposium on Multimedia Signal Processing (MMSP), "report on triangle soup decoding" by Nakagami O., m52279 in 2020 by ISO / IEC JTC1 / SC29-WG11, and US10,192,353.

[0029] However, for all lossy compression schemes, the reconstruction quality of the points in the point cloud is crucial. Summary of the invention

[0030] It is therefore an object of the present invention to provide a method for decoding the geometry of a 3D point cloud from a bitstream and encoding the 3D point cloud into a bitstream, which method improves the accuracy.

[0031] This problem is solved by a decoding method according to claim 1 , an encoding method according to claim 2 , an encoder according to claim 16 , a decoder according to claim 17 , a bitstream according to claim 18 and a software according to claim 19 .

[0032] In a first aspect, a method for decoding the geometry of a 3D point cloud from a bitstream is provided, preferably implemented in a decoder. The method comprises:

[0033] Receive and decode a bitstream, wherein the bitstream includes octree information and vertex information, the octree information includes information about an octree structure of a point cloud volume, and the vertex information includes information about the existence and position of vertices on edges of cubes of leaf nodes of the octree structure;

[0034] Determining triangles by connecting vertices of a cube associated with a leaf node of the octree structure;

[0035] voxelizing the triangles to determine points of the point cloud,

[0036] Characterized in that the method further comprises:

[0037] Determine whether the additional information included in the bitstream satisfies a predefined condition, wherein the additional information is determined based on the density of the point cloud, preferably by a sampling distance d of the point cloud. sampl Evaluate;

[0038] When the predefined condition is met, based on the sampling distance d sampl At least one triangle is expanded along at least one edge for voxelization.

[0039] Therefore, in the first step, a bitstream is received, and the bitstream contains information about the octree structure of the volume of the decoded point cloud. Preferably, the geometric structure of the point cloud is GPCC encoded. Therefore, by decoding from the bitstream, octree information about the volume of the point cloud can be provided. Further, the bitstream also includes vertex information, which includes information about the existence and position of vertices on the edges of the cube associated with the leaf nodes in the octree structure. Therefore, the vertex information is provided by decoding from the bitstream. Wherein, the bitstream is preferably encoded by the Trisoup encoding scheme at the encoder.

[0040] After decoding the octree information and vertex information from the bitstream described in the previous step, in a further step for reconstructing the point cloud geometry, triangles are determined for each cube by connecting the vertices on the edges of the cube. Therefore, the surface of the triangle is determined by the position of the vertices included in the bitstream. In order to reconstruct the points of the point cloud from the triangles, voxelization is performed by ray tracing, wherein, during the ray tracing process, the rays are emitted in three directions parallel to any of the three axes. Their origins are integer coordinate points corresponding to the sampling accuracy required for rendering. The intersection of the ray with one of the triangles (if any) is then determined, and the intersection is added to the list of rendered points, i.e., to the points of the point cloud. During the voxelization process, the surface of the triangle is sampled using rays to determine the points of the point cloud.

[0041] According to the present invention, how to determine the triangle is based on the additional information contained in the bit stream, and there are different schemes for determining the triangle:

[0042] Scheme 1. Adaptive halo: wherein, for / during voxelization, at least one triangle is extended along at least one edge to a distance d based on the sampling distance of the point cloud. sampl The surface of the triangle is extended along at least one direction. The sampling distance is a property of the initial point cloud data and is related to the distance between the actual sampled points of the point cloud in units of the sampling resolution if no points were lost during data collection. sampl Set by, for example, a device (such as LIDAR, etc.) that acquires the points of the point cloud. Therefore, by expanding the triangle during the voxelization process, the accuracy of the voxelization process can be improved, because additional points of the original point cloud can be reliably determined, otherwise these points will be ignored during the voxelization process. Since the triangle is sampled with a certain accuracy and sampling resolution, by expanding the triangle along at least one edge to expand the surface of the triangle, the points of the point cloud just outside the triangle can now be captured. In addition, since the extension of the triangle is based on the sampling distance of the point cloud, the extension will be applicable to any point cloud regardless of the sampling distance. Preferably, the extension is proportional to the sampling distance of the point cloud. Therefore, if the sampling distance of the point cloud becomes larger, the triangle will also be expanded to a greater extent. Further details of the adaptive halo scheme are described in the dependent claims. Therefore, in many cases, a higher accuracy of reconstructing the 3D point cloud is achieved, and the number of sampling errors in the voxelization process is reduced. In addition, the complexity of the encoding and / or decoding algorithm is maintained. However, the scheme does not perform well on all kinds of point clouds. In some cases, it may even result in a loss of compression results.

[0043] Solution 2. Halo with fixed or other values: Compared with adaptive halo, the expansion of triangles in this solution is based on a fixed value that is independent of the sampling distance of the point cloud. It should be understood that Solution 2 can also be other solutions for determining triangles, for example, not expanding triangles at all.

[0044] Therefore, according to the present invention, additional information included in the bitstream is used to select an adaptive halo or non-adaptive halo scheme. By introducing such an indicator, each scheme according to the present invention can be applied to the appropriate use case, thereby achieving an overall better compression performance compared to solutions that only implement a single scheme.

[0045] Preferably, at least one triangle is expanded at more than one edge to further expand the surface of the corresponding triangle. Thus, a triangle can be expanded at one edge, two edges, or all three edges to include points of the original point cloud that just exceed the triangle defined by the vertices on the edge of the cube.

[0046] Preferably, if a cube of a leaf node of an octree structure can contain more than one triangle, each triangle of the cube is extended along at least one edge for voxelization. Therefore, the extension of the triangular surface can be applied to all triangles in the cube. Alternatively or additionally, in each cube of the octree structure, at least one triangle is extended along at least one edge for voxelization. Alternatively, the extension of one or more edges of a triangle will be applied only to a subset of leaf nodes in the octree structure. Wherein, the subset can be determined by, for example, application, the density of points in the leaf nodes of the point cloud or the requirements for accuracy and decoding speed. More preferably, one or more edges of a triangle are extended based on a local sampling distance. Therefore, the triangles of each subset of leaf nodes can be extended in a manner that can achieve local optimal performance.

[0047] Preferably, the expansion is the same for each edge. Thus, the triangle is expanded by the same amount along at least two directions to enlarge the surface of the triangle. More preferably, the amount of expansion is the same for all three directions. Alternatively, the expansion along at least two directions is different. Thus, different directions can be treated differently to improve decoding accuracy.

[0048] Preferably, the extension is the same or different for each leaf node of the octree structure. If there is a different extension for more than one edge or each edge of a triangle in a leaf node of the octree structure, this may be the same or different in other leaf nodes of the octree structure. The extension may be pre-selected or may be determined, for example, by the application, the density of points in the leaf nodes of the point cloud, or the requirements for accuracy and decoding speed.

[0049] Preferably, by -Trumbore algorithm to perform voxelization.

[0050] Preferably, -In the Trumbore algorithm, the convex hull requirement is relaxed to -ε a ≤u,v,w, where ε a >0, u, v, w are the barycentric coordinates of the triangle, where ε a Sampling distance d based on point cloud sampl OK. In the original -In the Trumbore algorithm, the convex hull requirement is set to 0≤u,v,w. Therefore, by relaxing this requirement to -ε a ≤u,v,w, the surface of the considered triangle is enlarged and the voxelization of the points of the original point cloud will now be included, which otherwise would not have been considered in the reconstructed point cloud during sampling. In particular, due to ε a is the sampling distance d based on the point cloud sampl Determined, so the expansion will work for any point cloud regardless of the sampling distance. Preferably, the expansion is proportional to the sampling distance of the point cloud. Therefore, if the sampling distance of the point cloud becomes larger, the triangles will also expand to a greater extent. As a result, the reconstruction quality and appearance of the final reconstructed point cloud is improved.

[0051] Preferably, the convex hull is required to be set to -ε u_a ≤u, -ε v_a ≤v and -ε w_a ≤w, where ε u_a ,ε v_a ,ε w_a ≥0, and u, v, w are the barycentric coordinates of the triangle, where ε u_a , ε v_a , ε w_a At least one of the sampling distances d based on the point cloud sampl Therefore, for different directions, separate convex hull requirements can be provided to individually control the expansion of the considered triangle. u_a ≠ε w_a Alternatively or additionally, ε u_a ≠ε v_a Alternatively or additionally, ε v_a ≠ε w_a Thus, the expansion in one or more directions may be selected independently of the other directions to determine the expansion individually.

[0052] Preferably, the expansion is provided by an adaptive halo parameter. In the case of the Trumbore algorithm, the adaptive halo parameter is given by ε aProvided, and for different directions by ε u_a , ε v_a and ε w_a Thus, by adapting the halo parameters, the amount of extension can be determined and quantified based on the sampling distance of the point cloud.

[0053] Preferably, the adaptive halo parameter is set to be less than 1 / 4d sampl More preferably, the adaptive halo parameter is set to be less than 1 / 8d sampl Therefore, by choosing an adaptive halo parameter, the amount of expansion can be adjusted to achieve the best results, where larger values ​​will result in more points being determined during the voxelization process. The preferred range for the adaptive halo parameter will be between 0 and d sampl If the sampling distance is large, the adaptive halo parameter will also become larger, thereby increasing the expansion amount. Therefore, even if the sampling distance changes, the present invention also provides an adaptive solution for expanding the triangle, thereby ensuring that the expanded triangle always covers a reasonable number of points.

[0054] Preferably, the adaptive halo parameters are pre-set. Thus, the encoder and decoder may have agreed on the adaptive halo parameters, and thus the adaptive halo parameters are fixed for each point cloud generated by the encoder and reconstructed by the decoder. Information about the adaptive halo parameters does not need to be encoded into the bitstream.

[0055] Alternatively, the adaptive halo parameters are encoded into the bitstream, and preferably in a geometry parameter set (GPS) of the bitstream. In the case where the adaptive halo parameters are set for each subsequent point cloud to be decoded, this can be done only once. Alternatively, for each point cloud, the corresponding adaptive halo parameters or adaptive halo parameter sets can be encoded separately.

[0056] Alternatively, the halo parameters also depend on the size of the volume of the cube, ie the level of the octree of the current leaf node.

[0057] Preferably, the sampling distance d of the point cloud is sampl Depend on Determine, where N leaf is the number of leaf nodes, N total is the number of points in the point cloud, and N is the size of the corresponding cube of the leaf node, or the sampling distance d of the point cloud. sampl Determined by a cyclic method. In which, on the encoder side, N total It is known to the encoder. In addition, the number of leaf nodes N leaf It is known on the encoder side. Further, N defines the size of the leaf node in units of the sampling resolution of the original point cloud data acquired by the device. Therefore, d samplcan be determined from the point cloud data before voxelization and depends on the size of the leaf node cube. Thus, as the leaf node size N increases, d sampl will also increase, thereby increasing the adaptive halo parameter. Additionally or alternatively, during the voxelization process, the sampling distance can also be determined by a loop method to select the best sampling distance. Specifically, the loop method starts from 1 to N, tries different integer values ​​to estimate the sampling distance, and increases the sampling distance by 1 from this cycle to the next cycle. In each cycle k, by using the sampling distance d of the cycle k To estimate the number of points in the reconstructed point cloud generated during the voxelization process, and compare this number with the N of the original point cloud total Compare; if the number of points in the reconstructed point cloud is greater than N in the i-th cycle total , then the loop method ends and the estimated sampling distance for voxelization is equal to d i -1.

[0058] Preferably, at least one triangle is weighted based on a halo parameter ε a_t Extend along at least one edge to voxelize, where the weighted halo parameter ε a_t By ε a_t =ε a *t is determined, where ε a is the sampling distance d based on the point cloud sampl The adaptive halo parameter of the triangle is set, and the expansion of at least one triangle is provided, and t is the corresponding weight associated with the sampling distance, preferably set to 2. In some embodiments, t is selected to be between 1 and 4, more preferably between 1.5 and 2.5. Among them, a heuristic method can be used to determine the value of t. A heuristic method is an optimization method that attempts to find the global optimal feasible solution for the specific problem under consideration. Heuristic methods are iterative in nature. After each iteration, a feasible solution to a specific problem is determined. When the heuristic method terminates after a certain time or a certain number of iterations, the output solution is the best solution found in any iteration. More preferably, the weight to be tried in each iteration is an integer selected from the range of 1 to 4. Among them, the adaptive halo parameter is less than 1. If the weight is too large, the overall accuracy of the TriSoup model may be affected. Therefore, the upper limit can be set to 4. For example, if the adaptive halo parameter is 1 / 4 and it is determined that the best result can be obtained by assigning a weight of 2 to the sampling distance, then if the adaptive halo parameter is proportional to the sampling distance, the updated adaptive halo parameter can be 1 / 4*2=1 / 2. Thus, by providing an appropriate weight setting range, the efficiency and accuracy of the overall algorithm can be further improved.It is understandable that different weights can also be determined in different directions of the triangle.

[0059] Preferably, the additional information is a flag for enabling or disabling the functionality of the encoding or decoding method, preferably one bit. In the simplest case, the additional information can be a 1-bit flag indicating whether the adaptive halo scheme should be enabled. It will be understood that the additional information can also be multiple bits, as long as it can indicate the information required according to the present invention.

[0060] Preferably, the additional information is encoded into a Geometry Parameter Set (GPS) of the bitstream.

[0061] In another aspect of the present invention, a method for encoding a 3D point cloud into a bitstream is provided, preferably implemented in an encoder. The encoding method of the 3D point cloud comprises:

[0062] Obtaining octree information, the octree information including an octree structure of a body, the body including a plurality of cubes;

[0063] Acquire vertex information from a surface of a point cloud of each cube associated with a leaf node, wherein the vertex information includes information about the existence and position of vertices on an edge of the cube;

[0064] Encoding the octree information and the vertex information into a bitstream;

[0065] Reconstructing the geometric data of the point cloud by using the octree information and vertex information obtained in the aforementioned encoding process;

[0066] Wherein, reconstructing the geometric data of the point cloud includes:

[0067] Determining triangles by connecting vertices of a cube associated with a leaf node of the octree structure;

[0068] voxelizing the triangles to determine points of the point cloud;

[0069] Characterized in that the method further comprises:

[0070] Additional information is determined based on the density of the point cloud, preferably by a sampling distance d of the point cloud. sampl Evaluate;

[0071] encoding the additional information into the bitstream;

[0072] Determining whether the additional information satisfies a predefined condition;

[0073] When the predefined condition is met, based on the sampling distance d sampl At least one triangle is expanded along at least one edge for voxelization.

[0074] Therefore, through the encoding method, octree information and vertex information are generated. In addition, based on the density of the point cloud, for example, based on the sampling distance of the point cloud, additional information is determined and generated. It should be understood that the density can also be determined by other methods, which will not be described in detail here. The information is encoded into the bit stream. Subsequently, on the encoder side, a reconstruction step is performed. In this reconstruction step, the point cloud geometric information is reconstructed, wherein the reconstruction steps are the same as the steps in the above-mentioned decoding method. Then, the reconstructed geometric structure of the point cloud is used on the encoder side to encode the encoding attributes (color, reflectivity, ...) of the points of the point cloud, for example, by RAHT (Regional Adaptive Hierarchical Transform), prediction transform or lifting transform to encode the attributes of the points of the point cloud.

[0075] Preferably, the geometric structure of the point cloud is encoded into the bitstream via Geometry-based Point Cloud Compression (G-PCC).

[0076] Preferably, the bitstream is an MPEG G-PCC compliant bitstream.

[0077] Preferably, the encoding method is further constructed according to the features described above in conjunction with the decoding method.

[0078] In another aspect of the present invention, an encoder for encoding a 3D point cloud into a bit stream is provided. The encoder comprises a memory and a processor, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the steps of the aforementioned encoding method are performed.

[0079] In another aspect of the present invention, a decoder for decoding a 3D point cloud from a bitstream is provided. The decoder comprises a memory and a processor, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the steps of the above-mentioned decoding method are performed.

[0080] In another aspect of the present invention, a bit stream is provided, wherein the bit stream is encoded based on the steps of the aforementioned encoding method.

[0081] In another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium comprising instructions for executing the steps of the above method for encoding a 3D point cloud into a bitstream.

[0082] In another aspect of the present invention, a computer-readable storage medium is provided, comprising instructions for executing the steps of the above method for decoding a 3D point cloud from a bitstream.

[0083] In another aspect of the present invention, a computer-readable storage medium is provided, comprising instructions for executing the steps of the above-mentioned method for encoding a 3D point cloud into a bitstream, and further comprising a configuration file indicating the type of the point cloud, which indicates the density of the point cloud. Among them, the type of point cloud can be, for example, solid, dense, sparse, and scant. However, it should be understood that these types can essentially be distinguished by the sampling distance of the point cloud as described above. In any of the above embodiments, additional information can also be determined based on the type of point cloud (for example, by obtaining information from a configuration file). BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Hereinafter, the present invention is described in more detail with reference to the accompanying drawings.

[0085] These figures show:

[0086] Figure 1a A flowchart showing a method for decoding 3D point cloud geometry according to the present invention;

[0087] Figure 1b A simplified flow chart showing a decoding method according to the present invention;

[0088] Figure 2 An example of generating an octree structure is shown;

[0089] Figure 3 Show according to Figure 2 The octree of

[0090] Figure 4 An example of determining vertices on the edge of a cube is shown;

[0091] Figure 5 An example of generating a triangle is shown;

[0092] Figure 6 An example showing vertices on the edges of a cube;

[0093] Figure 7 Shows the generation of triangles from vertices;

[0094] Figure 8 Show the basis for determination Figure 7 An example of the order of triangles;

[0095] Fig. 9 A schematic diagram showing the voxelization step;

[0096] Fig.10 shows the 2D representation of triangles in the leaf nodes of the octree;

[0097] Fig.11 Show Fig.10Example of voxelization of a triangle;

[0098] Fig.12 Show Fig.10 The coordinates and definition of the center of gravity of the triangle;

[0099] Fig.13 shows the comparison between the vertex triangles and the original point cloud;

[0100] Fig.14a Show Fig.10 The triangle of is expanded in one direction based on a fixed halo parameter in barycentric coordinates;

[0101] Fig.14b Show Fig.10 The triangle is expanded in a direction based on the adaptive halo parameters in barycentric coordinates;

[0102] Fig.15a Show Fig.10 The triangle is expanded in all three directions based on a fixed halo parameter;

[0103] Fig.15b Show Fig.10 The triangle is expanded in all three directions based on the adaptive halo parameters;

[0104] Fig.16 The weighted halo parameter ε is shown a_t A representation of a triangle extending in three directions from the sampling distance of the point cloud;

[0105] Fig.17a shows a representation of a triangle expanded in three directions based on the sampling distance of the point cloud;

[0106] Fig.17b A representation showing a triangle expanded in three directions by a fixed amount;

[0107] Fig.17c shows a representation of a triangle expanded by a fixed amount in three directions when the sampling distance is 1;

[0108] Fig.18 A schematic flow chart showing an encoding method;

[0109] Fig.19a The performance of the longdress data based on different halo parameters is shown;

[0110] Fig.19b shows the performance of the house_without_roof data based on different halo parameters; and

[0111] Fig.19c Shows the performance of the ulb_unicorn data based on different halo parameters. DETAILED DESCRIPTION

[0112] Reference Figure 1a , which shows a schematic diagram of a method for decoding geometric information of a 3D point cloud from a bitstream.

[0113] A method for decoding the geometry of a 3D point cloud from a bitstream, preferably implemented in a decoder, comprises the following steps:

[0114] In step S01, a bit stream is received and decoded, wherein the bit stream includes octree information and vertex information, the octree information includes information about an octree structure of a point cloud volume, and the vertex information includes information about the existence and position of vertices on edges of cubes of leaf nodes of the octree structure;

[0115] In step S02, a triangle is determined by connecting the vertices of a cube associated with a leaf node of the octree structure;

[0116] In step S03, the triangles are voxelized to determine the points of the point cloud.

[0117] Determine whether the additional information included in the bitstream satisfies a predefined condition, the additional information being determined based on the density of the point cloud, preferably by the sampling distance d of the point cloud sampl Evaluation; when the predefined conditions are met, at least one triangle is based on the sampling distance d sampl Expand along at least one edge to voxelize.

[0118] To determine the octree information, the first step of the geometry encoding process is to construct and encode the octree, such as Figure 2 and Figure 3 As shown. The bounding box is the body 100 that contains all the points and is associated with a root node 112 (i.e., a single node at the top of the tree 110). The body 100 is first divided into eight sub-volumes 102, called octants, each of which is represented by a node 114 in the tree 110. Then, octants 106 are recursively split out of the sub-volumes 104 until the target level is reached, where an octant 106 is occupied by at least one point. Figure 2 and Figure 3 Indicated by shading.

[0119] Each octant (or node) is represented by an occupation byte containing one bit for each sub-octant, with the corresponding bit set to one if the sub-octant is occupied by at least one node, and set to zero otherwise. The occupation bytes 118 of all octants are serialized (in breadth-first order) and entropy encoded using a binary arithmetic encoder.

[0120] Figure 4A block representation of a 3D surface 210 is shown, along with an example of a block 220 in TriSoup. The surface 210 intersects the block 220, so the block 220 is an occupied block, and the block 220 exists between multiple blocks 200 in the 3D space. Within the block 220, the closed portion of the surface 210 intersects the edge of the block at the six illustrated vertices of the polygon 230. If an edge of the block 220 contains a vertex, the edge is said to be selected.

[0121] Figure 5 Block 220 in TriSoup is shown, with surface 210 omitted for clarity, and unselected edge 270, selected edge 260, and i-th edge 250 are shown. Assume that i-th edge 250 is selected. To specify vertex v on edge i i , specifies a scalar value indicating the corresponding fraction of the length of edge 250.

[0122] like Figure 4 and Figure 5 As shown, within each octant 220 of the target level of the octree, trisoup represents the original surface 210 as a set of triangles 245. The surface is encoded and used to obtain the position of the reconstructed (or decoded) points. First, the intersection of the surface represented by the original points with the edges of the octant is estimated by averaging the positions of the points closest to those edges within the octant among the original points used to represent the surface. Secondly, the twelve edges of all octants and their associated intersections (if any) are stored as segments and vertices respectively. Each (unique) segment is then encoded as follows. The first single bit is arithmetically encoded, and the bit is set to 1 if the segment is occupied by a vertex, otherwise it is set to 0. If it is occupied, the relative position of the vertex on the segment is also arithmetically encoded.

[0123] The vertices 310 of the triangles are encoded along the edges 320 of the volume associated with the leaf nodes 300 of the tree, as Figure 6 As shown. These vertices 310 on the edge 320 are shared between multiple leaf nodes 300 with a common edge 320. This means that each edge belonging to at least one leaf node encodes at most one vertex. In this way, the continuity of the model is ensured through the leaf nodes.

[0124] As mentioned above, the encoding of TriSoup vertices requires two pieces of information per edge:

[0125] A vertex flag indicating whether a TriSoup vertex exists on the edge, and

[0126] The vertex positions along the edge, when present.

[0127] Therefore, the encoded data includes Octree data as well as TriSoup data.

[0128] The vertex flags are encoded by an adaptive binary arithmetic encoder that uses a specific context to encode the vertex flags. The length is N = 2 s The positions of vertices on the edges of can be encoded with unit precision by pushing s bits into the (bypassing / non-entropy encoding) bitstream.

[0129] Within a leaf node, if there are at least three vertices 310 on an edge 320 of the leaf node 300, a triangle is constructed from the TriSoup vertices. Figure 7 Reconstructed triangles 330, 340 are depicted in FIG.

[0130] Obviously, other combinations of triangles 330, 340 are possible. The selection of triangles results from a three-step process:

[0131] 1. Determine the dominant direction along one of the three axes;

[0132] 2. Sort the TriSoup vertices based on the dominant direction;

[0133] 3. Build a triangle based on the ordered list of vertices.

[0134] Knowledge of the exact position of the triangle in the current leaf node is not required and can be inferred from the vertices.

[0135] Figure 8 will be used to explain this process. Each of the three axes is tested and the one that maximizes the total surface of the triangle is chosen as the principal axis. For simplicity of the diagram, Figure 8 Only tests on two axes are described.

[0136] The first test along the vertical axis (top) is performed by vertically projecting the cube and TriSoup vertices 310 on the 2D plane. The vertices 310 are then sorted in clockwise order relative to the center of the projection node (square). Then, based on the ordered vertices, triangles 330, 340 are constructed according to a fixed rule. Here, when 4 vertices are involved, triangle 123 as well as triangle 134 are systematically constructed. When there are 3 vertices, the only possible triangle is 123. When there are 5 vertices, the fixed rule can be to construct triangles 123, 134 and 451. And so on, up to a maximum of 12 vertices.

[0137] The second test (left side) along the horizontal vertical axis is performed by projecting the cube and trisoup vertices horizontally onto the 2D plane.

[0138] The vertical projection shows the maximum total 2D surface of the triangle, therefore, the principal axis is chosen as the vertical axis, and the constructed TriSoup triangles are obtained in the order of vertical projection, such as Figure 8 As shown, it is located inside the node. It should be noted that taking the horizontal axis as the main axis will result in another construction of the triangle.

[0139] The principal axes are appropriately selected by maximizing the projection surface, thus achieving continuous reconstruction of point clouds without holes.

[0140] Rendering of TriSoup triangles into points is performed by ray tracing. The set of all points rendered by ray tracing will form the decoded point cloud.

[0141] for Fig. 9 In the ray tracing shown, rays are emitted in three directions parallel to the axis. Their origin is a point with integer (voxelized) coordinates of a precision corresponding to the sampling precision required for rendering. The intersection with one of the trisoup triangles (dashed point if any) is then voxelized (= rounded to the nearest point of the required sampling precision) and added to the list of rendered points.

[0142] After applying Trisoup to all leaf nodes, i.e., building triangles and obtaining points by ray tracing, discard all copies of the same point in the list of rendered points (i.e., keep only one voxel among all voxels sharing the same position and volume) to obtain a set of decoded (unique) points.

[0143] To keep it simple, start here. Figures 10 to 16 Rather than 3D volumes (cubes), 2D volumes (squares) associated with leaf nodes will be depicted. The reader will remember that all methods described in this invention apply to 3D space.

[0144] refer to Fig.10 , which shows an example of an N×N×N volume, where N=2 s = 8. There are at least three vertices V1, V2, V3 on the edge 410 of the body (depicted as squares in the figure, but actually cubes).

[0145] The edges of the leaf nodes are located at positions -0.5 and N-0.5 to ensure the continuity of the TriSoup model when passing from a "volume" to an adjacent volume. In practice, this means that the faces of the cubes are shared between adjacent volumes. This way, the position of the vertices present on an edge does not depend on the cube to which the edge belongs.

[0146] The position p of the vertex along its corresponding edge k are quantized positions 400 and are encoded into the bitstream. These positions 400 can be quantized with a unit step size such that p k is an integer in the interval [0, N-1]. Fig.10 In the example, p1=4, p2=2 and p3=2.

[0147] The TriSoup triangle 440 is constructed from vertices V1, V2, V3, and the set of triangles belonging to a volume models the point cloud contained in the volume.

[0148] The process of recovering the points 430 (of the decoded point cloud) from the triangles 440 is called voxelization of the triangles. Fig.11 Shows Fig.10 Voxelization of the TriSoup triangle. Rays are fired along all integer coordinates 420 (white and black dots), and rays that intersect the triangle are directed to the partially decoded points (black dots). The origins of the rays have a spacing D that sets the sampling resolution for voxelization.

[0149] The intersection between the ray and the triangle is determined by using -Trumbore algorithm, which determines the location of the intersection point by using the barycentric coordinates, such as Fig.12 shown.

[0150] Any point P in 3D space can be uniquely represented by its barycentric coordinates relative to any non-degenerate 3D triangle ABC (equivalently any triangle V1V2V3 from the TriSoup model).

[0151] Any point P in 3D space can be uniquely represented as

[0152] P=uA+vB+wC

[0153] Among them, there are conditions

[0154] u+v+w=1.

[0155] The points of the triangle correspond to the convex hull. Therefore,

[0156] 0≤u,v,w.

[0157] -Trumbore determines the values ​​of u and v; then w can be simply found by v = 1-uv. Fig.12 , the light from point P start By direction Emit. Set the following signs for the 3D vector derived from the 3D point, such as Fig.12 As shown:

[0158] and

[0159] The intersection point P between the ray and the only plane passing through A, B, and C is obtained by the following calculation

[0160]

[0161] The intersection point P belongs to the triangle if and only if 0≤u,v,w.

[0162] There is a slight offset between the position of the Trisoup triangle V1V2V3 determined by the vertices in the bitstream and the natural position of the triangle 450 in the current volume, such as Fig.13 This position is natural because the encoder has inferred the vertex V from the nearest (with respect to the edge) point in the original point cloud. k Therefore, it is very likely that the vertex V k The voxelized points of are the points of the point cloud. These points are natural candidates for constructing "natural" triangles that model the point cloud.

[0163] This offset is caused by the continuity constraints of adjacent volumes. This causes the ray tracing to miss some points 460 ( Fig.13 P miss ), because these points do not belong to the Trisoup triangle (as opposed to the triangle formed by Fig.11 The direct result is a degradation of the quantized geometry metric and a reduction in the rate-distortion performance of the scheme.

[0164] Therefore, by slightly relaxing the convex hull condition 0≤u,v,w, a "halo" can be created around the TriSoup triangle. In this way, the size of the triangle is slightly increased so that the ray trace will intersect the increased size triangle and reduce the number of lost points P miss .

[0165] Let ε>0 be the halo parameter. Fig.14a As shown, the condition 0≤u is relaxed to −ε≤u, where u is the centroid weight associated with point A, and the size of the triangle is increased along the side BC opposite to point A, as shown by the dashed area 470.

[0166] The relaxation of the condition can be applied to the three centroid weights u, v and w by changing the convex hull 0≤u,v,w to -ε≤u,v,w.

[0167] The resulting halo 480 around the triangle 440 is as follows Fig.15a As shown in Figure 2. To a first order approximation, the size of the halo is proportional to the parameter ε.

[0168] The halo parameter can depend on the weight of each barycenter of the triangle, for example - ε u ≤u, -ε v ≤v and -ε w ≤w. Among them, ε u , ε v and ε w are the three halo parameters.

[0169] The impact on voxelization is as follows Fig.17b As shown, some points P miss is now part of the "halo" and is therefore included as a point of the decoded point cloud. Therefore, it is not lost as it was in the original algorithm.

[0170] Of course, the halo parameter ε (alternatively ε u , ε v and ε w ) must be set to have halos of sufficient size. In the case where ε is too small, the halos are very small and have almost no effect, thus falling back to the missing point problem of the prior art. In the case where ε is too large, the halos become larger and the overall accuracy of the Trisoup model suffers. In both cases, the distortion of the decoded point cloud is not optimal.

[0171] It has been observed that reasonable values ​​of the halo parameter ε are around ε≈1 / 4 or ε≈1 / 8.

[0172] However, setting the halo parameters to fixed values ​​has the disadvantage that the "halos" created may not always give optimal results.

[0173] In order to prove that an arbitrarily fixed halo value cannot obtain the best result on a set of data, the performance of different halo values ​​ε on three test point clouds named "longdress_viewdep_vox12", "house_without_roof_00057_vox12" and "ulb_unicorn_vox13" used in MPEG G-PCC are tested. In the test experiment, the halo parameter values ​​used in G-PCC code are obtained by multiplying ε by 256 to improve the calculation accuracy, and they are set to 16, 32, 64 and 128, so that the corresponding values ​​of ε are 1 / 16, 1 / 8, 1 / 4 and 1 / 2. For each data, the coding performance is obtained using these four halo values ​​ε at the same compression rate r02.

[0174] Fig.19a a, b, and c show the relationship between the quality (geometric PSNR) of the decoded point cloud and the halo ε value. Higher PSNR means better quality. It is observed that different data can achieve the maximum PSNR quality at different halo values ​​ε. For example, the best halo value for longdress data is 128, the best halo value for house_without_roof data is 128, and the best halo value for ulb_unicorn data is 32. Therefore, in order to achieve the best compression performance on various datasets, the value of the halo parameter ε may not be fixed.

[0175] like Fig.17cAs shown, there are two natural points (indicated by all black dots) near some edges of the volume, and TriSoup triangle V1V2V3 is derived from these points. And it is observed that there is an offset between the TriSoup triangle V1V2V3 and the natural position in the current volume. Fig.17c In the example, the sampling distance of the points is 1, and the enlarged triangle obtained using the current fixed halo parameter ε can cover all the points except P. miss However, if the sampling distance becomes larger, such as Fig.17b As shown in the figure, the natural point is farther away from the triangle than when the sampling distance is 1, so more natural points will be lost using the current fixed halo parameter. Therefore, in order to reduce the reconstruction error of the point, a larger halo parameter is required for point data with a larger sampling distance.

[0176] Therefore, by slightly relaxing the convex hull condition 0≤u,v,w, an adaptive "halo" is created based on the sampling distance of the point cloud around the TriSoup triangle. In this way, the size of the triangle is slightly increased so that the ray tracing will intersect the triangle with the increased size and reduce the number of missing points P. miss .

[0177] The advantages of the adaptive halo method are

[0178] Less distortion in the decoded point cloud. In fact, a quantitative metric (BDBR) shows that the proposed method can achieve a 2.6% bitrate gain (for equivalent quality) compared to a non-adaptive approach (i.e. fixed halo parameters)

[0179] Since the overall algorithm has not changed, the complexity is maintained.

[0180] Assume ε a >0 is the adaptive halo parameter determined based on the sampling distance of the point cloud. Fig.14b As shown, the condition 0≤u is relaxed to ε a ≤u, where u is the centroid weight associated with point A, increasing the size of the triangle along the side BC opposite point A, as shown by the dashed area 472.

[0181] By changing the convex hull 0≤u,v,w to -ε a ≤u,v,w, the relaxation of the condition can be applied to the three centroid weights u, v and w.

[0182] The resulting halo 482 around the triangle 440 is as follows Fig.15b As shown in the first order approximation, the size of the halo is related to the adaptive halo parameter ε a Therefore, it may also be proportional to the sampling distance of the point cloud.

[0183] In one embodiment, the adaptive halo parameter may additionally depend on each barycentric weight of the triangle, for example -ε u_a ≤u, -ε v_a ≤v and -ε w_a ≤w. Among them, ε u_a , ε v_a and ε w_a are the three adaptive halo parameters.

[0184] Fig.17a and Fig.17b The effect on voxelization is depicted, with Fig.17b In contrast, when the sampling distance becomes larger and the halo parameter remains fixed (applicable to smaller sampling distances), there are many missing points P miss ,exist Fig.17a In the example, due to the application of the adaptive halo parameters according to the present invention, several missing points P miss are now part of the halo and are therefore decoded as points of the decoded point cloud. Therefore, they are not lost as in the original algorithm.

[0185] refer to Fig.16 ,and Fig.17a Compared with miss is now part of the halo. Larger halos are provided by the weighted halo parameter. The weighted parameter not only takes into account the sampling distance of the point cloud, but also associates a weight t with the sampling distance. Fig.16 In the example, the weight t is set to 2. Thus, the accuracy of the decoding or reconstruction process of the 3D point cloud is further improved.

[0186] Of course, the adaptive halo parameters ε a(alternatively ε u_a , ε v_a and ε w_a ) must be set to give a halo of sufficient size. a If ε is too small, the halo is very small and has almost no effect, thus falling back to the missing point problem in the prior art. If ε is too large, the halo becomes larger and the overall accuracy of the Trisoup model suffers. In both cases, the distortion of the decoded point cloud is not optimal.

[0187] It has been observed that the halo parameter ε a A reasonable value of a ≈d sampl / 4 or ε a ≈d sampl / 8.

[0188] If the sampling distance is fixed, the adaptive halo parameter ε a (Alternatively ε u_a , ε v_a and εw_a ) can be a fixed value. In one variation, the halo parameter ε a (Alternatively ε u_a , ε v_a and ε w_a ) is encoded into the bitstream, for example in the geometry parameter set (GPS). In another variant, the halo parameter ε a (Alternatively ε u_a , ε v_a and ε w_a ) further depends on the size N of the volume. In yet another variant, for a set of volumes representing a point cloud, the adaptive halo parameter ε is locally expressed a (Alternatively ε u_a , ε v_a and ε w_a ).

[0189] Although the adaptive halo method has many advantages, it cannot perform well on all kinds of MPEG point cloud datasets. If used directly in the MPEG G-PCC software, it will lead to a loss of overall compression results in terms of the D2 (point-to-plane distortion) metric, and the overall performance gain in terms of the D1 (point-to-point distortion) metric will not be large (about 2%, less than 5%), so only implementing a single adaptive halo method cannot achieve the overall optimal coding performance on all kinds of MPEG point cloud datasets.

[0190] In particular, in the MPEG G-PCC standard, AR / VR datasets can be divided into 4 categories, namely solid, dense, sparse, and scant. The solid category is a voxelized point cloud with a continuous surface, the dense category is a less continuous voxel point cloud, the sparse category is not dense (sparser than the dense category), and the sparse category is very sparse data. The adaptive halo method of the Trisoup model has been tested on all the above categories of data, and two metrics are used in BDBR to evaluate the quality of the reconstructed point cloud, as mentioned above, one is the D1 (point-to-point distortion) metric, and the other is the D2 (point-to-plane distortion) metric. The detailed experimental results for each category are as follows:

[0191] Solid Class: The adaptive halo method has no impact on the compression efficiency of solid class data in terms of D1 and D2 metrics.

[0192] Dense categories: In terms of the D1 metric, the adaptive halo method works well in improving the compression efficiency of dense category data (in fact, it has a 5% gain). In terms of the D2 metric, the method has little impact on the compression efficiency of dense category data.

[0193] · Sparse and rare categories: The adaptive halo method can slightly improve the compression efficiency of sparse and rare categories in terms of the D1 metric, but it causes losses in terms of the D2 metric.

[0194] Therefore, in Figure 1a the last step S03 of Figure 1b , the adaptive halo method is selectively executed to achieve better overall performance of the encoding. Referring to

[0195] , which shows a simplified flow according to the proposed method. Among them, a flag (such as adaptive_halo_enabled_flag) can be included in the bitstream at the encoder, and the decoder can determine whether to enable the adaptive halo method based on this flag. Preferably, this flag can be included in the geometric parameter set (GPS) of the bitstream. The above GPS contains parameters that specify the features and activation tools used in the encoded geometric information bitstream of the point cloud slice, and the GPS is placed in the slice header of the geometric information stream. For example, if the flag is set to "true", the adaptive halo method is enabled, otherwise the adaptive halo method is not used for trisoup encoding. If the adaptive halo method is not used, the triangle for voxelization can be extended along at least one edge based on a fixed value. As described above, how to set the flag can be based on the category of the point cloud data, which can be evaluated by the sampling distance of the point cloud data. For example, if the sampling distance d of the point cloud satisfies the condition: 1 < d < 4 (i.e., dense), the value of the flag is set to true; otherwise, the value of the flag is set to false.

[0196] See Fig.18 , which shows a schematic flowchart of a method for encoding a 3D point cloud into a bitstream according to the present invention.

[0197] The method includes:

[0198] In step S11, determine the octree information, where the octree information includes the octree structure of the volume, and the volume includes a plurality of cubes;

[0199] In step S12, obtain vertex information from the surface of the point cloud of each cube associated with the leaf node, where the vertex information includes information about the existence and position of vertices on the edges of the cube;

[0200] In step S13, encode the octree information and the vertex information into a bitstream;

[0201] In step S14, the point cloud data is reconstructed by using the octree information and vertex information obtained in the above encoding process, wherein the point cloud data is reconstructed including:

[0202] In step 141, triangles are determined by connecting vertices of a cube associated with a leaf node of the octree structure;

[0203] In step 142, the triangles are voxelized to determine the points of the point cloud; additional information is determined based on the density of the point cloud (which can be determined by the sampling distance d of the point cloud). sampl Evaluation), encoding the additional information into the bit stream, determining whether the additional information satisfies a predefined condition, and when the predefined condition is met, based on the sampling distance d sampl At least one triangle is expanded along at least one edge for voxelization.

[0204] Among them, steps S11 to S13 involve the known Trisoup coding, which is known, for example, from "Adaptive multi-level triangle soup for geometry-based point cloud coding" by A. DRICOT et al. at the 21st International Symposium on Multimedia Signal Processing (MMSP) of the Institute of Electrical and Electronics Engineers (IEEE) in 2019, "report on triangle soup decoding" by Nakagami O., ISO / IEC JTC1 / SC29-WG11 in m52279 in 2020, and US10,192,353. In addition to the usual point cloud encoding, the method also includes a reconstruction step, which includes the same or similar steps as the decoding method previously described with particular reference to Figure 1. The reconstructed point cloud can then be used to interpolate attributes (such as color), and then the attributes of the points of the point cloud are encoded based on the reconstructed geometry.

Claims

1. A method for decoding the geometry of a 3D point cloud from a bitstream, preferably implemented in a decoder, the method comprising: Receive and decode a bitstream, wherein the bitstream includes octree information and vertex information, the octree information includes information about an octree structure of a point cloud volume, and the vertex information includes information about the existence and position of vertices on edges of cubes of leaf nodes of the octree structure; Determining triangles by connecting vertices of a cube associated with a leaf node of the octree structure; voxelizing the triangles to determine points of the point cloud, Characterized in that the method further comprises: Determine whether the additional information included in the bitstream satisfies a predefined condition, wherein the additional information is determined based on the density of the point cloud, preferably by a sampling distance d of the point cloud. sampl Evaluate; When the predefined condition is met, based on the sampling distance d sampl At least one triangle is expanded along at least one edge for voxelization.

2. A method for encoding a 3D point cloud into a bitstream, preferably implemented in an encoder, the method comprising: Obtaining octree information, the octree information including an octree structure of a body, the body including a plurality of cubes; Acquire vertex information from a surface of a point cloud of each cube associated with a leaf node, wherein the vertex information includes information about the existence and position of vertices on an edge of the cube; Encoding the octree information and the vertex information into a bitstream; Reconstructing the geometric data of the point cloud by using the octree information and vertex information obtained in the aforementioned encoding process, wherein reconstructing the geometric data of the point cloud includes: Determining triangles by connecting vertices of a cube associated with a leaf node of the octree structure; voxelizing the triangles to determine points of the point cloud; Characterized in that the method further comprises: Additional information is determined based on the density of the point cloud, preferably by a sampling distance d of the point cloud. sampl Evaluate; encoding the additional information into the bitstream; Determining whether the additional information satisfies a predefined condition; When the predefined condition is met, based on the sampling distance d sampl At least one triangle is expanded along at least one edge for voxelization.

3. The method according to claim 2, wherein: The encoding is Trisoup encoding.

4. The method according to any one of claims 1 to 3, wherein: The at least one triangle is expanded on two or three sides for voxelization.

5. The method according to any one of claims 1 to 4, wherein: Each triangle in a cube is expanded, and preferably at least one triangle in each cube of a point cloud having triangles is expanded.

6. The method according to any one of claims 1 to 5, wherein: The extension is the same for each edge or different for at least two edges.

7. The method according to any one of claims 1 to 6, wherein: For voxelization use algorithm, and / or obtain a voxelization of the point by rounding its coordinates to the nearest integer.

8. The method according to claim 7, wherein: The convex hull is required to be -ε a ≤u,v,w, where ε a > 0, and u, v, w are the barycentric coordinates of the triangle, where ε a Based on the sampling distance d of the point cloud sampl Sure.

9. The method according to claim 7, wherein: The convex hull is required to be -ε u_a ≤u, -ε v_a ≤v and -ε w_a ≤w, where ε u_a , ε v_a , ε w_a > 0, and u, v, w are the barycentric coordinates of the triangle, and ε u_a ≠ε w_a and / or ε u_a ≠ε v_a and / or ε v_a ≠ε w_a , where ε u_a , ε v_a , ε w_a At least one of the sampling distances d of the point cloud is based on sampl Sure.

10. The method according to any one of claims 1 to 9, wherein: The extension is provided by a halo parameter, and the halo parameter of the extension is less than d sampl / 4, and preferably less than d sampl / 8.

11. The method according to any one of claims 1 to 10, wherein: The extension is provided by the adaptive halo parameters and the extension is pre-set.

12. The method according to any one of claims 1 to 10, wherein: The extension is provided by adaptive halo parameters and the adaptive halo parameters are encoded in the bitstream, preferably in a geometry parameter set.

13. The method according to any one of claims 1 to 12, wherein: The sampling distance d of the point cloud sampl Depend on Determine, where N leaf is the number of leaf nodes, N total is the number of points in the point cloud, N is the size of the corresponding cube of the leaf node, or the sampling distance d of the point cloud sampl Determined by round-robin method.

14. The method according to any one of claims 1 to 13, wherein: The at least one triangle is based on a weighted halo parameter ε a_t and is extended along at least one side for voxelization, wherein the weighted halo parameter ε a_t is determined by ε a_t = ε a * t (1 < t < 4), where ε a is an adaptive halo parameter based on the sampling distance d sampl of the point cloud, and provides the extension of the at least one triangle, and t is a corresponding weight associated with the sampling distance, preferably set to 2.

15. The method according to any one of claims 1 to 14, wherein: The additional information is a flag for enabling or disabling the function of the encoding or decoding method, preferably one bit.

16. An encoder for encoding a 3D point cloud into a bitstream, the encoder comprising at least one processor and a memory, wherein: The memory stores instructions which, when executed by the processor, perform the steps of the method according to any one of claims 2 to 15.

17. A decoder for decoding a 3D point cloud from a bitstream, the decoder comprising at least one processor and a memory, wherein: The memory stores instructions which, when executed by the processor, perform the steps of the method according to any one of claim 1 and claims 3 to 15 dependent on claim 1.

18. A bit stream encoded by the method according to any one of claims 2 to 15.

19. A computer-readable storage medium comprising instructions, which, when executed by a processor, perform the steps of the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Method and apparatus for encoding / decoding geometry of point cloud representing 3D object

    CN111417985A

  • Encoding and decoding methods, encoder, decoder and software

    CN112438049A

  • Cloud rendering method and device in virtual environment, equipment and storage medium

    CN112907716A

  • Point cloud data transmission device, point cloud data transmission method, point cloud data reception device and point cloud data reception method

    WO2020262824A1

  • Hybrid-tree coding for inter and intra prediction for geometry coding

    WO2022147015A1