Method and apparatus for coding a geometry of a point cloud, and data stream having encoded therein a geometry of a point cloud

By modifying vertex positions near corners and calculating centroid drifts, the method addresses Trisoup coding artifacts, enhancing point cloud reconstruction quality without increased complexity.

WO2025152086A1PCT designated stage expired Publication Date: 2025-07-24BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/072875
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing point cloud compression methods, particularly Trisoup coding, produce undesired folds and artifacts in reconstructed surfaces due to non-optimal vertex positions near corners of leaf nodes, leading to suboptimal visual quality without significant computing complexity reduction.

Method used

Modify the positions of vertices near corners to a common position, such as the corner itself, and calculate a centroid drift based on these modified positions, encoding this drift for improved vertex positioning during encoding and decoding.

Benefits of technology

Reduces or removes undesired folds and artifacts in reconstructed point clouds without substantial computing overhead or bitstream modifications, maintaining visual quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024072875_24072025_PF_FP_ABST
    Figure CN2024072875_24072025_PF_FP_ABST
Patent Text Reader

Abstract

A method of encoding into a bitstream a geometry of a point cloud is described. The point cloud is represented by a plurality of cuboid volumes. The plurality of cuboid volumes includes at least one occupied cuboid volume being modelled by one or more triangles. The one or more triangles have vertices on edges of the occupied cuboid volume. The method includes encoding positions of vertices located on edges of the occupied cuboid volumes, and encoding a centroid drift per occupied cuboid volume. Encoding the centroid drift for at least one occupied cuboid volume includes modifying the position of at least one vertex on an edge of the occupied cuboid volume, and calculating the centroid drift, which is to be encoded, using the modified position of the at least one vertex.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR CODING A GEOMETRY OF A POINT CLOUD, AND DATA STREAM HAVING ENCODED THEREIN A GEOMETRY OF A POINT CLOUDTECHNICAL FIELD

[0001] The present invention generally relates to data compression, more specifically to methods and apparatus for coding a geometry of a point cloud. Embodiments of the present invention concern methods and apparatus for encoding / decoding a geometry of a point cloud using a centroid drift for an occupied cuboid volume which is calculated using a modified position of at least one vertex of the occupied cuboid volume, and a data or bitstream having encoded therein the centroid drift.BACKGROUND

[0002] Data compression is used in communications and computer networking to store, transmit, and reproduce information efficiently. For example, as a format for the representation of three-dimensional (3D) data, point clouds have recently gained attraction as they are versatile in their capability in representing all types of 3D objects or scenes. Therefore, many use cases can be addressed by point clouds, among which are

[0003] · movie post-production,

[0004] · real-time 3D immersive telepresence or VR  / AR (virtual reality  / augmented reality) applications,

[0005] · free viewpoint video, e.g., for sports viewing,

[0006] · geographical information systems, also known as cartography,

[0007] · culture heritage, e.g., the storage of scans of rare objects into a digital form,

[0008] · autonomous driving, including 3D mapping of the environment and real-time LiDAR data acquisition (LiDAR: Light Detection And Ranging = a method for measuring distances (ranging) by illuminating the target with laser light and measuring the reflection with a sensor) .

[0009] A point cloud is a set of points in a three-dimensional coordinate system. The points are often intended to represent an external surface of one or more objects. Each  point has a location or position in the three-dimensional coordinate system. The position may be represented by three coordinates (X, Y, Z) , which can be Cartesian or any other coordinate system. The points may have other associated attributes, such as color, which may also be a three-component value in some cases, such as R, G, B or Y, Cb, Cr. Other associated attributes may include transparency, reflectance, a normal vector, etc., depending on the desired application for the point cloud data.

[0010] Point clouds can be static or dynamic. For example, a detailed scan or mapping of an object or topography may be static point cloud data. The LiDAR-based scanning of an environment for machine-vision purposes may be dynamic in that the point cloud, at least potentially, changes over time, e.g., with each successive scan of a volume. The dynamic point cloud is therefore a time-ordered sequence of point clouds.

[0011] As mentioned above, point cloud data may be used in a number of applications or use cases, including conservation, like scanning of historical or cultural objects, mapping, machine vision, e.g., for autonomous or semi-autonomous cars, and virtual or augmented reality systems. Dynamic point cloud data for applications, like machine vision, can be quite different from static point cloud data, like that for conservation purposes. Automotive vision, for example, typically involves relatively small resolution, non-colored, highly dynamic point clouds obtained through LiDAR or similar sensors with a high frequency of capture. The objective of such point clouds is not for human consumption or viewing but rather for machine object detection / classification in a decision process. As an example, typical LiDAR frames contain in the order of tens of thousands of points, whereas high quality virtual reality applications require several millions of points. It may be expected that there is a demand for higher resolution data over time as computational speed increases and new applications or use cases are found.

[0012] Stated differently, a point cloud is a set of points located in a 3D space, optionally with additional values attached to each of the points. These additional values are usually called point attributes. Consequently, a point cloud may be considered a combination of a geometry (the 3D position of each point) and attributes. Attributes may be, for example, three-component colours, material properties, like reflectance, and / or two-component normal vectors to a surface associated with the point. Point clouds may be  captured by various types of devices like an array of cameras, depth sensors, the mentioned LiDARs, scanners, or they may be computer-generated, e.g., in movie post-production use cases. Depending on the use cases, points clouds may have from thousands to up to billions of points for cartography applications.

[0013] Raw representations of point clouds require a very high number of bits per point, with at least a dozen of bits per spatial component X, Y or Z, and optionally more bits for the one or more attributes, for instance three times 10 bits for the colours. Therefore, a practical deployment of point-cloud-based applications or use cases requires compression technologies that enable the storage and distribution of point clouds with reasonable storage and transmission infrastructures. In other words, while point cloud data is useful, a lack of effective and efficient compression, i.e., encoding and decoding processes, may hamper adoption and deployment. A particular challenge in coding point clouds that does not arise in the case of other data compression, like audio or video, is the coding of the geometry of the point cloud, and the tendency of point clouds to be sparsely populated makes efficiently coding the location of the points much more challenging.

[0014] Until recently, point cloud compression, also referred to as PCC, was not addressed by the mass market and there was no standardized point cloud codec available. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11, also known as Moving Picture Experts Group or MPEG, initiated work items on point cloud compression which have led to two standards, namely:

[0015] · MPEG-I Part 5 (ISO / IEC 23090-5) also referred to as Video-based Point Cloud Compression, V-PCC,

[0016] · MPEG-I Part 9 (ISO / IEC 23090-9) also referred to as Geometry-based Point Cloud Compression, G-PCC.

[0017] The first versions of the V-PCC standard and the G-PCC standard were finalized respectively in 2020 and 2022.

[0018] The V-PCC coding method compresses a point cloud by performing multiple projections of a 3D object to obtain two-dimensional (2D) patches that are packed into an image or into a video when dealing with moving point clouds. The images or videos are then compressed using existing image / video codecs, allowing for the leverage of  already deployed image and video solutions. By its very nature, V-PCC is efficient only on dense and continuous point clouds because image / video codecs are unable to compress non-smooth patches in case they are obtained from the projection of, for example, LiDAR acquired sparse geometry data.

[0019] The G-PCC coding method has two schemes for the compression of the geometry.

[0020] · The first scheme is based on an occupancy tree representation of the point cloud geometry, for example, by means of an octree representation, a quad tree representation or a binary tree representation. In a tree-based structure, the bounding three-dimensional volume for the point cloud is recursively divided into sub-volumes. Nodes of the tree correspond to sub-volumes. The decision of whether or not to further divide a sub-volume may be based on a resolution of the tree and / or whether there are any points contained in the sub-volume. A leaf node may have an occupancy flag that indicates whether its associated sub-volume contains a point or not. Splitting flags may signal whether a node has child nodes, i.e., whether a current volume has been further split into sub-volumes. A commonly used tree structure is an octree. In this structure, the volumes / sub-volumes are all cuboids and each split of a sub-volume results in eight further sub-volumes / sub-cuboids. Another commonly used tree structure is a KD-tree, in which a volume, like a cuboid, is recursively divided in two by a plane orthogonal to one of the axes. Octrees are a special case of KD-trees, where the volume is divided by three planes, each being orthogonal to one of the three axes.

[0021] In other words, occupied nodes are split down until a certain size is reached, and occupied leaf modes provide the location of points, typically at the center of these nodes. By using neighbor-based prediction techniques, a high level of compression may be obtained for dense point clouds. Sparse point clouds are also addressed by directly coding the position of a point within a node with a non-minimal size, by stopping the tree construction when only isolated points are present in a node. This stopping technique is also referred to as a direct coding mode (DCM) .

[0022] · The second scheme is based on a predictive tree, in which each node represents the 3D location of one point and the relation between nodes is a spatial prediction from the parent node to the child nodes. This method may only address sparse point clouds and offers the advantage of a lower latency and a simpler decoding when compared to using an occupancy tree. However, the compression performance is slightly better while, when compared to the first scheme, the encoding is complex due to the need to intensively look for a best predictor among a long list of potential predictors when constructing the predictive tree.

[0023] In both schemes attribute coding, i.e., attribute encoding and attribute decoding, is performed after coding the complete geometry which, in turn, leads to a two-pass coding process. A low latency may be obtained by using slices that decompose the 3D space into sub-volumes that are coded independently, without prediction between the sub-volumes. However, this may heavily impact the compression performance when many slices are used.

[0024] One use case of specific interest is the transmission of dynamic AR / VR point clouds, wherein dynamic means that the point cloud evolves over time. Also, AR / VR point clouds are typically locally 2D as, most of the time, they represent the surface of an object. As such, AR / VR point clouds are highly connected, also referred to as being dense, in the sense that a point is rarely isolated and, instead, has many neighbors. Thus, dense or solid point clouds represent continuous surfaces with a resolution such that volumes, also referred to as small cubes or voxels, associated with points touch each other without exhibiting any visual hole in the surface. Such point clouds, as mentioned above, are typically used in AR / VR environments and may be viewed by an end user through a device, like a TV, a smart phone or a headset including AR / VR glasses. The point clouds may be transmitted to the device or may be stored locally. Many AR / VR applications make use of moving point clouds which, as opposed to static point clouds, vary with time. Therefore, the volume of data may be huge and needs to be compressed. For example, when applying the above-mentioned octree representation of the geometry of a point cloud, a lossless compression may be achieved down to slightly less than 1 bit per point (or 1 bpp) . However, this may not be sufficient for real time transmissions that may involve several  millions of points per frame with a frame rate as high as 50 frames per second leading, in turn, to hundreds of megabytes of data per second.

[0025] Consequently, a lossy compression scheme may be used with the usual requirement of maintaining an acceptable visual quality by providing for a compression that is sufficient to fit the compressed data within a bandwidth available in the transmission channel while, at the same time, maintaining a real time transmission of the frames. In many applications, bit rates as low as 0.1 bpp may already allow for a real time transmission, meaning that by means of the lossy compression the point cloud is compressed ten times more than when applying a lossless coding scheme.

[0026] The codec based on MPEG-I part 5 (ISO / IEC 23090-5) or V-PCC may achieve such low bitrates by using the lossy compression of video codecs that compress 2D frames obtained from the projection of the point cloud on the planes. The geometry is represented by a series of projection patches assembled into a frame with each patch being a small local depth map. However, V-PCC is not versatile and is limited to a narrow type of point clouds that do not exhibit a locally complex geometry, like trees or hair or the like, because the obtained projected depth map may not be smooth enough to be efficiently compressed by video codecs.

[0027] On the other hand, pure 3D compression techniques may handle any type of point clouds. For example, G-PCC may provide in the future a lossy compression that also allows compressing dense point clouds as good as V-PCC intra while maintaining the versatility of G-PCC so as to handle any type of point clouds, like dense point clouds, point clouds obtained by LiDAR or point clouds representing 3D maps. For implementing such a G-PCC mechanism, the so-called Trisoup coding scheme may be applied over a first layer based on an octree. Currently, the Trisoup coding scheme is discussed in the standardization working group JTC1 / SC29 / WG7 of ISO / IEC. When considering the possibilities for obtaining a lossy scheme from G-PCC, there are basically three approaches for obtaining a lossy scheme over the octree representation as used by the 3-PCC codec, namely

[0028] · down-sampling + (lossless) coding + re-up-sampling

[0029] · modifying the voxels locally on the encoder side

[0030] · modelling the point cloud locally.

[0031] The first approach basically comprises down-sampling the entire point cloud to a smaller resolution, lossless coding of the down-sampled point cloud, and then up-sampling after decoding. There are many up-sampling schemes, e.g., super resolution, artificial intelligence, AI or learning-based 3D post-processing and the like, which may provide for good peak signal-to-noise ratio, PSNR, results when the down-sampling is not too aggressive, for example not more than a factor of two in each direction. However, even if the metrics show a good PSNR, the visual quality is still disputable and not well controlled.

[0032] The second approach allows the encoder to adjust the point cloud locally such that the coding of the octree requires a lesser bitrate. For this purpose, the points may be slightly moved so as to obtain occupancy information that may be better predicted by neighboring nodes, thereby leading to a lossless encoding of a modified octree with a lower bitrate. However, this approach, unfortunately, only leads to a small bitrate reduction.

[0033] The third approach is to code the geometry using a tree, like an octree, down to a certain resolution, for example down to N×N×N blocks, where N may be 4, 8 or 16, for example. This tree is then coded using a lossless scheme, like the G-PCC scheme. The tree itself does not require a high bitrate because it does not go down to the deepest depth and has only a small number of leaf nodes when compared to the number of points in the point cloud. Then, in each N×N×N block the point cloud is modelled by a local model. Such a model may be a mean plane or a set of triangles as in the above-mentioned Trisoup coding scheme which is described now in more detail.

[0034] The Trisoup coding scheme models a point cloud locally by using a set of triangles without explicitly providing connectivity information -that is why its name is derived from the term “soup of triangles” . As mentioned above, each N×N×N block defines a volume associated with a leaf node, and in each N×N×N block or volume the point cloud is modeled locally using a set of triangles wherein vertices of the triangles are coded along the edges of the volume associated with the leaf nodes of the tree. Fig. 1 illustrates a volume 100 associated with a leaf node which is a cuboid volume designed by twelve edges 1001 to 10012.

[0035] The part of the point cloud encompassed by the volume 100 is modeled by at least one triangle having at least one vertex on one of the edges 1001 to 10012. In the example of Fig. 1, five vertices 1 to 4 are illustrated among which vertices 1 to 4 are located on the edges 1002, 1001, 1008 and 1007 , respectively.

[0036] The vertices located on the edges are shared among those leaf nodes that have a common edge which means that at most one vertex is coded per edge that belongs to at least one leaf node, and by doing so the continuity of the model is ensured through the leaf nodes. The coding of the Trisoup vertices requires two information per edge:

[0037] · a vertex flag indicating if a Trisoup vertex is present on the edge, also referred to herein as the presence flag, and

[0038] · in case the vertex is present, the vertex position along the edge.

[0039] Consequently, the coded data comprises the octree data plus the Trisoup data. For example, the vertex flag may be coded by an adaptive binary arithmetic coder that uses one specific context for coding vertex flags, while the position of the vertex on the edge having a length N=2s is coded with unitary precision by pushing s bits into the bitstream, i.e., by bypassing / not entropy coding the s bits.

[0040] Fig. 2 illustrates a volume 100 associated with a leaf node including two Trisoup triangles 102, 104 having their respective vertices 1 to 4 on the edges (see Fig. 1) 1002, 1001, 1008 and 1004, respectively, of the volume 100. Triangle 102 comprises the vertices 1, 2 and 3, while triangle 104 comprises the vertices 1, 3 and 4. Thus, triangles may be constructed in case at least three vertices are present on the edges of the volume 100. Naturally, any other combination of triangles than those shown in Fig. 2 is possible inside the volume 100 associated with a leaf node. Also, the one or more triangles inside the volume 100 do not have necessarily all of their vertices on the edges of the volume 100, rather, one or two of the vertices of a triangle may be located anywhere inside the volume 100.

[0041] The triangles to be constructed inside the volume 100 is based on the following three-step process including:

[0042] 1. Determining a dominant direction along one of the three axes.

[0043] 2. Ordering the Trisoup vertices dependent on the dominant direction.

[0044] 3. Constructing the triangles based on the ordered list of vertices.

[0045] Fig. 3 illustrates the process for choosing triangles to be constructed inside the volume 100 associated with a leaf node of Fig. 2 which is illustrated again in Fig. 3 (a) without the triangles. Fig. 3 (b) and Fig. 3 (c) illustrate the process over two axes, namely the vertical or z-axis (Fig. 3 (b) ) and the horizontal axis or x-axis (Fig. 3 (c) ) .

[0046] The first test along the vertical axis, i.e., from the top, is performed by projecting the volume or cube 100 and the Trisoup vertices vertically onto a 2D plane as is illustrated in Fig. 3 (c) . The vertices are then ordered following a clockwise order relative to the center of the projected node 114 which, in the illustrated example, is a square. The triangles are constructed following a fixed rule based on the ordered vertices, and in the example of Fig. 3 four vertices are involved and the triangles 102, 104 are constructed systematically to include the vertices 1, 2 and 3 for the first triangle, and vertices 1, 3 and 4 for the second triangle, as illustrated in Fig. 3 (c) . In case only three vertices are present, the only possible triangle is a triangle including vertices 1, 2 and 3, and in case five vertices are present, a fixed rule may be used to construct triangles including the vertices (1, 2, 3) , (1, 3, 5) and (4, 5, 1) , and so on. This may be repeated up to 12 vertices.

[0047] A second test along the horizontal axis is performed by projecting the cube 100 and the Trisoup vertices horizontally on a 2D plane when looking from the left of Fig. 3 (a) yielding the projection 116 illustrated in Fig. 3 (b) . When ordering the vertices following the clockwise order relative to the center of the projected node 100, the triangles 102, 104 include the vertices 1, 2 and 3 for the first triangle, and vertices 1, 3 and 4 for the second triangle, as illustrated in Fig. 3 (b) .

[0048] As may be seen from Fig. 3, the vertical projection (Fig. 3 (c) ) exhibits a 2D total surface of triangles that is the maximum so that the dominant axis is selected to be the vertical or z axis, and the Trisoup triangles to be constructed are obtained from the order of the vertical projection as illustrated in Fig. 3 (c) , which, in turn, yields triangles inside the volume as depicted in Fig. 2. It is noted that when considering the horizontal axis as the dominant axis, this leads to a different construction of the triangles within the volume 100 as depicted in Fig. 4 illustrating the volume 100 in which the triangles 102, 104 are  constructed when assuming the dominant axis to be the horizontal axis and in order of the vertices as illustrated in Fig. 3 (b) .

[0049] The adequate selection of the dominant axis by maximizing the projected surface leads to a continuous reconstruction of the point cloud without holes.

[0050] In addition, one centroid vertex per volume or leaf node 100 is coded so as to characterize a surface curvature within each volume 100. Fig. 5 illustrates a Trisoup geometry representation of the volume 100 comprising four vertices V1 to V4. The centroid vertex C is encoded as a drift value of the gravity center Cmean of all vertices V1 to V4. The drift value may be coded as the centroid drift Cres along the normal vector of the surface. Fig. 5 illustrates the vector indicating the normal of the triangle surfaces. Optionally, also face vertices illustrated in Fig. 5 by the plurality of dots may be created and signaled.

[0051] The rendering of the Trisoup triangles is performed by ray tracing, and the set of all rendered points by ray tracing results in the decoded point cloud. Fig. 6 illustrates the ray tracing to render the Trisoup triangle 102 of Fig. 2 including the vertices 1, 2 and 3. Rays, like ray 118 in Fig. 6, are launched along directions parallel to an axis, like the z axis in Fig. 6. The origin of the rays is a point of integer, voxelized, coordinates of precision corresponding to the sampling position desired for the rendering. The intersection 120 of the ray 118 with triangle 102 is then voxelized, i.e., is rounded to the closest point at the desired sampling position, and is added to the list of rendered points. After applying the Trisoup coding scheme to all leaf nodes, i.e., after constructing the triangles and obtaining the intersections by ray tracing, copies of the same points in the list of all rendered points are discarded, i.e., only one voxel is kept among all voxels sharing the same 3D position, thereby obtaining a set of decoded, unique points.

[0052] In the above-described approach for Trisoup coding, the vertex determination along an edge within a leaf node depends on the points whose distance to the edge of the node is within a predefined distance, as is described, for example, in WO 2023 / 197122 A1. However, it has been found that, if the surface passes very close to one corner of the leaf node, significant visual artifacts may be produced when reconstructing the point cloud. This is illustrated in Fig. 7 showing in Fig. 7 (a) a leaf node  100 and the original points or surface 130 within the volume 100 for which the vertices 1 and 2 are created in the above-described way. As may be seen from Fig. 7 (a) , as is indicated at 132, the points or surface 130 passes very close to the corner 134 of the volume 100. For points in the region 132 of the original point cloud 130 within the leaf node 100, i.e., for points close to the corner 134, the above-described Trisoup coding approach may actually generate more than one vertex along more than one edge, for example instead of generating a single vertex 3, on the edges around the corner 134 actually three vertices V1, V2, and V3 are generated. In the depicted examples, the respective vertices are generated on each of the edges meeting at corner 134, however, it is noted that the plurality of vertices may also be generated on only one or two of the edges intersecting the corner 134.

[0053] When using the vertices illustrated in Fig. 7 (a) , namely vertices 1, 2, V1, V2 and V3, to construct triangles together with the centroid vertex C (Vc) so as to obtain a reconstructed surface within the node, an undesired fold is produced in the reconstructed surface although there are no folded areas in the original point cloud. Fig. 7 (b) illustrates the points reconstructed for the volume 100 of Fig. 7 (a) using the vertices 1, 2, V1, V2, and V3. As may be seen, for the vertices 1 and 2 located at a certain distance from the respective corners of the volume 100, the point cloud is correctly reconstructed, as is illustrated at 130 in Fig. 7 (b) . However, the three vertices V1 to V3 generated for the region 132 of the point cloud in Fig. 7 (a) which is close to the corner 134 of the volume 100 lead to the undesired fold in the reconstructed surface, as is illustrated in Fig. 7 (b) at 136. This, as mentioned above, leads to significant visual artifacts in the reconstructed point cloud, as may be seen from Fig. 7 (b) .SUMMARY

[0054] Accordingly, it is an object of the present invention to provide for methods and apparatus that allow for removing or reducing undesired folds or artefacts in a reconstructed surface of a Trisoup coding generated due to non-optimal vertices positions.

[0055] The present invention provides a method of encoding into a bitstream a geometry of a point cloud, the point cloud being represented by a plurality of cuboid  volumes, the plurality of cuboid volumes comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the method comprising: encoding positions of vertices located on edges of the occupied cuboid volumes, and encoding a centroid drift per occupied cuboid volume,

[0056] wherein encoding the centroid drift for at least one occupied cuboid volume comprises:

[0057] - modifying the position of at least one vertex on an edge of the occupied cuboid volume, and

[0058] - calculating the centroid drift, which is to be encoded, using the modified position of the at least one vertex.

[0059] The present invention provides a method of decoding from a bitstream a geometry of a point cloud, the point cloud being represented by a plurality of cuboid volumes, the plurality of cuboid volumes comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the method comprising:

[0060] decoding positions of vertices located on edges of the occupied cuboid volumes, and

[0061] decoding a centroid drift per occupied cuboid volume,

[0062] wherein for at least one occupied cuboid volume the method comprises:

[0063] - modifying the position of at least one vertex on an edge of the occupied cuboid volume,

[0064] - calculating an initial position of a centroid vertex of the at least one occupied cuboid volume using the modified position of the at least one vertex, and

[0065] - adapting the initial position of the centroid vertex using the decoded centroid drift.

[0066] Optionally, the method further comprises:

[0067] - constructing for the at least one occupied cuboid volume at least one triangle using the decoded positions of the vertices and the adapted initial position of the centroid vertex, and

[0068] - reconstructing the point cloud using ray tracing on the constructed triangle in the occupied cuboid volume.

[0069] Optionally, a position of the at least one vertex is modified when one or  more of the following conditions are met:

[0070] - a distance between the at least one vertex and a corner of the occupied cuboid volume is less than or equal to a distance threshold,

[0071] - a total number of vertices in the occupied cuboid volume is greater than a first minimum number,

[0072] - in the occupied cuboid volume, a total number of vertices, which have a distance to the corner being less than or equal to the distance threshold, is greater than a second minimum number.

[0073] Optionally, one or more of the distance threshold, the first minimum number and the second minimum number is signaled in the bitstream or is a fixed value known at both the encoding side and the decoding side.

[0074] Optionally, the distance threshold is determined using a width of the occupied cuboid volume and a scaling coefficient.

[0075] Optionally, the distance threshold is determined as follows: th= ω*blockWidth

[0076] where:

[0077] th = the distance threshold,

[0078] ω = the scaling coefficient ranging within (0, 1) ,

[0079] blockWidth = a width of the occupied cuboid volume, and

[0080] * = indicates a multiplication operation or a k-bits right shift operation.

[0081] Optionally, the scaling coefficient and / or k is signaled in the bitstream or is a fixed value known at both the encoding side and the decoding side.

[0082] Optionally, the distance threshold is a number of vertex sampling positions away from the corner.

[0083] Optionally, the bitstream has encoded therein a geometry parameter set, GPS, the GPS including one or more of the following:

[0084] - the distance threshold, the first minimum number and the second minimum number,

[0085] - the scaling coefficient and / or k,

[0086] - the number of vertex sampling positions.

[0087] Optionally, modifying the position of the at least one vertex comprises modifying the position of the at least one vertex to match a position of a corner of the occupied cuboid volume.

[0088] The present invention provides a non-transitory computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the inventive method.

[0089] The present invention provides an apparatus for encoding into a bitstream a geometry of a point cloud, the point cloud being represented by a plurality of cuboid volumes, the plurality of cuboid volumes comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the apparatus comprising:

[0090] an encoding module configured to encode positions of vertices located on edges of the occupied cuboid volumes and a centroid drift per occupied cuboid volume,

[0091] wherein the encoding module is configured to encode the centroid drift for at least one occupied cuboid volume by:

[0092] - modifying the position of at least one vertex on an edge of the occupied cuboid volume, and

[0093] - calculating the centroid drift, which is to be encoded, using the modified position of the at least one vertex.

[0094] The present invention provides an apparatus for decoding from a bitstream a geometry of a point cloud, the point cloud being represented by a plurality of cuboid volumes, the plurality of cuboid volumes comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the apparatus comprising:

[0095] a decoding module configured to decode positions of vertices located on edges of the occupied cuboid volumes and a centroid drift per occupied cuboid volume,

[0096] wherein for at least one occupied cuboid volume the decoding module configured to:

[0097] - modify the position of at least one vertex on an edge of the occupied cuboid volume,

[0098] - calculate an initial position of a centroid vertex of the at least one occupied cuboid volume using the modified position of the at least one vertex, and

[0099] - adapt the initial position of the centroid vertex using the decoded centroid drift.

[0100] The present invention provides a data stream having encoded thereinto a geometry of a point cloud, the point cloud being represented by a plurality of cuboid volumes, the plurality of cuboid volumes comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the data stream comprising:

[0101] encoded positions of vertices located on edges of the occupied cuboid volumes, and

[0102] an encoded centroid drift per occupied cuboid volume, the encoded centroid drift for at least one occupied cuboid volume being based on a modified position of at least one vertex of the at least one occupied cuboid volume.

[0103] The technical solutions provided according to embodiments of the present invention have the following beneficial effects. In accordance with the inventive approach, artifacts in a reconstructed point cloud are removed or reduced without adding substantial computing complexity and overhead at the encoder / decoder and without or only minor changes to the bitstream.

[0104] It should be understood that the content described in this section is not intended to identify key or critical features of embodiments of the present invention, nor is intended to limit the scope of the present invention. Other features of the present invention will become readily appreciated from the following descriptions.BRIEF DESCRIPTION OF THE DRAWINGS

[0105] The drawings are explanatory and serve to explain the present invention, and are not construed to limit the present invention to the illustrated embodiments.

[0106] Fig. 1 illustrates Trisoup vertices along edges of a volume associated with a leaf node of an octree representation of a point cloud geometry;

[0107] Fig. 2 illustrates a volume including two Trisoup triangles having their respective vertices on the edges of the volume;

[0108] Fig. 3 illustrates a process for choosing triangles to be constructed inside a leaf node, wherein Fig. 3 (a) illustrates the volume of Fig. 2 without triangles, Fig. 3 (b) illustrates a 2D surface of the triangles using a vertical projection of the volume, and  Fig. 3 (c) illustrates a 2D surface of the triangles using a vertical horizontal projection of the volume;

[0109] Fig. 4 illustrates the volume of Fig. 2 with two Trisoup triangles constructed under the assumption of the horizontal axis being the dominant axis;

[0110] Fig. 5 illustrates of a Trisoup geometry representation including a centroid vertex to be coded as a drift value of a gravity center of all vertices.;

[0111] Fig. 6 illustrates the ray tracing to render a Trisoup triangle as a decoded point cloud;

[0112] Fig. 7 illustrates in Fig. 7 (a) original points of a point cloud and generated vertices for one Trisoup node, and in Fig. 7 (b) reconstructed points using those generated vertices;

[0113] Fig. 8 illustrates a flow diagram of a method of encoding a geometry of a point cloud into a bitstream in accordance with embodiments of the present invention;

[0114] Fig. 9 illustrates a flow diagram of a method of decoding a geometry of a point cloud from a bitstream in accordance with embodiments of the present invention;

[0115] Fig. 10 illustrates a data stream in accordance with embodiments of the present invention;

[0116] Fig. 11 illustrates vertices on the six edges intersecting at one node corner of a leaf node;

[0117] Fig. 12 illustrates the modification of the positions of vertices from Fig. 11 satisfying conditions to be moved from their original positions to a corner of the leaf node in accordance with embodiments of the present invention;

[0118] Fig. 13 illustrates an embodiment for encoding geometry of a point cloud into a bitstream;

[0119] Fig. 14 illustrates an embodiment for decoding a bitstream generated by an encoder using a process as described with reference to Fig. 13;

[0120] Fig. 15 illustrates a block diagram of an apparatus / encoder for encoding into a bitstream a geometry of a point cloud in accordance with embodiments;

[0121] Fig. 16 illustrates a block diagram of an apparatus / decoder for decoding from a bitstream a geometry of a point cloud in accordance with embodiments;

[0122] Fig. 17 illustrates a block diagram illustrating an electronic device configured to implement an image processing method according to embodiments of the present disclosure.DETAILED DESCRIPTION

[0123] Illustrative embodiments of the present invention are described below with reference to the drawings, where various details of the embodiments of the present invention are included to facilitate understanding and should be considered as illustrative only. Accordingly, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope of the present invention. Also, descriptions of well-known functions and constructions are omitted from the following description for clarity and conciseness.

[0124] In the present invention, the terms "node" , "volume" , "sub-volume" and “occupied cubic volume” may be used interchangeably. It will be appreciated that a node is associated with a volume or sub-volume. The node is a particular point on the tree that may be an internal node or a leaf node. The volume or sub-volume is the bounded physical space that the node represents. The term "volume" may, in some cases, be used to refer to the largest bounded space defined for containing the point cloud. A volume may be recursively divided into sub-volumes for the purpose of building out a tree-structure of interconnected nodes for coding the point cloud data. A “occupied cubic volume” is a volume including one or more triangles that may be rendered.

[0125] In the present invention, the term "and / or" is intended to cover all possible combinations and sub-combinations of the listed elements, including any one of the listed elements alone, any sub-combination, or all of the elements, and without necessarily excluding additional elements.

[0126] In the present invention, the phrase "at least one of... or... " is intended to cover any one or more of the listed elements, including any one of the listed elements alone, any sub-combination, or all of the elements, without necessarily excluding any additional elements, and without necessarily requiring all of the elements.

[0127] In the present invention, the term “coding” refers to "encoding” or to “decoding” as becomes apparent from the context of the described embodiments concerning the coding of the geometrical information into / from a bitstream. Likewise, the term “coder” refers to " an encoder” or to “adecoder” .

[0128] When modeling a point cloud by applying the Trisoup coding scheme in a way as described above using a set of triangles for each leaf node or volume, Trisoup data is provided. The Trisoup data includes, for example, the information about the vertices of the respective triangles for a volume, also referred to as an occupied cuboid volume or leaf node. However, in the prior art, as outlined above, a point in the original point cloud which is close to a corner of the leaf node may cause the generation of more than one vertex along more than one edges and when using these vertices to construct triangles together with a centroid vertex to get a reconstructed surface within the leaf node, an undesired fold is produced in the reconstructed surface. The present invention is based on the finding that

[0129] The present invention is based on the inventors’ finding that undesired artifacts in a reconstructed point cloud may be removed or substantially reduced by modifying the position of one or more vertices, which are close to a corner, into a common position, for example the position of a corner of the volume or leaf node, and then using the vertices for calculating the centroid drift value which is encoded into the bitstream. At the decoder side, the decoded centroid drift obtained in accordance with the inventive approach is used for adapting an initial position of the centroid vertex which is determined at the decoder side also on the basis of vertices having their position modified so that multiple vertices created close to a corner of the leaf volume are merged into a common position, like a corner position of the volume. This process removes or reduces undesired folds or artifacts in the reconstructed point cloud or surface without adding significant calculating complexity at the decoder side and at the encoder side. At the encoder, the original vertices, like vertices 1, 2, V1, V2, V3 (see Fig. 7 (a) ) are encoded into the bitstream, as well as the centroid drift value which, however, is obtained by simply replacing vertices V1 to V3 by a single vertex located, for example, at the position of the corner. Thus, only simple shift operations are needed at the encoder side while the process for determining the centroid drift need not to be modified. Thus, no substantial processing complexity and overhead is  created by the inventive approach at the encoder side. The same is true for the decoder side which performs the conventional steps for decoding the vertices and the centroid drift from the bitstream. Only simple shift operations are needed at the decoder side for modifying the position of the one or more vertices into a common position on the basis of which, as is also done in conventional decoders, an initial position of the centroid vertex is determined or calculated which is then adapted by the decoded centroid drift. Thus, also at the decoder side, substantially no additional processing complexity or overhead is required.

[0130] A further advantage of embodiments of the present invention is that when hardcoding the required parameters for determining whether vertex positions need to be shifted or not, there is no change in the bitstream which still includes the encoded positions and the centroid drift value, which, however, is now determined in accordance with the inventive approach. In accordance with other embodiments, only minor additional information is to be included into the bitstream for signaling certain parameters needed for the inventive approach from the encoder side to the decoder side which, however, does not significantly add to the complexity of the bitstream.

[0131] Thus, in accordance with the inventive approach, the above-described artifacts in a reconstructed point cloud are avoided without adding substantial computing complexity and overhead at the encoder / decoder and without or only minor changes in the bitstream.

[0132] Fig. 8 illustrates a flow diagram of a method of encoding into a bitstream a geometry of a point cloud. The point cloud is represented by a plurality of cuboid volumes. The plurality of cuboid volumes includes at least one occupied cuboid volume being modelled by one or more triangles. The one or more triangles have vertices on edges of the occupied cuboid volume. In accordance with embodiments, as depicted in Fig. 8, the method includes the following steps:

[0133] S100: Encoding positions of vertices located on edges of the occupied cuboid volumes.

[0134] S102: Encoding a centroid drift per occupied cuboid volume.

[0135] Encoding the centroid drift for at least one occupied cuboid volume comprises:

[0136] · S102a: Modifying the position of at least one vertex on an edge of the occupied cuboid volume.

[0137] · S102b: Calculating the centroid drift, which is to be encoded, using the modified position of the at least one vertex.

[0138] In other words, the centroid drift to be encoded for the occupied cuboid volume is based on a modified position of at least one vertex of the occupied cuboid volume.

[0139] Fig. 9 illustrates a flow diagram of a method of decoding from a bitstream a geometry of a point cloud. The point cloud is represented by a plurality of cuboid volumes. The plurality of cuboid volumes includes at least one occupied cuboid volume being modelled by one or more triangles. The one or more one triangles have vertices on edges of the occupied cuboid volume. In accordance with embodiments, as depicted in Fig. 9, the method includes the following steps:

[0140] S200: Decoding positions of vertices located on edges of the occupied cuboid volumes.

[0141] S202: Decoding a centroid drift per occupied cuboid volume.

[0142] For at least one occupied cuboid volume

[0143] · S204: Modifying the position of at least one vertex on an edge of the occupied cuboid volume.

[0144] · S206: Calculating an initial position of a centroid vertex of the at least one occupied cuboid volume using the modified position of the at least one vertex.

[0145] · S208: Adapting the initial position of the centroid vertex using the decoded centroid drift.

[0146] Fig. 10 illustrates a data stream 300 in accordance with embodiments of the present invention, which has encoded thereinto a geometry of a point cloud. The point cloud is represented by a plurality of cuboid volumes. The plurality of cuboid volumes includes at least one occupied cuboid volume modelled by one or more triangles. The one or more triangles have vertices on edges of the occupied cuboid volume. For example, the data stream or bitstream 300 may be provided by an encoder 400 that performs the inventive method for encoding into a bitstream 300 the geometry of a point cloud. The data stream 300 is transmitted to a decoder 500 via a wired or wireless transmission medium  302, like a cable or a radio link, and the decoder 500 decodes from the data stream 300 the geometry of the point cloud. Thus, in accordance with embodiments, as depicted in Fig. 10, the data stream 300 includes encoded positions 304 of vertices located on edges of the occupied cuboid volumes and an encoded centroid drift per occupied cuboid volume. An encoded centroid drift 306 for at least one occupied cuboid volume is based on a modified position of at least one vertex of the at least one occupied cuboid volume.

[0147] Embodiments of the present invention are now described in more detail.

[0148] As described above, the problems associated with the creation of artifacts in a reconstructed point cloud are addressed by the inventive approach by modifying the position of Trisoup vertices which are close to a corner of a leaf mode to a common position so that only one position for the multiple vertices close to the corner exists. In accordance with embodiments of the present invention, the position to which the Trisoup vertices close to the corner are modified is the position directly at the corner, i.e., the position of the vertices close to the corner is modified in such a way that their position corresponds to the position of the corner. Using the corner position as the new position for the one or more modified vertices may be the most favorable situation in most of the cases because of its central position, however, the present invention is not limited to such embodiments. In accordance with other embodiments the problems associated with the creation of artifacts in a reconstructed point cloud are addressed by modifying the position of the one or more vertices to any common position which is within a distance threshold around the corner.

[0149] In the following, embodiments will be described in accordance with which the positions of vertices close to a corner are modified to match the position of the corner of the leaf node or volume, without limiting the inventive approach to such embodiments. When performing Trisoup coding of a plurality of volumes or leaf nodes, during the Trisoup vertices deriving process, for each volume or leaf node all vertices are determined which are considered to be close enough to a corner of the leaf node.

[0150] Fig. 11 illustrates a leaf node 100 and a plurality of vertices, illustrated by the dots in Fig. 11, on the respective six edges intersecting at the node corner 134. Fig. 11 illustrates also a threshold distance th. For each corner of the node 100, for example for  corner 134 in Fig. 11, the positions of the vertices, after reverse quantization, along all edges that intersect at the corner 134 are checked, more specifically the number of vertices Vn, denoted as NUM, having a distance to the corner 134 being less than or equal to the distance threshold th is determined. In Fig. 11, when applying the threshold th, among the vertices illustrated along the six edges, the vertices 138 are considered to be close enough to the corner 134. Thus, in the example of Fig. 11, the number of vertices close enough to the corner 134 is 6, i.e., NUM = 6.

[0151] In accordance with embodiments, the threshold th may be set as follows: th= ω*blockWidth        Equation 1

[0152] where:

[0153] th = the distance threshold,

[0154] ω = the scaling coefficient ranging within (0, 1) ,

[0155] blockWidth = a width of the occupied cuboid volume, and

[0156] * = indicates a multiplication operation.

[0157] Once the number of vertices close enough to the corner 138 is determined, within each leaf node 100, the positions of the vertices Vn, those close to the corner, are modified to the corresponding position of the corner 138 provided there are enough vertices Vn (close to the corner) in the current leaf node and provided that enough of these vertices are close to the corner 134. These conditions are met when the following applies:

[0158] (1) The number of vertices V within the current node is greater than a minimum value N which may be signaled in the bitstream, for example using the geometry parameter set, GPS, or which may be fixed and known at the encoder and decoder side. In accordance with embodiments, N may be equal to 6 meaning that the current leaf node or volume at least contains half of the possible vertices. In accordance with embodiments, this test is applied to determine whether there are enough vertices in the node. If there are not enough vertices in the node, then the folding behavior is likely to not occur and there is no need to have a process to prevent it. Thus, in accordance with embodiments, this test may be applied to speed up the processing

[0159] (2) The number NUM, i.e., the number of vertices Vn being close to the corner 134, is greater than a minimum value M which, again may be signaled in the bitstream, for example in the GPS, or may be fixed and known at the encoder side and at the decoder side. In accordance with embodiments, M may be equal to 4 meaning that the corner contains at least four vertices to be considered close enough to the corner.

[0160] For example, when considering the embodiment of Fig. 11, one can see that the leaf node 100, i.e., the current node, includes 14 vertices illustrated on the edges intersecting the corner 134, thereby fulfilling condition (1) . Moreover, when considering the surrounding of the corner 134, the number of vertices Vn being within the distance threshold th is 6, i.e., NUM = 6 so that condition (2) is also fulfilled which takes into consideration not only the close vertices Vn from the current node but all close vertices Vn from the surrounding of the corner 134 as is illustrated in Fig. 11.

[0161] This allows, as is illustrated in Fig. 12, to move the positions of vertices V1, V2, V3 from their original positions to the corner position 134.

[0162] On the basis of such modified vertex positions, at the encoder side, the centroid drift is determined which is then encoded into the bitstream. At the decoder side, the initial position of the centroid vertex is determined which is then modified in accordance with the centroid drift signaled in the bitstream.

[0163] Thus, applying the inventive approach improves the visual quality of the reconstructed point cloud, for example, undesired folds or other artifacts are removed or reduced when applying the inventive approach.

[0164] Fig. 13 illustrates an embodiment implementing the inventive approach at the encoder side. For a triangle surface reconstruction using a Trisoup coding method for point cloud data, at the encoder side, initially, as is illustrated at S300, for each Trisoup node or leaf node 100, the vertex presence flag and the quantified vertex position, if the vertex presence flag is true, is determined for each edge.

[0165] At step S302, all unique Trisoup edges are reordered according to a lexicographical order and then, at step S304, the context for the vertex presence flags and the vertex positions is determined using neighboring information. The vertex presence flags and the quantized vertex positions are encoded into a bitstream, for example by using  context-based adaptive binary arithmetic coding, CABC, following the coding order determined in step S302.

[0166] In step S306, the position of one or more vertices is modified so as to correspond to a corner position. More specifically, by applying the above-described approach, eligible vertices are determined and their positions are replaced with the position of the corresponding corner.

[0167] At step S308, an initial position of the centroid of the leaf node is calculated using the vertex coordinates which includes the coordinates of both modified and non-modified vertices. Using the original points of the point cloud around the calculated initial position of the centroid the centroid drift value Cres is calculated, quantized and encoded into the bitstream. Since this value depends on the updated vertices position according to the present invention, this is actually the only thing in the bitstream that is different whether or not the inventive approach is used (and if no signaling is used) . Stated differently, only the actual value for the parameter Cres is different, i.e., implementing the inventive approach in accordance with the described embodiment has the advantage that no modifications of the existing bitstream is required.

[0168] In accordance with embodiments, at step S310, the above-described face vertices (see Fig. 5) may be determined for each leaf node and the presence of the face vertices may be signaled via the bitstream.

[0169] Fig. 14 illustrates an embodiment for decoding a bitstream generated by an encoder using a process as described above with reference to Fig. 13. More specifically, Fig. 14 illustrates an embodiment of a method for a triangle surface reconstruction in a Trisoup decoding method for point cloud data.

[0170] At the step S400, neighbor information for each edge is determined to be later used as context in the subsequent entropy decoding process.

[0171] At step S402, the vertex presence flag and the quantized vertex positions (if the vertex presence flag is true) are entropy decoded for each edge using the neighbor information obtained in step S400.

[0172] In step S404, one or more positions of vertices are modified in accordance with the above-described approach so as to replace positions of eligible vertices with their corresponding corner positions based on the decoded and quantized vertex positions.

[0173] At step S406, the quantized centroid drift values are decoded for each of the nodes from the bitstream, and an initial position of a centroid vertex for a currently processed leaf node is calculated using the vertices or vertex coordinates for both the modified and non-modified vertices. The initial position of the centroid vertex is adapted using the quantized drift value decoded from the bitstream from the currently processed leaf node thereby obtaining a refined centroid vertex C.

[0174] At step S408, information about the presence of face vertices is decoded and, if present, the Trisoup face vertices for the currently processed leaf node are derived.

[0175] At step S410, the vertex coordinates and face vertex coordinates, if present, within each node are sorted and then the triangles are constructed using the edge vertices, the refined centroid vertex and the face vertices. The edge vertices used include those vertices modified in accordance with the inventive approach.

[0176] In step S412, the point cloud is reconstructed by performing ray tracing onto the constructed triangles.

[0177] In the embodiments described so far, reference has been made to a modification of the positions of the vertices which are eligible for a position replacement in accordance with the above-described conditions, to a position of a corner of the vertex. The present invention is not limited to such embodiments, rather, the vertices eligible for a replacement of their position may be modified so as to have a common position, for example when considering Fig. 7 (a) , the vertices V2 and V3 may have their positions modified so as to correspond or match with the position of the vertex V1 or any other vertex position within the distance threshold from the corner. Such an approach may be employed if the corner is not the best choice. For instance, three vertices are close to the corner but two of them are immediately close and one of them is at the limit of being close. Then choosing to replace all of them with a position between the corner and the furthest vertex may offer better results. In accordance with embodiments, if not choosing the corner, some extra processing may be necessary to find the best position.

[0178] Further, in the above-described embodiments, to find the vertices Vn whose distance to the corner is less than or equal to a distance threshold th, Equation 1included a multiplication. However, the present invention is not limited to such embodiments. In accordance with other embodiments, the multiplication in Equation 1 may be replaced by a k-bits right shift operation, or by a simple integer threshold T which indicates a maximum number of vertex sampling positions away from the corner a vertex is considered to be close to the corner. For example, if an edge is sampled into eight possible vertex positions, all vertices one and two sampling steps away from the corner may be considered to be close.

[0179] In accordance with further embodiments, one or more or all of the above-described parameters ω, k and T may be signaled in the bitstream, for example using the GPS. In accordance with other embodiments, one or more or all of the above-described parameters ω, k and T may be fixed at both the encoder side and at the decoder side so that they need not to be signaled in the bitstream.

[0180] So far, the inventive concept has been described with reference to embodiments concerning methods for encoding / decoding the geometry of a point cloud into / from a bitstream. In accordance with further embodiments, the present invention also provides apparatuses for encoding / decoding a geometry of a point cloud into / from a bitstream, e.g., an encoder and / or a decoder operating in accordance with the above-described embodiments.

[0181] Fig. 15 illustrates a block diagram of an apparatus or encoder 400 for encoding into a bitstream a geometry of a point cloud. The point cloud is represented by a plurality of cuboid volumes. The plurality of cuboid volumes includes at least one occupied cuboid volume being modelled by one or more triangles. The one or more triangles have vertices on edges of the occupied cuboid volume. In accordance with embodiments, the apparatus 400 includes the following modules:

[0182] An encoding module 402 for encoding positions of vertices located on edges of the occupied cuboid volumes and a centroid drift per occupied cuboid volume.

[0183] The encoding module 402, for encoding the centroid drift for at least one occupied cuboid volume, comprises a modifying module 404 for modifying the position of  at least one vertex on an edge of the occupied cuboid volume, and a calculating module 406 for calculating the centroid drift, which is to be encoded, using the modified position of the at least one vertex.

[0184] Fig. 16 illustrates a block diagram of an apparatus or decoder 500 for decoding from a bitstream a geometry of a point cloud. The point cloud is represented by a plurality of cuboid volumes. The plurality of cuboid volumes includes at least one occupied cuboid volume being modelled by one or more triangles. The one or more triangles have vertices on edges of the occupied cuboid volume. In accordance with embodiments, the apparatus 500 includes the following modules:

[0185] A decoding module 502 configured to decode positions of vertices located on edges of the occupied cuboid volumes and a centroid drift per occupied cuboid volume.

[0186] The decoding module 502, for operating on at least one occupied cuboid volume, comprises a modifying module 504 for modifying the position of at least one vertex on an edge of the occupied cuboid volume, a calculation module 506 for calculating an initial position of a centroid vertex of the at least one occupied cuboid volume using the modified position of the at least one vertex, and an adapting module 508 for adapting the initial position of the centroid vertex using the decoded centroid drift.

[0187] The present invention further provides in embodiments an electronic device, a computer-readable storage medium and a computer program product.

[0188] Although some aspects of the disclosed concept have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or a device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.

[0189] Fig. 17 is a block diagram illustrating an electronic device 600 according to embodiments of the present invention.

[0190] The electronic device is intended to represent various forms of digital computers, such as a laptop, a desktop, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device  may also represent various forms of mobile devices, such as a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are described as examples only, and are not intended to limit implementations of the present invention described and / or claimed herein.

[0191] Referring to Fig. 17, the device 600 includes a computing unit 601 to perform various appropriate actions and processes according to computer program instructions stored in a read only memory (ROM) 602, or loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data for the operation of the storage device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0192] Components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse; an output unit 607, such as various types of displays, speakers; a storage unit 608, such as a disk, an optical disk; and a communication unit 609, such as network cards, modems, wireless communication transceivers, and the like. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0193] The computing unit 601 may be formed of various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU) , graphics processing unit (GPU) , various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processor (DSP) , and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs various methods and processes described above, such as an image processing method. For example, in some embodiments, the image processing method may be implemented as computer software programs that are tangibly embodied on a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM  602 and / or the communication unit 609. When a computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the image processing method described above may be performed. In some embodiments, the computing unit 601 may be configured to perform the image processing method in any other suitable manner (e.g., by means of firmware) .

[0194] Various implementations of the systems and techniques described herein above may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA) , application specific integrated circuits (ASIC) , application specific standard products (ASSP) , system-on-chip (SOC) , complex programmable logic device (CPLD) , computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, and the programmable processor may be a special-purpose or general-purpose programmable processor, and may receive data and instructions from a storage system, at least one input device and at least one output device, and may transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0195] Program code for implementing the methods of the present invention may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general computer, a dedicated computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions and / or operations specified in the flowcharts and / or block diagrams is performed. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on a machine and partly on a remote machine or entirely on a remote machine or server.

[0196] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A  machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memories (RAM) , read-only memories (ROM) , erasable programmable read-only memories (EPROM or flash memory) , fiber optics, compact disc read-only memories (CD-ROM) , optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0197] To provide interaction with a user, the systems and techniques described herein may be implemented on a computer having a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) ) for displaying information for the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide an input to the computer. Other types of devices can also be used to provide interaction with the user, for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback) ; and may be in any form (including acoustic input, voice input, or tactile input) to receive the input from the user.

[0198] The systems and techniques described herein may be implemented on a computing system that includes back-end components (e.g., as a data server) , or a computing system that includes middleware components (e.g., an application server) , or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein) , or a computer system including such a backend components, middleware components, front-end components or any combination thereof. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network) . Examples of the communication network includes: Local Area Networks (LAN) , Wide Area Networks (WAN) , the Internet and blockchain networks.

[0199] The computer system may include a client and a server. The Client and server are generally remote from each other and usually interact through a communication  network. The relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business expansion in traditional physical hosts and virtual private servers ( "VPS" for short) . The server may also be a server of a distributed system, or a server combined with a blockchain.

[0200] It should be understood that the steps may be reordered, added or deleted by using the various forms of flows shown above. For example, the steps described in the present invention may be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions in the present invention can be achieved, and no limitation is imposed herein.

[0201] The above-mentioned specific embodiments do not limit the scope of protection of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and replacements may be made depending on design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1.A method of encoding into a bitstream a geometry of a point cloud, the point cloud being represented by a plurality of cuboid volumes, the plurality of cuboid volumes comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the method comprising:encoding positions of vertices located on edges of the occupied cuboid volumes, andencoding a centroid drift per occupied cuboid volume,wherein encoding the centroid drift for at least one occupied cuboid volume comprises:- modifying the position of at least one vertex on an edge of the occupied cuboid volume, and- calculating the centroid drift, which is to be encoded, using the modified position of the at least one vertex.2.A method of decoding from a bitstream a geometry of a point cloud, the point cloud being represented by a plurality of cuboid volumes, the plurality of cuboid volumes comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the method comprising:decoding positions of vertices located on edges of the occupied cuboid volumes, anddecoding a centroid drift per occupied cuboid volume,wherein for at least one occupied cuboid volume the method comprises:- modifying the position of at least one vertex on an edge of the occupied cuboid volume,- calculating an initial position of a centroid vertex of the at least one occupied cuboid volume using the modified position of the at least one vertex, and- adapting the initial position of the centroid vertex using the decoded centroid drift.3.The method of claim 2, further comprising:- constructing for the at least one occupied cuboid volume at least one triangle using the decoded positions of the vertices and the adapted initial position of the centroid vertex, and- reconstructing the point cloud using ray tracing on the constructed triangle in the occupied cuboid volume.4.The method of any one of the preceding claims, wherein a position of the at least one vertex is modified when one or more of the following conditions are met:- a distance between the at least one vertex and a corner of the occupied cuboid volume is less than or equal to a distance threshold,- a total number of vertices in the occupied cuboid volume is greater than a first minimum number,- in the occupied cuboid volume, a total number of vertices, which have a distance to the corner being less than or equal to the distance threshold, is greater than a second minimum number.5.The method of claim 4, wherein one or more of the distance threshold, the first minimum number and the second minimum number is signaled in the bitstream or is a fixed value known at both the encoding side and the decoding side.6.The method of claim 4, wherein the distance threshold is determined using a width of the occupied cuboid volume and a scaling coefficient.7.The method of claim 6, wherein the distance threshold is determined as follows: th=ω*blockWidthwhere:th = the distance threshold,ω = the scaling coefficient ranging within (0, 1) ,blockWidth = a width of the occupied cuboid volume, and* = indicates a multiplication operation or a k-bits right shift operation.8.The method of claim 7, wherein the scaling coefficient and / or k is signaled in the bitstream or is a fixed value known at both the encoding side and the decoding side.9.The method of claim 4 or 5, wherein the distance threshold is a number of vertex sampling positions away from the corner.10.The method of any one of claims 4 to 9, wherein the bitstream has encoded therein a geometry parameter set, GPS, the GPS including one or more of the following:- the distance threshold, the first minimum number and the second minimum number,- the scaling coefficient and / or k,- the number of vertex sampling positions.11.The method of any one of the preceding claims, wherein modifying the position of the at least one vertex comprises modifying the position of the at least one vertex to match a position of a corner of the occupied cuboid volume.12.A non-transitory computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of any one of the preceding claims.13.An apparatus for encoding into a bitstream a geometry of a point cloud, the point cloud being represented by a plurality of cuboid volumes, the plurality of cuboid volumes  comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the apparatus comprising:an encoding module configured to encode positions of vertices located on edges of the occupied cuboid volumes and a centroid drift per occupied cuboid volume,wherein the encoding module is configured to encode the centroid drift for at least one occupied cuboid volume by:- modifying the position of at least one vertex on an edge of the occupied cuboid volume, and- calculating the centroid drift, which is to be encoded, using the modified position of the at least one vertex.14.An apparatus for decoding from a bitstream a geometry of a point cloud, the point cloud being represented by a plurality of cuboid volumes, the plurality of cuboid volumes comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the apparatus comprising:a decoding module configured to decode positions of vertices located on edges of the occupied cuboid volumes and a centroid drift per occupied cuboid volume,wherein for at least one occupied cuboid volume the decoding module configured to:- modify the position of at least one vertex on an edge of the occupied cuboid volume,- calculate an initial position of a centroid vertex of the at least one occupied cuboid volume using the modified position of the at least one vertex, and- adapt the initial position of the centroid vertex using the decoded centroid drift.15.A data stream having encoded thereinto a geometry of a point cloud, the point cloud  being represented by a plurality of cuboid volumes, the plurality of cuboid volumes comprising at least one occupied cuboid volume being modelled by one or more triangles, the one or more triangles having vertices on edges of the occupied cuboid volume, the data stream comprising:encoded positions of vertices located on edges of the occupied cuboid volumes, andan encoded centroid drift per occupied cuboid volume, the encoded centroid drift for at least one occupied cuboid volume being based on a modified position of at least one vertex of the at least one occupied cuboid volume.

Citation Information

Patent Citations

  • Point cloud processing method and device, encoder, decoder and readable storage medium

    CN117223287A

  • Multiresolution surface representation and compression

    US10192353B1

  • Method for encoding and decoding a 3D point cloud, encoder, decoder

    WO2023184393A1

  • Apparatus for coding vertex position for point cloud, and data stream including vertex position

    WO2023193533A1

  • Methods and apparatus for coding presence flag for point cloud, and data stream including presence flag

    WO2023193534A1