Encoding / decoding positions of points of point cloud contained in cubic volume

By interleaving and encoding the occupancy information and Trisoup information of the octree structure in a single transmission channel, the problems of transmission delay and low efficiency in point cloud compression technology are solved, and efficient point cloud data transmission and display are achieved.

CN120266478APending Publication Date: 2025-07-04BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380068171.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-28
Filing Date
2023-04-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When existing point cloud compression technology deals with intensive point clouds, especially in dynamic AR/VR applications, there are problems of transmission delay and low compression efficiency, which is difficult to meet the needs of real-time transmission and high-quality display.

Method used

The interleaved encoding method is adopted to interleave the octree structure occupancy information and Trisoup information in a single transmission channel, reducing transmission delay and improving encoding efficiency.

Benefits of technology

Through the interleaved encoding method, the transmission delay between the encoder and the decoder is reduced, the transmission efficiency and quality of point cloud data are improved, and the needs of real-time transmission and high-quality display are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266478A_ABST
    Figure CN120266478A_ABST
Patent Text Reader

Abstract

A method and apparatus of encoding / decoding are provided that encode / decode position information of points of a point cloud into / from a bitstream containing at least one data unit. The points are included in a cubic volume associated with leaf nodes of an octree structure. The method comprises: encoding occupancy information of nodes of an octree structure, the occupancy information representing the presence of points of a point cloud contained in the cubic volume; and encoding, for each current cube volume containing at least one point of the point cloud, Trisound information representing the presence of vertices on an edge of the current cube volume and the position of the vertices along the edge. According to the invention, the occupancy information of a leaf node and the Trisup information of another leaf node are encoded in a data unit in an interleaved manner.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority based on European Patent Application No. “22306441.1” filed on September 28, 2022, the entire content of which is incorporated herein by reference. Technical field

[0003] This application generally relates to point cloud compression, and particularly to methods and apparatuses for encoding / decoding, which encode the positions of points of a point cloud in a bitstream including at least one data unit / decode the positions of points of a point cloud from a bitstream including at least one data unit. Background art

[0004] This section aims to introduce readers to various aspects of the field that may be relevant to at least one exemplary embodiment of the present application described and / or claimed below. This discussion is considered to help provide background information for readers to better understand the various aspects of the present application.

[0005] As a representation format of three - dimensional (3D) data, point clouds have recently gained increasing attention for their general use in representing all types of physical objects or scenes. Point clouds can be used for various purposes, such as cultural heritage / architecture, where objects (such as statues or buildings) are 3D scanned to share their spatial structure without sending or accessing the object. Additionally, it is a way to ensure the preservation of knowledge of an object in case it may be destroyed (e.g., a temple is damaged by an earthquake). Such point clouds are usually static, colored, and huge.

[0006] Another use case is in topography and cartography, where the use of 3D representation makes it possible to create maps that are not limited to a plane and can include terrain. Currently, Google Maps is an excellent example of a 3D map, but it uses meshes instead of point clouds. Nevertheless, point clouds can be a suitable data format for 3D maps, and such point clouds are usually static, colored, and huge.

[0007] Virtual reality (VR), augmented reality (AR), and immersive worlds have recently become popular topics, and many people foresee them as the future of two - dimensional (2D) flat - screen videos. The basic idea is to immerse the viewer in the surrounding environment, while ordinary TV only allows the viewer to watch the virtual world in front of her / him. Depending on the degree of freedom of the viewer in the environment, the immersion can be divided into several levels. Point clouds are excellent format candidates for distributing VR / AR worlds.

[0008] The automotive industry, more particularly autonomous vehicles, is also an area where point clouds can be used extensively. Autonomous vehicles should be able to “detect” their surrounding environment to make excellent driving decisions based on the presence and nature of nearby objects detected immediately and the road structure.

[0009] A point cloud is a set of points located in 3D space, optionally each point being attached with additional values. These additional values are typically referred to as attributes. Attributes can be, for example, three-component color, material properties (such as reflectivity), and / or a two-component normal vector of the surface associated with the point.

[0010] Thus, a point cloud is a combination of geometric data (the positions of the points in 3D space, usually represented by three-dimensional Cartesian coordinates x, y, and z) and attributes.

[0011] A point cloud can be sensed by various types of devices (such as an array of the following: cameras, depth sensors, lasers (light detection and ranging, also known as lidar), radars), or can be computer-generated (for example, in post-production of movies). Depending on the use case, for cartography applications, a point cloud can have thousands to billions of points. The raw representation of a point cloud requires each point to have a very large number of bits, at least a dozen bits for each Cartesian coordinate x, y, or z, and optionally more bits for one or more attributes, for example, 3 times 10 bits for color.

[0012] In many applications, it is very important to distribute the point cloud to end-users or store it in a server while consuming only a reasonable amount of bitrate or storage space, while maintaining an acceptable (or preferably very excellent) quality of experience. The efficient compression of these point clouds is a key point to make the distribution chain of many immersive worlds practical.

[0013] Compression can be lossy (as in video compression), for distribution to end-users and for their visualization, such as on AR / VR glasses or any other 3D-capable device. Other use cases do require lossless compression, such as medical applications or autonomous driving, to avoid changing the decision results obtained from subsequent analysis of the compressed and transmitted point cloud.

[0014] Until recently, point cloud compression (PCC) has only received mass-market attention, and there has been no available standardized point cloud codec. In 2017, the standardization working group ISO / JCT1 / SC29 / WG11, also known as the Moving Picture Experts Group or MPEG, initiated a work item on point cloud compression. This has resulted in two standards, namely:

[0015] · Part 5 of MPEG-I (ISO / IEC 23090-5), also known as Video-based Point Cloud Compression (V-PCC).

[0016] · Part 9 of MPEG-I (ISO / IEC 23090-9), also known as Geometry-based Point Cloud Compression (G-PCC).

[0017] The V-PCC encoding / decoding method compresses a point cloud by performing multiple projections on a 3D object to obtain 2D patches that are packed into an image (or in a video when processing dynamic point clouds). Then, an existing image / video codec is used to compress the obtained image or video, thus allowing the full utilization of the deployed image and video solutions. By its nature, V-PCC is only efficient on dense and continuous point clouds because if non-smooth patches are obtained from projections of sparse geometric data sensed by, for example, lidar, then the image / video codec cannot compress them.

[0018] The G-PCC encoding / decoding method has two schemes for the compression of sensed sparse geometric data.

[0019] The first scheme is based on an occupancy tree representing the geometry of the point cloud, which can locally be any of an octree, a quadtree, or a binary tree. Occupied nodes (i.e., nodes associated with a cubic volume containing at least one point of the point cloud) are split until a specific size is reached, and the occupied leaf nodes provide the 3D locations of the points, e.g., at the centers of these nodes. The occupancy information is carried by occupancy data (binary data, flags) that signal the occupancy status of each child node of a node. By using neighbor-based prediction techniques, a high level of compression can be obtained for the occupancy data of dense point clouds. For sparse point clouds, it can also be solved by directly encoding / decoding the positions of the points within nodes of non-minimal size by stopping the tree construction when there are only isolated points in a node; this technique is called the direct coding / decoding mode (DCM).

[0020] The second scheme is based on a prediction tree where each node represents the 3D location of a point and the parent / child relationship between nodes represents a spatial prediction from the parent node to the child node. Compared with using an occupancy tree, this method can only solve sparse point clouds and offers the advantages of lower latency and simpler decoding. However, compared with the first occupancy-based method, the compression performance is only slightly better, and the encoding is complex because the encoder needs to densely search for the best predictor among a large number of potential predictors when constructing the prediction tree.

[0021] In both schemes, the attribute encoding / decoding is performed after the complete geometry encoding / decoding, which actually results in two-pass encoding / decoding. Therefore, the low latency of geometric / attribute joint is obtained by using a partitioning of the point cloud data that decomposes the three-dimensional space into independently encoded / decoded sub-volumes without prediction between sub-volumes. When using many partitions of the point cloud data, this severely affects the compression performance.

[0022] An important use case is the transmission of dynamic AR / VR point clouds. Dynamic means that the point cloud evolves over time. Additionally, AR / VR point clouds are typically locally two-dimensional as they represent the surface of an object most of the time. Thus, AR / VR point clouds are highly connected (or dense), meaning that points are rarely isolated but have many neighboring points.

[0023] A dense (or solid) point cloud represents a continuous surface with a certain resolution such that the cubic volumes (voxels) associated with the points touch each other without showing any visible holes on the surface.

[0024] Such point clouds are typically used in AR / VR environments and can be viewed by end-users through devices such as TVs, smartphones, or headsets. They are transmitted to the devices or are usually stored locally. Many AR / VR applications use moving point clouds, which, in contrast to static point clouds, change over time. Thus, the point cloud data is very large and must be compressed.

[0025] Now, lossless compression based on the octree representation of the geometry of point cloud frames (temporal instances of the point cloud) can achieve slightly less than one bit per point (1 bpp). For real-time transmission that may involve millions of points per point cloud frame and a frame rate of up to 50 frames per second (fps), resulting in hundreds of megabits of data per second, this may not be sufficient.

[0026] Therefore, lossy compression that typically requires sufficient compression while maintaining an acceptable visual quality to fit the bandwidth provided by the transmission channel can be used. In most applications, real-time transmission of point cloud frames must be ensured, which imposes limitations on the complexity and latency of the compression scheme. In many applications, bitrates as low as 0.1 bpp (a 10-fold compression compared to lossless codecs) have made real-time transmission possible.

[0027] V-PCC can achieve such low bitrates through lossy compression using a video codec that compresses the two-dimensional frames (pictures) obtained by projecting (multiple) parts of the point cloud onto a set of planes to obtain projection patches. The geometry of the point cloud frame is represented by assembling these projection patches into a frame, where each patch is a small local depth map. However, V-PCC is not general and is limited to a narrow type of point cloud that does not show locally complex geometries (such as trees, hair) because the obtained projection depth maps are not smooth enough to be effectively compressed by video codecs such as HEVC or VVC.

[0028] Pure 3D compression techniques can handle any type of point cloud. Regarding dense point clouds, it is still an open question whether 3D compression techniques can compete with V-PCC (or any projection + image codec scheme). Standardization work to provide an extension (revision) to the current G-PCC codec to provide competitive lossy compression that is as good as V-PCC for dense point clouds while maintaining the generality of G-PCC (able to handle any type of point cloud (dense, sparse (e.g., point clouds captured based on Lidar), 3D maps, etc.)) is still ongoing.

[0029] Basically, there are three main approaches to obtain lossy schemes based on the octree representation of point clouds as accomplished by the current G-PCC codec.

[0030] The first approach consists of downsampling the entire point cloud to a smaller resolution, losslessly encoding the downsampled point cloud, decoding the encoded downsampled point cloud, and then upsampling the decoded downsampled point cloud to the original resolution. Many upsampling schemes have been proposed (such as super-resolution, AI-based upsampling, learning-based 3D post-processing, etc.). When the downsampling is not too aggressive, e.g., the downsampling in each spatial direction does not exceed a factor of 2, the first approach can provide good Peak Signal to Noise Ratio (PSNR) results. However, even though the first approach can provide good PSNR results, the visual quality of the upsampled point cloud is controversial and cannot be well controlled.

[0031] The second approach is to have the encoder perform local "adjustments" to the point cloud so that the bitrate required for encoding and decoding the octree is lower. For this purpose, the points of the point cloud can be slightly moved, e.g., to obtain occupancy information that is better predicted by neighboring nodes, thus resulting in lossless encoding of the modified octree at a reduced bitrate. Unfortunately, this approach only results in a small reduction in bitrate.

[0032] The third approach is based on the so-called TriSoup codec scheme, which is being explored by the ISO / IEC JTC1 / SC29 / WG7 standardization working group. The TriSoup codec scheme is the most promising of the three approaches but still requires some work before it matures.

[0033] The TriSoup codec scheme provides a local model of the geometry of the point cloud by using a set of triangles (without explicitly providing connectivity information), hence its name from "soup of triangles".

[0034] Point cloud geometry is represented in a three-dimensional coordinate system by a three-dimensional tree (usually an octree) with a resolution reduced to a certain extent, for example, reduced to a cubic volume (cube) of NxNxN associated with a leaf node, where N can be 4, 8, or 16. A lossless encoding and decoding scheme (such as G-PCC) is used to encode and decode this tree. The tree itself does not require much bitrate because it does not go to the deepest depth and has a smaller number of leaf nodes (cubic volumes) compared to the number of points in the point cloud. Then, within each NxNxN cubic volume, the point cloud geometry is modeled by a certain local model consisting of a set of triangles. As Figure 1 shown, the vertices of the triangles are encoded and decoded along the edges of the cubic volume associated with the leaf nodes of the tree. Four vertices (dark black dots) along the edge of the cube are shown. These four vertices are used to form two triangles, as discussed subsequently.

[0035] The TriSoup encoding and decoding scheme is not limited to a specific number of vertices distributed along the edges of the cubic volume, but stipulates that there is at most one vertex per edge. The TriSoup encoding and decoding scheme is not limited to a specific number of triangles in a specific cubic volume. The vertices are shared between cubic volumes that have a common edge, and the point belongs to this common edge. This means that: for each edge belonging to at least one cubic volume, at most one vertex is encoded and decoded. Thus, the continuity of the local model is ensured through the cubic volumes associated with the leaf nodes of the tree.

[0036] Encoding and decoding the vertices located on the edges of the cubic volume requires encoding the vertex information for each vertex, where the vertex information includes a vertex flag indicating whether the vertex exists on the edge of the cubic volume and the vertex position along the edge (if it exists). Subsequently, the position of the vertex in the three-dimensional coordinate system is obviously derived from the vertex position along the edge and the positions of the starting vertex and the ending vertex of the cubic volume that defines the edge.

[0037] The vertex flag can be encoded and decoded by an adaptive binary arithmetic codec that uses a specific context to encode and decode the vertex flag. The vertex position along an edge of length N = 2 s can be encoded with unit precision by pushing (bypassing / non-entropy encoding) s bits into the bitstream.

[0038] In summary, using the TriSoup encoding and decoding scheme to encode and decode point cloud geometry requires encoding octree data (such as through G-PCC), occupancy information, and Trisoup information that represents the existence of vertices on the edges of the cubic volume and represents the vertex positions along the said edges. The increase in the encoded data caused by encoding the Trisoup information is greatly compensated by the improvement in reconstructing the point cloud with triangles, as discussed below.

[0039] Within the cubic volume associated with the leaf node of the tree, if there are at least three vertices on the edge of the cubic volume, at least one triangle is constructed based on the vertices. When there are more than three vertices on the edge of the cubic volume, more than one triangle can be constructed.

[0040] There are multiple construction processes to construct such triangles.

[0041] In the current TriSoup encoding and decoding scheme, triangles are constructed as follows: First, determine the dominant direction along one of the three axes of the three-dimensional coordinate system representing the cubic volume; then, sort at least three vertices according to the determined dominant direction; finally, construct at least one triangle based on the sorted list of at least three vertices.

[0042] The dominant direction is determined by testing each of the three directions along the three axes, and the direction that maximizes the total surface of the triangle (as seen along the test direction) is retained as the dominant direction.

[0043] Figure 2 and Figure 3 An illustrative example showing the process of determining the dominant direction according to the current TriSoup is presented.

[0044] For simplicity of the figure, Figure 2 and Figure 3 only the cases of testing for two axes are depicted.

[0045] Figure 2 An illustrative example is shown when the cubic volume and four vertices (lower right) are projected onto a two-dimensional plane (upper right) along the vertical axis. Then, the four vertices are sorted in clockwise order relative to the center of the projected cubic volume (square). Then, based on the sorted vertices, two triangles (upper left) are constructed according to a fixed rule. Here, the fixed rule is that when there are 4 vertices, triangles 123 and 134 are constructed. When there are 3 vertices, the only possible triangle 123 can be constructed. When there are 5 vertices, triangles 123, 134, and 451 can be constructed. And so on, up to 12 vertices. Other fixed rules can also be used to construct triangles from vertices.

[0046] Figure 3 An illustrative example is shown when the cubic volume and four vertices (lower right) are projected onto a two-dimensional plane (lower left) along the horizontal axis. Then, the four vertices are sorted in clockwise order relative to the center of the projected cubic volume (square). Then, based on the sorted vertices, two triangles (upper left) are constructed according to the above fixed rule.

[0047] In Figure 2 andFigure 3 In the example of Figure 2 ), the vertical projection ( Figure 3 ) shows the (maximum) two-dimensional total projection surface of the triangle on the two-dimensional plane (greater than the two-dimensional total projection surface of the triangle on the two-dimensional plane produced by the horizontal projection,

[0048] ) on the two-dimensional plane. Therefore, the vertical direction is chosen as the dominant direction, and two triangles are constructed from those vertices sorted according to the numbers of the vertices produced by the vertical projection.

[0049] In the current TriSoup encoding and decoding scheme, ray tracing is performed on each cubic volume of the tree to render the triangles defined within each cubic volume as rendered points representing (modeling) the points of the point cloud contained in the cubic volume. The positions of all the rendered points obtained by ray tracing will be the positions of the decoded points of the point cloud. This process is called "voxelization" of the triangles.

[0050] As Figure 4 shown, ray tracing performed on the triangles of the cubic volume essentially consists of emitting rays in three directions parallel to the axes of the three-dimensional coordinate system. The origin of the emitted rays is a point with integer (voxelized) coordinates, whose precision corresponds to the sampling precision required for rendering. Then the intersections of the rays with the triangles (if any, shown as dotted points) are voxelized ( = rounded to the nearest point with the required sampling precision) and added to the list of rendered points.

[0051] After applying the TriSoup encoding and decoding scheme to the cubic volumes associated with all the leaf nodes of the tree representing the geometry of the point cloud (i.e., constructing triangles for each cubic volume of the cubic volumes and obtaining rendered points through ray tracing), copies of the same rendered points in the list of all the rendered points are discarded (i.e., only one voxel (rendered point) is retained among all the voxels sharing the same three-dimensional position) to obtain the set of decoded (unique) points of the point cloud.

[0052] When there are more than 4 vertices involved, constructing triangles according to fixed rules following the sorted vertices may result in poor visual effects. In fact, for a given set of vertices, other triangles different from those constructed by the above fixed rules can be constructed, leading to different sets of rendered points (different visual aspects of the decoded points of the point cloud). Moreover, it turns out that reconstructing triangles according to the above fixed rules is not the most ideal in terms of the PSNR (peak signal-to-noise ratio) results.

[0053] Occupancy information and Trisoup information of the point cloud data partitioning can be carried by the bitstream in multiple separate data units of the bitstream or in two separate parts of the same data unit DU. After encoding the bitstream, the compressed point cloud data is transmitted through a transmission system (such as the Internet). If there is only one transmission channel in the transmission system, the entire bitstream in the data unit carrying the octree information (including all information for reconstructing the octree structure, which includes occupancy information) of the octree structure is transmitted first, and then the data unit (or part of the same data unit) carrying the Trisoup information of the leaf nodes of the octree structure is transmitted. Since the octree information of the partitioning of the point cloud data is large, it introduces a latency in the transmission of the encoded point cloud geometry data, thus affecting use cases of real-time applications (such as autonomous vehicles). A straightforward solution is to use two transmission channels in the transmission system, one channel can be used to transmit occupancy information, and the other channel can be used to transmit Trisoup information. This allows for parallel transmission of occupancy information and Trisoup information, thus avoiding the introduction of transmission latency. However, in this case, the capacity of the channels is not fully utilized, and in some cases, it causes waste of capacity. In addition, there are many use cases that use a single channel.

[0054] The problem to be solved is to reduce the transmission latency of the bitstream for transmitting occupancy information and Trisoup information of the octree structure when they are transmitted through a single transmission channel.

[0055] Based on the above considerations, at least one exemplary embodiment of the present application is designed. Summary of the Invention

[0056] The following part presents a simplified overview of at least one exemplary embodiment to provide a basic understanding of certain aspects of the present application. This overview is not a comprehensive review of the exemplary embodiments. It is not intended to identify the key or core elements of the embodiments. The following overview only presents certain aspects of at least one exemplary embodiment in a simplified form as a prelude to the more detailed description provided in other parts of this document.

[0057] According to a first aspect of the present application, there is provided a method of encoding position information of points of a point cloud into a bitstream including at least one data unit, the points being included in a cubic volume associated with a leaf node of an octree structure having a maximum depth d, at least three vertices being located on the edges of each of the cubic volumes, at most one vertex on each edge, and the position information of the points included in the cubic volume being represented by a triangle connecting the at least three vertices; the method includes: encoding occupancy information of nodes of the octree structure, the occupancy information representing the presence of points of the point cloud included in the cubic volume; and for each current cubic volume including at least one point of the point cloud, encoding Trisoup information, the Trisoup information representing the presence of vertices on the edges of the current cubic volume and the vertex positions along the edges, wherein the occupancy information of a leaf node and the Trisoup information of another leaf node are encoded in the data unit (DU) in an interleaved manner.

[0058] According to a second aspect of the present application, there is provided a method of decoding position information of points of a point cloud from a bitstream including at least one data unit, the points being included in a cubic volume associated with a leaf node of an octree structure having a maximum depth d, at least three vertices being located on the edges of each of the cubic volumes, at most one vertex on each edge, and the position information of the points included in the cubic volume being represented by a triangle connecting the at least three vertices, the method includes: decoding occupancy information of nodes of the octree structure, the occupancy information representing the presence of points of the point cloud included in the cubic volume; and for each current cubic volume including at least one point of the point cloud, decoding Trisoup information, the Trisoup information representing the presence of vertices on the edges of the current cubic volume and the vertex positions along the edges, wherein the occupancy information of a leaf node and the Trisoup information of another leaf node are encoded in the data unit (DU) in an interleaved manner.

[0059] In an exemplary embodiment, the method further includes a first list storing nodes of depth d - 1 sorted in encoding order or decoding order according to the coordinates of the nodes in a three-dimensional system; and encoding occupancy information of p child nodes of depth d of the nodes in the first list into a bit sequence BSoct1 of the data unit / decoding occupancy information of p child nodes of depth d of the nodes in the first list from the bit sequence BSoct1 of the data unit.

[0060] In an exemplary embodiment, the method further includes, for a first child node of a first node in the first list, encoding Trisoup information into a bit sequence BStris1 of the data unit / decoding Trisoup information from the bit sequence BStris1 of the data unit.

[0061] In an exemplary embodiment, the method further includes, for child nodes of a node in the first list, alternately: encoding occupancy information into a bit sequence BSoct2 of the data unit / decoding occupancy information from the bit sequence BSoct2 of the data unit; and encoding Trisoup information into a bit sequence BStris2 of the data unit / decoding Trisoup information from the bit sequence BStris2 of the data unit, until the encoding / decoding of the occupancy information of the last child node of the node in the first list.

[0062] In an exemplary embodiment, the method further includes encoding / decoding Trisoup information of remaining child nodes added to a bit sequence BStris3 in the data unit.

[0063] In an exemplary embodiment, the method further includes obtaining a second list that stores child nodes of a node at depth d - 1, where the node has the same coordinates of the three - dimensional system, the nodes are sorted in a first encoding order or a first decoding order, and the child nodes are sorted in a second encoding order or a second decoding order.

[0064] In an exemplary embodiment, the method further includes, for a first child node of a node in the second list in the bit sequence BStris1 of the data unit, encoding / decoding Trisoup information; and encoding / decoding octree information of a first child node of a node with updated same coordinates in a bit sequence BSoct1 of the data unit, where the updated same coordinates are equal to the same coordinates plus 1.

[0065] In an exemplary embodiment, the method further includes, in an interleaved manner: encoding Trisoup information of the next child node of the node with the updated same coordinates into a bit sequence BStris2 of the data unit / decoding Trisoup information of the next child node of the node with the updated same coordinates from the bit sequence BStris2 of the data unit; and encoding occupancy information of the next child node into a bit sequence BSoct2 of the data unit (DU) / decoding occupancy information of the next child node from the bit sequence BSoct2 of the data unit (DU), until the encoding / decoding of the occupancy information of the last child node of the node in the first list.

[0066] In an exemplary embodiment, the method further includes encoding the Trisoup information of the remaining child nodes of the second list into a bit sequence BStris3 of the data unit (DU) / decoding the Trisoup information of the remaining child nodes of the second list from the bit sequence BStris3 of the data unit (DU).

[0067] In an exemplary embodiment, the encoding order and / or the decoding order and / or the first encoding order and / or the first decoding order and / or the second encoding order and / or the second decoding order is a Morton order or a raster scan order.

[0068] According to a third aspect of the present application, there is provided a bitstream formatted to include encoded point cloud geometry data obtained by the method according to the first aspect of the present application.

[0069] According to a fourth aspect of the present application, there is provided a device including means for performing one of the methods according to the first or second aspect of the present application.

[0070] According to a fifth aspect of the present application, there is provided a computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the methods according to the first or second aspect of the present application.

[0071] According to a sixth aspect of the present application, there is provided a non-transitory storage medium carrying instructions of program code for performing the methods according to the first or second aspect of the present application.

[0072] The specific nature of at least one exemplary embodiment and other objects, advantages, features and uses of the at least one exemplary embodiment will become apparent from the following description of the examples in conjunction with the drawings. Description of the Drawings

[0073] Reference will now be made to the drawings to illustrate exemplary embodiments of the present application by way of example, where:

[0074] Figure 1 A schematic diagram of a TriSoup cube volume including four vertices is shown;

[0075] Figure 2 An illustrative example of the process of determining the dominant direction according to the current TriSoup is shown;

[0076] Figure 3 An illustrative example of the process of determining the dominant direction according to the current Trisoup is shown;

[0077] Figure 4 Schematically shows an example of ray tracing for points of a point cloud to be rendered according to an exemplary embodiment;

[0078] Figure 5 Schematically illustrates, according to one exemplary embodiment, the data unit structure of a bitstream B generated by Figure 6 an encoding method;

[0079] Figure 6 Schematically illustrates a block diagram of a method 300 for encoding position information of points of a point cloud included in a cubic volume associated with a leaf node of an octree structure into a bitstream according to an exemplary embodiment;

[0080] Figure 7 Schematically illustrates a block diagram of a method 400 for decoding position information of points of a point cloud included in a cubic volume associated with a leaf node of an octree structure from a bitstream according to an exemplary embodiment;

[0081] Figure 8 Schematically illustrates a block diagram of a variant of method 300;

[0082] Figure 9 Schematically illustrates a block diagram of a variant of method 400;

[0083] Figure 10 Illustrates an example of a raster scan order of a complete octree structure in a global cube containing all points of a point cloud;

[0084] Figure 11 Illustrates an example of a delay introduced by p child nodes according to an exemplary embodiment;

[0085] Figure 12 Illustrates an example of a delay introduced by p child nodes according to an exemplary embodiment;

[0086] Figure 13 Illustrates a schematic block diagram of an example of a system implementing various aspects and exemplary embodiments.

[0087] Similar reference numerals may be used in different figures to denote similar components. Detailed Description

[0088] Hereinafter, at least one exemplary embodiment will be described more fully with reference to the accompanying drawings, in which examples of at least one exemplary embodiment are depicted. However, the exemplary embodiments may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Accordingly, it should be understood that the exemplary embodiments are not intended to be limited to the specific forms disclosed. Instead, this disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of this application.

[0089] One aspect generally relates to the encoding and decoding of point clouds, another aspect generally relates to the transmission of the generated or encoded bitstream, and another aspect relates to the reception / access of the decoded bitstream.

[0090] Furthermore, these aspects are not limited to the MPEG standards (such as MPEG-I Part 5 or Part 9 related to point cloud compression), but can also be applied to, for example, other standards and recommendations, whether existing or to be developed in the future, and extensions of any such standards and recommendations (including MPEG-I Part 5 and Part 9). Unless otherwise stated or technically excluded, the various aspects described in this application can be used alone or in combination.

[0091] In short, the present invention provides a method for encoding the position information of points of a point cloud into a bitstream containing at least one data unit / a method for decoding the position information of points of a point cloud from a bitstream containing at least one data unit. The points are contained in a cubic volume associated with a leaf node of an octree structure having a maximum depth d. At least three vertices are located on the edges of each such cubic volume, with at most one vertex on each edge, and the position information of the points contained in the cubic volume is represented by a triangle connecting the at least three vertices. The method includes encoding the occupancy information of the nodes of the octree structure, the occupancy information representing the presence of the points of the point cloud contained in the cubic volume; and for each current cubic volume containing at least one point of the point cloud, encoding Trisoup information, the Trisoup information representing the presence of vertices on the edges of the current cubic volume and the vertex positions along the edges. According to the present invention, the occupancy information of the leaf nodes and the Trisoup information of other leaf nodes are encoded in the data unit in an interleaved manner.

[0092] The present invention is beneficial because if a single transmission channel is used, the transmission of the Trisoup information does not need to wait for the occupancy information of all the nodes of the octree structure to be transmitted first and then the Trisoup information of the leaf nodes to be transmitted. Therefore, compared with the prior art, the present invention reduces the transmission delay between the encoder and the decoder, so that the bitstream generated by the encoding method of the present invention can be transmitted as quickly as possible.

[0093] Figure 5 Schematically illustrates the data unit structure of the bitstream B generated by the encoding method according to an exemplary embodiment through Figure 6 the encoding method.

[0094] Assume that the octree structure of the point cloud has a depth of (d + 1), starting from the 0th layer depth (the root node of the octree structure) and ending at the dth layer depth (the last layer depth). The leaf nodes are located at the last layer depth d of the octree structure, and the parent nodes of the child nodes are located at depth (d - 1).

[0095] The bitstream B includes a plurality of data units DU, for example, one data unit DU for each point cloud data partition. Each data unit DU is dedicated to a point cloud data partition, and the point cloud data can be divided into a plurality of point cloud data partitions.

[0096] Figure 5 The DU structure includes a DU header (header) usually formed by two bytes and a Raw Byte Sequence Payload (RBSP) formed by a number of bit sequences.

[0097] The DU header provides information about the type of information carried by the RBSP.

[0098] The RBSP includes a bit sequence BSoct1 and at least one bit sequence BSoct2, representing the occupancy information of the nodes of the octree structure, where these nodes are either located at a position with a depth strictly less than depth d or at depth d.

[0099] The RBSP also includes a bit sequence BStris1, at least one bit sequence BStris2, and a third bit sequence BStris3, representing the Trisoup information of the leaf nodes of the octree structure. According to the present invention, the bit sequences BSoct2 and BStris2 are interleaved in the data unit DU.

[0100] Figure 6 A block diagram schematically illustrates a method 300 for encoding the position information of the points of the point cloud contained in the cubic volume associated with the leaf nodes of the octree structure according to an exemplary embodiment.

[0101] Basically, the method 300 encodes the occupancy information of the nodes of the octree structure, and the occupancy information represents the presence of the points of the point cloud contained in the cubic volume associated with the nodes of the octree structure. For each current cubic volume containing at least one point of the point cloud, the method 300 includes encoding the Trisoup information, which represents the presence of the vertices on the edges of the current cubic volume and the vertex positions along the edges.

[0102] In step 310, the method 300 encodes the occupancy information of the nodes at the first d layer depths (from depth 0 to depth d - 1) in the bit sequence BSoct1.

[0103] In step 320, obtain list L d-1 . The list L d-1 Stores the nodes at depth (d - 1) sorted in encoding order according to the coordinates of these nodes in the three-dimensional system.

[0104] In step 330, initialize a first child node index m for encoding occupancy information and a second child node index k for encoding Trisoup information. For example, m = k = 0.

[0105] In step 340, the occupancy information of p (an integer) child nodes at depth d of the nodes in list L d-1 (i.e., the occupancy information of p leaf nodes) is encoded in bit sequence BSoct1, starting from the first child node (m = 0) and ending at the p-th child node (m = p - 1). The child nodes are considered in encoding order. Then, k = 0, m = p - 1.

[0106] Encoding the occupancy information of p child nodes before encoding the Trisoup information of the leaf nodes introduces a small delay between encoding the occupancy information and the Trisoup information. This delay is unavoidable because encoding the Trisoup information of the leaf nodes requires pre-encoding the occupancy information of adjacent nodes (here, p child nodes). The reason that the Trisoup encoding of a node requires the occupancy information of adjacent nodes (p child nodes) is due to the characteristics of the Trisoup encoding of each child node.

[0107] In step 350, for the first child node (k = 0) of the first node in list (L d-1 ), encode the Trisoup information into bit sequence BStris1.

[0108] In step 360, for the child nodes of the nodes in list L d-1 , encode the occupancy information in bit sequence BSoct2 and the Trisoup information in bit sequence BStris2 in an interleaved manner (i.e., alternately) until the encoding of the occupancy of the last child node (m = Nmax) of the last node in L d-1 . Then, m = Nmax, k = Nmax - p.

[0109] In sub-step 361 of step 360, increment the first sub-index m by 1 and also increment the second sub-index by 1 to encode the octree information and Trisoup information of the next child node of the nodes in list L d-1 .

[0110] Note that the child node index used for occupancy information encoding and the child node index used for Trisoup information encoding are not the same. There is a slight time delay (latency) between these two child node indexes.

[0111] In sub-step 362, the occupancy information of the m-th child node is encoded in the bit sequence BSoct2; and in sub-step 363, the Trisoup information of the k-th child node is encoded in the bit sequence BStris2. As Figure 5 shown, the bit sequences BSoct2 and BStris2 are alternately added to the data unit, that is, the occupancy information of the m-th child node and the Trisoup information of the k-th child node are encoded in an interleaved manner.

[0112] Sub-steps 361 - 363 are iterated, and in each iteration, the bit sequences BSoct2 and BStris2 are added to the data unit DU in an interleaved manner.

[0113] In step 370, the Trisoup information of the remaining child nodes (p = Nmax - k) is encoded in the bit sequence BStris3.

[0114] The data unit DU ends with a bit sequence (e.g., 10000).

[0115] Figure 7 A block diagram schematically illustrates a method 400 for decoding the position information of points in a point cloud contained in a cubic volume associated with a leaf node of an octree structure from a bitstream according to an exemplary embodiment.

[0116] Basically, method 400 decodes the occupancy information of the nodes of the octree structure, and the occupancy information represents the presence of points in the point cloud contained in the cubic volume associated with the nodes of the octree structure. For each current cubic volume containing at least one point of the point cloud, method 400 includes decoding the Trisoup information, which represents the presence of vertices on the edges of the current cubic volume and the vertex positions along the edges.

[0117] After obtaining the data unit DU of the bitstream B related to the information of the partition of the point cloud data, in step 410, method 400 decodes the occupancy information of the nodes at the first d levels of depth (from depth 0 to depth d - 1) from the bit sequence BSoct1.

[0118] In step 420, the list L is obtained d-1 . The list L d-1 stores the nodes at depth (d - 1) sorted in the decoding order according to the coordinates of these nodes in the three-dimensional system.

[0119] In step 430, a first child node index m is initialized for decoding occupancy information, and a second child node index k is initialized for decoding Trisoup information. For example, m = k = 0.

[0120] In step 440, the occupancy information of p child nodes (i.e., the occupancy information of p leaf nodes) with depth d of the nodes in list L is decoded from the bit sequence BSoct1 of data unit DU, starting from the first child node (m = 0) and ending at the p-th child node (m = p - 1). Then, k = 0, m = p - 1. d-1 the occupancy information of p child nodes (i.e., the occupancy information of p leaf nodes) with depth d of the nodes in list L is decoded from the bit sequence BSoct1 of data unit DU, starting from the first child node (m = 0) and ending at the p-th child node (m = p - 1). Then, k = 0, m = p - 1.

[0121] In step 450, for the first child node (k = 0) of the first node in list L d-1 Trisoup information is decoded from the bit sequence BStris1 of data unit DU.

[0122] In step 460, for the child nodes of the nodes in list L d-1 it is done in an interleaved manner (i.e., alternately): the occupancy information is decoded from the bit sequence BSoct2 of data unit DU and the Trisoup information is decoded from the bit sequence BStris2 of data unit DU, and the above interleaved decoding is iterated until the decoding of the occupancy information of the last child node (m = Nmax) of the last node in L d-1 is completed. Then, m = Nmax, k = Nmax - p.

[0123] In sub-step 461 of step 460, the first sub-index m is incremented by 1, and the second sub-index is also incremented by 1 to decode the octree information and Trisoup information of the next child node of the nodes in list L d-1 is completed. Then, m = Nmax, k = Nmax - p.

[0124] In sub-step 462, the occupancy information of the m-th child node is decoded from the bit sequence BSoct2 of data unit DU; and in sub-step 463, the Trisoup information of the k-th child node is decoded from the second bit sequence BStris2 of data unit DU. As Figure 5 shown, the bit sequences BSoct2 and BStris2 can be alternately decoded from data unit DU, that is, the occupancy information of the m-th child node and the Trisoup information of the k-th child node are decoded in an interleaved manner.

[0125] Sub-steps 461 - 463 are iterated, and in each iteration, the bit sequences BSoct2 and BStris2 are considered to be in an interleaved manner in data unit DU.

[0126] In step 470, the Trisoup information of the remaining child nodes (p = Nmax - k) is decoded from the bit sequence BStris3 of the data unit DU.

[0127] The data unit DU ends with a bit sequence (e.g., 10000).

[0128] In one exemplary embodiment, by leveraging the characteristics of a linear memory structure (i.e., each element in the memory retains information on how to locate the next and previous elements), the iterations of steps 361 - 363 and steps 461 - 463 can directly locate the next child node in the manner of a linear memory structure such as a list.

[0129] This exemplary embodiment is beneficial because it avoids using the child node indices m and k to locate and then iterate over each child node.

[0130] In one exemplary embodiment, encoding / decoding the occupancy information of a leaf node and encoding / decoding the Trisoup information of the leaf node are performed in parallel.

[0131] This embodiment reduces the latency between encoding / decoding the occupancy information and encoding / decoding the Trisoup.

[0132] In one exemplary embodiment, the p child nodes in steps 340 and 440 include the child nodes of a point cloud slice (a slice is formed by a cubic volume with the same coordinates), the child nodes of a tube, and the child nodes at the first child node of another tube located in another point cloud slice. The number p can be:

[0133] p = N slice + N tube + 1,

[0134] where, N slice represents the number of child nodes in a point cloud slice, and N tube represents the number of child nodes in a tube.

[0135] Figure 11 Shows an example of 43 child nodes, including 36 child nodes of the point cloud slice PCS (gray cubic volume), 6 child nodes of the tube TU, and 1 child node CN at the first child node of another tube located in another point cloud slice.

[0136] In one exemplary embodiment, the p child nodes in steps 540 and 640 include the child nodes of nodes that have the same x in the list L d-1 as shown in Figure 12 and as shown in Figure 12As shown, the light gray cubes represent the children of the nodes with x = 0, the medium gray cubes represent the children of the nodes with x = 1, the cube CU represents the child node currently undergoing Trisoup encoding / decoding, and the dark gray cube represents the child node currently undergoing octree encoding / decoding. Therefore, the delay is for two slices of the children of the nodes with the same x.

[0137] In Figure 6 and Figure 7 In the exemplary embodiments of method 300 and method 400, the encoding / decoding of Trisoup information uses the same list L as used for encoding occupancy information d-1 to iterate through list L d-1 for the children of the nodes in

[0138] In one variant, the node information of the child nodes for which occupancy information is encoded / decoded is stored in list (buffer) L tris and can subsequently be used for encoding / decoding the Trisoup information of the child nodes by iterating through the children of the nodes in list L tris

[0139] Figure 8 Schematically illustrates a block diagram of a variant of method 300.

[0140] In step 510, method 300 encodes the occupancy information of the nodes at the first d levels of depth (from depth 0 to depth d - 1) in the first bit sequence BSoct1.

[0141] In step 520, list L d-1 is obtained. List L d-1 stores the nodes at depth (d - 1) sorted in the encoding order according to the coordinates of these nodes in the three - dimensional system.

[0142] In step 530, a child node index k is initialized for encoding the Trisoup information. For example, k = 0.

[0143] In step 540, consider the nodes at depth (d - 1) with the same coordinates in the three - dimensional system, for example, the nodes at depth (d - 1) with the same x - coordinate (e.g., x = 0). These nodes are sorted in the first encoding order, and their children (at depth d) are sorted in the second encoding order. The occupancy information of these children is encoded in the bit sequence BSoct1 ( Figure 5 ). The children of the nodes with the same coordinates are added to list L tris . These children are sorted in the second encoding order in list L tris .

[0144] In step 550, for list L​tris For the first child node of the node (k = 0), the Trisoup information is encoded into the bit sequence BStris1. At the same time, the octree information of the first child node of the node with the same coordinate x = 1 (the coordinate x of the node with a depth of d - 1 plus 1) is encoded in the bit sequence BSoct2.

[0145] In step 560, it is done in an interleaved manner: the Trisoup information of the next child node of the node with the updated same coordinate x = 1 is encoded into the second bit sequence BStris2; and the occupancy information of the next child node (k = k + 1) of the list L tris is encoded into the bit sequence BSoct2, and the above interleaved encoding is iterated until the encoding of the occupancy information of the last child node (the Nmax-th child node) of the nodes in the list L d-1 is completed.

[0146] In each iteration, when the encoding of the occupancy information of the child nodes of the nodes with the same coordinates ends, the coordinate is incremented by 1 (updated) to encode the occupancy information of the child nodes of the nodes with this updated coordinate.

[0147] In step 570, the Trisoup information of the remaining child nodes (Nmax - k) of the list L tris is encoded in the bit sequence BStris3.

[0148] The data unit DU ends with the bit sequence (10000).

[0149] Figure 9 A block diagram schematically illustrates a variant of method 400.

[0150] After obtaining the data unit DU of the bitstream B related to the information of the partition of the point cloud data, in step 610, method 400 decodes the occupancy information of the nodes at the first d layers of depth (from depth 0 to depth d - 1) from the bit sequence BSoct1.

[0151] In step 620, the list L d-1 is obtained. The list L d-1 stores the nodes at depth (d - 1) sorted in the decoding order according to the coordinates of these nodes in the three - dimensional system.

[0152] In step 630, a child node index k is initialized for the decoding of the Trisoup information. For example, k = 0.

[0153] In step 640, consider the nodes with the same coordinates in a three-dimensional system at depth (d - 1), for example, the nodes at depth (d - 1) with the same x coordinate (e.g., x = 0). These nodes are sorted in a first decoding order, and their child nodes (at depth d) are sorted in a second decoding order. Decode the occupancy information ( Figure 5 ) of these child nodes from the bit sequence BSoct1. Add the child nodes of the nodes with the same x coordinate to the list L tris . These child nodes are sorted in the second decoding order in the list L tris .

[0154] In step 650, for the first child node (k = 0) of the nodes in the list L tris , decode the Trisoup information from the bit sequence BStris1. At the same time, decode the octree information of the first child node of the nodes with the updated same coordinate x = 1 from the bit sequence BSoct2.

[0155] In step 660, proceed in an interleaved manner: decode the Trisoup information of the next child node of the nodes with the updated same coordinate x = 1 from the bit sequence BStris2 of the data unit DU; and decode the occupancy information of the next child node (k = k + 1) from the bit sequence BSoct2 of the data unit DU until the decoding of the occupancy information of the last child node (Nmax-th) of the nodes in the list L d-1 .

[0156] In each iteration, when the decoding of the occupancy information of the nodes with the same coordinate ends, the coordinate is incremented by 1 (updated) to decode the occupancy information of the nodes with this updated coordinate.

[0157] In step 670, decode the Trisoup information of the remaining child nodes (Nmax - k tris ) of the list L from the bit sequence BStris3. )

[0158] The data unit DU ends with a bit sequence (e.g., 10000).

[0159] In an exemplary embodiment, entropy encoding / decoding is performed on the occupancy information and / or the Trisoup information.

[0160] In an exemplary embodiment, the encoding order and / or the decoding order is the Morton order or the raster scan order.

[0161] In an exemplary embodiment, the first encoding order and / or the first decoding order and / or the second encoding order and / or the second decoding order is the Morton order or the raster scan order.

[0162] The last two exemplary embodiments include: the encoding order and the decoding order are the same.

[0163] Figure 10 An example of a raster scan order of a complete octree structure in a global cube that includes all points of a point cloud is illustrated.

[0164] The cube that includes all points of the point cloud is divided into (a plurality of) cube volumes. A point cloud slice s is formed by cube volumes having the same coordinate (e.g., y coordinate) in a three-dimensional coordinate system (x, y, z). Each point cloud slice s is divided into a set of cube volumes defined as having the same coordinate (e.g., x). Scanning the nodes of the octree structure according to the raster scan order includes independently scanning each pipeline of each point cloud slice s. Raster scanning a point cloud slice includes: considering a first cube volume cs of a first pipeline T1 of the point cloud slice, considering subsequent cube volumes of pipeline T1 until a last cube volume ce. Next, considering a first cube volume of another pipeline T2 of the point cloud slice, and next considering subsequent cube volumes of pipeline T2 until the last cube volume of pipeline T2, and so on.

[0165] Figure 13 A schematic block diagram showing an example of a system in which various aspects and exemplary embodiments are implemented is shown.

[0166] System 500 may be embedded as one or more devices, including various components described below. In various embodiments, system 500 may be configured to implement one or more aspects described in the present application.

[0167] Examples of equipment that may form all or part of system 500 include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, set-top boxes, and any other device for processing point clouds, videos, or images or other communication devices. The elements of system 500 may be implemented singly or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 500 may be distributed across multiple ICs and / or discrete components. In various embodiments, system 500 may be communicatively connected to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.

[0168] System 500 may include at least one processor 510 configured to execute instructions loaded therein for implementing various aspects described, for example, in the present application. The processor 510 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 500 may include at least one memory 520 (e.g., volatile memory devices and / or non-volatile memory devices). System 500 may include a storage device 540, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 540 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0169] System 500 may include an encoder / decoder module 530 configured to, for example, process data to provide encoded / decoded point cloud geometry data, and the encoder / decoder module 530 may include its own processor and memory. The encoder / decoder module 530 may represent the (one or more) modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 530 may be implemented as a separate element of system 500 or may be incorporated within the processor 510 as a combination of hardware and software known to those skilled in the art.

[0170] The program code to be loaded into the processor 510 or encoder / decoder 530 to execute the various aspects described in the present application may be stored in the storage device 540 and subsequently loaded into the memory 520 for execution by the processor 510. According to various embodiments, during the execution of the processes described in the present application, one or more of the processor 510, memory 520, storage device 540, and encoder / decoder module 530 may store one or more of various items. Such stored items may include but are not limited to point cloud frames, encoded / decoded geometry / attribute video / images or portions thereof, bitstreams, matrices, variables, and intermediate or final results of equations, formulas, operations, and operational logic processing.

[0171] In several embodiments, the memory internal to the processor 510 and / or encoder / decoder module 530 may be used to store instructions and provide working memory for the processing that may be performed during encoding or decoding.

[0172] However, in other embodiments, a memory external to the processing device (e.g., the processing device can be the processor 510 or the encoder / decoder module 530) is used for one or more of these functions. The external memory can be the memory 520 and / or the storage device 540, e.g., dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM can be used as a working memory for video encoding and decoding operations, e.g., for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), or MPEG-I Part 5 or Part 9.

[0173] As indicated in block 590, input to the elements of the system 500 can be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section that can receive RF signals transmitted, for example, by a broadcast device over the air, (ii) composite input terminals, (iii) USB input terminals, and / or (iv) HDMI input terminals.

[0174] In various embodiments, the input devices of block 590 have associated respective input processing elements, as known in the art. For example, the RF section can be associated with elements necessary for (i) selecting a desired frequency (also known as selecting a signal or band-limiting a signal to a band), (ii) down-converting the selected signal, (iii) band-limiting the signal again to a narrower band to select a signal band that can be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments can include one or more elements that perform these functions, e.g., a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section can include a tuner that performs various functions among these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband.

[0175] In one set-top box embodiment, the RF section and its associated input processing elements can receive an RF signal transmitted over a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band.

[0176] Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0177] Adding elements can include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter, for example. In various embodiments, the RF section can include an antenna.

[0178] In addition, the USB and / or HDMI terminals can include corresponding interface processors for connecting the system 500 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented, for example, within a separate input processing IC or within the processor 510 when necessary. Similarly, various aspects of USB or HDMI interface processing can be implemented within a separate interface IC or within the processor 510 when necessary. The demodulated, error-corrected, and demultiplexed stream can be provided to various processing elements, including, for example, the processor 510 and the encoder / decoder 530, which operate in conjunction with memory and storage elements to process the data stream as necessary for presentation on an output device.

[0179] Various elements of the system 500 can be provided within an integrated housing. Within the integrated housing, a suitable connection arrangement 590, such as internal buses (including I2C buses), wiring, and printed circuit boards known in the art, can be used to interconnect the various elements and transfer data between them.

[0180] The system 500 can include a communication interface 550 that enables communication with other devices via a communication channel 900. The communication interface 550 can include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 900. The communication interface 550 can include, but is not limited to, a modem or a network card, and the communication channel 900 can be implemented, for example, within a wired and / or wireless medium.

[0181] In various embodiments, data can be streamed to the system 500 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments can be received via the communication channel 900 and the communication interface 550 suitable for Wi-Fi communication. The communication channel 900 of these embodiments can generally be connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-top communications.

[0182] Other embodiments can use a set-top box to provide streamed data to the system 500, which delivers the data via an HDMI connection to the input block 590.

[0183] There are other embodiments that can use the RF connection of input box 590 to provide the streamed data to system 500.

[0184] The streamed data can be used as a way of signaling information used by system 500. The signaling information can include the bitstream B and / or information such as the number of points, coordinates, and / or sensor setting parameters of a point cloud.

[0185] It should be recognized that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. can be used to signal information to the corresponding decoder.

[0186] System 500 can provide output signals to various output devices, including display 600, speaker 700, and other peripheral devices 800. In various examples of the embodiments, other peripheral devices 800 can include one or more of a stand-alone DVR, disc player, stereo system, lighting system, and other devices based on the output providing function of system 500.

[0187] In various embodiments, control signals can be communicated between system 500 and display 600, speaker 700, or other peripheral devices 800 using signaling of communication protocols such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention.

[0188] The output devices can be communicatively connected to system 500 via dedicated connections through corresponding interfaces 560, 570, and 580.

[0189] Alternatively, the output devices can be connected to system 500 via communication interface 550 using communication channel 900. Display 600 and speaker 700 can be integrated into a single unit in an electronic device (such as, for example, a television) with other components of system 500.

[0190] In various embodiments, display interface 560 can include a display driver, such as, for example, a timing controller (TCon) chip.

[0191] For example, if the RF part of input terminal 390 is part of a separate set-top box, then display 600 and speaker 700 can optionally be separate from one or more of the other components. In various embodiments where display 600 and speaker 700 can be external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0192] In Figure 1-13In this context, various methods are described herein, and each method includes one or more steps or actions to implement the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.

[0193] Some examples are described with respect to block diagrams and / or operational flowcharts. Each block represents a circuit element, module, or portion of code that includes one or more executable instructions for implementing the specified logical function(s). It should also be noted that in other embodiments, the function(s) labeled in a block may not occur in the order indicated. For example, depending on the functions involved, two consecutive blocks shown may actually be executed substantially concurrently, or sometimes the blocks may be executed in the reverse order.

[0194] The embodiments and aspects described herein can be implemented, for example, in a method or process, apparatus, computer program, data stream, bit stream, or signal. Even if discussed only in the context of a single form of embodiment (e.g., only as a method), the embodiments of the features discussed can be implemented in other forms (e.g., an apparatus or computer program).

[0195] A method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor also includes a communication device.

[0196] Furthermore, a method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the embodiments) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product implemented in one or more computer-readable media and having computer-readable program code executable by a computer implemented thereon. Considering the inherent ability to store information therein and the inherent ability to retrieve information therefrom, a computer-readable storage medium as used herein can be considered a non-transitory storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. It should be recognized that the following, while providing more specific examples of computer-readable storage media to which this embodiment can be applied, is merely illustrative and not exhaustive as would be readily recognized by a person of ordinary skill in the art: a portable computer floppy disk; a hard disk; a read-only memory (ROM); an erasable programmable read-only memory (EPROM or flash memory); a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing.

[0197] Instructions can form an application program tangibly embodied on a processor-readable medium.

[0198] For example, the instructions can be in hardware, firmware, software, or a combination. For example, the instructions can be found in an operating system, a separate application, or a combination of both. Thus, a processor can be characterized as, for example, a device configured to execute a process and a device that includes a processor-readable medium (such as a storage device) having instructions for executing the process. Additionally, in addition to or instead of the instructions, the processor-readable medium can store data values generated by an implementation.

[0199] The apparatus can be implemented in, for example, suitable hardware, software, and firmware. Examples of such an apparatus include a personal computer, a laptop computer, a smart phone, a tablet computer, a digital multimedia set-top box, a digital television receiver, a personal video recording system, a connected household appliance, a head-mounted display device (HMD, see-through glasses), a projector (projector), a "cave" (a system including multiple displays), a server, a video encoder, a video decoder, a post-processor that processes the output from the video decoder, a pre-processor that provides input to the video encoder, a web server, a set-top box, and any other device for processing point clouds, videos, or images, or other communication devices. It should be clear that the equipment can be mobile and even installed in a moving vehicle.

[0200] Computer software can be implemented by the processor 510 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can also be implemented by one or more integrated circuits. The memory 520 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). The processor 510 can be of any type suitable for the technical environment and can encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.

[0201] As will be apparent to those of ordinary skill in the art, the embodiments can generate various signals that are formatted to carry information such as can be stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, the signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted over various different wired or wireless links. The signal can be stored on a processor-readable medium.

[0202] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" may also be intended to include the plural forms, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "include / comprise" and / or "including / comprising" may specify the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. Also, when an element is referred to as being "responsive to" or "connected to" another element, it can be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly responsive to" or "directly connected to" another element, no intervening elements are present.

[0203] It should be recognized that, for example, in the case of "A / B", "A and / or B" and "at least one of A and B", the use of any of the symbols / terms " / ", "and / or" and "at least one of" can be intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B and / or C" and "at least one of A, B and C", such wording is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.

[0204] In this application, various numerical values can be used. Specific values can be used for illustrative purposes and the aspects described are not limited to these specific values.

[0205] It will be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of this application, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. No ordering is implied between the first element and the second element.

[0206] References to "an exemplary embodiment" or "exemplary embodiments" or "an embodiment" or "embodiments" and other variations thereof are frequently used to convey that a particular feature, structure, characteristic, etc. (described in connection with the embodiment / embodiment) is included in at least one embodiment / embodiment. Thus, the appearances of the phrases "in an exemplary embodiment" or "in exemplary embodiments" or "in an embodiment" or "in embodiments" and any other variations thereof throughout this application do not necessarily all refer to the same embodiment.

[0207] Similarly, references herein to "according to an exemplary embodiment / example / embodiment" or "in an exemplary embodiment / example / embodiment" and other variations thereof are frequently used to convey that a particular feature, structure, or characteristic (described in connection with the exemplary embodiment / example / embodiment) may be included in at least one exemplary embodiment / example / embodiment. Thus, the expressions "according to an exemplary embodiment / example / embodiment" or "in an exemplary embodiment / example / embodiment" appearing throughout this specification do not necessarily all refer to the same exemplary embodiment / example / embodiment, nor are the individual or alternative exemplary embodiments / example / embodiments necessarily mutually exclusive of other exemplary embodiments / example / embodiments.

[0208] The reference numerals appearing in the claims are for illustrative purposes only and have no limiting effect on the scope of the claims. Although not explicitly described, the embodiments / examples and variations can be employed in any combination or sub - combination.

[0209] When a figure is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0210] Although some figures include arrows on communication paths to indicate the main direction of communication, it should be understood that communication can occur in a direction opposite to that depicted by the arrows.

[0211] Various embodiments relate to decoding. As used in this application, "decoding" can cover, for example, all or part of a process performed on a received point cloud frame (which may include a received bitstream encoding one or more point cloud frames) to produce a final output suitable for display or further processing in a reconstructed point cloud domain. In various embodiments, such processes include one or more of the processes typically performed by a decoder. In various embodiments, for example, such processes also include or optionally include processes performed by the decoders of the various embodiments described in this application.

[0212] As a further example, in one embodiment "decoding" can refer only to dequantization, in one embodiment "decoding" can refer to entropy decoding, in another embodiment, "decoding" can refer only to differential decoding, and in another embodiment, "decoding" can refer to a combination of dequantization, entropy decoding, and differential decoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process, and it is believed that this will be well understood by those skilled in the art.

[0213] Various embodiments also relate to encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application can cover, for example, all or part of a process performed on an input point cloud frame to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder. In various embodiments, such processes also include or optionally include processes performed by the encoders of the various embodiments described in this application.

[0214] As a further example, in one embodiment "encoding" can refer only to quantization, in one embodiment "encoding" can refer only to entropy encoding, in another embodiment, "encoding" can refer only to differential encoding, and in another embodiment, "encoding" can refer to a combination of quantization, differential encoding, and entropy encoding. Based on the context of the specific description, it will be clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process, and it is believed that this will be well understood by those skilled in the art.

[0215] In addition, this application may refer to "acquiring" each piece of information. Acquiring information can include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from a memory.

[0216] In addition, this application may refer to "accessing" each piece of information. Accessing information can include, for example, one or more of the following: receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0217] In addition, this application may refer to "receiving" each piece of information. Like "access", "receiving" is intended to be a broad term. Receiving information may include, for example, one or more of the following: accessing information or retrieving information (e.g., from a memory). In addition, "receiving" is typically involved in processes such as storing information, processing information, transmitting information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information in one way or another.

[0218] Moreover, as used herein, the word "signal" particularly refers to indicating something to a corresponding decoder, etc. For example, in certain embodiments, an encoder signals specific information, such as the number of points or coordinates of points in a point cloud or sensor setting parameters, etc. In this way, in an embodiment, the same parameter can be used on the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicitly signal) a specific parameter to a decoder such that the decoder can use the same specific parameter. Conversely, if the decoder already has a specific parameter and other parameters, then signaling can be used without transmission (implicitly signal) to simply allow the decoder to know and select the specific parameter. By avoiding transmitting any actual functionality, bit savings are achieved in various embodiments. It should be recognized that signaling can be accomplished in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing refers to the verb form of the word "signal", the word "signal" can also be used as a noun herein.

[0219] Multiple embodiments have been described. However, it should be understood that various modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to produce other embodiments. In addition, those of ordinary skill in the art will understand that other structures and processes can replace the disclosed structures and processes, and the resulting embodiments will perform at least substantially the same (one or more) functions in at least substantially the same (one or more) ways to achieve at least substantially the same (one or more) results as the disclosed embodiments. Thus, this application contemplates these and other embodiments.

Claims

1. A method for encoding the position information of points in a point cloud into a bitstream comprising at least one data unit, wherein the points are contained in a cubic volume associated with a leaf node of an octree structure having a maximum depth d, - at least three vertices are located on the edges of each said cubic volume, at most one vertex on each edge and the position information of the points contained in the cubic volume is represented by a triangle connecting the at least three vertices; The method comprises: - encoding the occupancy information of the nodes of the octree structure, the occupancy information representing the presence of the points of the point cloud contained in the cubic volume; and for each current cubic volume containing at least one point of the point cloud, encoding Trisoup information, the Trisoup information representing the presence of vertices on the edges of the current cubic volume and the vertex positions along the edges; wherein the occupancy information of a leaf node and the Trisoup information of another leaf node are encoded in an interleaved manner in the data unit (DU).

2. A method for decoding the position information of points in a point cloud from a bitstream comprising at least one data unit, wherein the points are contained in a cubic volume associated with a leaf node of an octree structure having a maximum depth d, - at least three vertices are located on the edges of each said cubic volume, at most one vertex on each edge and the position information of the points contained in the cubic volume is represented by a triangle connecting the at least three vertices; The method comprises: - decoding the occupancy information of the nodes of the octree structure, the occupancy information representing the presence of the points of the point cloud contained in the cubic volume; and for each current cubic volume containing at least one point of the point cloud, decoding Trisoup information, the Trisoup information representing the presence of vertices on the edges of the current cubic volume and the vertex positions along the edges; wherein the occupancy information of a leaf node and the Trisoup information of other leaf nodes are encoded in an interleaved manner in the data unit (DU).

3. The method according to claim 1 or 2, wherein the method further comprises obtaining (320, 420, 520, 620) a first list (L d-1 ), the first list (L d-1 ) storing nodes of depth d-1 sorted in encoding order or decoding order according to the coordinates of the nodes in the three-dimensional system; and encoding (340) occupancy information of p child nodes of depth d of the nodes in the first list (L d-1 ) into the bit sequence BSoct1 of the data unit (DU) / decoding (440) occupancy information of p child nodes of depth d of the nodes in the first list (L d-1 ) from the bit sequence BSoct1 of the data unit (DU).

4. The method according to claim 3, wherein the method further comprises, for a first child node (k = 0) of a first node in the first list (L d-1 ), encoding (350) Trisoup information into a bit sequence BStris1 of the data unit (DU) / decoding (450) Trisoup information from the bit sequence BStris1 of the data unit (DU).

5. The method according to claim 4, wherein the method further comprises, for the child nodes of the nodes of the first list (L d-1 ), alternately: Encoding (360) the occupancy information into the bit sequence BSoct2 of the data unit (DU) / decoding (460) the occupancy information from the bit sequence BSoct2 of the data unit (DU); and Encoding (360) the Trisoup information into the bit sequence BStris2 of the data unit (DU) / decoding (460) the Trisoup information from the bit sequence BStris2 of the data unit (DU), up to the encoding / decoding of the occupancy information of the last child node (Nmax-th child node) of the nodes in the first list (L d-1 ).

6. The method according to claim 5, wherein the method further comprises encoding / decoding (370, 470) the Trisoup information of the remaining child nodes (p = Nmax - k) added to the bit sequence BStris3 of the data unit (DU).

7. The method according to claim 3, wherein the method further comprises: Obtain (540, 640) a second list (L tris ), where the second list (L tris ) stores the child nodes of the nodes with a depth of d - 1, the nodes have the same coordinates of the three-dimensional system, the nodes are sorted in a first encoding order or a first decoding order, and the child nodes are sorted in a second encoding order or a second decoding order.

8. The method according to claim 7, wherein the method further comprises, for a first child node (k = 0) of a node in the second list (L tris ) in the bit sequence BStris1 of the data unit (DU), encoding / decoding (550, 650) Trisoup information; and encoding / decoding octree information of a first child node of a node having updated same coordinates (x = 1) in the bit sequence BSoct1 of the data unit (DU), the updated same coordinates being equal to the same coordinates plus 1.

9. The method according to claim 8, wherein the method further comprises, in an interleaved manner: Encode (560) the Trisoup information of the next child node of the node with the same coordinates having the update into the bit sequence BStris2 of the data unit (DU) / decode (660) from the bit sequence BStris2 of the data unit (DU) the Trisoup information of the next child node of the node with the same coordinates having the update; and Encode (560) the occupancy information of the next child node (k = k + 1) into the bit sequence BSoct2 of the data unit (DU) / decode (660) from the bit sequence BSoct2 of the data unit (DU) the occupancy information of the next child node (k = k + 1), until the encoding / decoding of the occupancy information of the last child node (Nmax-th) of the nodes in the first list (L d-1 ).

10. The method according to claim 9, wherein the method further comprises encoding (570) the Trisoup information of the remaining child nodes (Nmax-k) of the second list (L tris ) into the bit sequence BStris3 of the data unit (DU) / decoding (670) the Trisoup information of the remaining child nodes (Nmax-k) of the second list (L tris ) from the bit sequence BStris3 of the data unit (DU).

11. The method according to any one of claims 3 to 7, wherein the encoding order and / or the decoding order and / or the first encoding order and / or the first decoding order and / or the second encoding order and / or the second decoding order is a Morton order or a raster scan order.

12. A bitstream formatted to contain the encoded positions of the points of a point cloud obtained from the method according to any one of claims 1, 3 to 11.

13. An apparatus comprising means for performing one of the methods according to any one of claims 1 to 11.

14. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 11.

15. A non-transitory storage medium carrying instructions for a program code for performing the method according to any one of claims 1 to 11.