Point cloud compression using space-filling curves for level-of-detail generation

By using spatial fill curves and attribute prediction correction technology in point cloud encoder, the problem of high storage and transmission costs of point cloud data is solved, and the effect of fast storage and real-time transmission is achieved.

CN113272866BActive Publication Date: 2025-05-13APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080008331.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-07
Filing Date
2020-01-08
Publication Date
2025-05-13
Estimated Expiration
2040-01-08

AI Technical Summary

Technical Problem

The prior art is costly and time-consuming when storing and transmitting large amounts of point cloud data, making it difficult to achieve real-time or almost real-time data transmission.

Method used

By configuring the attribute information of the compressed point cloud in the encoder of the point cloud, the detailed level structure is determined using the spatial fill curve, the attribute information is predicted and corrected, the compressed attribute information is generated, and the decompression is performed in the decoder.

Benefits of technology

It effectively reduces the storage and transmission cost of point cloud data, realizes the rapid storage and transmission of point cloud data, and supports real-time or almost real-time data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113272866B_ABST
    Figure CN113272866B_ABST
Patent Text Reader

Abstract

The present invention discloses a system, which includes an encoder configured to compress attribute information for a point cloud and / or a decoder configured to decompress the compressed attribute information. Attribute values ​​for at least one starting point are included in a compressed attribute information file, and attribute correction values ​​are included in the compressed attribute information file. The order of these points is determined based on a space filling curve, wherein the encoder and the decoder determine the same order of these points based on the space filling curve. The level of detail is determined by sampling these ordered points according to different sampling parameters, and the attribute values ​​for these points in these levels of detail are predicted using the determined order. The encoder determines the attribute correction values ​​based on a comparison of these predicted values ​​with the initial values ​​before compression. The decoder corrects the predicted attribute values ​​based on the received attribute correction values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to compression and decompression of point clouds, which include a plurality of points, each point having associated attribute information.

[0002] Related technical description

[0003] Various types of sensors (such as light detection and ranging (LIDAR) systems, 3D cameras, 3D scanners, etc.) can capture data indicating the location of points in three-dimensional space (e.g., locations in the X, Y, and Z planes). In addition, such systems can capture attribute information in addition to spatial information for corresponding points, such as color information (e.g., RGB values), intensity attributes, reflectivity attributes, motion-related attributes, modal attributes, or various other attributes. In some cases, additional attributes can be assigned to corresponding points, such as a timestamp when the point was captured. The points captured by such sensors can constitute a "point cloud", which includes a set of points each having associated spatial information and one or more associated attributes. In some cases, a point cloud can include thousands of points, hundreds of thousands of points, millions of points, or even more points. In addition, in some cases, a point cloud can be generated, for example, in software, unlike a point cloud being captured by one or more sensors. In either case, such a point cloud can include a large amount of data, and storing and transmitting these point clouds can be costly and time-consuming. Summary of the invention

[0004] In some embodiments, a system includes one or more sensors configured to capture points that collectively constitute a point cloud, wherein each of the points includes spatial information identifying a spatial position of the corresponding point in 3D space and attribute information defining one or more attributes associated with the corresponding point. The system also includes an encoder configured to compress the attribute information for the points. To compress the attribute information, the encoder is configured to assign attribute values ​​to at least one point of the point cloud based on the attribute information included in the captured point cloud. In addition, the encoder is configured to: for each of the corresponding other points of the point cloud, identify a set of neighboring points; determine a predicted attribute value for the corresponding point based at least in part on the predicted or assigned attribute values ​​for the neighboring points; and determine an attribute correction value for the point based at least in part on comparing the predicted attribute value for the corresponding point with the attribute information for the point included in the captured point cloud. The encoder is further configured to encode the compressed attribute information for the point cloud, wherein the compressed attribute information includes the assigned attribute value for the at least one point and data indicating the corresponding determined attribute correction value for the corresponding other points.

[0005] In some embodiments, in order to compress the attribute information, the encoder is configured to construct a hierarchical level of detail (LOD) structure for the point cloud. For example, the encoder may be configured to determine points to be included in a first level of detail of the compressed attribute information for the point cloud, and to determine points to be included in one or more additional levels of detail of the compressed attribute information for the point cloud. In order to determine the points to be included in the first level of detail or the points to be included in the one or more additional levels of detail, the encoder is configured to determine the order of these points of the point cloud based on a space filling curve, wherein the corresponding points of the point cloud are assigned to indices that index these corresponding points based on the proximity of these corresponding points to the position along the space filling curve. Further as part of determining the points to be included in the first level of detail or the points to be included in one or more additional levels of detail, the encoder is configured to sample the index according to one or more sampling rates to determine the points of the point cloud to be included in the first level of detail or the one or more additional levels of detail. In addition, the encoder is configured to compress the attribute information of the points determined to be included in the first level of detail, and to compress the attribute information of the points determined to be included in the one or more additional levels of detail. For example, the encoder may generate attribute correction values ​​for points determined to be included in a first level of detail and for points determined to be included in one or more additional levels of detail. In some embodiments, the prediction and correction process described above may be used to determine attribute correction values ​​for corresponding sets of points in corresponding levels of detail determined to be included in a level of detail.

[0006] In some embodiments, the decoder is configured to receive compressed attribute information for a point cloud, the compressed attribute information comprising at least one assigned attribute value for at least one point of the point cloud and data indicating attribute correction values ​​for attributes of other points of the point cloud. In some embodiments, the attribute correction values ​​may be sorted at multiple levels of detail for multiple subsets of points of the point cloud. For example, the decoder may receive a compressed point cloud compressed by an encoder as described above. The decoder may be further configured to provide decompressed attribute information for a first level of detail, and to update a decompressed version of the point cloud to include attribute information for additional subsets of points at one or more other levels of detail in the multiple levels of detail.

[0007] In some embodiments, in order to decompress the attribute information, the decoder is configured to receive spatial information for points of a point cloud, and to receive compressed attribute information for one or more levels of detail of the point cloud. The decoder is further configured to determine the points to be included in one or more levels of detail for the point cloud based on the order of the points of the point cloud determined according to the space filling curve, wherein the corresponding points of the point cloud are assigned to indices, which index the corresponding points based on the proximity of the corresponding points to the position along the space filling curve, and the indexes are sampled according to one or more sampling rates to determine the points of the point cloud to be included in the one or more levels of detail. In addition, the decoder is configured to determine the attribute values ​​of the points determined to be included in the one or more levels of detail based on the compressed attribute information received for the points of the corresponding one or more levels of detail of the point cloud. For example, the decoder may predict the attribute values ​​for the points included in a given level of detail, and then apply the attribute correction values ​​to the predicted values, wherein the attribute correction values ​​are included in the compressed attribute information for the given level of detail.

[0008] In some embodiments, a non-transitory computer readable medium may store program instructions that, when executed by one or more processors, cause the one or more processors to determine a level of detail of a level of detail structure based on applying a space filling curve as described herein. Additionally, the program instructions may cause the one or more processors to encode or decode attribute information of a point cloud using the determined level of detail structure as described herein.

[0009] In some embodiments, the method includes determining a level of detail based on applying a space filling curve as described herein. The method may also include encoding or decoding attribute information of the point cloud using the level of detail structure as described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1A A system including a sensor that captures information for points of a point cloud and an encoder that compresses attribute information and / or spatial information of the point cloud, wherein the compressed point cloud information is sent to a decoder, is shown according to some embodiments.

[0011] Figure 1B A process for determining points to be included in a level of detail (LOD) for compressing / encoding attribute information of a point cloud is shown according to some embodiments.

[0012] Figure 1C to Figure 1D A process for compressing and encoding attribute information for points of a point cloud, the points of which are selected to be included in a corresponding level of detail (LOD) for the point cloud, is shown according to some embodiments.

[0013] Figure 2AA process of determining and signaling a sampling rate to be used to determine points to be included in a level of detail (LOD) for a point cloud is shown in accordance with some embodiments.

[0014] Figure 2B An exemplary index generated by applying a space-filling curve to a point cloud is shown, and sampling of the index to determine points to be included in a corresponding level of detail (LOD) for the point cloud is also shown, according to some embodiments.

[0015] Figure 3A A process of determining and signaling a sampling order to be used to determine points to be included in a level of detail (LOD) for a point cloud is shown in accordance with some embodiments.

[0016] Figure 3B An exemplary index generated by applying a space-filling curve to a point cloud is shown according to some embodiments, and sampling of the index in a forward sampling order to determine points of the point cloud to be included in a corresponding level of detail (LOD) for the point cloud is also shown.

[0017] Figure 3C An exemplary index generated by applying a space-filling curve to a point cloud is shown according to some embodiments, and sampling of the index in a reverse sampling order to determine points of the point cloud to be included in a corresponding level of detail (LOD) for the point cloud is also shown.

[0018] Figure 3D An exemplary index generated by applying a space filling curve to a point cloud is shown according to some embodiments, and the index is sampled in an interior out order to determine points of the point cloud to be included in a corresponding level of detail (LOD) for the point cloud.

[0019] Figure 4A A process of determining and signaling sampling offset values ​​to be used to determine points to be included in a level of detail (LOD) for a point cloud is shown in accordance with some embodiments.

[0020] Figure 4B An exemplary index generated by applying a space filling curve to a point cloud is shown according to some embodiments, and sampling of the index using different sampling offset values ​​to determine points to be included in a corresponding level of detail (LOD) for the point cloud is also shown.

[0021] Figure 5 A process for determining sampling parameters for points to be used in determining a point cloud to be included in a corresponding level of detail (LOD) for the point cloud is shown in accordance with some embodiments.

[0022] Figure 6 An example of interleaving attribute correction values ​​between different levels of detail (LODs) for a point cloud is shown according to some embodiments.

[0023] Figure 7 An exemplary process for determining points to be included in a level of detail (LOD) for a point cloud and for compressing attribute information for points determined to be included in the LOD is shown according to some embodiments.

[0024] Figure 8 An exemplary process for decompressing compressed attribute information for points included in a level of detail (LOD) of a point cloud is illustrated in accordance with some embodiments.

[0025] Fig.9A Exemplary components of an encoder according to some embodiments are shown.

[0026] Fig. 9B Exemplary components of a decoder according to some embodiments are shown.

[0027] Fig.10 An example compression properties file including compression property information for multiple levels of detail (LODs) of a compressed point cloud is shown according to some embodiments.

[0028] Fig.11 An exemplary process for compressing attribute information for level of detail (LOD) based on nearest neighbor prediction and prediction-correction techniques is shown in accordance with some embodiments.

[0029] FIG. 12A to FIG. 12B An exemplary process for compressing spatial information of a point cloud is shown according to some embodiments.

[0030] Fig.13 Another exemplary process for compressing spatial information of a point cloud is shown in accordance with some embodiments.

[0031] Fig.14 An exemplary process for decompressing attribute information for one or more levels of detail (LOD) based on nearest neighbor prediction and prediction-correction techniques is shown in accordance with some embodiments.

[0032] Fig.15 An exemplary encoder for generating compressed attribute information for multiple levels of detail (LODs) of a compressed point cloud is shown according to some embodiments.

[0033] Fig.16 An exemplary level of detail (LOD) structure is shown according to some embodiments.

[0034] Fig.17 Compressed point cloud information is shown being used in a 3D application according to some embodiments.

[0035] Fig.18 Compressed point cloud information is shown being used in a virtual reality application according to some embodiments.

[0036] Fig.19 An exemplary computer system is shown that may implement an encoder or decoder according to some embodiments.

[0037] This specification includes references to "one embodiment" or "an embodiment." The appearance of the phrase "in one embodiment" or "in an embodiment" does not necessarily refer to the same embodiment. The particular features, structures or characteristics may be combined in any suitable manner consistent with the present disclosure.

[0038] The term "comprising" is open ended. As used in the appended claims, the term does not exclude additional structures or steps. Consider the following cited claim: "an apparatus comprising one or more processor units..." Such a claim does not exclude the apparatus from including additional components (e.g., a network interface unit, a graphics circuit, etc.).

[0039] "Configured to", various units, circuits or other components may be described or described as "configured to" perform one or more tasks. In such contexts, "configured to" is used to imply a structure (e.g., a circuit) that performs the one or more tasks during operation by indicating that the unit / circuit / component includes the structure. In this way, the unit / circuit / component is said to be configured to perform the task even when the specified unit / circuit / component is currently inoperable (e.g., not turned on). The units / circuits / components used with the "configured to" language include hardware-such as circuits, memories storing executable program instructions to implement operations, etc. Reference to a unit / circuit / component "configured to" perform one or more tasks is explicitly intended not to invoke 35 U.S.C. §112 (f) for the unit / circuit / component. In addition, "configured to" may include a general structure (e.g., a general circuit) manipulated by software and / or firmware (e.g., an FPGA or a general processor executing software) to operate in a manner capable of performing one or more tasks to be solved. "Configured to" may also include adapting a manufacturing process (eg, a semiconductor fabrication facility) to produce a device (eg, an integrated circuit) suitable for implementing or performing one or more tasks.

[0040] "First," "second," etc. As used herein, these terms act as labels for the nouns that precede them, and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.). For example, a buffer circuit may be described herein as performing a write operation of a "first" value and a "second" value. The terms "first" and "second" do not necessarily imply that the first value must be written before the second value.

[0041] "Based on". As used herein, this term is used to describe one or more factors that influence a determination. This term does not exclude additional factors that influence a determination. That is, a determination may be based solely on these factors or at least in part on these factors. Consider the phrase "A is determined based on B". In this case, B is a factor that influences the determination of A, and such a phrase does not exclude that the determination of A may also be based on C. In other examples, A may be determined based solely on B. DETAILED DESCRIPTION

[0042] As data acquisition and display technologies become more advanced, the ability to capture point clouds comprising thousands or tens of thousands of points in 2D or 3D space has increased (such as via LIDAR systems). Moreover, the development of advanced display technologies (such as virtual reality or augmented reality systems) has increased the potential uses of point clouds. However, point cloud files are typically very large, and storing and transmitting these point cloud files can be costly and time consuming. For example, communication of point clouds over private networks or public networks (such as the Internet) can require considerable amounts of time and / or network resources, such that some uses of the point cloud data (such as real-time uses) may be limited. In addition, the storage requirements of the point cloud files may consume a significant amount of storage capacity of the device storing the point cloud files, which may also limit potential applications for using the point cloud data.

[0043] In some embodiments, an encoder may be used to generate a compressed point cloud to reduce the cost and time associated with storing and transmitting large point cloud files. In some embodiments, the system may include an encoder that compresses the attribute information and / or spatial information (also referred to herein as geometric information) of a point cloud file so that the point cloud file can be stored and transmitted faster than a non-compressed point cloud, and in a manner that the point cloud file occupies less storage space than a non-compressed point cloud. In some embodiments, compression of the spatial information and / or attributes of points in a point cloud enables the point cloud to be transmitted over a network in real time or near real time. For example, a system may include a sensor that captures spatial information and / or attribute information about points in an environment where the sensor is located, wherein the captured points and corresponding attributes constitute a point cloud. The system may also include an encoder that compresses the attribute information of the captured point cloud. The compressed attribute information of the point cloud may be sent to a decoder that decompresses the compressed attribute information of the point cloud in real time or near real time over a network. The decompressed point cloud may be further processed, for example, to make control decisions based on the surrounding environment at the sensor location. The control decision can then be transmitted back to a device at or near the sensor location, where the device receiving the control decision implements the control decision in real time or almost in real time. In some embodiments, the decoder can be associated with an augmented reality system, and the decompressed attribute information can be displayed or otherwise used by the augmented reality system. In some embodiments, compressed attribute information about point clouds can be sent together with compressed spatial information about the points of point clouds. In other embodiments, spatial information and attribute information can be encoded and / or sent to decoders respectively. In some embodiments, an encoder or decoder can be implemented in hardware via a processing circuit specially designed to perform encoding / decoding. In some embodiments, a device or system may include one or more processors and a memory storing program instructions, which enable the one or more processors to implement an encoder or decoder.

[0044] In some embodiments, the system may include a decoder that receives one or more point cloud files including compressed attribute information from a remote server or other storage device storing one or more point cloud files via a network. For example, a 3D display, a holographic display, or a head-mounted display may be manipulated in real time or near real time to display different parts of a virtual world represented by a point cloud. To update the 3D display, holographic display, or head-mounted display, a system associated with the decoder may request point cloud files from a remote server based on user manipulation of the display, and these point cloud files may be transmitted from the remote server to the decoder and decoded by the decoder in real time or near real time. The display may then be updated with updated point cloud data (such as updated point attributes) in response to the user manipulation.

[0045] In some embodiments, a system may include one or more LIDAR systems, 3D cameras, 3D scanners, etc., and such sensor devices may capture spatial information, such as the X, Y, and Z coordinates of points in the sensor device's field of view. In some embodiments, the spatial information may be relative to a local coordinate system or may be relative to a global coordinate system (e.g., a Cartesian coordinate system may have fixed reference points such as fixed points on the earth, or may have non-fixed local reference points such as sensor locations).

[0046] In some embodiments, such sensors may also capture attribute information about one or more points, such as color attributes, reflectivity attributes, velocity attributes, acceleration attributes, time attributes, modalities, and / or various other attributes. In some embodiments, in addition to LIDAR systems, 3D cameras, 3D scanners, etc., other sensors may capture attribute information to be included in the point cloud. For example, in some embodiments, a gyroscope or accelerometer may capture motion information to be included in the point cloud as an attribute associated with one or more points of the point cloud. For example, a vehicle equipped with a LIDAR system, a 3D camera, or a 3D scanner may include the direction and rate of the vehicle in the point cloud captured by the LIDAR system, the 3D camera, or the 3D scanner. For example, when points in the field of view of the vehicle are captured, these points may be included in a point cloud, wherein the point cloud includes the captured points and the associated motion information corresponding to the state of the vehicle when the points are captured.

[0047] In some embodiments, the attribute information may include string values, such as different modalities. For example, the attribute information may include string values ​​indicating modalities such as "walking", "running", "driving", etc. In some embodiments, the encoder may include a "string value" to an integer index, where certain strings are associated with certain corresponding integer values. In some embodiments, the point cloud may indicate the string value for the point by including an integer associated with the string value as an attribute of the point. Both the encoder and the decoder may store a common string value as an integer index, so that the decoder can determine the string value for the point based on looking up the integer value of the string attribute of the point in the decoder's string value to integer index that matches or is similar to the decoder's string value to integer index.

[0048] In some embodiments, in addition to compressing the attribute information of the attributes of the points of the point cloud, the encoder also compresses and encodes the spatial information of the point cloud to compress the spatial information. For example, in order to compress the spatial information, a KD tree can be generated, wherein the corresponding number of points in each cell of the cell included in the KD tree is encoded. This sequence of coded point counts can encode the spatial information of the points of the point cloud. In addition, in some embodiments, subsampling and prediction methods can be used to compress and encode the spatial information for the point cloud. In some embodiments, the spatial information can be quantized before being compressed and encoded. In addition, in some embodiments, the compression of the spatial information can be lossless. Therefore, the decoder may be able to determine the same spatial information view as the encoder. In addition, once the compressed spatial information is decoded, the encoder may be able to determine the spatial information view that the decoder will encounter. Because both the encoder and the decoder may have or be able to reconstruct the same spatial information for the point cloud, the spatial relationship can be used to compress the attribute information for the point cloud.

[0049] For example, in many point clouds, the attribute information between adjacent points or points located relatively short distances from each other may have a high level of correlation between the attributes, and therefore, the differences in the point attribute values ​​are relatively small. For example, adjacent points in a point cloud may have relatively small differences in color when considered relative to points that are further apart in the point cloud.

[0050] In some embodiments, the encoder may include a predictor that determines a predicted attribute value for an attribute of a point in the point cloud based on attribute values ​​of similar attributes for adjacent points in the point cloud and based on the corresponding distance between the point being evaluated and the adjacent points. In some embodiments, the attribute values ​​of the attributes of adjacent points closer to the point being evaluated can be given a higher weight than the attribute values ​​of the attributes of adjacent points farther away from the point being evaluated. In addition, the encoder can compare the predicted attribute value with the attribute value of the attribute for the point in the original point cloud before compression. A residual difference, also referred to as an "attribute correction value" in this article, can be determined based on the comparison. The attribute correction value can be encoded and included in the compressed attribute information for the point cloud, wherein the decoder uses the encoded attribute correction value to correct the predicted attribute value for the point, wherein the attribute value is predicted at the decoder using the same or similar prediction method as the prediction method used at the encoder.

[0051] In some embodiments, the encoder can assign attribute values ​​to the starting point of the point cloud to be used as the starting point in the evaluation order. The encoder can predict the attribute value for the next closest point to the starting point based on the attribute value of the starting point and the distance between the starting point and the next closest point. The encoder can then determine the difference between the predicted attribute value for the next closest point and the actual attribute value for the next closest point included in the uncompressed original point cloud. The difference can be encoded in a compressed attribute information file as an attribute correction value for the next closest point. The encoder can then repeat a similar process for each point in the evaluation order. In order to predict the attribute values ​​for subsequent points in the evaluation order, the encoder can identify the K nearest neighboring points as the specific points being evaluated, wherein the identified K nearest neighboring points have been assigned or predicted attribute values. In some embodiments, "K" can be a configurable parameter transmitted from the encoder to the decoder.

[0052] The encoder can determine the distances in X, Y, and Z space between the point being evaluated and each identified neighboring point. For example, the encoder can determine the corresponding Euclidean distances from the point being evaluated to each of the neighboring points. The encoder can then predict attribute values ​​for the attributes of the point being evaluated based on the attribute values ​​of the neighboring points, where the attribute values ​​of the neighboring points are weighted according to the inverse of the distances from the point being evaluated to corresponding ones of the neighboring points. Attribute values ​​of neighboring points that are closer to the point being evaluated can be given greater weight than attribute values ​​of neighboring points that are farther away from the point being evaluated.

[0053] In a similar manner as described for the first neighboring point, the encoder can compare the predicted value for each of the other points in the point cloud with the actual attribute value in the original non-compressed point cloud (e.g., the captured point cloud). The difference can be encoded as an attribute correction value for the attribute of one of the other points being evaluated. In some embodiments, the attribute correction values ​​can be encoded in a compressed attribute information file in order according to an evaluation order determined based on a space filling curve. Because the encoder and decoder can determine the same evaluation order based on spatial information for the point cloud, the decoder can determine which attribute correction value corresponds to which attribute of which point based on the order in which the attribute correction values ​​are encoded in the compressed attribute information file. In addition, the starting point and one or more attribute values ​​of the starting point can be explicitly encoded in the compressed attribute information file so that the decoder can determine the evaluation order starting from the same point used to start the evaluation order at the encoder. In addition, the one or more attribute values ​​of the starting point can provide the value of the neighboring point, and the decoder uses the value of the neighboring point to determine the predicted attribute value for the point being evaluated, which is a neighboring point of the starting point.

[0054] In some embodiments, the encoder may determine predicted values ​​for attributes of points based on temporal considerations. For example, in addition to or instead of determining predicted values ​​based on neighboring points in the same "frame" (e.g., the same time point as the point being evaluated), the encoder may consider attribute values ​​of points in adjacent and subsequent time frames.

[0055] Figure 1A A system is shown according to some embodiments including a sensor that captures information for points of a point cloud and an encoder that compresses attribute information of the point cloud, wherein the compressed attribute information is sent to a decoder.

[0056] The system 100 includes a sensor 102 and an encoder 104. The sensor 102 captures a point cloud 110 that includes points representing a structure 106 in a view 108 of the sensor 102. For example, in some embodiments, the structure 106 may be a mountain, a building, a sign, the environment surrounding a street, or any other type of structure. In some embodiments, a captured point cloud (such as captured point cloud 110) may include spatial information and attribute information about the points included in the point cloud. For example, point A of the captured point cloud 110 includes X, Y, Z coordinates and attributes 1, 2, and 3. In some embodiments, the attributes of the point may include attributes such as R, G, B color values, speed at the point, acceleration at the point, reflectivity of the structure at the point, a timestamp indicating when the point was captured, a string value indicating the modality when the point was captured, such as "walking", or other attributes. The captured point cloud 110 may be provided to the encoder 104, where the encoder 104 generates a compressed version of the point cloud (compressed attribute information 112), which is transmitted to the decoder 116 via the network 114. In some embodiments, a compressed version of a point cloud (such as compressed attribute information 112) may be included in a common compressed point cloud that also includes compressed spatial information for points of the point cloud, or in some embodiments, the compressed spatial information and compressed attribute information may be transmitted as separate files.

[0057] In some embodiments, the encoder 104 can be integral to the sensor 102. For example, the encoder 104 can be implemented in hardware or software included in a sensor device such as the sensor 102. In other embodiments, the encoder 104 can be implemented on a separate computing device adjacent to the sensor 102.

[0058] Low complexity level of detail generation process

[0059] In some embodiments, an encoder (such as encoder 104) may utilize a level of detail generation process to determine multiple levels of detail for a point cloud (such as point cloud 106). Instead of predicting attribute values ​​and determining attribute correction values ​​for each point of the point cloud before encoding the determined attribute correction values, the encoder may instead predict attribute values ​​and determine attribute correction values ​​for points included in a first level of detail, and encode the determined attribute correction values ​​before determining attribute correction values ​​for all other points of the point cloud. The encoder may then predict attribute values ​​and determine attribute correction values ​​for other points included in other additional levels of detail of the point cloud, and encode these determined attribute correction values ​​after encoding the attribute correction values ​​for points included in lower levels of detail. This approach may allow for faster compression, encoding, and transmission of representations of lower levels of detail of a point cloud than is possible with a full point cloud. The representations of the lower levels of detail may then be supplemented with compressed attribute information for the additional levels of detail. This in turn may enable a decoder to more quickly reconstruct representations of lower levels of detail of the point cloud, and later supplement the representations of the lower levels of detail to include more detail by adding attribute values ​​for points included in other additional levels of detail.

[0060] In some embodiments, a level of detail (LOD) structure divides a point cloud into non-overlapping subsets of points called refinement levels, e.g., l ) l=0…L-1 In some embodiments where a distance-based approach is used to determine the level of detail refinement, the level of refinement may be determined based on a set of Euclidean distances (d l ) l=0…L-1 To determine, to some extent, that the entire point cloud is represented by the union of all refinement levels. The current level of detail l, or (LOD) l , by taking the refinement levels R0, R1, ..., R l The union of is used to obtain:

[0061] LOD0 = R0

[0062] LOD1=LOD0 U R1…

[0063] LOD l =LOD (l-1) UR l …

[0064] LOD (L-1) =LOD (L-2) UR (L) Represents the entire point cloud

[0065] In some embodiments, where points to be included in a particular level of detail are selected based on the distance between the points, the points in each level of refinement are extracted in such a way that the Euclidean distance between these points in that particular LOD is greater than or equal to a user defined threshold D. As the level of detail l increases, D decreases and more points are included between points in lower LODs, thereby increasing the point cloud reconstruction detail. (l-1) The k nearest neighbor points in predict R l Finally, an entropy encoder (eg, an arithmetic encoder) is used to encode the prediction residual (eg, attribute correction value) (ie, the difference between the actual value and the predicted value of the attribute).

[0066] The distance-based LOD generation process as described above attempts to ensure uniform sampling across different LODs. This strategy provides valid prediction results for smooth attribute signals defined on uniformly or nearly uniformly sampled point clouds. However, this may lead to poor prediction results for non-smooth attribute signals defined on irregularly sampled point clouds. In addition, the distance-based LOD generation process requires calculating the distance between each point and its neighboring points, which may be complex in practice for certain use case scenarios.

[0067] In some embodiments, as an alternative to the distance-based LOD generation process for encoding attribute values, a low-complexity LOD generation process that uses a space-filling curve to sort points and determine the level of refinement can be used. The low-complexity LOD generation process using a space-filling curve can achieve more effective prediction of non-smooth attribute signals defined on irregularly sampled point clouds. In some embodiments, various additional features can be combined with a low-complexity LOD generation process that uses a space-filling curve to sort points and determine the level of refinement, such as a combined sorting / sampling LOD method, an adaptive scanning mode, an adaptive offset mode, an attribute interleaving mode, an attribute inter / cross-component prediction mode, a prediction adaptation mode, and / or a related LOD encoder optimization mode.

[0068] LOD generation using space filling curves

[0069] In some embodiments, a low-complexity LOD generation process that uses a space-filling curve to sort points and determine the level of refinement can be used. Spatial information can be encoded using any technique for encoding spatial information, such as KD tree, octree encoding, subsampling, and inter-point prediction. In this way, both the encoder and the decoder can know the spatial position of the points of the point cloud. However, instead of determining which points are to be included in the corresponding level of refinement based on the distance between the points as described above, the points to be included in the corresponding level of detail can be determined by sorting them according to the position of the points along the space-filling curve. For example, these points can be organized according to their Morton codes. Alternatively, other space-filling curves can be used. For example, a technique that maps a position (e.g., in the form of X, Y, Z coordinates) to a space-filling curve such as Morion order (or Z order), Hillbert curve, Peano curve, etc. can be used. In this way, all points of the point cloud encoded and decoded using spatial information can be organized into indexes in the same order at the encoder and decoder.

[0070] For example, Figure 1B A process for determining points to be included in a level of detail (LOD) for encoding attribute information for a point cloud is shown according to some embodiments.

[0071] At 152, the encoder receives a point cloud to be compressed, such as point cloud 162. The received point cloud may be a captured point cloud, such as a point cloud captured by sensor 102, or may be a point cloud generated in software, such as a 3D software environment.

[0072] At 154, the encoder generates a 3D space-filling curve and applies the 3D space-filling curve to the received point cloud to determine corresponding positions of points of the received point cloud along the space-filling curve. For example, space-filling curve 164 may be generated, and points of point cloud 162 may be mapped to the closest position of space-filling curve 164 to a corresponding point of the point cloud.

[0073] At 156 , the encoder determines an index for the point cloud that indexes points of the point cloud based on proximity of the corresponding points to a position along the space-filling curve. For example, index 166 shows points of point cloud 162 ordered in index positions 1-6 to N of index 166 .

[0074] At 158, the determined indices are sampled at a specified or known sampling rate to determine points of the point cloud 162 to be included in the first level of detail (e.g., the first refinement level). For example, the sampled indices 168 illustrate that the indices 166 are sampled at a rate of "every 2 positions" to determine points of the point cloud 162 to be included in the first refinement level / level of detail.

[0075] At 160, the determined indices are sampled at a second specified or known sampling rate to determine points of the point cloud 162 to be included in a second level of refinement, which, when combined with points of a lower level of detail (e.g., the first level of detail), constitute a second level of detail that is more detailed than the lower level of detail. For example, the sampled indices 170 illustrate the indices 166 sampled at a rate of "every 3 positions" to determine points of the point cloud 162 to be included in the second level of refinement / level of detail.

[0076] Figure 1C to Figure 1D A process for compressing and encoding attribute information for points of a point cloud, the points of which are selected to be included in a corresponding level of detail (LOD) for the point cloud, is shown according to some embodiments.

[0077] At 180, an encoder (such as encoder 104) selects a determined level of detail for which attribute information is to be compressed. For example, 192 shows an exemplary set of points of point cloud 162 that have been selected to be included in a first level of detail of point cloud 162. Each point in the point cloud shown in 192 may have one or more attributes associated with the point. Note that for ease of illustration, point cloud 192 is shown in 2D, but the point cloud may include points in 3D space.

[0078] At 182, an evaluation order for the points of the point cloud is determined. The evaluation order may be determined based on the corresponding index positions of the points in the index 166 or the sample index 168. For example, the evaluation order may be based on the Morton order.

[0079] In addition, in some embodiments, the minimum spanning tree can be determined based on the spatial information of the point cloud received by the encoder. In order to determine the minimum spanning tree, the minimum spanning tree generator of the encoder can select the starting point for the minimum spanning tree. The minimum spanning tree generator can then identify the points adjacent to the starting point. The adjacent points can then be sorted based on the corresponding distances between the corresponding identified adjacent points and the starting point. The adjacent point with the shortest distance from the starting point can be selected as the next point to be visited. The "weight" of the "edge" can be determined for the edge between the starting point and the adjacent point selected as the next to be visited, for example, the distance between the points in the point cloud, where a longer distance is given a greater weight than a shorter distance. After the adjacent point closest to the starting point is added to the minimum spanning tree, the adjacent point can then be evaluated, and the points adjacent to the point currently being evaluated (e.g., previously selected as the next point to be visited) can be identified. The identified adjacent points can be sorted based on the corresponding distances between the point currently being evaluated and the identified adjacent points. The neighboring point (e.g., an "edge") with the shortest distance to the point currently being evaluated can be selected as the next point to be included in the minimum spanning tree. The weight of the edge between the point currently being evaluated and the next selected neighboring point can be determined and added to the minimum spanning tree. A similar process can be repeated for each of the other points in the point cloud to generate a minimum spanning tree for the point cloud.

[0080] For example, 194 shows an illustration of a minimum spanning tree. In the minimum spanning tree shown in 194, each vertex can represent a point in a point cloud, and the edge weights (e.g., 1, 2, 3, 4, 7, 8, etc.) between vertices can represent the distance between points in the point cloud. For example, the distance between vertex 191 and vertex 193 can have a weight of 7, while the distance between vertex 193 and 195 can have a weight of 8. This can indicate that the distance in the point cloud between the point corresponding to vertex 193 and the point corresponding to vertex 195 is greater than the distance in the point cloud between the point corresponding to vertex 193 and the point corresponding to vertex 191. In some embodiments, the weights shown in the minimum spanning tree can be based on vector distances in 3D space, such as Euclidean distances.

[0081] At 184, attribute values ​​of one or more attributes for a starting point (such as a starting point for generating a minimum spanning tree or a starting point determined by a space filling curve index) may be assigned to be encoded and included in the compressed attribute information for the point cloud. As discussed above, predicted attribute values ​​for points of a point cloud may be determined based on attribute values ​​of neighboring points. However, the initial attribute values ​​for at least one point are provided to the decoder so that the decoder may determine attribute values ​​for other points using at least the initial attribute values ​​and attribute correction values ​​for correcting predicted attribute values ​​predicted based on the initial attribute values. Thus, one or more attribute values ​​for at least one starting point are explicitly encoded in the compressed attribute information file. Additionally, spatial information for the starting point may be explicitly encoded so that a minimum spanning tree generator of a space filling curve or decoder may determine which of the points of the point cloud is to be used as a starting point. In some embodiments, the starting point may be indicated in other ways besides explicitly encoding spatial information for the starting point, such as marking the starting point or other point identification methods.

[0082] Because the decoder will receive an indication of the starting point and will encounter the same or similar spatial information for points in the point cloud space as the encoder, the decoder can determine the same evaluation order from the same starting point as determined by the encoder.

[0083] At 186, for the current point being evaluated, the prediction / correction evaluator of the encoder determines the predicted attribute value for the attribute of the point currently being evaluated. In some embodiments, the point currently being evaluated may have more than one attribute. Therefore, the prediction / correction evaluator of the encoder can predict more than one attribute value for the point. For each point being evaluated, the prediction / correction evaluator can identify a group of nearest neighboring points that have been assigned or predicted attribute values. In some embodiments, the number "K" of the identified neighboring points can be a configurable parameter of the encoder, and the encoder can include configuration information indicating the parameter "K" in the compressed attribute information file so that when performing attribute prediction, the decoder can identify the same number of neighboring points. The prediction / correction evaluator can then use the weight from the minimum spanning tree, or can otherwise determine the distance between the point being evaluated and the corresponding ones of the identified neighboring points. The prediction / correction evaluator can use the inverse distance interpolation method to predict the attribute value for each attribute of the point being evaluated. The prediction / correction evaluator can then predict the attribute value of the point being evaluated based on the average value of the inverse distance weighted attribute values ​​of the identified neighboring points.

[0084] For example, 196 shows a point (X, Y, Z) being evaluated, where attribute A1 is determined based on the inverse distance weighted attribute values ​​of eight identified neighboring points.

[0085] At 188, an attribute correction value is determined for each point. The attribute correction value is determined based on comparing the predicted attribute value for each attribute of the point with the corresponding attribute value of the point in the original non-compressed point cloud (such as the point included in the selected LOD). For example, 198 shows a formula for determining the attribute correction value, where the captured value is subtracted from the predicted value to determine the attribute correction value. Note that although Figure 1C The attribute value predicted at 186 and the attribute correction value determined at 188 are shown, but in some embodiments, the attribute correction value for the point can be determined after the attribute value for the point is predicted. The next point can then be evaluated, wherein the predicted attribute value for the point is determined, and the attribute correction value for the point is determined. Therefore, 186 and 188 can be repeated for each point being evaluated. In other embodiments, the predicted values ​​for multiple points can be determined, and the attribute correction value can then be determined. In some embodiments, the prediction for the subsequent point being evaluated can be based on the predicted attribute value or can be based on the corrected attribute value or based on both. In some embodiments, both the encoder and the decoder can follow the same rules about whether to determine the predicted value of the subsequent point based on the predicted or corrected attribute value.

[0086] Note also that for higher levels of detail (e.g., levels of detail greater than the first level of detail), attribute values ​​determined for points in lower levels of detail may be used to predict attribute values ​​for points in higher levels of detail. For example, when predicting attribute values ​​for points included in level of detail two (which includes points in level of detail one plus points in an additional level of detail), points in level of detail one for which attribute values ​​have been determined may be selected as neighboring points for points in level of detail two for which attribute values ​​are being predicted.

[0087] At 190, the determined attribute correction values ​​for the points at the selected level of detail (or refinement level) are encoded. Additionally, in some embodiments, one or more assigned attribute values ​​for the starting point, spatial information or other markings for the starting point, and any configuration information to be included in the compressed attribute information file may be encoded. In some embodiments, various encoding methods such as arithmetic coding and / or Golomb coding may be used to encode the attribute correction values, assigned attribute values, and configuration information.

[0088] Changing sampling rate

[0089] To determine the various refinement levels, a sampling rate for the ordered index of points may be defined. For example, to divide a point cloud into four levels of detail, the index that maps Morton values ​​to corresponding points may be sampled, for example, at a rate of four, where every third index point is included in the lowest level of refinement. For each additional refinement level, the remaining points in the index that have not yet been sampled may be sampled (e.g., every third index point, etc.) until all points are sampled to obtain the highest level of detail. For example, a low complexity LOD generation process that uses a space filling curve to sort the points and determine the refinement level may be performed as follows:

[0090] First, the points of the point cloud (points with known spatial information) can be sorted according to the space filling curve. For example, the points can be sorted according to their Morton codes.

[0091] Then, assuming I L-1 is a set of ordered indexes, and LOD L-1 is the associated LOD representing the entire point cloud.

[0092] Next, define (k l ) l=0...L-1 A set of sampling rates, where k l An integer describing the sampling rate used for LOD1.

[0093] OK l It may be automatically determined based on characteristics of the signal and / or point cloud distribution, previous statistics, or may be fixed.

[0094] OK l Can be provided as a user-defined parameter (e.g., 4).

[0095] oSampling rate k l can be further updated within the LOD to better adapt to the point cloud distribution. More precisely, the encoder can be updated for the latest available k l A predefined set of points of values ​​(eg, each consecutive H=1024 points) explicitly encodes different values ​​or updates in the bitstream.

[0096] Next, through I l+1 Perform subsampling and at every k l An index is reserved from the indexes to calculate the ordered array of indices associated with LOD1=L-2, L-3, ..., 0, denoted as I l .

[0097] In some embodiments, different subsampling rates may be defined per attribute (eg, color, reflectance), per channel (eg, Y and U / V), etc.

[0098] Figure 2A A process of determining and signaling a sampling rate to be used to determine points to be included in a level of detail (LOD) of a point cloud is shown in accordance with some embodiments.

[0099] At 202, the encoder may determine one or more sampling rates to be used to determine points to be included in corresponding levels of refinement / levels of detail. In some embodiments, a rate-distortion optimization process may be used to determine the sampling rates to be used for corresponding levels of refinement / levels of detail. In some embodiments, the encoder may utilize known sampling rates known to the encoder, or may signal one or more selected sampling rates in the compressed bitstream at 204. In some embodiments, the encoder may utilize known or implied sampling rates that are known or can be inferred by the decoder, and may only signal sampling rates for specific LODs that deviate from the known or implied sampling rates.

[0100] Figure 2B An exemplary index generated by applying a space-filling curve to a point cloud is shown, and sampling of the index to determine points to be included in a corresponding level of detail (LOD) of the point cloud is also shown, according to some embodiments.

[0101] For example, in Figure 2B In the example, index 250 is sampled at a rate of every 2 positions to determine points to be included in a first level of refinement corresponding to a first level of detail. Additionally, index 250 is sampled at a rate of every 3 positions to determine points to be included in a second level of refinement, which points in the second level of refinement, when combined with the points included in the first level of refinement, constitute the second level of detail. Additionally, index 250 is sampled at a rate of every N-1 positions to determine points to be included in an Nth level of refinement, which points in the level of refinement, when combined with the points included in the previous level of refinement, constitute the Nth level of detail.

[0102] Note that in some embodiments, the index sampled according to the sampling rate for a given level of detail may be a raw index determined based on a space filling curve, or may be an index that includes points not yet included in a refinement level and excludes points already included in a refinement level. For example, in some embodiments, every N-1 positions corresponding to positions of points not yet included in level of detail 1 or level of detail 2 may be sampled to determine points to be included in level of detail N. Conversely, in some embodiments, when index 250 is sampled, all points / positions of that index may be retained, and the sampling rate / sampling offset value may be selected so that points included in a lower level of detail are not repeated in a later level of refinement.

[0103] In some embodiments, the prediction between detail levels can also be used to determine predicted attribute values ​​for points at various detail levels.As previously described, attribute correction values ​​can be encoded for points, where these attribute correction values ​​represent the difference between the predicted value and the original or pre-compression value of the attribute.

[0104] Adaptive scanning mode

[0105] In some embodiments, as described above, instead of using a single prediction order, the prediction mode in the low complexity LOD generation process using the space filling curve can achieve improved coding efficiency by selecting the prediction direction that gives improved rate distortion performance for a given refinement level. For example, in one case, the first level of detail can use the Morton order, and in another case, the reverse Morton order can be used. In another case, the traversal of these points can start from the center or any point explicitly signaled by the encoder. The scanning order can also be explicitly encoded in the bitstream or agreed upon between the encoder and the decoder. For example, the encoder can explicitly signal to the decoder that a point should be skipped and processed at a later time. The signaling of the mode and its associated parameters can be completed at the sequence / frame / tile / slice / LOD / point group level. In some embodiments, different sampling orders can be selected for different levels of detail.

[0106] Figure 3A A process of determining and signaling a sampling order to be used to determine points to be included in a level of detail (LOD) of a point cloud is shown in accordance with some embodiments.

[0107] At 302, the encoder determines one or more sampling orders to be used for determining points to be included in one or more levels of detail / refinement. In some embodiments, a rate-distortion optimization process may be used to determine the sampling order to be used for the corresponding level of refinement / level of detail. In some embodiments, the encoder may utilize a known sampling order known to the encoder, or may signal one or more selected sampling orders in the compressed bitstream at 304. In some embodiments, the encoder may utilize a known or implied sampling order that is known or can be inferred by the decoder, and may only signal a sampling order for a particular LOD that deviates from the known or implied sampling order.

[0108] Figure 3B An exemplary index generated by applying a space-filling curve to a point cloud is shown according to some embodiments, and sampling of the index in a forward sampling order to determine points of the point cloud to be included in a corresponding level of detail (LOD) of the point cloud is also shown.

[0109] For example, Figure 3BAn index 350 is shown sampled in a forward sampling order at a sampling rate of every 2 positions.

[0110] Figure 3C An exemplary index generated by applying a space-filling curve to a point cloud is shown according to some embodiments, and sampling of the index in a reverse sampling order to determine points of the point cloud to be included in a corresponding level of detail (LOD) of the point cloud is also shown.

[0111] For example, Figure 3C The index 370 is shown to be sampled in reverse sampling order at a sampling rate of every 2 positions. Figure 3B The reverse sampling order samples the index positions starting from the opposite end of the index compared to the case of the forward sampling order shown. Note also that in some embodiments, sampling the same index at the same sampling rate according to different sampling orders may result in different points being selected for inclusion in a given refinement level / level of detail. For example, the point determined to be included in the first level of detail based on sampling index 370 in the reverse sampling order is different from the point selected when index 350 is sampled in the forward sampling order.

[0112] Figure 3D An exemplary index generated by applying a space-filling curve to a point cloud is shown according to some embodiments, and sampling of the index in an inside-outside sampling order to determine points of the point cloud to be included in a corresponding level of detail (LOD) of the point cloud is also shown.

[0113] For example, Figure 3D An index 390 is shown sampled in an inside-outside sampling order at a sampling rate of every 2 positions. Note that the inside-outside sampling order samples the index position starting at the inside position of the index and continues to sample the index at every 2 positions in either direction from the inside position. This also results in a different set of points selected to be included in the first level of detail than the points selected according to the forward sampling order and the reverse sampling order. In some embodiments, the inside position starting point may be the center position in the index or may be offset from the center.

[0114] In some embodiments, an adaptive scan mode may be used, where a higher LOD allows its samples to be predicted from the current lower LOD samples. Of course, this may affect the decoding process for the higher level LOD (e.g., limiting its parallelization capabilities). However, parallelization can still be achieved by defining "independent" decoding groups within the LOD. Such groups can allow parallel decoding by not allowing predictions across them. However, predictions may be allowed using decoded samples within the lower level LOD as well as the current LOD group.

[0115] In some embodiments, the encoder can select an appropriate adaptive scan mode by utilizing a rate-distortion optimization (RDO) strategy. In some embodiments, the encoder can also consider various additional criteria, such as computational complexity, battery life, memory requirements, delays, pre-analysis, collected statistics (history) of past frames, user feedback, etc.

[0116] Adaptive Scan Shift Mode

[0117] In some embodiments, another mode that can be applied in the low complexity LOD generation process using space filling curves can be an "alternating sampling phase / offset" mode. For example, the encoder can signal a sampling offset value to the decoder, which is used to select which points should be sampled from the index when generating the next LOD level. For example, a sampling offset value of 1 can provide better rate distortion (RD) performance than an offset value of 0. For example, instead of sampling ordered Morton codes starting with the first Morton code of the first point, sampling can start from an offset value, such as the second Morton code, the third Morton code, etc. This may have an impact on performance because it may change the encoding and prediction process for each LOD. The signaling of the sampling offset can be done at the sequence / frame / tile / slice / LOD / point group level.

[0118] Figure 4A A process of determining and signaling sampling offset values ​​to be used to determine points to be included in a level of detail (LOD) of a point cloud is shown in accordance with some embodiments.

[0119] At 402, the encoder determines one or more sampling offset values ​​to be used to determine points to be included in one or more levels of level of detail / level of refinement. In some embodiments, a rate-distortion optimization process may be used to determine the sampling offset values ​​to be used for the corresponding level of refinement / level of detail. In some embodiments, the encoder may utilize known sampling offset values ​​known to the encoder, or may signal one or more selected sampling offset values ​​in the compressed bitstream at 304. In some embodiments, the encoder may utilize known or implied sampling offset values ​​that are known or can be inferred by the decoder, and may only signal sampling offset values ​​for a specific LOD that deviate from the known or implied sampling offset values. In some embodiments, in the absence of an implied or signaled sampling offset value, no offset may be applied.

[0120] Figure 4B An exemplary index generated by applying a space-filling curve to a point cloud is shown, and sampling of the index using different sampling offset values ​​to determine points to be included in a corresponding level of detail (LOD) of the point cloud is also shown in accordance with some embodiments.

[0121] For example, Figure 4BThe index 450 is shown sampled using a sampling rate of every third position and a sampling offset value of zero to determine points to be included in a first level of detail. To determine points to be included in a second level of refinement / level of detail, the same sampling rate of every third position is used, but a sampling offset value of one is applied. To determine points to be included in a third level of refinement / level of detail, the same sampling rate of every third position is used, but a sampling offset value of two is applied. It can be seen that even when the same sampling rate is applied, the sampling offset value can be used to shift which points at a particular index position of the index are selected to be included in a given level of detail.

[0122] In some embodiments, various combinations of sampling rates, sampling orders, and / or sampling offset values ​​may be selected to apply to determining points to be included in various levels of detail of the point cloud being compressed.

[0123] In some embodiments, a rate-distortion optimization process may be followed to test various combinations of sampling rates, sampling orders, and / or sampling offset values ​​to determine a set of sampling parameters that balances the desire to maximize compression efficiency and the desire to minimize distortion.

[0124] Figure 5 A process for determining sampling parameters to be used to determine points to be included in corresponding levels of detail (LOD) of a point cloud is shown in accordance with some embodiments.

[0125] At 502, a rate-distortion optimization module of an encoder selects an initial combination of sampling parameters to be applied to select points to be included in a corresponding level of detail for a compressed point cloud. For example, various combinations of sampling rates, sampling orders, sampling offset values, etc. may be selected for determining points to be included in a corresponding level of detail for a compressed point cloud.

[0126] At 504, the rate-distortion optimization module may determine the compression efficiency of the selected sampling parameter combination. For example, when the corresponding points are organized into detail levels as determined according to the initial combination of sampling parameters, the rate-distortion optimization module may determine the number of bits required to compress the property values ​​for these points. In addition, the rate-distortion optimization module may determine the number of bits required to signal the sampling parameters. For example, some sampling parameters may be known to the decoder, so that only exceptions need to be signaled, so selecting sampling parameters that need to be signaled (rather than default or known sampling parameters) may require more bits.

[0127] At 506 , the rate-distortion optimization module determines one or more distortion metrics for a reconstructed version of the point cloud, where the reconstructed version of the point cloud is reconstructed using compressed attribute information determined for the points organized into levels of detail based on the initial combination of sampling parameters (e.g., the same sampling parameters applied at 504 ).

[0128] At 508, the rate-distortion optimization module selects additional combinations of sampling parameters to be applied to select points to be included in the corresponding level of detail for the compressed point cloud. For example, various combinations of sampling rates, sampling orders, sampling offset values, etc. may be selected for determining points to be included in the corresponding level of detail for the compressed point cloud. The combination of sampling parameters selected at 508 includes at least one or more sampling parameter values ​​that are different from the initial combination of sampling parameters selected at 502.

[0129] At 510, the rate-distortion optimization module determines a compression efficiency when applying the combination of sampling parameters selected at 508. Additionally, at 512, the rate-distortion optimization module determines one or more distortion metrics when applying the combination of sampling parameters selected at 508.

[0130] Although Figure 5 The rate-distortion optimization is shown testing at least two combinations of sampling parameters to perform the optimization, but in some embodiments, any number of combinations of sampling parameters may be tested.

[0131] In some embodiments, a minimum rate-distortion optimization threshold may also be applied, such as a minimum compression efficiency or a maximum acceptable distortion level. For example, at 514, it may be determined whether the RDO threshold is met. If not, the process may return to 508, and the next combination of sampling parameters may be tested. However, if the RDO threshold is met at 514, then at 516, the best performing combination of sampling parameters (e.g., lowest distortion / maximum compression efficiency) of the tested combinations may be selected for determining the points to be included in the corresponding LOD for compressing the attribute information of the compressed point cloud.

[0132] Attribute interleaving mode

[0133] In some embodiments, when predicting and encoding / decoding residual data (e.g., attribute correction values) for each LOD level, the attributes / attribute channels of the point cloud may be interleaved at different levels. Specifically, the interleaving may be done at the following locations:

[0134] Point level,

[0135] ·Peer group level,

[0136] LOD level,

[0137] Slice level, and / or

[0138] Frame level.

[0139] In some embodiments, interleaving can be completed at the attribute channel level. For example, color channels can be interwoven, and other attributes are only interwoven at the LOD level. However, in some embodiments, different combinations can be used for different types of attribute data. Interleaving methods can be fixed and known between encoder and decoder, but can also be adaptive and can be signaled at different levels of bit streams. The decision-making of the method used can be based on rate-distortion (RD) standards, pre-analysis, past coding statistics, encoding / decoding complexity or some other standards determined by users or systems.

[0140] Figure 6 An example of interleaving attribute correction values ​​between different levels of detail (LODs) of a point cloud is shown in accordance with some embodiments.

[0141] Index 600 is shown as being sampled at a rate of every 2 positions, with the first, second, and third detail levels being generated with sampling offset values ​​of 0, 1, and 2, respectively. However, instead of encoding the red, green, and blue channel values ​​for each point included in the corresponding detail level, the red channel value is encoded for the first detail level, the green channel value is encoded for the second detail level, and the blue channel value is encoded for the third detail level. In some embodiments, the decoder may predict attribute correction values ​​not included in a given detail level based on attribute correction values ​​interleaved for other detail levels. For example, for a point included in detail level two, the decoder may directly decode the green channel attribute correction value, and may predict the red attribute correction value based on the red channel attribute correction value ended for an adjacent point in detail level one. In addition, based on the known green channel attribute correction value and the predicted red channel attribute correction value, the encoder may imply that the blue attribute channel correction value is to be applied to the point included in the second detail level.

[0142] Inter-attribute / cross-component prediction

[0143] In some embodiments, different interleaving methods may also allow inter-attribute / attribute channel prediction, which may result in more coding benefits. For example, for YCbCr data, the Cb and Cr color components may be predicted by their luma components. Such prediction modes may be selected at various levels, such as:

[0144] Point level,

[0145] ·Peer group level,

[0146] LOD level,

[0147] Slice level, and / or

[0148] Frame level.

[0149] In some embodiments, different prediction methods can be used. For example, the prediction can use a linear or nonlinear prediction model, wherein the parameters of the model (e.g., scale and offset in the model of chrominance=a*Y+b) are also signaled to the decoder. Such parameters can be estimated in the encoder using different methods (e.g., using the least square method). An alternative mode would be to combine the prediction based on brightness with the value generated by a conventional prediction method. For example, we can consider a weighted average of a brightness-based predictor and a distance-based predictor. The weighted parameters can be explicitly encoded in the bitstream, or can be implicitly determined by the decoder. The encoder can use, for example, a standard based on rate distortion (RD) or other methods that can consider pre-analysis of the encoding process and past statistical values ​​to determine such parameters.

[0150] In some embodiments, the prediction strategy for the current point attribute can also be adjusted based on the attribute / attribute channel value that has been encoded / decoded for the same point. For example, if the brightness (or other X attribute) value for a given point is known, neighboring points can be excluded from the prediction process for the given point, where these neighboring points have brightness (or X attribute) values ​​that are very different from the brightness (or X attribute) value of the given point. For example, based on the variance between the brightness (or X attribute) values, the chrominance components or other attributes of the neighboring points can be excluded from the prediction for the given point. In some embodiments, the prediction can be completed only with points with similar characteristics, for example, satisfying multiple attribute thresholds, not just distance thresholds. The encoder can explicitly signal which attributes or attribute channels that have been encoded / decoded should be used for prediction adaptation. Similarity can be determined based on a threshold T or a set of thresholds, assuming that multiple attributes are considered for the prediction selection process, T can be signaled in the bitstream. Such thresholds / threshold sets can be signaled at different levels of the encoding process, for example, point groups, LODs, slices, tiles, frames, or sequences.

[0151] Combined Sorting / Sampling LOD Method

[0152] In some embodiments, the low-complexity LOD generation process that uses a space-filling curve to sort the points and determine the level of refinement can be combined with the distance-based refinement process described above. For example, for portions of a point cloud with regularly sampled smooth attribute signals, a distance-based refinement strategy can be used. However, for portions of a point cloud that include irregularly sampled non-smooth attribute signals, a low-complexity LOD generation process that uses a space-filling curve to sort the points and determine the level of refinement can be used. In some embodiments, switching between a distance-based LOD generation process and a low-complexity LOD generation process using a space-filling curve can be operated at the following locations:

[0153] ·Peer group level,

[0154] LOD level,

[0155] Slice level, and / or

[0156] Frame level.

[0157] Figure 7 An exemplary process for determining points to be included in a level of detail (LOD) for a point cloud and for compressing attribute information for points determined to be included in the LOD according to some embodiments is shown.

[0158] At 702, an encoder receives a point cloud to be compressed. At 704, the encoder determines points of the point cloud to be included in various levels of detail using a space filling curve index sampling process as described above.

[0159] At 704, the encoder selects a first (or next) level of detail / refinement for which to compress attribute information.

[0160] At 708, the encoder performs rate distortion optimization to determine the compression technique to be used for compressing the attribute value. For example, the encoder may determine that the nearest neighbor inverse distance prediction technique is to be used. Alternatively, the encoder may determine that the space filling curve neighborhood prediction technique is to be used. For example, in the space filling curve neighborhood technique, the point in a given LOD at the position in the space filling curve index on either side of the point being evaluated can be selected as the nearest neighbor without calculating the Euclidean distance. Such techniques can weight the attribute values ​​of adjacent points based on the corresponding distances of the adjacent points in the index. The nearest neighbor inverse distance prediction technique may need to calculate the Euclidean distance, and the Euclidean distance can be used to weight the attribute values ​​of adjacent points when predicting the attribute value of a given point being evaluated.

[0161] At 710, the encoder then compresses the property values ​​for the selected level of detail using the selected property value compression technique (selected at 708).

[0162] At 712, the encoder determines whether there are additional detail levels to be compressed. If so, the process is repeated for the next detail level at 706. If not, the process ends at 714.

[0163] Figure 8 An exemplary process for decompressing compressed attribute information for points included in a level of detail (LOD) of a point cloud is illustrated in accordance with some embodiments.

[0164] At 802, the decoder receives compressed spatial information for a point cloud. The spatial information may have been compressed using a KD tree, an octree, a subsampling and prediction process, etc. At 804, the decoder reconstructs the geometry of the point cloud based on the compressed spatial information. Additionally, at 806, the decoder receives compressed attribute information for one or more levels of detail of the compressed point cloud.

[0165] At 808, the decoder uses a space filling curve index sampling technique to determine the level of detail for the point cloud. For example, the decoder may apply the same space filling curve to the reconstructed geometry of the point cloud as the space filling curve applied by the encoder. This may result in the same position index, such as that shown in FIG. Figure 7 Additionally, the decoder may apply the same set of sampling parameters to the index to sample the index in the same manner as performed at the encoder to generate a level of detail that includes the same points as those determined to be included in a corresponding level of detail at the encoder.

[0166] At 810 , the decoder selects a first (or next) level of detail for which to reconstruct the property value based on the compressed property information received at 806 .

[0167] At 812, the decoder determines an implied or signaled attribute prediction technique to be used for the selected level of detail, such as a nearest neighbor inverse distance prediction technique, a space filling curve neighborhood prediction technique, or the like.

[0168] At 814, the decoder decompresses the attribute information for the selected level of detail using the selected attribute prediction technique.

[0169] At 816, the decoder determines whether additional levels of detail are to be decompressed, and if so, returns to 810 and selects the next level of detail to decompress. If not, at 818, a reconstructed point cloud is generated.

[0170] Predictive Adaptation

[0171] In some embodiments, the prediction adaptation process can also utilize various statistics (e.g., spatial information) related to the point cloud geometry. For example, it can only include points with the same x value and / or y value and / or z value, and exclude other points (not only based on distance, but also based on other geometric criteria such as angle). The prediction can also include points with limited distance in one or more dimensions. For example, points within the total distance D can be included, but these points are also within the distance along the coordinate axis, such as the X distance (dx), Y distance (dy), or Z distance (dz) from the current point. Such prediction modes can be selected / signaled at various levels.

[0172] Related LOD encoder optimization

[0173] In some embodiments, the encoder can select the encoding parameters for the current LOD by considering not only its own distortion / coding performance but also the distortion introduced when predicting the next LOD. Such distortion can be calculated for all attributes or a subset of attributes (e.g., only the luminance component). Such a subset can be predetermined by a user or some other device, such as by analyzing the data and determining which attribute is the most active / has the most energy. For example, which attribute has the greatest impact on the prediction of other attributes and the prediction across multiple refinement levels or LODs. The distortion evaluated for the next LOD can be based on the mean square error (MSE), or some other distortion criteria can be considered. It is also possible to consider subsampling the points in the next LOD to reduce complexity. For example, only half of the samples affected in the LOD can be considered in the calculation. Sampling can be random, fixed based on some defined sampling processes, or can also be based on the characteristics of the signal, such as the LOD can be analyzed, and the most "important" point in the LOD can be considered. Importance can be determined, for example, based on the magnitude of the attribute.

[0174] Example Encoder

[0175] Fig.9A Components of an encoder according to some embodiments are shown.

[0176] Encoder 902 may be Figure 1A The encoder 902 includes a spatial encoder 904, a space filling curve LOD generator 910, a prediction / correction evaluator 906, an input data interface 914, and an output data interface 908. The encoder 902 also includes a context memory 916 and a configuration memory 918.

[0177] In some embodiments, a spatial encoder (such as spatial encoder 904) can compress spatial information associated with points of a point cloud so that the spatial information can be stored or transmitted in a compressed format. Fig.13 As discussed in more detail, the spatial encoder can utilize a KD tree to compress the spatial information of points in a point cloud. In addition, in some embodiments, a spatial encoder (such as spatial encoder 904) can utilize a KD tree to compress the spatial information of points in a point cloud. FIG. 12A to FIG. 12B Subsampling and prediction techniques are discussed in more detail. In some embodiments, a spatial encoder (such as spatial encoder 904) can utilize an octree to compress spatial information for points of a point cloud.

[0178] In some embodiments, the compressed spatial information may be stored or transmitted together with the compressed attribute information, or may be stored or transmitted separately. In either case, a decoder that receives compressed attribute information for points of a point cloud may also receive compressed spatial information for these points of the point cloud, or may obtain spatial information for these points of the point cloud.

[0179] A space filling curve LOD generator, such as space filling curve LOD generator 910, may utilize spatial information for points of a point cloud to determine points of the point cloud to be included in a corresponding LOD. Figure 8 Because the decoder is provided or otherwise obtains spatial information for points of the point cloud that is available at the encoder, the LOD determined by the space filling curve LOD generator of the encoder (such as the space filling curve LOD generator 910 of the encoder 902) can be the same or similar to the LOD generated by the space filling curve LOD generator of the decoder (such as the space filling curve LOD generator 928 of the decoder 920).

[0180] A prediction / correction evaluator (such as prediction / correction evaluator 906 of encoder 902) can determine a predicted attribute value for a point in a point cloud based on an inverse distance interpolation method using the attribute values ​​of the K nearest neighboring points of the point for which the attribute value is being predicted. The prediction / correction evaluator can also compare the predicted attribute value of the point being evaluated with the original attribute value of the point in the non-compressed point cloud to determine an attribute correction value. Alternatively, the prediction / correction evaluator 906 can utilize other prediction techniques, such as a space-filling curve neighborhood prediction technique.

[0181] The outgoing data encoder (such as the outgoing data encoder 908 of the encoder 902) can encode the attribute correction value and the assigned attribute value included in the compressed attribute information file for the point cloud. In some embodiments, the outgoing data encoder 908 can also encode the sampling parameters to be used to generate the LOD. In some embodiments, the outgoing data encoder (such as the outgoing data encoder 908) can select the encoding context for encoding the value based on the number of symbols included in the value, such as the assigned attribute value or the attribute correction value. In some embodiments, the encoding context including the Golomb index encoding can be used to encode the value with more symbols, while the arithmetic coding can be used to encode the value with fewer symbols. In some embodiments, the encoding context may include more than one encoding technique. For example, arithmetic coding can be used to encode a part of the value, while Golomb index encoding can be used to encode another part of the value. In some embodiments, the encoder (such as the encoder 902) may include a context memory such as a context memory 916, which stores the encoding context used by the outgoing data encoder such as the outgoing data encoder 908 to encode the attribute correction value and the assigned attribute value.

[0182] In some embodiments, an encoder such as encoder 902 may also include an incoming data interface such as incoming data interface 914. In some embodiments, the encoder may receive incoming data from one or more sensors that capture points of a point cloud or capture attribute information associated with points of a point cloud. For example, in some embodiments, the encoder may receive data from a LIDAR system, a 3D camera, a 3D scanner, etc., and may also receive data from other sensors such as gyroscopes, accelerometers, etc. In addition, the encoder may receive other data, such as the current time, from a system clock, etc. In some embodiments, such different types of data may be received by the encoder via an incoming data interface such as the incoming data interface 914 of encoder 902.

[0183] In some embodiments, an encoder (such as encoder 902) may also include a configuration interface (such as configuration interface 912), wherein one or more parameters used by the encoder to compress the point cloud can be adjusted via the configuration interface. In some embodiments, the configuration interface (such as configuration interface 912) can be a programmatic interface, such as an API. The configuration used by the encoder (such as encoder 902) can be stored in a configuration memory (such as configuration memory 918).

[0184] In some embodiments, an encoder (such as encoder 902) may include a Fig.9A More or fewer components as shown.

[0185] Fig. 9BComponents of a decoder according to some embodiments are shown.

[0186] Decoder 920 may be a Figure 1A Decoder 920 includes an encoding data interface 926 , a spatial decoder 922 , a space filling curve LOD generator 928 , a prediction evaluator 924 , a context memory 932 , a configuration memory 934 , and a decoding data interface 920 .

[0187] A decoder (such as decoder 920) may receive an encoded compressed point cloud and / or an encoded compressed attribute information file for points of the point cloud. For example, a decoder (such as decoder 920) may receive a compressed attribute information file such as Figure 1A The compressed attribute information 112 or Fig.10 The compressed attribute information file 1000 shown. The decoder can receive the compressed attribute information file via an encoding data interface (such as encoding data interface 926). The decoder can use the encoded compressed point cloud to determine the spatial information for the points of the point cloud. For example, the spatial information of the points of the point cloud included in the compressed point cloud can be generated by a spatial information generator (such as spatial information generator 922). In some embodiments, the compressed point cloud can be received from a storage device or other intermediate source via an encoded data interface (such as encoded data interface 926), where the compressed point cloud was previously encoded by an encoder (such as encoder 902). In some embodiments, the encoding data interface (such as encoding data interface 926) can decode the spatial information. For example, the spatial information may have been encoded using various encoding techniques (such as arithmetic coding, Golomb coding, etc.). The spatial information generator (such as spatial information generator 922) can receive the decoded spatial information from the encoding data interface (such as encoding data interface 926) and can use the decoded spatial information to generate a representation of the geometric structure of the point cloud being decompressed. For example, the decoded spatial information may be formatted as residual values ​​to be used in a sub-sampled prediction method to reconstruct the geometry of the point cloud to be decompressed. In such cases, the spatial information generator 922 may use the decoded spatial information from the encoded data interface 926 to reconstruct the geometry of the point cloud being decompressed, and the space filling curve LOD generator 928 may determine the LOD for the point cloud being decompressed based on the reconstructed geometry generated by the spatial information generator 922 for the point cloud being decompressed.

[0188] A predictive evaluator of the decoder (such as predictive evaluator 924) can select a starting point of the LOD based on an assigned starting point included in a compressed attribute information file. In some embodiments, the compressed attribute information file may include one or more assigned values ​​for one or more corresponding attributes of the starting point. In some embodiments, a predictive evaluator (such as predictive evaluator 924) can assign values ​​to one or more attributes of the starting point in the decompressed model of the point cloud being decompressed based on the assigned values ​​for the starting point included in the compressed attribute information file. The predictive evaluator (such as predictive evaluator 924) can also use the assigned values ​​of the attributes of the starting point to determine the attribute values ​​of the neighboring points. For example, the predictive evaluator can select the neighboring point closest to the starting point as the next point to be evaluated, where the next nearest neighboring point is selected based on the shortest distance from the starting point to the neighboring point.

[0189] Once the predictive evaluator has identified the "K" nearest neighboring points to the point being evaluated, the predictive evaluator may predict one or more attribute values ​​for one or more attributes of the point being evaluated based on the attribute values ​​of the corresponding attributes of the "K" nearest neighboring points. In some embodiments, an inverse distance interpolation technique may be used to predict the attribute values ​​of the point being evaluated based on the attribute values ​​of the neighboring points, wherein the attribute values ​​of the neighboring points at a closer distance to the point being evaluated are weighted more heavily than the attribute values ​​of the neighboring points at a greater distance from the point being evaluated.

[0190] A prediction evaluator (such as prediction evaluator 924) may apply the attribute correction value to the predicted attribute value to determine the attribute value to be included for the point in the decompressed point cloud. In some embodiments, the attribute correction value for the attribute of the point may be included in the compressed attribute information file. In some embodiments, the attribute correction value may be encoded using one of a plurality of supported coding contexts, wherein different coding contexts are selected based on the number of symbols included in the attribute correction value to encode different attribute correction values. In some embodiments, a decoder (such as decoder 920) may include a context memory (such as context memory 932), wherein the context memory stores a plurality of coding contexts, which may be used to decode the assigned attribute values ​​or attribute correction values ​​that have been encoded at the encoder using the corresponding coding context.

[0191] A decoder (such as decoder 920) may provide a decompressed point cloud LOD generated based on a received compressed point cloud and / or a received compressed attribute information file to a receiving device or application via a decoded data interface (such as decoded data interface 930). The decompressed LOD of the point cloud may include attribute values ​​of the points of the point cloud selected to be included in a given LOD and the attributes of the points for the given LOD. In some embodiments, the decoder may decode some attribute values ​​for the attributes of the point cloud without decoding other attribute values ​​for other attributes of the point cloud. For example, a point cloud may include color attributes for points of the point cloud, and may also include other attributes for the points of the point cloud, such as speed. In such cases, the decoder may decode one or more attributes of the points of the point cloud (such as a speed attribute) without decoding other attributes of the points of the point cloud (such as a color attribute).

[0192] In some embodiments, the decompressed point cloud and / or decompressed attribute information file for a given LOD can be used to generate a visual display, such as for a head mounted display. Additionally, in some embodiments, the decompressed point cloud and / or decompressed attribute information file can be provided to a decision engine that uses the decompressed point cloud and / or decompressed attribute information file to make one or more control decisions. In some embodiments, the decompressed point cloud and / or decompressed attribute information file can be used in various other applications or for various other purposes.

[0193] Fig.10 An exemplary compressed attribute information file according to some embodiments is shown. The attribute information file 1000 includes configuration information 1002, point cloud sampling parameters 1004 for a first LOD, point cloud data 1006 for the first LOD, and point attribute correction values ​​1008 for the first LOD. The attribute file 1000 also includes point cloud sampling parameters 1010 for another LOD, point cloud data 1012 for a second LOD, and point attribute correction values ​​1014 for points included in the second LOD. In some embodiments, the attribute information file 1000 may include information for any number of LODs. In some embodiments, the point cloud file 1000 may be transmitted in batches via multiple data packets. In some embodiments, not all of the parts shown in the attribute information file 1000 may be included in each data packet that transmits the compressed attribute information. In some embodiments, an attribute information file (such as the attribute information file 1000) may be stored in a storage device (such as a server that implements an encoder or decoder) or other computing device.

[0194] Fig.11 A process for compressing attribute information of a point cloud according to some embodiments is shown.

[0195] At 1102, an encoder receives a point cloud including attribute information for at least some of the points of the point cloud. The point cloud may be received from one or more sensors that capture the point cloud, or the point cloud may be generated in software. For example, a virtual reality or augmented reality system may have generated the point cloud.

[0196] At 1104, spatial information of the point cloud, such as the X, Y, and Z coordinates of the points of the point cloud, may be quantized. In some embodiments, the coordinates may be rounded to the nearest unit of measurement, such as meters, centimeters, millimeters, and the like.

[0197] At 1106, the quantized spatial information is compressed. In some embodiments, the quantized spatial information may be compressed using FIG. 12A to FIG. 12B The subsampling and subdivision prediction techniques discussed in more detail are used to compress spatial information. In addition, in some embodiments, the Fig.13 The KD tree compression technique discussed in more detail above can be used to compress the spatial information, or the octree compression technique can be used to compress the spatial information. In some embodiments, other suitable compression techniques can be used to compress the spatial information of the point cloud.

[0198] At 1108, the compressed spatial information for the point cloud is encoded as a compressed point cloud file or a portion of a compressed point cloud file. In some embodiments, the compressed spatial information and the compressed attribute information may be included in a common compressed point cloud file, or may be transmitted or stored as separate files.

[0199] At 1112, the spatial information of the received point cloud is used to generate a level of detail for the point cloud. In some embodiments, the spatial information of the point cloud can be quantized before the level of detail is generated. Additionally, in some embodiments where a lossy compression technique is used to compress the spatial information of the point cloud, the spatial information can be lossy encoded and lossy decoded before the level of detail is generated. In embodiments where lossy compression is used for the spatial information, encoding and decoding the spatial information at the encoder can ensure that the level of detail generated at the encoder will match the level of detail that will be generated at the decoder using the previously lossy encoded decoded spatial information.

[0200] Additionally, in some embodiments, attribute information for points of the point cloud may be quantized at 1110. For example, attribute values ​​may be rounded to integers or specific measurement increments. In some embodiments where attribute values ​​are integers, such as when integers are used to convey string values ​​such as "walking", "running", "driving", etc., quantization at 1110 may be omitted.

[0201] At 1114, attribute values ​​for the starting point are assigned. The assigned attribute values ​​for the starting point are encoded in a compressed attribute information file together with the attribute correction value. Because the decoder predicts the attribute value based on the distance to the adjacent point and the attribute value of the adjacent point, at least one attribute value for at least one point is explicitly encoded in the compressed attribute file. In some embodiments, the point of the point cloud may include multiple attributes, and in such embodiments, at least one attribute value for each type of attribute may be encoded for at least one point of the point cloud. In some embodiments, the starting point may be the first point evaluated for the first LOD determined at 1112. In some embodiments, the encoder may encode data indicating other marks of which point of the spatial information for the starting point and / or the point cloud is one or more starting points. In addition, the encoder may encode the attribute values ​​for one or more attributes of the starting point. In some embodiments, the starting point may be encoded for each LOD, or a single starting point may be encoded for the first LOD, and the predicted value from the lower level LOD may be used to predict the higher level LOD.

[0202] At 1116, the encoder determines an evaluation order for predicting attribute values ​​for other points in the point cloud other than the starting point, and the prediction and determination of attribute correction values ​​may be referred to herein as "evaluating" the attributes of the point. The evaluation order may be determined based on the shortest distance from the starting point to the adjacent neighboring points, wherein the nearest neighboring point is selected as the next point in the evaluation order. In some embodiments, the evaluation order may be determined only for the next point to be evaluated. In other embodiments, the evaluation order for all or multiple points in the point cloud may be determined at 1116. In some embodiments, the evaluation order may be determined dynamically, for example, one point at a time when evaluating the point. In some embodiments, the points may be evaluated in an evaluation order that corresponds to the order of the points in an index generated based on a space-filling curve, such as a Morton order.

[0203] At 1118, select the adjacent points of the starting point or subsequent point being evaluated. In some embodiments, the next adjacent point to be evaluated can be selected based on the adjacent point with the shortest distance to the most recently evaluated point compared to other adjacent points of the point evaluated last time. In some embodiments, the point selected at 1118 can be selected based on the evaluation order determined at 1116. In some embodiments, the evaluation order can be determined dynamically, for example, one point is determined at a time when evaluating the point. For example, each time the next point to be evaluated is selected at 1118, the next point in the evaluation order can be determined. In such embodiments, 1116 can be omitted. Because the points are evaluated in order, where the distance between each next point to be evaluated and the point evaluated last time is the shortest, the entropy between the attribute values ​​of the point being evaluated can be minimized. This is because points adjacent to each other are most likely to have similar attributes. Although in some cases, the similarity between the attributes of adjacent points may be different.

[0204] At 1120, the "K" nearest neighbor points to the point currently being evaluated are determined. The parameter "K" may be a configurable parameter selected by the encoder, or may be provided to the encoder as a user-configurable parameter. In order to select the "K" nearest neighbor points, the encoder may identify the first "K" closest points to the point being evaluated according to the minimum spanning tree. In some embodiments, points having only assigned attribute values ​​or points for which predicted attribute values ​​have been determined may be included in these "K" nearest neighbor points. In some embodiments, various numbers of points may be identified. For example, in some embodiments, "K" may be 5 points, 10 points, 16 points, etc. Because the point cloud includes points in 3D space, a particular point may have multiple neighbor points in multiple planes. In some embodiments, the encoder and decoder may be configured to identify a point as the "K" nearest neighbor points, regardless of whether a value has been predicted for the point. In addition, in some embodiments, the attribute value for a point used in the prediction may be a previous predicted attribute value or a corrected predicted attribute value that has been corrected based on an applied attribute correction value. In either case, the encoder and decoder may be configured to apply the same rules when identifying the "K" nearest neighboring points and when predicting the attribute value of a point based on the attribute values ​​of the "K" nearest neighboring points.

[0205] At 1122, one or more attribute values ​​are determined for each attribute of the point currently being evaluated. The attribute values ​​may be determined based on inverse distance interpolation. Inverse distance interpolation may interpolate predicted attribute values ​​based on the attribute values ​​of the "K" nearest neighboring points. The attribute values ​​of the "K" nearest neighboring points may be weighted based on respective distances between respective ones of the "K" nearest neighboring points and the point currently being evaluated. Attribute values ​​of neighboring points that are closer to the point currently being evaluated may be weighted more heavily than attribute values ​​of neighboring points that are farther away from the point currently being evaluated.

[0206] At 1124, attribute correction values ​​are determined for one or more predicted attribute values ​​of the point currently being evaluated. The attribute correction values ​​may be determined based on comparing the predicted attribute values ​​with corresponding attribute values ​​for the same point (or similar points) in the point cloud before compression of the attribute information. In some embodiments, quantized attribute information (such as the quantized attribute information generated at 1110) may be used to determine the attribute correction values. In some embodiments, the attribute correction values ​​may also be referred to as "residual errors," where the residual errors indicate the difference between the predicted attribute values ​​and the actual attribute values.

[0207] At 1126, it is determined whether there are additional points in the point cloud for which attribute correction values ​​are to be determined. If there are additional points to be evaluated, the process returns to 1118 and the next point to be evaluated in the evaluation order is selected, or the next LOD is selected. As discussed above, in some embodiments, the evaluation order can be determined dynamically, for example, one point at a time when evaluating the points. Therefore, in such embodiments, the minimum spanning tree can be consulted to select the next point to be evaluated based on the shortest distance between the next point and the last evaluated point. The process can repeat steps 1118 to 1126 until all or a portion of all points in the point cloud have been evaluated to determine predicted attribute values ​​and attribute correction values ​​for the predicted attribute values.

[0208] At 1128, the determination of the property correction value, the assigned property value, and any configuration information (such as parameter "K") for decoding the compressed property information file are encoded.

[0209] Various encoding techniques may be used to encode the attribute correction values, the assigned attribute values, and any configuration information.

[0210] FIG. 12A to FIG. 12B An exemplary process for compressing spatial information of a point cloud is shown according to some embodiments.

[0211] At 1202, an encoder receives a point cloud. The point cloud may be a captured point cloud from one or more sensors, or may be a generated point cloud (such as a point cloud generated by a graphics application). For example, 1204 shows points of an uncompressed point cloud.

[0212] At 1206, the encoder subsamples the received point cloud to generate a subsampled point cloud. The subsampled point cloud may include fewer points than the received point cloud. For example, the received point cloud may include hundreds of points, thousands of points, or millions of points, and the subsampled point cloud may include tens of points, hundreds of points, or thousands of points. For example, 1208 shows subsampled points of the point cloud received at 1202, for example, subsampling the points of the point cloud in 1204.

[0213] In some embodiments, the encoder may encode and decode the sub-sampled point cloud to generate a representative sub-sampled point cloud that the decoder will encounter when decoding the compressed point cloud. In some embodiments, the encoder and decoder may perform a lossy compression / decompression algorithm to generate a representative sub-sampled point cloud. In some embodiments, the spatial information for the points of the sub-sampled point cloud may be quantized as part of generating the representative sub-sampled point cloud. In some embodiments, the encoder may utilize lossless compression techniques and may omit encoding and decoding of the sub-sampled point cloud. For example, when using lossless compression techniques, the original sub-sampled point cloud may represent the sub-sampled point cloud that the decoder will encounter because in lossless compression, data may not be lost during compression and decompression.

[0214] At 1210, the encoder identifies subdivision locations between points of the subsampled point cloud according to configuration parameters selected for the compressed point cloud or according to fixed configuration parameters. Configuration parameters used by the encoder that are not fixed configuration parameters are communicated to the encoder by including values ​​for these configuration parameters in the compressed point cloud. Thus, the decoder can determine the same subdivision locations as evaluated by the encoder based on the subdivision configuration parameters included in the compressed point cloud. For example, 1212 shows identified subdivision locations between adjacent points of the subsampled point cloud.

[0215] At 1214, the encoder determines whether to include or exclude points at the subdivided positions in the decompressed point cloud for corresponding ones of the subdivided positions. Data indicating this determination is encoded in the compressed point cloud. In some embodiments, the data indicating the determination may be a single bit, which means that the point is to be included if "true", and which means that the point is not to be included if "false". In addition, the encoder may determine that the points to be included in the decompressed point cloud will be repositioned relative to the subdivided positions in the decompressed point cloud. For example, 1216 shows some points to be repositioned relative to the subdivided positions. For such points, the encoder may also encode data indicating how to reposition the points relative to the subdivided positions. In some embodiments, the position correction information may be quantized and entropy encoded. In some embodiments, the position correction information may include ΔX, ΔY, and / or ΔZ values ​​that indicate how to reposition the point relative to the subdivided positions. In other embodiments, the position correction information may include a single scalar value corresponding to the normal component of the position correction information, which is calculated as follows:

[0216] ΔN=[X A ,Y A ,Z A ]-[X,Y,Z])·[Normal vector]

[0217] In the above formula, ΔN is a scalar value indicating position correction information, which is the point position repositioned or adjusted relative to the subdivision position (e.g. [X A , Y A , Z A ]) and the original subdivided position (e.g., [X, Y, Z]). The vector product of this vector difference and the normal vector at the subdivided position produces a scalar value ΔN. Because the decoder can determine the normal vector at the subdivided position and can determine the coordinates of the subdivided position (e.g., [X, Y, Z]), by solving the above formula for the adjusted position, the decoder can also determine the coordinates of the adjusted position (e.g., [X A , Y A , Z A ]), which coordinates represent the repositioned position of the point relative to the subdivided position. In some embodiments, the position correction information can be further decomposed into a vertical component and one or more additional tangential components. In such embodiments, the vertical component (e.g., ΔN) and the one or more tangential components can be quantized and encoded for inclusion in the compressed point cloud.

[0218] In some embodiments, the encoder may determine whether to include one or more additional points (in addition to the points included at the subdivision positions or the points included at the positions relocated relative to the subdivision positions) in the decompressed point cloud. For example, if the original point cloud has an irregular surface or shape, so that the subdivision positions between the points in the subsampled point cloud do not adequately represent the irregular surface or shape, the encoder may determine to include one or more additional points in addition to the points to be included at the subdivision positions or the points relocated relative to the subdivision positions in the decompressed point cloud. In addition, the encoder may determine whether to include one or more additional points in the decompressed point cloud based on system constraints such as target bit rate, target compression ratio, quality target metric, etc. In some embodiments, the bit budget may change due to changing conditions such as network conditions, processor load, etc. In such embodiments, the encoder may adjust the amount of additional points encoded to be included in the decompressed point cloud based on the changed bit budget. In some embodiments, the encoder may include additional points so that the bit budget is consumed without exceeding the bit budget. For example, when the bit budget is higher, the encoder may include more additional points to consume the bit budget (and improve quality), and when the bit budget is smaller, the encoder may include fewer additional points so that the bit budget is consumed without exceeding the bit budget.

[0219] In some embodiments, the encoder may also determine whether to perform additional subdivision iterations. If so, the points determined to be included, repositioned, or additionally included in the decompressed point cloud are taken into account, and the process returns to 1210 to identify new subdivision locations for an updated subsampled point cloud, which includes the points determined to be included, repositioned, or additionally included in the decompressed point cloud. In some embodiments, the number of subdivision iterations (N) to be performed may be a fixed or configurable parameter of the encoder. In some embodiments, different subdivision iteration values ​​may be assigned to different portions of the point cloud. For example, the encoder may consider the point view from which the point cloud is being viewed, and may perform more subdivision iterations for points of the point cloud in the foreground of the point cloud viewed from the point view, and may perform fewer subdivision iterations for points in the background of the point cloud viewed from the point view.

[0220] At 1218, spatial information for subsampled points of the point cloud is encoded. Additionally, subdivided locations include and repositioning data is encoded. Additionally, any configurable parameters selected by the encoder or provided to the encoder from a user are encoded. The compressed point cloud may then be sent to a receiving entity as one compressed point cloud file, multiple compressed point cloud files, or the compressed point cloud may be packaged and transmitted to a receiving entity, such as a decoder or storage device, via multiple data packets. In some embodiments, the compressed point cloud may include both compressed spatial information and compressed attribute information. In other embodiments, the compressed spatial information and compressed attribute information may be included in separate compressed point cloud files.

[0221] Fig.13 Another exemplary process for compressing spatial information of a point cloud is shown in accordance with some embodiments.

[0222] In some embodiments, in addition to FIG. 12A to FIG. 12B Other spatial information compression techniques other than the subsampling and prediction spatial information techniques described in . For example, a spatial encoder (such as spatial encoder 904) or a spatial decoder (such as spatial decoder 922) may utilize other spatial information compression techniques (such as KD tree spatial information compression techniques). For example, in Fig.11 The compressed spatial information at 1106 can be used similar to FIG. 12A to FIG. 12B The subsampling and prediction techniques described in , can be performed using something like Fig.13 The KD tree spatial information compression technology described in , or another suitable spatial information compression technology can be used to perform it.

[0223] In the KD tree spatial information compression technique, a point cloud including spatial information may be received at 1302. In some embodiments, the spatial information may have been pre-quantized or may be further quantized after being received. For example, 1318 shows a captured point cloud that may be received at 1302. For simplicity, 1318 shows the point cloud in two dimensions. However, in some embodiments, the received point cloud may include points in 3D space.

[0224] At 1304, a K-dimensional tree or KD tree is constructed using the spatial information of the received point cloud. In some embodiments, a KD tree can be constructed by dividing a space such as a 1D, 2D, or 3D space of a point cloud into two halves in a predetermined order. For example, a 3D space including points of a point cloud can be initially divided into two halves via a plane intersecting one of the three axes (such as the X-axis). Then, a subsequent division can divide the resulting space along another of the three axes (such as the Y-axis). Then, another division can divide the resulting space along another of the axes (such as the Z-axis). Each time a division is performed, the number of points included in the sub-cell created by the division can be recorded. In some embodiments, the number of points in only one of the two sub-cells caused by the division can be recorded. This is because the number of points included in other sub-cells can be determined by subtracting the number of points in the recorded sub-cell from the total number of points in the parent cell before the division.

[0225] The KD tree may include a sequence of the number of points included in the cell, resulting from a sequential partitioning of the space including the points of the point cloud. In some embodiments, constructing the KD tree may include continuing to subdivide the space until only a single point is included in each lowest level sub-cell. The KD tree may be transmitted as a sequence of the number of points in the sequential cells resulting from the sequential partitioning. The decoder may be configured with information indicating the subdivision sequence followed by the encoder. For example, the encoder may follow a predefined partitioning sequence until only a single point is left in each lowest level sub-cell. Because the decoder may know the partitioning order followed to construct the KD tree and the number of points generated by each subdivision (transmitted to the decoder as compressed spatial information), the decoder may be able to reconstruct the point cloud.

[0226] For example, 1320 shows a simplified example of KD compression in a two-dimensional space. The initial space includes seven points. This can be considered as the first parent cell, and the KD tree can be encoded with the number of points "7" as the first number of the KD tree, which indicates that there are a total of seven points in the KD tree. The next step can be to divide the space along the X axis, thereby obtaining two sub-cells, the left sub-cell has three points and the right sub-cell has four points. The KD tree may include the number of points in the left sub-cell, for example, including "3" as the next number of the KD tree. Remember that the number of points in the right sub-cell can be determined based on the number of points in the left sub-cell minus the number of points in the parent cell. Further, it can be a time to add the space division along the Y axis so that each of the left and right sub-cells is divided into two halves and becomes a lower-level sub-cell. Similarly, the number of points included in the left lower-level sub-cell may be included in the KD tree, such as "0" and "1". Then, the next step may be to divide the non-zero lower-level sub-cells along the X axis and record the number of points of each of the lower-level left sub-cells in the KD tree. This process can continue until only a single point remains in the lowest level sub-cell. The decoder can utilize the inverse process to reconstruct the point cloud based on the sequence of the total number of points for each left sub-cell of the received KD tree.

[0227] At 1306, a coding context is selected for encoding the number of points for the first cell (e.g., a parent cell including seven points) of the KD tree. In some embodiments, the context memory may store hundreds or thousands of coding contexts. In some embodiments, the highest number of point coding contexts may be used to encode cells that include more points than the highest number of point coding contexts. In some embodiments, the coding contexts may include arithmetic coding, Golomb exponential coding, or a combination of the two. In some embodiments, other coding techniques may be used. In some embodiments, the arithmetic coding context may include probabilities for specific symbols, wherein different arithmetic coding contexts include different symbol probabilities.

[0228] At 1308, the number of points for the first cell is encoded according to the selected coding context.

[0229] At 1310, a coding context for encoding the child cell is selected based on the number of points included in the parent cell. The coding context for the child cell may be selected in a similar manner as at 1306 for the parent cell.

[0230] At 1312, the number of points included in the sub-cell is encoded according to the selected coding context selected at 1310. At 1314, it is determined whether there are additional lower-level sub-cells to be encoded in the KD tree. If so, the process returns to 1310. If not, at 1316, the number of encoded points in the parent cell and the sub-cell is included in a compressed spatial information file (such as a compressed point cloud). The encoded values ​​are sorted in the compressed spatial information file so that the decoder can reconstruct the point cloud based on the number of points for each parent cell and sub-cell and the order of the number of points of the corresponding cell included in the compressed spatial information file.

[0231] In some embodiments, the number of points in each cell may be determined and then encoded into groups at 1316. Alternatively, in some embodiments, the number of points in a cell may be encoded after determination without waiting to determine the total number of points for all sub-cells.

[0232] Fig.14 An example process for decompressing compressed attribute information of a point cloud is shown according to some embodiments.

[0233] At 1402, the decoder receives compressed attribute information for a point cloud, and at 1404, the decoder receives compressed spatial information for the point cloud. In some embodiments, the compressed attribute information and the compressed spatial information may be included in one or more common files or separate files.

[0234] At 1406, the decoder decompresses the compressed spatial information. The compressed spatial information may have been compressed according to subsampling and prediction techniques, and the decoder may perform similar subsampling, prediction, and prediction correction actions as performed at the encoder and further apply the correction values ​​to the predicted point positions to generate a non-compressed point cloud based on the compressed spatial information. In some embodiments, the compressed spatial information may be compressed in a KD tree format, and the decoder may generate a decompressed point cloud based on the encoded KD tree included in the received spatial information. In some embodiments, the compressed spatial information may have been compressed using octree technology, and octree decoding technology may be used to generate decompressed spatial information for the point cloud. In some embodiments, other spatial information compression techniques may have been used, and these techniques may be decompressed via a decoder.

[0235] At 1408, the decoder may determine points to be included in various levels of detail based on the decompressed spatial information. For example, the decoder may be provided via an encoded data interface of the decoder such as Fig. 9B926) of the decoder 920 shown in to receive the compressed spatial information and / or the compressed attribute information. A spatial decoder (such as the spatial decoder 922) can decompress the compressed spatial information, and a space filling curve generator (such as the space filling curve generator 928) can generate the LOD based on the decompressed spatial information.

[0236] At 1410, a prediction evaluator of a decoder, such as prediction evaluator 924 of decoder 920, may assign attribute values ​​to the starting point based on the assigned attribute values ​​included in the compressed attribute information. In some embodiments, the compressed attribute information may identify the point as a starting point to be used to generate a minimum spanning tree and to predict attribute values ​​for the point according to an evaluation order based on a minimum spanning tree. The one or more assigned attribute values ​​for the starting point may be included in the decompressed attribute information for the decompressed point cloud.

[0237] At 1412, the predictive evaluator of the decoder or another decoder component determines the evaluation order for at least the next point to be evaluated after the starting point. In some embodiments, the evaluation order for all or multiple points in the point can be determined, or in other embodiments, the evaluation order can be confirmed point by point when determining the attribute value for the point. The points can be evaluated in order based on the minimum distance between the continuous points being evaluated. For example, the adjacent point with the shortest distance from the starting point compared to other adjacent points can be selected as the next point to be evaluated after the starting point. In a similar manner, other points to be evaluated can then be selected based on the shortest distance from the most recently evaluated point. At 1414, the next point to be evaluated is selected. In some embodiments, 1412 and 1414 can be performed together.

[0238] At 1416, the predictive evaluator of the decoder determines the "K" nearest neighbors to the point being evaluated. In some embodiments, neighbors may be included in the "K" nearest neighbors only if they have been assigned or predicted attribute values. In other embodiments, neighbors may be included in the "K" nearest neighbors regardless of whether they have been assigned or predicted attribute values. In such embodiments, the encoder may follow similar rules as the decoder regarding whether to include points without predicted values ​​as neighbors when identifying the "K" nearest neighbors.

[0239] At 1418, predicted attribute values ​​are determined for one or more attributes of the point being evaluated based on the attribute values ​​of the "K" nearest neighboring points and the distances between the point being evaluated and corresponding ones of the "K" nearest neighboring points. In some embodiments, inverse distance interpolation techniques may be used to predict attribute values, where attribute values ​​of points closer to the point being evaluated are weighted more heavily than attribute values ​​of points farther from the point being evaluated. The attribute prediction techniques used by the decoder may be the same as the attribute prediction techniques used by the encoder that compressed the attribute information.

[0240] At 1420, the predictive evaluator of the decoder may apply the attribute correction value to the predicted attribute value of the point to correct the attribute value. The attribute correction value may make the attribute value match or nearly match the attribute value of the original point cloud before compression. In some embodiments where a point has more than one attribute, 1418 and 1420 may be repeated for each attribute of the point. In some embodiments, some attribute information may be decompressed without decompressing all attribute information for the point cloud or point. For example, a point may include speed attribute information and color attribute information. Speed ​​attribute information may be decoded without decoding color attribute information, and vice versa. In some embodiments, an application utilizing compressed attribute information may indicate which attributes to decompress for a point cloud.

[0241] At 1422, it is determined whether there are additional points or LODs to be evaluated. If so, the process returns to 1414 and the next point to be evaluated is selected. If there are no additional points to be evaluated, at 1424, decompressed attribute information is provided, for example as a decompressed point cloud, where each point includes spatial information and one or more attributes.

[0242] In some cases, the number of bits required to encode attribute information for a point cloud may constitute a significant portion of the bitstream for the point cloud. For example, the attribute information may constitute a larger portion of the bitstream than that used to transmit compressed spatial information for the point cloud.

[0243] In some embodiments, the spatial information may be used to construct a hierarchical level of detail (LOD) structure. The LOD structure may be used to compress attributes associated with a point cloud. The LOD structure may also enable advanced features such as progressive / view-related streaming and scalable rendering. For example, in some embodiments, compressed attribute information may be sent (or decoded) only for a portion of a point cloud (e.g., level of detail) without sending (or decoding) all attribute information for the entire point cloud.

[0244] Fig.15 An exemplary encoding process for generating a hierarchical LOD structure according to some embodiments is shown. For example, in some embodiments, an encoder (such as encoder 902) may use Fig.15A similar process is shown to generate compressed attribute information in the LOD structure.

[0245] In some embodiments, geometric information (also referred to herein as "spatial information") can be used to effectively predict attribute information. Fig.15 The compression of color information is shown in . However, the LOD structure can be applied to the compression of any type of attribute associated with the points of the point cloud (e.g., reflectivity, texture, modality, etc.). It should be noted that the pre-encoding step of applying color space conversion or updating the data to make the data more suitable for compression can be performed depending on the attribute to be compressed.

[0246] In some embodiments, compression of attribute information according to the LOD process is performed as described below.

[0247] For example, assume that geometry (G) = {point - P (0), P (1), ... P (N-1)} is a reconstructed point cloud position generated by a spatial decoder (geometry decoder GD 1502) included in the encoder after decoding a compressed geometry bitstream generated by a geometry encoder also included in an encoder (geometry encoder GE1514) (such as spatial encoder 904). For example, in some embodiments, an encoder (such as encoder 902) may include both a geometry encoder (such as geometry encoder 1514) and a geometry decoder (such as geometry decoder 1514). In some embodiments, the geometry encoder can be part of the spatial encoder 914, and the geometry decoder can be part of the prediction / correction evaluator 906.

[0248] In some embodiments, the decompressed spatial information may describe the locations of points in 3D space, such as the X, Y, and Z coordinates of the points that make up the mug 1500. Note that the spatial information may be available to both an encoder (such as encoder 902) and a decoder (such as decoder 920). For example, various techniques (such as KD-tree compression, octree compression, nearest neighbor prediction, etc.) may be used to compress and / or encode the spatial information for the mug 1500, and the spatial information may be sent to the decoder along with or in addition to compressed attribute information for the attributes of the points that make up the point cloud of the mug 1500.

[0249] In some embodiments, a deterministic reordering process may be applied at both the encoder side (such as at encoder 902) and the decoder side (such as at decoder 920) to organize the points of a point cloud (such as the points representing mug 1500) in a set of levels of detail (LODs). For example, the levels of detail may be generated by a level of detail generator 1504, which may be included in a prediction / correction evaluator of an encoder, such as prediction / correction evaluator 906 of encoder 902. In some embodiments, the level of detail generator 1504 may be a separate component of an encoder (such as encoder 902). For example, the level of detail generator 1504 may be a separate component of encoder 902.

[0250] In some embodiments, the encoder described above may further include a quantization module (not shown) that quantizes the geometry information included in the "position (x, y, z)" being provided to the geometry encoder 1514. In addition, in some embodiments, the encoder described above may additionally include a module for removing duplicate points after quantization and before the geometry encoder 1514.

[0251] In some embodiments, quantization may also be applied to compress attribute information, such as attribute correction values ​​and / or one or more attribute value starting points. For example, quantization is performed at 1510 to attribute correction values ​​determined by interpolation-based prediction module 1508. Quantization techniques may include uniform quantization, uniform quantization with dead zones, non-uniform / non-linear quantization, grid quantization, or other suitable quantization techniques.

[0252] Example Level of Detail Grading

[0253] Fig.16 1604 may include more details than 1602, and 1606 may include more details than 1604. In addition, 1608 may include more details than 1602, 1604, and 1606.

[0254] The hierarchical LOD structure can be used to construct an attribute prediction strategy. For example, in some embodiments, the points can be encoded in the same order as the order in which the points are visited during the LOD generation phase. The attributes of each point can be predicted by using the K nearest neighbors that have been previously encoded. In some embodiments, "K" is a parameter that can be defined by the user or can be determined by using an optimization strategy. "K" can be static or adaptive. In the latter case where "K" is adaptive, additional information describing the parameter can be included in the bitstream.

[0255] In some embodiments, different prediction strategies can be used. For example, one of the following interpolation strategies and a combination of the following interpolation strategies can be used, or the encoder / decoder can adaptively switch between different interpolation strategies. Different interpolation strategies may include interpolation strategies, such as: inverse distance interpolation, centroid interpolation, natural neighbor interpolation, moving least squares interpolation or other suitable interpolation techniques. For example, interpolation-based prediction can be performed at an interpolation-based prediction module 1508 included in a prediction / correction value evaluator (such as, prediction / correction value evaluator 906 of encoder 902) of the encoder. In addition, interpolation-based prediction can be performed at an interpolation-based prediction module 1508 included in a prediction evaluator (such as, prediction evaluator 924 of decoder 920) of the decoder. In some embodiments, before performing interpolation-based prediction, the color space can also be converted at a color space conversion module 1506. In some embodiments, the color space conversion module 1506 may be included in an encoder (such as encoder 902). In some embodiments, the decoder may also include a module that converts the converted color space back to the original color space.

[0256] In some embodiments, quantization may also be applied to the attribute information. For example, quantization may be performed at quantization module 1510. In some embodiments, an encoder (such as encoder 902) may also include quantization module 1510. The quantization techniques employed by quantization module 1510 may include uniform quantization, uniform quantization with dead zones, non-uniform / non-linear quantization, grid quantization, or other suitable quantization techniques.

[0257] In some embodiments, LOD attribute compression may be used to compress dynamic point clouds as follows:

[0258] oAssume FC is the current point cloud frame, and assume RF is the reference point cloud.

[0259] o Assume that M is the motion field that deforms RF into the shape of FC.

[0260] • M may be calculated at the decoder side, and in this case the information may not be encoded in the bitstream.

[0261] • M can be calculated by the encoder and explicitly encoded in the bitstream.

[0262] • M can be encoded by applying the hierarchical compression techniques described herein to the motion vectors associated with each point of the RF (eg, the motion of the RF can be considered an additional attribute).

[0263] M can be encoded as a skeleton / skin-based model with associated local and global transformations.

[0264] M can be encoded as a motion field defined based on an octree structure, which is adaptively improved to adapt to the complexity of the motion field.

[0265] M may be described using any suitable animation technique, such as keyframe-based animation, deformation techniques, free-form deformation, keypoint-based deformation, etc.

[0266] Assume RF' is the point cloud obtained after applying the motion field M to RF. Then, not only the 'K' nearest neighboring points of FC but also the 'K' nearest neighboring points of RF' are considered, and the points of RF' can also be used for attribute prediction strategy.

[0267] In addition, attribute correction values ​​may be determined based on comparing the interpolation-based prediction values ​​determined at the interpolation-based prediction module 1508 with the original uncompressed attribute values. The attribute correction values ​​may be further quantized at the quantization module 1510, and the quantized attribute correction values, the encoded spatial information (output from the geometry encoder 1502), and any configuration parameters used in the prediction may be encoded at the arithmetic coding module 1512. In some embodiments, the arithmetic coding module may use context-adaptive arithmetic coding techniques. The compressed point cloud may then be provided to a decoder (such as decoder 920), and the decoder may determine a similar level of detail and perform interpolation-based prediction based on the quantized attribute correction values, the encoded spatial information (output from the geometry encoder 1502), and the configuration parameters used in the encoder prediction to reconstruct the original point cloud.

[0268] Example application for point cloud compression and decompression

[0269] Fig.17 Illustration of a compressed point cloud being used in a 3D remote display application according to some embodiments.

[0270] In some embodiments, a sensor (such as sensor 102), an encoder (such as encoder 104 or encoder 902), and a decoder (such as decoder 116 or decoder 920) may be used to transmit a point cloud in a 3D application. For example, at 1702, a sensor (such as sensor 102) may capture a 3D image, and at 1704, the sensor or a processor associated with the sensor may perform 3D reconstruction based on the sensed data to generate a point cloud.

[0271] At 1706, an encoder (such as encoder 104 or 902) may compress the point cloud, and at 1708, the encoder or post-processor may package the compressed point cloud and transmit the compressed point cloud via network 1710. At 1712, the data packet may be received at a target location including a decoder (such as decoder 116 or decoder 920). At 1714, the decoder may decompress the point cloud, and at 1716, the decompressed point cloud may be rendered. In some embodiments, the 3D application may transmit the point cloud data in real time so that the display at 1716 may represent the image being observed at 1702. For example, at 1716, a camera in a canyon may allow a remote user to experience walking through a virtual canyon.

[0272] Fig.18 Illustration of a compressed point cloud being used in a virtual reality (VR) or augmented reality (AR) application according to some embodiments.

[0273] In some embodiments, the point cloud may be generated in software (e.g., as opposed to being captured by a sensor). For example, at 1802, virtual reality or augmented reality content is generated. The virtual reality or augmented reality content may include point cloud data and non-point cloud data. For example, as an example, a non-point cloud character may traverse a terrain represented by a point cloud. At 1804, the point cloud data may be compressed, and at 1806, the compressed point cloud data and the non-point cloud data may be packaged and transmitted via a network 1808. For example, the virtual reality or augmented reality content generated at 1802 may be generated at a remote server and transmitted to a VR or AR content consumer via a network 1808. At 1810, a data packet may be received and synchronized at a device of a VR or AR consumer. At 1812, a decoder operating at a device of a VR or AR consumer may decompress the compressed point cloud, and the point cloud and non-point cloud data may be rendered in real time, for example, in a head-mounted display of a device of a VR or AR consumer. In some embodiments, point cloud data may be generated, compressed, decompressed, and rendered in response to a VR or AR consumer manipulating a head mounted display to look in different directions.

[0274] In some embodiments, point cloud compression as described herein can be used in various other applications such as geographic information systems, live sports events, museum displays, autonomous navigation, etc.

[0275] Exemplary Computer System

[0276] Fig.19 An exemplary computer system 1900 is shown that can implement an encoder or decoder or any other of the components described herein (e.g., as described above with reference to FIGS. 1 to 2). Fig.18 1900). The computer system 1900 may be configured to perform any or all of the embodiments described above. In various embodiments, the computer system 1900 may be any of various types of devices, including, but not limited to, a personal computer system, a desktop computer, a laptop computer, a notebook computer, a tablet computer, an all-in-one computer, a tablet or netbook computer, a mainframe computer system, a handheld computer, a workstation, a network computer, a camera, a set-top box, a mobile device, a consumer device, a video game controller, a handheld video game device, an application server, a storage device, a television, a video recording device, a peripheral device (such as a switch, a modem, a router), or generally any type of computing or electronic device.

[0277] The various embodiments of the point cloud encoder or decoder described herein may be executed on one or more computer systems 1900, which may interact with various other devices. Fig.18 Any component, action or functionality described may be implemented in a configuration Fig.19 1900. In the illustrated embodiment, the computer system 1900 includes one or more processors 1910 coupled to a system memory 1920 via an input / output (EO) interface 1930. The computer system 1900 also includes a network interface 1940 coupled to the I / O interface 1930, and one or more input / output devices 1950, such as a cursor control device 1960, a keyboard 1970, and a display 1980. In some cases, it is contemplated that the embodiments may be implemented using a single instance of the computer system 1900, while in other embodiments, multiple such systems or multiple nodes making up the computer system 1900 may be configured to host different parts or instances of the embodiments. For example, in one embodiment, some elements may be implemented via one or more nodes of the computer system 1900 that are different from those nodes that implement other elements.

[0278] In various embodiments, computer system 1900 may be a uniprocessor system including one processor 1910, or a multiprocessor system including a number of processors 1910 (e.g., two, four, eight, or another suitable number). Processor 1910 may be any suitable processor capable of executing instructions. For example, in various embodiments, processor 1910 may be a general-purpose or embedded processor that implements any of a variety of instruction set architectures (ISAs), such as x86, PowerPC, SPARC, or MIPS ISAs, or any other suitable ISAs. In a multiprocessor system, each of processors 1910 may typically, but not necessarily, implement the same ISA.

[0279] The system memory 1920 may be configured to store point cloud compression or point cloud decompression program instructions 1922 and / or sensor data accessible by the processor 1910. In various embodiments, the system memory 1920 may be implemented using any suitable memory technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash type memory, or any other type of memory. In the illustrated embodiment, the program instructions 1922 may be configured to implement an image sensor control application incorporating any of the above functionalities. In some embodiments, the program instructions and / or data may be received, sent, or stored on different types of computer accessible media or similar media separate from the system memory 1920 or the computer system 1900. Although the computer system 1900 is described as implementing the functionality of the functional blocks of the preceding figures, any functionality described herein may be implemented via such a computer system.

[0280] In one embodiment, the I / O interface 1930 may be configured to coordinate I / O communications between the processor 1910, the system memory 1920, and any peripheral devices in the device (including the network interface 1940 or other peripheral device interfaces, such as the input / output device 1950). In some embodiments, the I / O interface 1930 may perform any necessary protocol, timing, or other data conversion to convert data signals from one component (e.g., the system memory 1920) into a format suitable for use by another component (e.g., the processor 1910). In some embodiments, the I / O interface 1930 may include support for devices attached, for example, via various types of peripheral device buses (such as, variations of the peripheral component interconnect (PCI) bus standard or the universal serial bus (USB) standard). In some embodiments, the functionality of the I / O interface 1930 may be divided into two or more separate components, such as a north bridge and a south bridge, for example. In addition, in some embodiments, some or all of the functionality of the I / O interface 1930 (such as an interface to the system memory 1920) may be directly incorporated into the processor 1910.

[0281] The network interface 1940 may be configured to allow data to be exchanged between the computer system 1900 and other devices (e.g., carriers or proxy devices) attached to the network 1985 or between nodes of the computer system 1900. In various embodiments, the network 1985 may include one or more networks, including, but not limited to, a local area network (LAN) (e.g., an Ethernet or an enterprise network), a wide area network (WAN) (e.g., the Internet), a wireless data network, some other electronic data network, or some combination thereof. In various embodiments, the network interface 1940 may support communication via a wired or wireless general data network (such as any suitable type of Ethernet network), for example; communication via a telecommunications / telephone network (such as an analog voice network or a digital fiber optic communication network); communication via a storage area network (such as Fibre Channel SANs), or communication via any other suitable type of network and / or protocol.

[0282] In some embodiments, input / output devices 1950 may include one or more display terminals, keyboards, keypads, touch pads, scanning devices, voice or optical recognition devices, or any other device suitable for inputting or accessing data by one or more computer systems 1900. Multiple input / output devices 1950 may be present in computer system 1900, or may be distributed across various nodes of computer system 1900. In some embodiments, similar input / output devices may be separate from computer system 1900 and may interact with one or more nodes of computer system 1900 via a wired or wireless connection (such as via network interface 1940).

[0283] like Fig.19 As shown, memory 1920 may include program instructions 1922, which may be executable by a processor to implement any element or action described above. In one embodiment, the program instructions may execute the method described above. In other embodiments, different elements and data may be included. It should be noted that the data may include any data or information described above.

[0284] It will be appreciated by those skilled in the art that computer system 1900 is merely illustrative, and is not intended to limit the scope of the embodiments. Specifically, computer systems and devices may include any combination of hardware or software that can perform the indicated functions, including computers, network equipment, Internet equipment, personal digital assistants, wireless telephones, pagers, etc. Computer system 1900 may also be connected to other devices not shown, or may otherwise be operated as an independent system. In addition, the functions provided by the components shown may be combined in fewer components or distributed in additional components in some embodiments. Similarly, in some embodiments, the functions of some components in the components shown may not be provided, and / or other additional functions may be available.

[0285] Those skilled in the art will also recognize that, although various items are shown as being stored in memory or on storage devices during use, for the purpose of memory management and data integrity, these items or parts thereof may be transmitted between memory and other storage devices. Alternatively, in other embodiments, some or all of these software components may be executed in a memory on another device and communicated with the illustrated computer system via inter-computer communication. Some or all of the system components or data structures may also be stored (e.g., as instructions or structured data) on a computer-accessible medium or portable article to be read by a suitable driver, and various examples thereof are described above. In some embodiments, instructions stored on a computer-accessible medium separated from the computer system 1900 may be transmitted to the computer system 1900 via a transmission medium or signal (such as an electrical signal, an electromagnetic signal, or a digital signal transmitted via a communication medium such as a network and / or a wireless link). Various embodiments may also include receiving, sending, or storing instructions and / or data implemented according to the above description on a computer-accessible medium. Generally speaking, a computer-accessible medium may include a non-transitory computer-readable storage medium or memory medium, such as a magnetic medium or an optical medium, for example a disk or a DVD / CD-ROM, a volatile or non-volatile medium, such as a RAM (e.g., SDRAM, DDR, RDRAM, SRAM, etc.), a ROM, etc. In some embodiments, a computer-accessible medium may include a transmission medium or signal, such as an electrical signal, an electromagnetic signal, or a digital signal transmitted via a communication medium such as a network and / or a wireless link.

[0286] In different embodiments, the methods described herein can be implemented in software, hardware or a combination thereof. In addition, the order of the frames of the method can be changed, and various elements can be added, reordered, combined, omitted, modified, etc. For those skilled in the art who benefit from the present disclosure, various modifications and changes can obviously be made. The various embodiments described herein are intended to be illustrative and not restrictive. Many variations, modifications, additions and improvements are possible. Therefore, multiple examples can be provided for the components described as a single example in this article. The boundaries between various components, operations and data repositories are arbitrary to a certain extent, and specific operations are shown in the context of a specific exemplary configuration. Other allocations of functions are expected, and they may fall within the scope of the appended claims. Finally, the structure and function presented as discrete components in the exemplary configuration may be implemented as a combined structure or component. These and other variations, modifications, additions and improvements may fall within the scope of the embodiments defined in the following claims.

Claims

1. A non-transitory computer-readable medium storing program instructions that, when executed by one or more processors, cause the one or more processors to: determining points to be included in a first level of detail of the point cloud for which attribute information is compressed; and determining points to be included in one or more additional levels of detail of the point cloud for which attribute information is compressed, Wherein to determine the points to be included in the first level of detail and the one or more additional levels of detail, the program instructions cause the one or more processors to: determining an order of the points of the point cloud based on a space-filling curve, wherein corresponding points of the point cloud are assigned indices that index the corresponding points based on a proximity of the corresponding points to a position along the space-filling curve; sampling the index according to a first sampling rate to determine points of the point cloud to be included in the first level of detail; as well as sampling the index according to one or more other sampling rates to determine points of the point cloud to be included in the one or more additional levels of detail; compressing attribute information for the points determined to be included in the first level of detail; and Attribute information for the points determined to be included in the one or more additional levels of detail is compressed.

2. The non-transitory computer readable medium of claim 1 , wherein the program instructions, when executed by the one or more processors, cause the one or more processors to: The first sampling rate and the one or more other sampling rates are signaled in a bitstream along with the compression attribute information for the first level of detail and the compression attribute information for the one or more additional levels of detail.

3. The non-transitory computer readable medium of claim 1 , wherein the program instructions, when executed by the one or more processors, cause the one or more processors to: determining a first sampling order to be used for sampling the indexes according to the first sampling rate to determine the points to be included in the first level of detail; determining one or more other sampling orders to be used for sampling the indexes according to the one or more other sampling rates to determine the points to be included in the one or more additional levels of detail; as well as The first sampling order and the one or more other sampling orders are signaled in a bitstream along with the compression attribute information for the first level of detail and the compression attribute information for the one or more additional levels of detail.

4. A non-transitory computer-readable medium according to claim 3, wherein the respective sampling rates indicate respective frequencies of the index positions in the index to be skipped or selected when traversing the index positions according to one of the respective sampling orders to sample the index to determine which points to include in a given level of detail.

5. The non-transitory computer readable medium of claim 4, wherein the first sampling order and the one or more other sampling orders include one or more of the following: Forward sampling order; a reverse sampling order, wherein the index positions are sampled in a direction opposite to the forward sampling order; or An inside-outside sampling order, where the index is sampled from the inside index position in a forward direction toward the end of the index, and from the inside index position in a reverse direction toward the beginning of the index.

6. The non-transitory computer readable medium of claim 5, wherein the program instructions, when executed on the one or more processors, cause the one or more processors to: A rate-distortion optimization analysis is performed on a respective level of detail to select the respective sampling order to be used to determine points to be included in the respective level of detail.

7. The non-transitory computer readable medium of claim 6, wherein the program instructions, when executed on the one or more processors, cause the one or more processors to: determining, based on a rate-distortion optimization analysis, a first sampling offset to be used as an offset when sampling the index positions according to the first sampling order to determine the points to be included in the first level of detail; determining, based on a rate-distortion optimization analysis, using the one or more other sampling rates to determine one or more other sampling offsets to be used as offsets when sampling the index positions according to the one or more other sampling orders to determine the points to be included in the one or more additional levels of detail; as well as The first sampling offset and the one or more further sampling offsets are signaled in a bitstream along with the compression attribute information for the first level of detail and the compression attribute information for the one or more additional levels of detail.

8. The non-transitory computer-readable medium of claim 7, wherein for the first level of detail and the one or more additional levels of detail, the signaled bitstream comprises one or more of the following different respective items: Sampling rate; Sampling order; or Sampling offset, for determining the point in the respective level of detail to be included in the level of detail.

9. The non-transitory computer readable medium of claim 8, wherein the program instructions, when executed by the one or more processors, cause the one or more processors to: Compressed attribute value information is interleaved between different ones of the levels of detail such that attribute value types not included in a given level of detail are included in adjacent levels of detail.

10. The non-transitory computer readable medium of claim 1, wherein the program instructions, when executed on the one or more processors, cause the one or more processors to: Points to be included in another level of detail are determined based on a rate-distortion optimization analysis that selects between a nearest neighbor traversal technique and a space filling curve sampling technique to determine which points to include in the another level of detail. 11 . The non-transitory computer readable medium of claim 1 , wherein the space filling curve is a Morton space filling curve and the order of the points in the index is according to a Morton order.

12. A non-transitory computer readable medium storing program instructions which, when executed by one or more processors, cause the one or more processors to: determining points to be included in one or more levels of detail for the point cloud, Wherein, in order to determine the points to be included in the one or more levels of detail, the program instructions cause the one or more processors to: determining an order of the points of the point cloud based on a space-filling curve, wherein corresponding points of the point cloud are assigned indices that index the corresponding points based on a proximity of the corresponding points to a position along the space-filling curve; as well as sampling the index according to one or more sampling rates to determine points of the point cloud to be included in the one or more levels of detail; as well as Attribute values ​​of the points determined to be included in the one or more levels of detail are determined based on compressed attribute information for corresponding one or more levels of detail of the point cloud.

13. The non-transitory computer readable medium of claim 12, wherein the program instructions, when executed by the one or more processors, cause the one or more processors to receive one or more of: Sampling rate; Sampling order; or Sampling offset, for determining the points to be included in the corresponding one or more levels of detail.

14. A non-transitory computer-readable medium according to claim 13, wherein for at least one of the levels of detail, a sampling rate, sampling order, or sampling offset to be used for the at least one level of detail is signaled to be different from the sampling rate, sampling order, or sampling offset to be used for other levels of the one or more levels of detail.

15. The non-transitory computer readable medium of claim 12, wherein to determine the attribute value of the point determined to be included in the one or more levels of detail, the program instructions cause the one or more processors to: determining a set of neighboring points of the given point being evaluated based on the determined order of the given point, wherein the set of neighboring points includes a plurality of points in the order included at previous or subsequent index positions in the order that precede or follow the given point being evaluated; Determining, from the set of neighborhood points, a subset of the neighborhood points whose locations in three-dimensional space are closest to the given point being evaluated; predicting a property value for the given point being evaluated based on corresponding property values ​​of points included in the subset of neighborhood points closest to the given point being evaluated; as well as Applying an attribute correction value for the given point being evaluated to the predicted attribute value to determine a decompressed attribute value for the given point being evaluated, wherein the attribute correction value for the corresponding point for a given level of detail is included in the compressed attribute information for the given level of detail.

16. The non-transitory computer readable medium of claim 15, wherein the compressed attribute information for two or more levels of detail comprises interleaved compressed attribute value information for the two or more levels of detail, and Wherein to predict the property value for the given point at a given level of detail being evaluated, one or more points in neighboring levels of detail are considered in making the prediction, the one or more points having interleaved property information that is not included for the given point in the given level of detail.

17. The non-transitory computer readable medium of claim 15, wherein to predict the property value for the given point being evaluated, the program instructions cause the one or more processors to: Predict the value of the attribute based on: the corresponding attribute values ​​of the points included in the subset of the neighborhood points closest to the given point being evaluated; and The position of the given point is being evaluated in terms of the shape of the point cloud formed together with other points of the point cloud.

18. An electronic device, comprising: Memory for storing program instructions; and one or more processors, wherein the program instructions, when executed by the one or more processors, cause the one or more processors to: Receiving spatial information for a point of the point cloud; receiving compressed attribute information for one or more levels of detail of the point cloud; determining points to be included in the one or more levels of detail for the point cloud, Wherein, in order to determine the points to be included in the one or more levels of detail, the program instructions cause the one or more processors to: determining an order of the points of the point cloud based on a space-filling curve, wherein corresponding points of the point cloud are assigned indices that index the corresponding points based on a proximity of the corresponding points to a position along the space-filling curve; as well as sampling the index according to one or more sampling rates to determine points of the point cloud to be included in the one or more levels of detail; as well as Attribute values ​​for the points determined to be included in the one or more levels of detail are determined based on the received compressed attribute information for corresponding one or more levels of detail of the point cloud.

19. The electronic device of claim 18, wherein the program instructions, when executed by the one or more processors, cause the one or more processors to: receiving a first set of sampling parameters for a first level of detail of the one or more levels of detail; and receiving a second set of sampling parameters for a second level of detail of the one or more levels of detail, wherein the first set of sampling parameters is applied to sample the index to determine the points of the point cloud to be included in the first level of detail; and wherein the second set of sampling parameters is applied to sample the index to determine the points of the point cloud to be included in the second level of detail, The first group of sampling parameters and the second group of sampling parameters include one or more parameters that are different between the groups.

20. The electronic device of claim 19, wherein the first set of sampling parameters or the second set of sampling parameters comprises one or more of the following: Sampling rate; Sampling order; or Sampling offset.