Trimming the Search Space for Nearest Neighbor Determination in Point Cloud Compression

By trimming the search space using space fill curves and Morton code technology, combining the adaptive prediction methods of encoder and decoder, the problem of high storage and transmission costs of point cloud data is solved, and efficient compression and real-time application of point cloud data is achieved.

CN114631118BActive Publication Date: 2025-07-08APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080076377.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-01
Filing Date
2020-10-02
Publication Date
2025-07-08
Estimated Expiration
2040-10-02

AI Technical Summary

Technical Problem

The storage and transmission of point cloud data is expensive and time-consuming, and the prior art is difficult to effectively compress and decompress large amounts of point cloud data, limiting its use in real-time applications.

Method used

By trimming the search space using spatial fill curves such as Morton sequences, combining the search of Morton code and adjacent voxels, identifying and compressing the nearest neighbor points in point cloud data, encoder and decoder are used to encode and decode the spatial information and attribute information of the point cloud, and compressing using adaptive prediction and prediction correction values.

Benefits of technology

It realizes efficient compression and real-time transmission of point cloud data, reduces storage requirements and transmission time, and supports real-time or almost real-time point cloud data applications, such as fast display and control decisions of augmented reality and virtual reality systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114631118B_ABST
    Figure CN114631118B_ABST
Patent Text Reader

Abstract

The search space for performing nearest neighbor search for encoding point cloud data can be trimmed. The ranges of a space filling curve can be used to identify points to be excluded from the search space. Additionally, adjacent voxels can be searched based on these ranges of the space filling curve to identify any adjacent points that were missed during the trimmed search.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art Technical Field

[0002] The present disclosure generally relates to the compression and decompression of point clouds, which include a plurality of points, each point having associated attribute information.

[0003] Description of Related Technologies

[0004] Various types of sensors (such as light detection and ranging (LIDAR) systems, 3D cameras, 3D scanners, etc.) can capture data indicating the position of points in three-dimensional space (e.g., positions in the X, Y, and Z planes). Additionally, such systems can capture attribute information in addition to the spatial information for the corresponding points, such as color information (e.g., RGB values), intensity attributes, reflectivity attributes, motion-related attributes, modality attributes, or various other attributes. In some cases, additional attributes can be assigned to the corresponding points, such as a timestamp when the point was captured. The points captured by such sensors can constitute a "point cloud", which includes a set of points each having associated spatial information and one or more associated attributes. In some cases, a point cloud can include thousands of points, hundreds of thousands of points, millions of points, or even more points. Additionally, in some cases, different from the point cloud being captured by one or more sensors, a point cloud can be generated, for example, in software. In either case, such point clouds can include a large amount of data, and storing and transmitting these point clouds can be costly and time-consuming. Summary of the Invention

[0005] In some embodiments, the search space for performing a nearest neighbor search for encoding point cloud data can be trimmed. Instead of generating nearest neighbor search results for at least some of the points in the point cloud that are within some range of a space filling curve, the range of the space filling curve can be used to identify the search space to be excluded or reused.

[0006] In some embodiments, the space filling curve used is the Morton order, where Morton codes are determined for the points in the point cloud that fall along the space filling curve. Additionally, in some embodiments, in addition to using the trimmed search space generated by searching using the range of Morton codes on either side of the point being evaluated for nearest neighboring points, the Morton codes of one or more adjacent voxels adjacent to the point being evaluated are also determined, and the Morton codes determined for the points in the point cloud are searched to see if the Morton codes of the adjacent voxels include points in the point cloud. This can identify the nearest neighboring points included in adjacent voxels whose Morton codes are outside the trimmed search range.

[0007] Additionally, in some embodiments, as an initial step, the Morton codes of adjacent voxels can be determined and a search can be performed in the index of the Morton codes of the point cloud being compressed. In such embodiments, if the number of adjacent points found in the adjacent voxels is less than the desired number of nearest neighbor points to be used for prediction / detail level generation, an additional search can be performed within a trimmed Morton code search range to identify additional nearest neighbor points. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1A A system is shown that includes a sensor that captures information for points of a point cloud and an encoder that compresses attribute information and / or spatial information of the point cloud, where the compressed point cloud information is sent to a decoder, in accordance with some embodiments.

[0009] Figure 1B A process for encoding attribute information of a point cloud is shown, in accordance with some embodiments.

[0010] Figure 1C Representative views of point cloud information at different stages of an encoding process are shown, in accordance with some embodiments.

[0011] Figure 2A Components of an encoder are shown, in accordance with some embodiments.

[0012] Figure 2B Components of a decoder are shown, in accordance with some embodiments.

[0013] Figure 3 An example compressed attribute file is shown, in accordance with some embodiments.

[0014] Figure 4A A process for compressing attribute information of a point cloud is shown, in accordance with some embodiments.

[0015] Figure 4B Using adaptive distance-based prediction to predict attribute values as part of compressing attribute information of a point cloud is shown, in accordance with some embodiments.

[0016] Figures 4C to 4E Parameters that can be determined or selected by an encoder and signaled via compressed attribute information of a point cloud are shown, in accordance with some embodiments.

[0017] Figure 5 A process for encoding an attribute correction value is shown, in accordance with some embodiments.

[0018] Figures 6A to 6B An exemplary process for compressing spatial information of a point cloud is shown, in accordance with some embodiments.

[0019] Figure 7Shows another example process for compressing the spatial information of a point cloud according to some embodiments.

[0020] Figure 8A Shows an exemplary process for decompressing the compressed attribute information of a point cloud according to some embodiments.

[0021] Figure 8B Shows using adaptive distance-based prediction to predict attribute values as part of decompressing the attribute information of a point cloud according to some embodiments.

[0022] Figure 9 Shows components of an example encoder for generating a hierarchical level of detail (LOD) structure according to some embodiments.

[0023] Figure 10 Shows an example process for determining points to be included at different refinement layers of a level of detail (LOD) structure according to some embodiments.

[0024] Figure 11A Shows an example level of detail (LOD) structure according to some embodiments.

[0025] Figure 11B Shows an example compressed point cloud file including a level of detail (LOD) for a point cloud according to some embodiments.

[0026] Figure 12A Shows a method for encoding the attribute information of a point cloud according to some embodiments.

[0027] Figure 12B Shows a method for decoding the attribute information of a point cloud according to some embodiments.

[0028] Figure 12C Shows an example neighborhood configuration of a cube of an octree according to some embodiments.

[0029] Figure 12D Shows an example look-ahead cube according to some embodiments.

[0030] Figure 12E Shows an example of 31 contexts that can be used to adaptively encode the index value of symbol S using a binary arithmetic encoder according to some embodiments.

[0031] Figure 12F Shows an example octree compression technique using a binary arithmetic encoder, cache, and look-ahead table according to some embodiments.

[0032] Figure 13AShows a direct transformation that can be applied at an encoder to encode attribute information of a point cloud according to some embodiments.

[0033] Figure 13B Shows an inverse transformation that can be applied at a decoder to decode attribute information of a point cloud according to some embodiments.

[0034] Figure 14 Shows an assignment of a bounding box to a space-filling curve range for determining the minimum distance to points within the space-filling curve range according to some embodiments.

[0035] Figure 15 Shows a high-level flowchart for applying a boundary shape to trim a search space for nearest neighbor search according to some embodiments.

[0036] Figure 16 Shows an example of reusing nearest neighbor search results according to the range of a space-filling curve according to some embodiments.

[0037] Figure 17 Shows a high-level flowchart for applying a boundary shape to trim a search space for nearest neighbor search according to some embodiments.

[0038] Figure 18 Shows points in a discrete space for improving nearest neighbor search according to some embodiments.

[0039] Figure 19 Illustrates a set of exemplary points and Morton codes of a space-filling curve, where two points fall on the space-filling curve, according to some embodiments.

[0040] Figure 20 Shows compressed point cloud information being used in a 3D remote display application according to some embodiments.

[0041] Figure 21 Shows compressed point cloud information being used in a virtual reality application according to some embodiments.

[0042] Figure 22 Shows an exemplary computer system that can implement an encoder or a decoder according to some embodiments.

[0043] This specification includes references to "one embodiment" or "embodiments". The appearances of the phrase "in one embodiment" or "in embodiments" do not necessarily refer to the same embodiment. Specific features, structures, or characteristics may be combined in any suitable manner consistent with this disclosure.

[0044] "Comprising", this term is open-ended. As used in the appended claims, this term does not exclude additional structures or steps. Consider the following cited claim: "An apparatus comprising one or more processor units...". Such a claim does not exclude the apparatus from including additional components (e.g., a network interface unit, graphics circuitry, etc.).

[0045] "Configured to", various units, circuits, or other components may be described or recited as "configured to" perform one or more tasks. In such contexts, "configured to" is used to imply, by indicating that the unit / circuit / component includes the structure (e.g., circuitry) that performs the one or more tasks during operation. Thus, the unit / circuit / component is alleged to be configured to perform the task even when the specified unit / circuit / component is currently inoperable (e.g., not powered on). Units / circuits / components used in conjunction with "configured to" language include hardware - such as circuitry, memory storing program instructions executable to implement the operation, etc. Referring to a unit / circuit / component "configured to" perform one or more tasks is specifically intended not to invoke 35 U.S.C. § 112(f) for that unit / circuit / component. Additionally, "configured to" can include a general structure (e.g., general circuitry) manipulated by software and / or firmware (e.g., an FPGA or a general-purpose processor executing software) to operate in a manner capable of performing the one or more tasks to be solved. "Configured to" can also include adjusting a manufacturing process (e.g., a semiconductor fabrication facility) to fabricate a device (e.g., an integrated circuit) suitable for implementing or performing the one or more tasks.

[0046] "First", "second", etc. As used herein, these terms serve as labels for the nouns preceding them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.). For example, a buffer circuit may be described herein as performing write operations on a "first" value and a "second" value. The terms "first" and "second" do not necessarily imply that the first value must be written before the second value.

[0047] "Based on". As used herein, this term is used to describe one or more factors that affect a determination. This term does not exclude additional factors that affect the determination. That is, the determination may be based solely on these factors or at least partially on these factors. Consider the phrase "determine A based on B". In this case, B is a factor that affects the determination of A, and such a phrase does not exclude the determination of A from also being based on C. In other instances, A may be determined based solely on B. Detailed Description

[0048] As data acquisition and display technologies become more advanced, the ability to capture point clouds including thousands of points in 2D or 3D space (such as via a LIDAR system) is enhanced. Moreover, the development of advanced display technologies (such as virtual reality or augmented reality systems) increases the potential uses of point clouds. However, point cloud files are typically very large, and storing and transmitting these point cloud files can be costly and time-consuming. For example, the communication of point clouds over a private network or a public network (such as the Internet) may require a significant amount of time and / or network resources, such that some uses of point cloud data (such as real-time use) may be restricted. Additionally, the storage requirements for point cloud files may consume a significant amount of the storage capacity of the device storing the point cloud files, which may also limit the potential applications using point cloud data.

[0049] In some embodiments, an encoder can be used to generate compressed point clouds to reduce the costs and time associated with storing and transmitting large point cloud files. In some embodiments, the system can include an encoder that compresses the attribute information and / or the spatial information (also referred to herein as geometric information) of the point cloud file, such that the point cloud file can be stored and transmitted more quickly than an uncompressed point cloud and can be stored and transmitted in a manner that the point cloud file occupies less storage space than an uncompressed point cloud. In some embodiments, the compression of the spatial information and / or attributes of the points in the point cloud can enable the point cloud to be transmitted over a network in real time or almost in real time. For example, the system can include a sensor that captures the spatial information and / or attribute information of points in the environment where the sensor is located, wherein the captured points and the corresponding attributes constitute a point cloud. The system can also include an encoder that compresses the attribute information of the captured point cloud. The compressed attribute information of the point cloud can be sent over the network in real time or almost in real time to a decoder that decompresses the compressed attribute information of the point cloud. The decompressed point cloud can be further processed, for example, to make a control decision based on the surrounding environment at the sensor location. The control decision can then be transmitted back to a device at or near the sensor location, where the device receiving the control decision implements the control decision in real time or almost in real time. In some embodiments, the decoder can be associated with an augmented reality system, and the decompressed attribute information can be displayed or otherwise used by the augmented reality system. In some embodiments, the compressed attribute information about the point cloud can be sent together with the compressed spatial information of the points of the point cloud. In other embodiments, the spatial information and the attribute information can be separately encoded and / or separately sent to the decoder.

[0050] In some embodiments, the system may include a decoder that receives, via a network, one or more point cloud files including compressed attribute information from a remote server or other storage device storing the one or more point cloud files. For example, a 3D display, a holographic display, or a head-mounted display may be manipulated in real-time or near real-time to display different portions of a virtual world represented by the point cloud. To update the 3D display, the holographic display, or the head-mounted display, the system associated with the decoder may request point cloud files from the remote server based on user manipulations of the display, and the point cloud files may be transmitted from the remote server to the decoder and decoded by the decoder in real-time or near real-time. The display may then be updated with updated point cloud data (such as updated point attributes) in response to the user manipulations.

[0051] In some embodiments, a system may include one or more LIDAR systems, 3D cameras, 3D scanners, etc., and such sensor devices may capture spatial information, such as the X, Y, and Z coordinates of points in the view of the sensor device. In some embodiments, the spatial information may be relative to a local coordinate system or may be relative to a global coordinate system (e.g., a Cartesian coordinate system may have a fixed reference point such as a fixed point on the Earth, or may have a non-fixed local reference point such as the sensor location).

[0052] In some embodiments, such sensors may also capture attribute information about one or more points, such as color attributes, reflectivity attributes, velocity attributes, acceleration attributes, time attributes, modality, and / or various other attributes. In some embodiments, in addition to LIDAR systems, 3D cameras, 3D scanners, etc., other sensors may capture attribute information to be included in the point cloud. For example, in some embodiments, a gyroscope or an accelerometer may capture motion information to be included in the point cloud as an attribute associated with one or more points of the point cloud. For example, a vehicle equipped with a LIDAR system, a 3D camera, or a 3D scanner may include the direction and rate of the vehicle in the point cloud captured by the LIDAR system, the 3D camera, or the 3D scanner. For example, when points in the field of view of the vehicle are captured, the points may be included in the point cloud, where the point cloud includes the captured points and associated motion information corresponding to the state of the vehicle when the points were captured.

[0053] In some embodiments, the attribute information may include string values, such as different modalities. For example, the attribute information may include string values indicating modalities such as "walking", "running", "driving", etc. In some embodiments, the encoder may include a mapping from "string values" to integer indices, where certain strings are associated with certain corresponding integer values. In some embodiments, the point cloud may indicate the string value for a point by including the integer associated with the string value as an attribute of the point. Both the encoder and the decoder may store the common string values as integer indices, such that the decoder can determine the string value for a point based on looking up the integer value of the string attribute of the point in a string value to integer index that matches or is similar to the decoder's string value to integer index.

[0054] In some embodiments, in addition to compressing the attribute information for the attributes of the points of the point cloud, the encoder also compresses and encodes the spatial information of the point cloud to compress the spatial information. For example, to compress the spatial information, a K-D tree may be generated, where the corresponding number of points included in each cell of the cells of the K-D tree are encoded. This sequence of encoded point counts may encode the spatial information for the points of the point cloud. Additionally, in some embodiments, subsampling and prediction methods may be used to compress and encode the spatial information for the point cloud. In some embodiments, the spatial information may be quantized before being compressed and encoded. Additionally, in some embodiments, the compression of the spatial information may be lossless. Thus, the decoder may be able to determine the same view of the spatial information as the encoder. Additionally, once the compressed spatial information is decoded, the encoder may be able to determine the view of the spatial information that the decoder will encounter. Because both the encoder and the decoder may have or be able to reconstruct the same spatial information for the point cloud, the spatial relationships may be used to compress the attribute information for the point cloud.

[0055] For example, in many point clouds, the attribute information between adjacent points or points located relatively close to each other may have a high level of correlation between the attributes, and thus, the difference in the point attribute values is relatively small. For example, when considered relative to points that are more separated in the point cloud, adjacent points in the point cloud may have relatively small differences in color.

[0056] In some embodiments, the encoder may include a predictor that determines a predicted attribute value for an attribute of a point in a point cloud based on attribute values of similar attributes for adjacent points in the point cloud and based on a corresponding distance between the point being evaluated and the adjacent points. In some embodiments, a higher weight may be given to the attribute value of an adjacent point that is closer to the point being evaluated compared to the attribute value of an adjacent point that is farther from the point being evaluated. Additionally, the encoder may compare the predicted attribute value with the attribute value of the attribute of the point in the original point cloud before compression. A residual, also referred to herein as an “attribute correction value,” may be determined based on this comparison. The attribute correction value may be encoded and included in the compressed attribute information for the point cloud, where the decoder uses the encoded attribute correction value to correct the predicted attribute value for the point, where the same or a similar prediction method used at the encoder is used at the decoder to predict the attribute value.

[0057] In some embodiments, to encode the attribute values, the encoder may generate an ordering of the points of the point cloud based on the spatial information of the points of the point cloud. For example, the points may be ordered according to a space-filling curve. In some embodiments, this ordering may represent a Morton ordering of the points. The encoder may select a first point as a starting point and may determine an evaluation order of the other points of the point cloud based on the minimum distance from the starting point to the nearest neighbor and subsequent minimum distances from the neighbor to the next nearest neighbor, etc. Additionally, in some embodiments, the adjacent points may be determined from a set of points within a user-defined search range of the index values of the given point being evaluated, where the index values and the search range values are values in the indices of the points of the point cloud organized according to the space-filling curve. In this way, an evaluation order for determining the predicted attribute values of the points of the point cloud may be determined. Since the decoder may receive or reconstruct the same spatial information as used by the encoder, the decoder may generate the same ordering of the points of the point cloud and may determine the same evaluation order of the points of the point cloud.

[0058] In some embodiments, the encoder may assign an attribute value to a starting point of the point cloud for predicting the attribute values of the other points of the point cloud. The encoder may predict the attribute value of an adjacent point based on the attribute value of the starting point and the distance between the starting point and the adjacent point of the starting point. Then, the encoder may determine the difference between the predicted attribute value of the adjacent point and the actual attribute value of the adjacent point included in the uncompressed original point cloud. This difference may be encoded in the compressed attribute information file as the attribute correction value of the adjacent point. Then, the encoder may repeat a similar process for each point in the evaluation order. To predict the attribute value of a subsequent point in the evaluation order, the encoder may identify K nearest neighbor points as the particular point being evaluated, where the identified K nearest neighbor points have been assigned or predicted attribute values. In some embodiments, “K” may be a configurable parameter transmitted from the encoder to the decoder.

[0059] The encoder can determine the distances in the X, Y, and Z spaces between the point being evaluated and each identified neighboring point. For example, the encoder can determine the respective Euclidean distances from the point being evaluated to each of the neighboring points. The encoder can then predict the value of the attribute for the point being evaluated based on the attribute values of the neighboring points, where the attribute values of the neighboring points are weighted according to the reciprocals of the distances from the point being evaluated to some of the respective neighboring points. Greater weight can be given to the attribute values of the neighboring points that are closer to the point being evaluated compared to the attribute values of the neighboring points that are farther from the point being evaluated.

[0060] In a similar manner as described for the first neighboring point, the encoder can compare the predicted value for each of the other points in the point cloud with the actual attribute value in the original uncompressed point cloud (e.g., the captured point cloud). The difference can be encoded as an attribute correction value for the attribute of the point being evaluated among these other points. In some embodiments, the attribute correction values can be encoded in the compressed attribute information file in order according to an evaluation order determined based on a space-filling curve order. Since the encoder and the decoder can determine the same evaluation order based on the spatial information of the point cloud, the decoder can determine which attribute correction value corresponds to which attribute of which point based on the order in which the attribute correction values are encoded in the compressed attribute information file. Additionally, the starting point and one or more attribute values of the starting point can be explicitly encoded in the compressed attribute information file such that the decoder can determine the evaluation order starting from the same point as used at the encoder to start the evaluation order. Additionally, the one or more attribute values of the starting point can provide the values of neighboring points that the decoder uses to determine the predicted attribute value for the point being evaluated, which is a neighboring point of the starting point.

[0061] In some embodiments, the encoder can determine the predicted value for the attribute of a point based on temporal considerations. For example, in addition to or instead of determining the predicted value based on neighboring points in the same "frame" (e.g., the same time point as the point being evaluated), the encoder can consider the attribute values of points in neighboring and subsequent time frames.

[0062] Figure 1A A system is shown that includes a sensor that captures information for points in a point cloud and an encoder that compresses the attribute information of the point cloud, where the compressed attribute information is sent to a decoder.

[0063] System 100 includes sensor 102 and encoder 104. Sensor 102 captures point cloud 110, which includes points representing structure 106 in view 108 of sensor 102. For example, in some embodiments, structure 106 can be a mountain, a building, a sign, the environment around a street, or any other type of structure. In some embodiments, capturing a point cloud (such as capturing point cloud 110) can include spatial information and attribute information about the points included in the point cloud. For example, point A of captured point cloud 110 includes X, Y, Z coordinates and attributes 1, 2, and 3. In some embodiments, the attributes of a point can include attributes such as R, G, B color values, the velocity at that point, the acceleration at that point, the reflectivity of the structure at that point, a timestamp indicating when the point was captured, a string value indicating the modality when the point was captured such as "walking", or other attributes. Captured point cloud 110 can be provided to encoder 104, where encoder 104 generates a compressed version of the point cloud (compressed attribute information 112), and this compressed version is transmitted via network 114 to decoder 116. In some embodiments, the compressed version of the point cloud (such as compressed attribute information 112) can be included in a common compressed point cloud that also includes compressed spatial information for the points of the point cloud, or in some embodiments, the compressed spatial information and compressed attribute information can be transmitted as separate files.

[0064] In some embodiments, encoder 104 can be integrated with sensor 102. For example, encoder 104 can be implemented in hardware or software included in a sensor device (such as sensor 102). In other embodiments, encoder 104 can be implemented on a separate computing device adjacent to sensor 102.

[0065] Figure 1B A process for encoding compressed attribute information of a point cloud according to some embodiments is shown. Additionally, Figure 1C A representative view of point cloud information at different stages of the encoding process according to some embodiments is shown.

[0066] At 152, an encoder such as encoder 104 receives a captured point cloud or a generated point cloud. For example, in some embodiments, a point cloud can be captured via one or more sensors (such as sensor 102), or a point cloud can be generated in software (such as in a virtual reality or augmented reality system). For example, 164 shows an example captured or generated point cloud. Each point in the point cloud shown in 164 can have one or more attributes associated with that point. Note that for ease of illustration, point cloud 164 is shown in 2D, but the point cloud can include points in 3D space.

[0067] At 154, the ordering of the points of the point cloud is determined according to a space-filling curve. For example, the space-filling curve can fill a three-dimensional space, and the points of the point cloud can be ordered based on their positions relative to the space-filling curve. For example, a Morton code can be used to represent multi-dimensional data in one dimension, where a "Z-order function" is applied to the multi-dimensional data to produce a one-dimensional representation. In some embodiments, as discussed in more detail herein, the points can also be sorted into multiple levels of detail (LOD). In some embodiments, the points included in the corresponding level of detail (LOD) can be determined by sorting the points according to their positions along the space-filling curve. For example, these points can be organized according to their Morton codes.

[0068] In some embodiments, other space-filling curves can be used. For example, techniques can be used that map positions (e.g., in the form of X, Y, Z coordinates) to space-filling curves such as Morton order (or Z-order), Hilbert curve, Peano curve, etc. In this way, all the points of the point cloud that are encoded and decoded using spatial information can be organized into an index in the same order on the encoder and decoder. To determine various refinement levels, sampling rates, etc., the sorted index of the points can be used. For example, to divide the point cloud into four levels of detail, the index that maps Morton values to the corresponding points can be sampled, for example, at a rate of four, where every third index point is included in the lowest level of refinement. For each additional level of refinement, the remaining points in the index that have not been sampled can be sampled, for example, every second index point, etc., until all points have been sampled to obtain the highest level of detail.

[0069] At 156, an attribute value for one or more attributes of the starting point can be assigned for encoding and included in the compressed attribute information of the point cloud. As discussed above, the predicted attribute value for the points of the point cloud can be determined based on the attribute values of adjacent points. However, an initial attribute value for at least one point is provided to the decoder so that the decoder can use at least the initial attribute value and an attribute correction value for correcting the predicted attribute value predicted based on the initial attribute value to determine the attribute values of other points. Therefore, one or more attribute values for at least one starting point are explicitly encoded in the compressed attribute information file. Additionally, the spatial information for the starting point can be explicitly encoded so that the decoder can identify the starting point to determine which point among the points of the point cloud is to be used as the starting point for generating the order according to the space-filling curve. In some embodiments, the starting point can be indicated in other ways besides explicitly encoding the spatial information for the starting point, such as marking the starting point or other point identification methods.

[0070] Because the decoder will receive an indication of the starting point and will encounter the same or similar spatial information for the points of the point cloud as the encoder, the decoder can determine the same spatial filling curve order starting from the same starting point as determined by the encoder. Additionally, the decoder can determine the same processing order as the encoder based on the spatial filling curve order determined by the decoder.

[0071] At 158, for the current point being evaluated, the prediction / correction evaluator of the encoder determines a predicted attribute value for the attribute of the current point being evaluated. In some embodiments, the current point being evaluated can have more than one attribute. Thus, the prediction / correction evaluator of the encoder can predict more than one attribute value for the point. For each point being evaluated, the prediction / correction evaluator can identify a set of nearest neighbor points for which attribute values have been assigned or predicted. In some embodiments, the number "K" of the identified neighboring points can be a configurable parameter of the encoder, and the encoder can include configuration information indicating the parameter "K" in the compressed attribute information file such that when performing attribute prediction, the decoder can identify the same number of neighboring points. Then, the prediction / correction evaluator can determine the distance between the point being evaluated and the corresponding neighboring points among the identified neighboring points. The prediction / correction evaluator can use an inverse distance interpolation method to predict the attribute value for each attribute of the point being evaluated. Then the prediction / correction evaluator can predict the attribute value of the point being evaluated based on the average of the inverse distance weighted attribute values of the identified neighboring points.

[0072] For example, 166 shows the point (X, Y, Z) being evaluated, where attribute A1 is determined based on the inverse distance weighted attribute values of eight identified neighboring points.

[0073] At 160, an attribute correction value for each point is determined. The attribute correction value is determined based on comparing the predicted attribute value for each attribute of the point with the corresponding attribute value of the point in the original uncompressed point cloud (such as the captured point cloud). For example, 168 shows the formula for determining the attribute correction value, where the captured value is subtracted from the predicted value to determine the attribute correction value. Note that although Figure 1BThe predicted property value at 158 and the property correction value determined at 160 are shown, but in some embodiments, the property correction value for a point may be determined after predicting the property value for that point. Then the next point can be evaluated, where the predicted property value for that point is determined and the property correction value for that point is determined. Thus, 158 and 160 can be repeated for each point being evaluated. In other embodiments, the predicted values for multiple points can be determined and then the property correction values can be determined. In some embodiments, the prediction for subsequent points being evaluated can be based on the predicted property value or can be based on the corrected property value or based on both. In some embodiments, both the encoder and the decoder can follow the same rules regarding whether to determine the predicted value for subsequent points based on the predicted or corrected property value.

[0074] At 162, the determined property correction values for the points of the point cloud, one or more assigned property values for the starting point, the spatial information or other markers of the starting point, and any configuration information to be included in the compressed property information file are encoded. As discussed more fully in Figure 5 various encoding methods such as arithmetic coding and / or Golomb coding can be used to encode the property correction values, the assigned property values, and the configuration information.

[0075] Figure 2A Shows components of an encoder according to some embodiments.

[0076] Encoder 202 can be an encoder similar to encoder 104 shown in Figure 1A Encoder 202 includes a spatial encoder 204, a space filling curve order generator 210, a prediction / correction evaluator 206, an incoming data interface 214, and an outgoing data interface 208. Encoder 202 also includes a context memory 216 and a configuration memory 218.

[0077] In some embodiments, a spatial encoder (such as spatial encoder 204) can compress the spatial information associated with the points of the point cloud such that the spatial information can be stored or transmitted in a compressed format. In some embodiments, as discussed more fully with respect to Figure 7 a spatial encoder can utilize a K-D tree to compress the spatial information for the points of the point cloud. Additionally, in some embodiments, a spatial encoder (such as spatial encoder 204) can utilize subsampling and prediction techniques as discussed more fully with respect to Figures 6A to 6B In some embodiments, as discussed more fully with respect to Figures 12C to 12F a spatial encoder (such as spatial encoder 204) can utilize an octree to compress the spatial information for the points of the point cloud.

[0078] In some embodiments, the compressed spatial information may be stored or transmitted together with the compressed attribute information, or may be stored or transmitted separately. In either case, a decoder that receives the compressed attribute information for the points of a point cloud may also receive the compressed spatial information for those points of the point cloud, or may obtain the spatial information for those points of the point cloud.

[0079] A space-filling curve order generator (such as space-filling curve order generator 210) may utilize the spatial information for the points of a point cloud to generate an index order for the points based on the positions at which the points fall along the space-filling curve. For example, Morton codes may be generated for the points of the point cloud. Because the decoder is provided with or otherwise obtains the same spatial information for the points of the point cloud that was available at the encoder, the space-filling curve order determined by the space-filling curve order generator of the encoder (such as space-filling curve order generator 210 of encoder 202) may be the same or similar to the space-filling curve order generated by the space-filling curve order generator of the decoder (such as space-filling curve order generator 228 of decoder 220).

[0080] A prediction / correction evaluator (such as prediction / correction evaluator 206 of encoder 202) may determine a predicted attribute value for a point of a point cloud based on an inverse distance interpolation method using the attribute values of the K nearest neighbor points of the point for which the attribute value is being predicted. The prediction / correction evaluator may also compare the predicted attribute value of the point being evaluated with the original attribute value of that point in the uncompressed point cloud to determine an attribute correction value. In some embodiments, the prediction / correction evaluator (such as prediction / correction evaluator 206 of encoder 202) may adaptively adjust the prediction strategy used to predict the attribute values of points in a neighborhood based on a measure of the variability of the attribute values of points in a given point neighborhood.

[0081] An outgoing data encoder, such as the outgoing data encoder 208 of encoder 202, may encode the attribute correction values and assigned attribute values included in the compressed attribute information file for the point cloud. In some embodiments, the outgoing data encoder (such as the outgoing data encoder 208) may select an encoding context for encoding a value based on the number of symbols included in the value, such as an assigned attribute value or an attribute correction value. In some embodiments, an encoding context including Golomb exponential coding may be used to encode a value with more symbols, while arithmetic coding may be used to encode a value with fewer symbols. In some embodiments, the encoding context may include more than one encoding technique. For example, arithmetic coding may be used to encode a part of the value, while Golomb exponential coding may be used to encode another part of the value. In some embodiments, the encoder (such as encoder 202) may include a context memory such as context memory 216 that stores the encoding context used by the outgoing data encoder such as the outgoing data encoder 208 to encode the attribute correction values and assigned attribute values.

[0082] In some embodiments, the encoder such as encoder 202 may further include an incoming data interface such as incoming data interface 214. In some embodiments, the encoder may receive incoming data from one or more sensors that capture points of the point cloud or capture attribute information associated with the points of the point cloud. For example, in some embodiments, the encoder may receive data from a LIDAR system, a 3D camera, a 3D scanner, etc., and may also receive data from other sensors such as gyroscopes, accelerometers, etc. Additionally, the encoder may receive other data such as the current time from a system clock, etc. In some embodiments, such different types of data may be received by the encoder via the incoming data interface (such as the incoming data interface 214 of encoder 202).

[0083] In some embodiments, the encoder (such as encoder 202) may further include a configuration interface (such as configuration interface 212), where one or more parameters used by the encoder to compress the point cloud may be adjusted via the configuration interface. In some embodiments, the configuration interface (such as configuration interface 212) may be a programmable interface such as an API. The configuration used by the encoder (such as encoder 202) may be stored in a configuration memory (such as configuration memory 218).

[0084] In some embodiments, the encoder (such as encoder 202) may include more or fewer components than Figure 2A shown.

[0085] Figure 2B Components of a decoder according to some embodiments are shown.

[0086] The decoder 220 can be a decoder similar to the decoder 116 shown in Figure 1A The decoder 220 includes an encoded data interface 226, a spatial decoder 222, a spatial filling curve order generator 228, a prediction evaluator 224, a context memory 232, a configuration memory 234, and a decoded data interface 220.

[0087] A decoder (such as decoder 220) can receive an encoded compressed point cloud and / or an encoded compressed attribute information file for points of a point cloud. For example, a decoder (such as decoder 220) can receive a compressed attribute information file, such as Figure 1A the compressed attribute information 112 shown in Figure 3 the compressed attribute information file 300 shown in. The decoder can receive the compressed attribute information file via an encoded data interface (such as encoded data interface 226). The decoder can use the encoded compressed point cloud to determine spatial information for points of the point cloud. For example, the spatial information for points of the point cloud included in the compressed point cloud can be generated by a spatial information generator (such as spatial information generator 222). In some embodiments, the compressed point cloud can be received from a storage device or other intermediate source via an encoded data interface (such as encoded data interface 226), where the compressed point cloud was previously encoded by an encoder (such as encoder 104).

[0088] In some embodiments, the encoded data interface (such as encoded data interface 226) can decode the spatial information. For example, the spatial information may have been encoded using various encoding techniques (such as arithmetic coding, Golomb coding, etc.). The spatial information generator (such as spatial information generator 222) can receive the decoded spatial information from the encoded data interface (such as encoded data interface 226), and can use the decoded spatial information to generate a representation of the geometry of the point cloud being decompressed. For example, the decoded spatial information can be formatted as residual values for use in a subsampled prediction method to reconstruct the geometry of the point cloud to be decompressed. In such cases, the spatial information generator 222 can use the decoded spatial information from the encoded data interface 226 to reconstruct the geometry of the point cloud being decompressed, and the spatial filling curve order generator 228 can determine the spatial filling curve order for the point cloud being decompressed based on the reconstructed geometry of the point cloud being decompressed generated by the spatial information generator 222.

[0089] Once the spatial information for the point cloud is determined and the spatial filling curve order has been determined, the spatial filling curve order can be used by the prediction evaluator of the decoder (such as prediction evaluator 224 of decoder 220) to determine the evaluation order for determining the attribute values of the points of the point cloud. Additionally, the spatial filling curve order can be used by the prediction evaluator (such as prediction evaluator 224) to identify the nearest neighbor points to the point being evaluated.

[0090] The prediction evaluator of the decoder (such as prediction evaluator 224) may select a starting point, such as based on an allocation starting point included in the compressed attribute information file. In some embodiments, the compressed attribute information file may include one or more allocation values for one or more corresponding attributes of the starting point. In some embodiments, the prediction evaluator (such as prediction evaluator 224) may assign values to one or more attributes of the starting point in the decompression model of the point cloud being decompressed based on the allocation values for the starting point included in the compressed attribute information file. The prediction evaluator (such as prediction evaluator 224) may also utilize the assigned values of the attributes of the starting point to determine the attribute values of adjacent points. For example, the prediction evaluator may select an adjacent point to the starting point as the next point to be evaluated, where the adjacent point is selected based on the index order of the points according to the space-filling curve order. Note that since the space-filling curve order is generated at the decoder based on the same or similar spatial information used to generate the space-filling curve order at the encoder, the decoder can determine the same order of points for evaluating the point cloud being decompressed as determined at the encoder by identifying the next nearest neighbor in the index according to the space-filling curve order.

[0091] Once the prediction evaluator has identified the "K" nearest neighbors to the point being evaluated, the prediction evaluator may predict one or more attribute values for one or more attributes of the point being evaluated based on the attribute values of the corresponding attributes of the "K" nearest neighbors. In some embodiments, inverse distance interpolation techniques may be used to predict the attribute values of the point being evaluated based on the attribute values of the adjacent points, where the attribute values of the adjacent points that are closer to the point being evaluated are weighted more heavily than the attribute values of the adjacent points that are farther from the point being evaluated. In some embodiments, the prediction evaluator of the decoder (such as prediction evaluator 224 of decoder 220) may adaptively adjust the prediction strategy for predicting the attribute values of points in a neighborhood based on a measure of the variability of the attribute values of the points in the neighborhood of a given point. For example, in embodiments where adaptive prediction is used, the decoder may mirror the prediction adaptation decisions made at the encoder. In some embodiments, the adaptive prediction parameters may be included in the compressed attribute information received by the decoder, where these parameters are signaled by the encoder that generated the compressed attribute information. In some embodiments, the decoder may utilize one or more default parameters in the absence of signaled parameters, or may infer the parameters based on the received compressed attribute information.

[0092] A prediction evaluator, such as prediction evaluator 224, may apply an attribute correction value to a predicted attribute value to determine an attribute value to include for a point in a decompressed point cloud. In some embodiments, the attribute correction value for an attribute of a point may be included in a compressed attribute information file. In some embodiments, one of a plurality of supported coding contexts may be used to code the attribute correction value, where different coding contexts are selected based on the number of symbols included in the attribute correction value to code different attribute correction values. In some embodiments, a decoder, such as decoder 220, may include a context memory, such as context memory 232, where the context memory stores a plurality of coding contexts that may be used to decode assigned attribute values or attribute correction values that have been coded at the encoder using corresponding coding contexts.

[0093] A decoder, such as decoder 220, may provide a decompressed point cloud generated based on the received compressed point cloud and / or the received compressed attribute information file to a receiving device or application via a decoded data interface, such as decoded data interface 230. The decompressed point cloud may include points of the point cloud and attribute values for attributes of the points of the point cloud. In some embodiments, the decoder may decode some attribute values for attributes of the point cloud without decoding other attribute values for other attributes of the point cloud. For example, the point cloud may include a color attribute for points of the point cloud and may also include other attributes for points of the point cloud, such as velocity. In such cases, the decoder may decode one or more attributes (such as the velocity attribute) of the points of the point cloud without decoding other attributes (such as the color attribute) of the points of the point cloud.

[0094] In some embodiments, the decompressed point cloud and / or the decompressed attribute information file may be used to generate a visual display, such as for a head-mounted display. Additionally, in some embodiments, the decompressed point cloud and / or the decompressed attribute information file may be provided to a decision engine that uses the decompressed point cloud and / or the decompressed attribute information file to make one or more control decisions. In some embodiments, the decompressed point cloud and / or the decompressed attribute information file may be used for various other applications or for various other purposes.

[0095] Figure 3Illustrates an example compressed attribute information file according to some embodiments. The attribute information file 300 includes configuration information 302, point cloud data 304, and point attribute correction values 306. In some embodiments, the point cloud file 300 may be transmitted in batches via multiple data packets. In some embodiments, not all parts shown in the attribute information file 300 may be included in each data packet for transmitting compressed attribute information. In some embodiments, an attribute information file (such as the attribute information file 300) may be stored in a storage device (such as a server implementing an encoder or decoder) or other computing devices. In some embodiments, additional configuration information may include adaptive prediction parameters, such as variability measurement techniques for determining variability measurement results for a point neighborhood, threshold variability values for triggering the use of a specific prediction procedure, one or more parameters for determining the size of the point neighborhood for which variability is to be determined, and so on.

[0096] Figure 4A Illustrates a process for compressing attribute information of a point cloud according to some embodiments.

[0097] At 402, the encoder receives a point cloud that includes attribute information for at least some of the points in the point cloud. The point cloud may be received from one or more sensors that capture the point cloud, or the point cloud may be generated in software. For example, a virtual reality or augmented reality system may have generated the point cloud.

[0098] At 404, the spatial information of the point cloud may be quantified. For example, the X, Y, and Z coordinates of the points in the point cloud. In some embodiments, the coordinates may be rounded to the nearest measurement unit, such as meters, centimeters, millimeters, etc.

[0099] At 406, the quantized spatial information is compressed. In some embodiments, subsampling and subdivision prediction techniques discussed in Figures 6A to 6B more detail may be used to compress the spatial information. Additionally, in some embodiments, K-D tree compression techniques discussed in Figure 7 more detail may be used to compress the spatial information, or octree compression techniques may be used to compress the spatial information. In some embodiments, other suitable compression techniques may be used to compress the spatial information of the point cloud.

[0100] At 408, the compressed spatial information for the point cloud is encoded as a compressed point cloud file or a part of a compressed point cloud file. In some embodiments, the compressed spatial information and the compressed attribute information may be included in a common compressed point cloud file, or may be transmitted or stored as separate files.

[0101] At 412, the spatial information of the received point cloud is used to generate an indexed point order according to a space-filling curve. In some embodiments, the spatial information of the point cloud may be quantified before generating the order according to the space-filling curve. Additionally, in some embodiments where lossy compression techniques are used to compress the spatial information of the point cloud, the spatial information may be lossily encoded and lossily decoded before generating the order according to the space-filling curve. In embodiments where lossy compression is used for the spatial information, encoding and decoding the spatial information at the encoder ensures that the order according to the space-filling curve generated at the encoder will match the order according to the space-filling curve that will be generated at the decoder using the decoded spatial information that was previously lossily encoded.

[0102] Additionally, in some embodiments, at 410, the attribute information for the points of the point cloud may be quantified. For example, the attribute values may be rounded to integers or to a specific measurement increment. In some embodiments where the attribute values are integers, such as when integers are used to convey string values such as "walking", "running", "driving", etc., the quantification at 410 may be omitted.

[0103] At 414, an attribute value for the starting point is assigned. The assigned attribute value for the starting point, along with an attribute correction value, is encoded in the compressed attribute information file. Since the decoder predicts the attribute values based on the distance to adjacent points and the attribute values of the adjacent points, at least one attribute value for at least one point is explicitly encoded in the compressed attribute file. In some embodiments, the points of the point cloud may include multiple attributes, and in such embodiments, at least one attribute value for each type of attribute may be encoded for at least one point of the point cloud. In some embodiments, the starting point may be the first point evaluated when determining the order according to the space-filling curve at 412. In some embodiments, the encoder may encode data indicating the spatial information for the starting point and / or other markers as to which point of the point cloud is one or more starting points. Additionally, the encoder may encode the attribute values for one or more attributes of the starting point.

[0104] At 416, the encoder determines an evaluation order for predicting the attribute values for points in the point cloud other than the starting point, and the prediction and determination of the attribute correction value may be referred to herein as "evaluating" the attributes of the points. The evaluation order may be determined based on the order according to the space-filling curve.

[0105] At 418, an adjacent point of the starting point or a subsequent point being evaluated is selected. In some embodiments, the adjacent point may be selected based on the next adjacent point to be evaluated being the next point in the point index order according to the space-filling curve.

[0106] At 420, the "K" nearest neighbor points to the point currently being evaluated are determined. The parameter "K" can be a configurable parameter selected by the encoder or can be provided to the encoder as a user-configurable parameter. To select the "K" nearest neighbor points, the encoder can identify the first "K" nearest points to the point being evaluated based on the point index order determined at 412 and the corresponding distances between the points. For example, instead of determining the absolute nearest neighbor point to the point being evaluated, the encoder can select a set of points in the point cloud that have index values in the user-defined search range (e.g., 8, 16, 32, 64, etc.) of the index value of the particular point being evaluated in the index according to the space-filling curve. Then, the encoder can utilize the distances within the set of points to select the "K" nearest neighbor points for prediction. In some embodiments, only points with assigned attribute values or for which predicted attribute values have been determined can be included in these "K" nearest neighbor points. In some embodiments, various numbers of points can be identified. For example, in some embodiments, "K" can be 5 points, 10 points, 16 points, etc. Since the point cloud includes points in 3D space, a particular point may have multiple adjacent points in multiple planes. In some embodiments, the encoder and decoder can be configured to identify points as the "K" nearest neighbor points regardless of whether a value has been predicted for that point. Additionally, in some embodiments, the attribute value for a point used in prediction can be a previous predicted attribute value or a corrected predicted attribute value that has been corrected based on an applied attribute correction value. In either case, the encoder and decoder can be configured to apply the same rules when identifying the "K" nearest neighbor points and when predicting the attribute value of a point based on the attribute values of the "K" nearest neighbor points.

[0107] At 422, one or more attribute values are determined for each attribute of the point currently being evaluated. The attribute values can be determined based on inverse distance interpolation. Inverse distance interpolation can interpolate the predicted attribute value based on the attribute values of the "K" nearest neighbor points. The attribute values of the "K" nearest neighbor points can be weighted based on the corresponding distances between the respective ones of the "K" nearest neighbor points and the point being evaluated. The attribute values of the adjacent points at a closer distance to the point currently being evaluated can be weighted more heavily than the attribute values of the adjacent points at a farther distance from the point currently being evaluated.

[0108] At 424, an attribute correction value for one or more predicted attribute values for the point currently being evaluated is determined. The attribute correction value can be determined based on comparing the predicted attribute value with the corresponding attribute value for the same (or similar) point in the point cloud before attribute information compression. In some embodiments, the quantized attribute information (such as the quantized attribute information generated at 410) can be used to determine the attribute correction value. In some embodiments, the attribute correction value can also be referred to as a "residual error", where the residual error indicates the difference between the predicted attribute value and the actual attribute value.

[0109] At 426, it is determined whether there are additional points in the point cloud for which an attribute correction value is to be determined. If there are additional points to be evaluated, the process returns to 418 and the next point to be evaluated in the evaluation order is selected. The process can repeat steps 418 to 426 until all or a portion of all the points in the point cloud have been evaluated to determine the predicted attribute values and the attribute correction values for the predicted attribute values.

[0110] At 428, the determined attribute correction values, the assigned attribute values, and any configuration information (such as the parameter "K") used to decode the compressed attribute information file are encoded.

[0111] Adaptive Prediction Attribute

[0112] In some embodiments, the encoder as described above can also adaptively change the prediction strategy and / or the number of points used in a given prediction strategy based on the attribute values of neighboring points. Additionally, the decoder can similarly adaptively change the prediction strategy and / or the number of points used in a given prediction strategy based on the reconstructed attribute values of neighboring points.

[0113] For example, the point cloud can include points representing a road, where the road is black and there are white stripes on the road. The default nearest neighbor prediction strategy can be adaptively changed to account for the variability in the attribute values of the points representing the white lines and the black road. Since the difference in the attribute values of these points is large, the default nearest neighbor prediction strategy may result in blurring of the white lines and / or high residual values that reduce compression efficiency. However, the updated prediction strategy can account for this variability by selecting a more suitable prediction strategy and / or by using fewer points in the K-nearest neighbor prediction. For example, for the black road, the white line points are not used in the K-nearest neighbor prediction.

[0114] In some embodiments, before predicting the attribute value of point P, the encoder or decoder may calculate the variability of the attribute values of the points in the neighborhood of point P, e.g., the K nearest neighbor points. In some embodiments, the variability may be calculated based on the variance, the maximum difference between any two attribute values (or reconstructed attribute values) of the points adjacent to point P. In some embodiments, the variability may be calculated based on the weighted average of the adjacent points, where the weighted average takes into account the distance of the adjacent points to point P. In some embodiments, the variability of a set of adjacent points may be calculated based on the weighted average of the attributes of the adjacent points and taking into account the distances of the adjacent points. For example,

[0115] Variability = E[(X - weighted average(X)) 2

[0116] In the above formula, E is the average attribute value of the points in the neighborhood of point P, and weighted average(X) is the weighted average of the attribute values of the points in the neighborhood of point P, where the distances of the adjacent points to point P are taken into account. In some embodiments, the variability may be calculated as the maximum difference compared to the average value E(X) of the attribute, the weighted average of the attribute, weighted average(X), or the median of the attribute (median(X)). In some embodiments, the variability may be calculated using the average value of the values corresponding to x percentage (e.g., x = 10), which has the largest difference compared to the average value E(X) of the attribute, the weighted average of the attribute, weighted average(X), or the median of the attribute (median(X)).

[0117] In some embodiments, if the calculated variability of the attributes of the points in the neighborhood of point P is greater than a threshold, rate-distortion optimization may be applied. For example, rate-distortion optimization may reduce the number of adjacent points used in the prediction or switch to a different prediction technique. In some embodiments, the threshold may be explicitly written into the bitstream. Additionally, in some embodiments, the threshold may be adaptively adjusted for each point cloud or sub-block of the point cloud or for the number of points to be encoded. For example, the threshold may be included in the compressed attribute information file 350 as additional configuration information included in the configuration information 302, as Figure 3 described, or may be included in the compressed attribute file 1150 as additional configuration information included in the configuration information 1152, as described below with respect to Figure 11B described.

[0118] In some embodiments, different distortion metrics such as sum of squared errors, weighted sum of squared errors, sum of absolute differences, or weighted sum of absolute differences may be used in the rate-distortion optimization procedure.

[0119] ​In some embodiments, the distortion can be calculated independently for each attribute or for multiple attributes corresponding to the same sample, and can be considered and weighted appropriately. For example, the distortion values for R, G, B or Y, U, V can be calculated and then linearly or non-linearly combined together to generate a total distortion value.

[0120] In some embodiments, advanced techniques for rate-distortion quantization can also be considered, such as grid-based quantization, where instead of considering individual points in isolation, multiple points are jointly encoded. For example, the encoding process can choose a method that uses a cost function of the form J = D + λ * rate to encode all these multiple points, where D is the overall distortion of all these points and rate is the overall rate cost of encoding these points.

[0121] In some embodiments, the encoder (such as encoder 202) can explicitly encode the index value of the selected prediction strategy for the point cloud, the level of detail of the point cloud, or a set of points within the level of detail of the point cloud, where the decoder can access an instance of the index and can determine the selected prediction strategy based on the received index value. The decoder can apply the selected prediction strategy to the set of points to which the rate-distortion optimization process is being applied. In some embodiments, if no rate-distortion optimization procedure is specified in the encoded bitstream, there can be a default prediction strategy and the decoder can apply the default prediction strategy. Additionally, in some embodiments, if a variability threshold is not met, the default prediction strategy can be applied.

[0122] For example, Figure 4B illustrates the use of adaptive distance-based prediction to predict attribute values as part of compressing the attribute information of a point cloud, according to some embodiments.

[0123] In some embodiments that employ adaptive distance-based prediction, as Figure 4A described in elements 420 and 422 of, predicting the attribute value of a point can also include the step of selecting a prediction procedure to be used for predicting the attribute value of the point, such as 450 - 456. In some embodiments, the selected prediction procedure can be a K-nearest neighbor prediction procedure, as described herein and with respect to Figure 4Aas described by element 420 therein. In some embodiments, the selected prediction procedure may be a modified K-nearest neighbor prediction procedure, where the number of nearest neighbors used for performing adaptive prediction includes fewer points compared to the number of points for which property values for a less variable portion of the point cloud are predicted. In some embodiments, if the variability of adjacent points exceeds a threshold associated with the selected prediction procedure, the prediction procedure may be: using only the property values of the nearest points to the point for which the property value is being predicted for the point for which the property value is being predicted. In some embodiments, other prediction procedures may be used depending on the variability of points in the neighborhood of the point for which the property value is being predicted. For example, in some embodiments, other prediction procedures may be used, such as non-distance-based interpolation procedures, such as barycentric interpolation, natural neighbor interpolation, moving least squares interpolation, or other suitable interpolation techniques.

[0124] At 450, the encoder identifies a set of adjacent points in the neighborhood of the point for which the property value of the point cloud is being predicted. In some embodiments, the K-nearest neighbor technique as described herein may be used to identify the set of adjacent points in the neighborhood. In some embodiments, the points to be used for determining variability may be identified in other ways. For example, in some embodiments, the neighborhood of points used for variability analysis may be defined to include more or fewer points or points within a greater or lesser distance from a given point compared to using the K nearest adjacent points for predicting property values based on inverse distance-based interpolation. In some embodiments where the parameters for identifying the neighborhood points used for determining variability are different from the parameters for K-nearest neighbor prediction, different parameters or data from which different parameters may be determined are signaled in the bitstream encoded by the encoder.

[0125] At 452, the variability of the property values of the adjacent points is determined. In some embodiments, each property value variability may be determined separately. For example, for a point having R, G, B property values, each property value (e.g., each of R, G, and B) may have its respective variability determined separately. Additionally, in some embodiments, grid quantization may be used, where a set of properties having related values such as RGB may be determined as a common variability. For example, in the example discussed above regarding the white stripe on the black road, the large variability of R may also apply to B and G, so it is not necessary to determine the variability of each of R, G, and B separately. Instead, the related property values may be considered as a group, and the common variability of the related properties may be determined.

[0126] In some embodiments, the following variability techniques for determining the variability of attributes in the neighborhood of point P may be used: sum of squared error variability technique, sum of squared error distance weighted variability technique, sum of absolute differences variability technique, sum of absolute differences distance weighted variability technique, or other suitable variability techniques. In some embodiments, the encoder may select the variability technique to be used for a given point P and may encode the index value of the index of the variability technique in the bitstream encoded by the encoder, where the decoder includes the same index and may determine the variability technique to be used for point P based on the encoded index value.

[0127] At 454 to 456, it is determined whether the variability determined at 452 exceeds one or more variability thresholds. If so, the corresponding prediction technique corresponding to the exceeded variability threshold is used to predict one or more attribute values for point P. In some embodiments, multiple prediction procedures may be supported. For example, if the first variability threshold is exceeded, element 458 indicates using the first prediction procedure, and if another variability threshold is exceeded, element 460 indicates using another prediction procedure. Further, if the variability thresholds 1 to N are not exceeded, 462 indicates using a default prediction procedure, such as an unmodified k-nearest neighbor prediction procedure. In some embodiments, in addition to the default prediction procedure, a single variability threshold and a single alternative prediction procedure may be used. In some embodiments, any number "N" of variability thresholds and corresponding prediction procedures may be used.

[0128] For example, in some embodiments, if the first variability threshold is exceeded, the first prediction procedure may use fewer neighboring points than those used in the default k-nearest neighbor prediction procedure. Further, if the second variability threshold is exceeded, the second prediction procedure may use only the nearest point to determine the attribute value of point P. Thus, in such embodiments, medium variability may result in omitting some outliers under the first prediction procedure, and high variability may result in omitting all points except the nearest neighbor from the prediction procedure, while if the variability is low, k nearest neighbors are used in the default prediction procedure.

[0129] Figures 4C to 4E Parameters that may be determined or selected by an encoder and signaled via compressed attribute information of a point cloud are shown, according to some embodiments.

[0130] At Figure 4C At 470, the encoder may select the variability measurement technique to be used for determining the attribute variability of points in the neighborhood of point P for which the attribute value is being predicted. In some embodiments, the encoder may utilize a rate-distortion optimization framework to determine which variability measurement technique to use. At 472, the encoder may include in the bitstream encoded by the encoder a signal indicating which variability technique has been selected.

[0131] AtFigure 4D In [description], at 480, the encoder may determine a variability threshold for points in the neighborhood of point P for which it is predicting an attribute value. In some embodiments, the encoder may utilize a rate-distortion optimization framework to determine the variability threshold. At 482, the encoder may include in the bitstream encoded by the encoder a signal indicating which variability threshold the encoder used to perform the prediction.

[0132] In Figure 4E In [description], at 490, the encoder may determine or select a neighborhood size to be used for determining variability. For example, the encoder may use rate-distortion optimization techniques to determine the size of the point neighborhood to be used for determining the variability of point P. At 492, the encoder may include in the bitstream encoded by the encoder one or more values for defining the neighborhood size. For example, the encoder may signal the minimum distance from point P, the maximum distance from point P, the total number of adjacent points to be included, etc., and these parameters may define which points are included among the neighborhood points considered for determining the variability of point P.

[0133] In some embodiments, one or more of the variability technique, the variability threshold, or the neighborhood size may not be signaled and instead may be determined at the decoder using predetermined parameters known to both the encoder and the decoder. In some embodiments, the decoder may infer one or more of the variability technique, the variability threshold, or the neighborhood size based on other data such as the spatial information of the point cloud.

[0134] Once the attribute value is predicted using the appropriate corresponding prediction procedure at 858 - 862, the decoder may proceed to 820 and apply the attribute correction value received in the encoded bitstream to adjust the predicted attribute value. In some embodiments, using adaptive prediction as described herein at the encoder and the decoder may reduce the number of bits required to encode the attribute correction value and may also reduce the distortion of the reconstructed point cloud using the prediction procedure and the signaled attribute correction value at the decoder.

[0135] Exemplary Process for Encoding Attribute Values and / or Attribute Correction Values

[0136] Various coding techniques may be used to encode the attribute correction value, the assigned attribute value, and any configuration information.

[0137] For example, Figure 5illustrates a process for encoding an attribute correction value according to some embodiments. At 502, the attribute correction value for the point whose value (e.g., the attribute correction value) is being encoded is converted to an unsigned value. For example, in some embodiments, an odd number can be assigned to an attribute correction value that can be negative, and an even number can be assigned to an attribute correction value that can be positive. Thus, whether the attribute correction value is positive or negative can be implied based on whether the value of the attribute correction value is even or odd. In some embodiments, the assigned attribute value can also be converted to an unsigned value. In some embodiments, the attribute values can all be positive, such as in the case of being assigned as an integer representing a string value, such as "walking", "running", "driving", etc. In such cases, 502 can be omitted.

[0138] At 504, an encoding context is selected to encode the first value for the point. For example, the value can be an assigned attribute value, or it can be an attribute correction value. The encoding context can be selected from multiple supported encoding contexts. For example, the context memory (such as context memory 216) of an encoder (such as encoder 202) as Figure 2A shown can store multiple supported encoding contexts for encoding the attribute values or attribute correction values of the points of a point cloud. In some embodiments, the encoding context can be selected based on the characteristics of the value to be encoded. For example, some encoding contexts can be optimized for encoding values with certain characteristics, while other encoding contexts can be optimized for encoding values with other characteristics.

[0139] In some embodiments, the encoding context can be selected based on the number or type of symbols included in the value to be encoded. For example, arithmetic coding techniques can be used to encode values with fewer or less diverse symbols, while exponential Golomb coding techniques can be used to encode values with more symbols or more diverse symbols. In some embodiments, the encoding context can use more than one encoding technique to encode a part of the value. For example, in some embodiments, the encoding context can indicate that arithmetic coding techniques are to be used to encode a part of the value, while Golomb coding techniques are to be used to encode another part of the value. In some embodiments, the encoding context can indicate that a first encoding technique (such as arithmetic coding) is to be used to encode a part of the value below a threshold, while another encoding technique (such as exponential Golomb coding) is to be used to encode another part of the value above the threshold. In some embodiments, the context memory can store multiple encoding contexts, where each encoding context is suitable for a value with specific characteristics.

[0140] At 506, the encoding context selected at 504 can be used to encode the first value (or additional value) for that point. At 508, it is determined whether there is an additional value for that point to be encoded. If there is an additional value for that point to be encoded, then at 506, the same selected encoding technique as selected at 504 can be used to encode the additional value. For example, a point may have "red", "green", and "blue" color attributes. Since the differences between adjacent points in the R, G, B color space may be similar, the attribute correction values for the red attribute, green attribute, and blue attribute may be similar. Thus, in some embodiments, the encoder can select an encoding context for encoding the attribute correction value for the first of the color attributes (e.g., the red attribute), and can use the same encoding context to encode the attribute correction values for other color attributes (such as the green attribute and the blue attribute).

[0141] At 510, the encoded values (such as the encoded assigned attribute values and the encoded attribute correction values) can be included in the compressed attribute information file. In some embodiments, the encoded values can be included in the compressed attribute information file according to an evaluation order determined based on the space-filling curve order for the point cloud. Thus, the decoder may be able to determine which encoded value corresponds to which attribute of which point based on the order in which the encoded values are included in the compressed attribute information file. Additionally, in some embodiments, data can be included in the compressed attribute information file that indicates the corresponding ones of the encoding contexts selected to encode the corresponding ones of the values for the points.

[0142] Exemplary Process for Encoding Spatial Information

[0143] Figures 6A to 6B An exemplary process for compressing spatial information of a point cloud according to some embodiments is shown.

[0144] At 602, the encoder receives a point cloud. The point cloud can be a captured point cloud from one or more sensors, or can be a generated point cloud (such as a point cloud generated by a graphics application). For example, 604 shows the points of an uncompressed point cloud.

[0145] At 606, the encoder subsamples the received point cloud to generate a subsampled point cloud. The subsampled point cloud can include fewer points than the received point cloud. For example, the received point cloud can include hundreds of points, thousands of points, or millions of points, and the subsampled point cloud can include dozens of points, hundreds of points, or thousands of points. For example, 608 shows the subsampled points of the point cloud received at 602, e.g., subsampling the points of the point cloud in 604.

[0146] In some embodiments, an encoder may encode and decode a subsampled point cloud to generate a representative subsampled point cloud that a decoder will encounter when decoding the compressed point cloud. In some embodiments, the encoder and decoder may perform a lossy compression / decompression algorithm to generate the representative subsampled point cloud. In some embodiments, the spatial information of the points of the subsampled point cloud may be quantized as part of generating the representative subsampled point cloud. In some embodiments, the encoder may utilize lossless compression techniques and may omit encoding and decoding of the subsampled point cloud. For example, when using lossless compression techniques, the original subsampled point cloud may represent the subsampled point cloud that the decoder will encounter because, in lossless compression, data may not be lost during compression and decompression.

[0147] At 610, the encoder identifies subdivision locations between points of the subsampled point cloud based on configuration parameters selected for the compressed point cloud or based on fixed configuration parameters. Configuration parameters of the non-fixed configuration parameters used by the encoder are communicated to the encoder by including the values of these configuration parameters in the compressed point cloud. Thus, the decoder can determine the same subdivision locations as evaluated by the encoder based on the subdivision configuration parameters included in the compressed point cloud. For example, 612 shows the identified subdivision locations between adjacent points of the subsampled point cloud.

[0148] At 614, the encoder determines for some of the subdivision locations whether to include a point or not at the subdivision location in the decompressed point cloud. Data indicating this determination is encoded in the compressed point cloud. In some embodiments, the data indicating this determination can be a single bit, which means that a point will be included if "true" and that a point will not be included if "false". Additionally, the encoder may determine that the points to be included in the decompressed point cloud will be repositioned relative to the subdivision locations in the decompressed point cloud. For example 616 shows some points to be repositioned relative to the subdivision locations. For such points, the encoder may also encode data indicating how to reposition the points relative to the subdivision locations. In some embodiments, the position correction information may be quantized and entropy encoded. In some embodiments, the position correction information may include ΔX, ΔY, and / or ΔZ values that indicate how to reposition the point relative to the subdivision location. In other embodiments, the position correction information may include a single scalar value that corresponds to the normal component of the position correction information and is calculated as follows:

[0149] ΔN = ([X A ,Y A ,Z A - [X,Y,Z]) · [normal vector]

[0150] In the above formula, ΔN is a scalar value indicating the position correction information, which is the point position repositioned or adjusted relative to the subdivision location (e.g., [XA , Y A , Z A ) and the original subdivision position (e.g., [X, Y, Z]). The vector product of this vector difference and the normal vector at the subdivision position produces a scalar value ΔN. Since the decoder can determine the normal vector at the subdivision position and can determine the coordinates of the subdivision position (e.g., [X, Y, Z]), by solving the above formula for the adjusted position, the decoder can also determine the coordinates of the adjusted position (e.g., [X A , Y A , Z A ), which coordinates represent the repositioned position of the point relative to the subdivision position. In some embodiments, the position correction information can be further decomposed into a vertical component and one or more additional tangential components. In such embodiments, the vertical component (e.g., ΔN) and the one or more tangential components can be quantized and encoded to be included in the compressed point cloud.

[0151] In some embodiments, the encoder can determine whether to include one or more additional points (in addition to the points included at the subdivision position or the points included at the repositioned positions relative to the subdivision position) in the decompressed point cloud. For example, if the original point cloud has an irregular surface or shape such that the subdivision positions between the points in the subsampled point cloud do not adequately represent the irregular surface or shape, the encoder can determine to include one or more additional points in addition to the points determined to be included at the subdivision position or the points repositioned relative to the subdivision position in the decompressed point cloud. Additionally, the encoder can determine whether to include one or more additional points in the decompressed point cloud based on system constraints such as a target bit rate, a target compression ratio, a quality target metric, etc. In some embodiments, the bit budget can change due to changing conditions such as network conditions, processor load, etc. In such embodiments, the encoder can adjust the amount of additional points encoded to be included in the decompressed point cloud based on the changed bit budget. In some embodiments, the encoder can include additional points such that the bit budget is consumed without exceeding the bit budget. For example, when the bit budget is high, the encoder can include more additional points to consume the bit budget (and improve quality), and when the bit budget is low, the encoder can include fewer additional points such that the bit budget is consumed without exceeding the bit budget.

[0152] In some embodiments, the encoder may also determine whether to perform additional subdivision iterations. If so, points determined to be included, repositioned, or additionally included in the decompressed point cloud are considered, and the process returns to 610 to identify new subdivision positions for an updated subsampled point cloud that includes the points determined to be included, repositioned, or additionally included in the decompressed point cloud. In some embodiments, the number of subdivision iterations (N) to be performed may be a fixed or configurable parameter of the encoder. In some embodiments, different subdivision iteration values may be assigned to different portions of the point cloud. For example, the encoder may consider the point view from which it is viewing the point cloud, and may perform more subdivision iterations on points in the foreground of the point cloud as viewed from that point view, and may perform fewer subdivision iterations on points in the background of the point cloud as viewed from that point view.

[0153] At 618, the spatial information for the subsampled points of the point cloud is encoded. Additionally, the subdivision position inclusion and repositioning data is encoded. Additionally, any configurable parameters selected by the encoder or provided to the encoder by the user are encoded. The compressed point cloud may then be sent to a receiving entity as one compressed point cloud file, multiple compressed point cloud files, or the compressed point cloud may be packaged and transmitted via multiple data packets to a receiving entity such as a decoder or storage device. In some embodiments, the compressed point cloud may include both compressed spatial information and compressed attribute information. In other embodiments, the compressed spatial information and compressed attribute information may be included in separate compressed point cloud files.

[0154] Figure 7 Another example process for compressing spatial information of a point cloud in accordance with some embodiments is shown.

[0155] In some embodiments, other spatial information compression techniques may be used in addition to Figures 6A to 6B the subsampling and predictive spatial information techniques described in. For example, a spatial encoder (such as spatial encoder 204) or a spatial decoder (such as spatial decoder 222) may utilize other spatial information compression techniques (such as K-D tree spatial information compression techniques). For example, the compressed spatial information at 406 in FIG. 4 may be performed using subsampling and predictive techniques similar to those described in Figure 6A -B, may be performed using K-D tree spatial information compression techniques similar to those described in Figure 7 or may be performed using another suitable spatial information compression technique.

[0156] In the K-D tree spatial information compression technique, a point cloud including spatial information can be received at 702. In some embodiments, the spatial information may already be pre-quantized or may be further quantized after being received. For example, 718 shows a captured point cloud that can be received at 702. For simplicity, 718 shows the point cloud in two dimensions. However, in some embodiments, the received point cloud may include points in 3D space.

[0157] At 704, the spatial information of the received point cloud is used to construct a K-dimensional tree or K-D tree. In some embodiments, the K-D tree can be constructed by dividing a space such as a 1D, 2D, or 3D space of a point cloud into two halves in a predetermined order. For example, a 3D space including the points of a point cloud can be initially divided into two halves via a plane intersecting one of the three axes (such as the X axis). Then, subsequent divisions can divide the resulting space along another one of the three axes (such as the Y axis). Then, another division can divide the resulting space along another one of the axes (such as the Z axis). Each time a division is performed, the number of points included in the sub-cell created by the division can be recorded. In some embodiments, the number of points in only one of the two sub-cells resulting from the division can be recorded. This is because the number of points included in the other sub-cell can be determined by subtracting the number of points in the recorded sub-cell from the total number of points in the parent cell before the division.

[0158] The K-D tree can include a sequence of the number of points included in the cells resulting from the sequential division of the space including the points of the point cloud. In some embodiments, constructing the K-D tree can include continuing to subdivide the space until only a single point is included in each lowest-level sub-cell. The K-D tree can be transmitted as a sequence of the number of points in the sequential cells resulting from the sequential division. The decoder can be configured with information indicating the subdivision sequence followed by the encoder. For example, the encoder can follow a predefined division sequence until only a single point remains in each lowest-level sub-cell. Because the decoder may know the division order followed in constructing the K-D tree and the number of points generated by each subdivision (transmitted to the decoder as compressed spatial information), the decoder may be able to reconstruct the point cloud.

[0159] For example, 720 shows a simplified example of K-D compression in a two-dimensional space. The initial space includes seven points. This can be considered as the first parent cell, and the K-D tree can be encoded with the number of points "7" as the first number of the K-D tree, which indicates that there are a total of seven points in the K-D tree. The next step can be to divide the space along the X-axis, resulting in two sub-cells, with the left sub-cell having three points and the right sub-cell having four points. The K-D tree can include the number of points in the left sub-cell, for example, including "3" as the next number of the K-D tree. Remember that the number of points in the right sub-cell can be determined by subtracting the number of points in the left sub-cell from the number of points in the parent cell. Further, the space can be divided along the Y-axis an additional time, such that each of the left and right sub-cells is divided into two halves, becoming lower-level sub-cells. Similarly, the number of points included in the left lower-level sub-cells can be included in the K-D tree, for example, "0" and "1". Then, the next step can be to divide the non-zero lower-level sub-cells along the X-axis, and record the number of points in each of the lower-level left sub-cells in the K-D tree. This process can continue until only a single point remains in the lowest-level sub-cells. The decoder can use the inverse process to reconstruct the point cloud based on the sequence of the total number of points in each left sub-cell of the received K-D tree.

[0160] At 706, an encoding context is selected for encoding the number of points for the first cell of the K-D tree (e.g., the parent cell including seven points). In some embodiments, the context memory can store hundreds or thousands of encoding contexts. In some embodiments, the highest number of points encoding context can be used to encode a cell including more points than the highest number of points encoding context. In some embodiments, the encoding context can include arithmetic coding, Golomb exponential coding, or a combination of both. In some embodiments, other coding techniques can be used. In some embodiments, the arithmetic coding context can include probabilities for specific symbols, where different arithmetic coding contexts include different symbol probabilities.

[0161] At 708, the number of points for the first cell is encoded according to the selected encoding context.

[0162] At 710, an encoding context for encoding the sub-cells is selected based on the number of points included in the parent cell. The encoding context for the sub-cells can be selected in a similar manner as for the parent cell at 706.

[0163] At 712, the number of points included in the sub-cell is encoded according to the selected coding context selected at 710. At 714, it is determined whether there is an additional lower-level sub-cell to be encoded in the K-D tree. If so, the process returns to 710. If not, at 716, the number of encoded points in the parent cell and the sub-cell is included in a compressed spatial information file (such as a compressed point cloud). The encoded values are sorted in the compressed spatial information file such that the decoder can reconstruct the point cloud based on the number of points in each parent cell and sub-cell and the order of the number of points in the corresponding cells included in the compressed spatial information file.

[0164] In some embodiments, the number of points in each cell can be determined and then encoded as a group at 716. Alternatively, in some embodiments, the number of points in a cell can be encoded after determination without waiting to determine the total number of points in all sub-cells.

[0165] Exemplary Decoding Process

[0166] Figure 8 illustrates an example process for decompressing compressed attribute information of a point cloud according to some embodiments.

[0167] At 802, the decoder receives compressed attribute information for the point cloud, and at 804, the decoder receives compressed spatial information for the point cloud. In some embodiments, the compressed attribute information and the compressed spatial information can be included in one or more common files or separate files.

[0168] At 806, the decoder decompresses the compressed spatial information. The compressed spatial information may have been compressed according to subsampling and prediction techniques, and the decoder can perform similar subsampling, prediction, and prediction correction actions as those performed at the encoder and further apply the correction values to the predicted point positions to generate an uncompressed point cloud from the compressed spatial information. In some embodiments, the compressed spatial information can be compressed in a K-D tree format, and the decoder can generate a decompressed point cloud based on the encoded K-D tree included in the received spatial information. In some embodiments, the compressed spatial information may have been compressed using octree techniques, and octree decoding techniques can be used to generate the decompressed spatial information for the point cloud. In some embodiments, other spatial information compression techniques may have been used, and these techniques can be decompressed via the decoder.

[0169] At 808, the decoder generates an order of points of the point cloud based on a space-filling curve. For example, it can be via the encoding data interface of the decoder (such as Figure 2Breceive compressed spatial information and / or compressed attribute information via the encoding data interface 226 of the decoder 220 as shown. A spatial decoder (such as the spatial decoder 222) may decompress the compressed spatial information, and a space filling curve order generator (such as the space filling curve order generator 228) may generate a space filling curve order based on the decompressed spatial information.

[0170] At 810, a prediction evaluator of the decoder (such as the prediction evaluator 224 of the decoder 220) may assign an attribute value to the starting point based on the assigned attribute value included in the compressed attribute information. In some embodiments, the compressed attribute information may determine a point as the starting point for generating the space filling curve order and for predicting the attribute value of a point according to the evaluation order based on the space filling curve order. The one or more assigned attribute values for the starting point may be included in the decompressed attribute information for the decompressed point cloud.

[0171] At 812, the prediction evaluator of the decoder or another decoder component determines an evaluation order for at least the next point to be evaluated after the starting point. In some embodiments, an evaluation order may be determined for all or multiple points among the points, or in other embodiments, the evaluation order may be confirmed point by point when determining the attribute value of a point. Points may be evaluated in order based on the minimum distance between consecutive points being evaluated. For example, the adjacent point with the shortest distance from the starting point compared to other adjacent points may be selected as the next point to be evaluated after the starting point. In a similar manner, other points to be evaluated may then be selected based on the shortest distance from the most recently evaluated point. At 814, the next point to be evaluated is selected. In some embodiments, 812 and 814 may be performed together.

[0172] At 816, the prediction evaluator of the decoder determines the "K" nearest neighbors of the point being evaluated. In some embodiments, only when the adjacent points have been assigned or predicted attribute values can they be included in the "K" nearest neighbors. In other embodiments, adjacent points may be included in the "K" nearest neighbors regardless of whether they have been assigned or have had their attribute values predicted. In such embodiments, the encoder may follow similar rules as the decoder, i.e., whether to include points without predicted values as adjacent points when identifying the "K" nearest neighbors.

[0173] At 818, predictive attribute values are determined for one or more attributes of the point being evaluated, based on the attribute values of the "K" nearest neighbor points and the distances between the point being evaluated and the respective ones of the "K" nearest neighbor points. In some embodiments, inverse distance interpolation techniques can be used to predict the attribute values, where the attribute values of points closer to the point being evaluated are weighted more heavily than the attribute values of points farther from the point being evaluated. The attribute prediction techniques used by the decoder can be the same as the attribute prediction techniques used by the encoder that compresses the attribute information.

[0174] At 820, a prediction evaluator of the decoder can apply an attribute correction value to the predicted attribute value of the point to correct the attribute value. The attribute correction value can cause the attribute value to match or nearly match the attribute value of the original point cloud before compression. In some embodiments where the point has more than one attribute, 818 and 820 can be repeated for each attribute of the point. In some embodiments, some of the attribute information can be decompressed without decompressing all of the attribute information for the point cloud or the point. For example, the point can include velocity attribute information and color attribute information. The velocity attribute information can be decoded without decoding the color attribute information, and vice versa. In some embodiments, the application using the compressed attribute information can indicate which attributes are to be decompressed for the point cloud.

[0175] At 822, it is determined whether there are additional points to be evaluated. If so, the process returns to 814 and the next point to be evaluated is selected. If there are no additional points to be evaluated, then at 824, decompressed attribute information is provided, such as a decompressed point cloud, where each point includes spatial information and one or more attributes.

[0176] In some embodiments, the decoder can perform a complementary adaptive prediction process as described above for the encoder in Figure 4B For example, Figure 8B illustrates using adaptive distance-based prediction to predict attribute values as part of decompressing the attribute information of a point cloud, according to some embodiments.

[0177] At 850, the decoder identifies a set of neighboring points of a neighborhood of points in the point cloud that are positive for its predicted attribute values. In some embodiments, a K-nearest neighbor technique as described herein may be used to identify the set of neighboring points of the neighborhood. In some embodiments, the points to be used for determining variability may be identified in other ways. For example, in some embodiments, the neighborhood of points used for variability analysis may be defined to include more or fewer points or points within a greater or lesser distance from a given point than those used to predict attribute values based on inverse distance-based interpolation using the K nearest neighboring points. In some embodiments where the parameters for identifying the neighborhood points used for determining variability are different from the parameters for K-nearest neighbor prediction, different parameters or data from which different parameters may be determined are signaled in the bitstream encoded by the encoder and received at the decoder.

[0178] At 852, the variability of the attribute values of the neighboring points is determined. In some embodiments, each attribute value variability may be determined separately. For example, for points having R, G, B attribute values, each attribute value (e.g., each of R, G, and B) may have its corresponding variability determined separately. Additionally, in some embodiments, grid quantization may be used, where a set of attributes having related values such as RGB may be determined as a common variability. For example, in the example discussed above regarding white stripes on a black road, a large variability in R may also apply to B and G, so it is not necessary to determine the variability of each of R, G, and B separately. Instead, the related attribute values may be considered as a group, and the common variability of the related attributes may be determined.

[0179] In some embodiments, the following may be used to determine the variability of attributes in the neighborhood of point P: sum of squared error variability technique, sum of squared error distance weighted variability technique, sum of absolute differences variability technique, sum of absolute differences distance weighted variability technique, or other suitable variability techniques. In some embodiments, the decoder may utilize the variability technique to be used for a given point P that is signaled. In some embodiments, the decoder may determine which variability technique to use based on an index value encoded in the bitstream, where the index value is an index for the variability technique, where the decoder includes the same index as the encoder, and may determine the variability technique to be used for point P based on the encoded index value.

[0180] At 854 to 856, it is determined whether the variability determined at 852 exceeds one or more variability thresholds. If so, the corresponding prediction technique corresponding to the exceeded variability threshold is used to predict one or more attribute values for point P. In some embodiments, multiple prediction programs may be supported. For example, if the first variability threshold is exceeded, element 858 indicates using the first prediction program, and if another variability threshold is exceeded, element 860 indicates using another prediction program. Additionally, if the variability thresholds 1 to N are not exceeded, 862 indicates using a default prediction program, such as an unmodified k-nearest neighbor prediction program. In some embodiments, in addition to the default prediction program, a single variability threshold and a single alternative prediction program may be used. In some embodiments, any number "N" of variability thresholds and corresponding prediction programs may be used.

[0181] Level of Detail Attribute Compression

[0182] In some cases, the number of bits required to encode the attribute information for a point cloud may constitute a significant portion of the bitstream for that point cloud. For example, the attribute information may constitute a larger portion than the bitstream used to transmit the compressed spatial information for the point cloud.

[0183] In some embodiments, the spatial information may be used to construct a hierarchical level of detail (LOD) structure. The LOD structure may be used to compress the attributes associated with the point cloud. The LOD structure may also enable advanced features, such as progressive / view-dependent streaming and scalable rendering. For example, in some embodiments, the compressed attribute information may be sent (or decoded) only for a portion of the point cloud (e.g., a level of detail), rather than sending (or decoding) all the attribute information for the entire point cloud.

[0184] Figure 9 An example encoding process for generating a hierarchical LOD structure according to some embodiments is shown. For example, in some embodiments, an encoder (such as encoder 202) may use a similar process as shown in Figure 9 to generate compressed attribute information in an LOD structure.

[0185] In some embodiments, geometric information (also referred to herein as "spatial information") may be used to effectively predict the attribute information. For example, the compression of color information is shown in Figure 9 . However, the LOD structure may be applied to the compression of any type of attribute (e.g., reflectance, texture, modality, etc.) associated with the points of the point cloud. Note that a precoding step of applying a color space transformation or updating the data to make the data more suitable for compression may be performed according to the attribute to be compressed.

[0186] In some embodiments, the compression of the attribute information according to the LOD process is performed as described below.

[0187] For example, assume that Geometry (G) = {points - P(0), P(1), … P(N - 1)} is the reconstructed point cloud position generated by a spatial decoder (Geometry Decoder GD 902) included in an encoder after decoding a compressed geometry bitstream generated by a geometry encoder (such as Geometry Encoder GE914) also included in the encoder (such as Spatial Encoder 204 (shown in Figure 2A ). For example, in some embodiments, an encoder (such as Encoder 202 (shown in Figure 2A )) may include both a geometry encoder (such as Geometry Encoder 914) and a geometry decoder (such as Geometry Decoder 902). In some embodiments, the geometry encoder may be part of the spatial encoder 214, and the geometry decoder may be part of the prediction / correction evaluator 206, both as shown in Figure 2A .

[0188] In some embodiments, the decompressed spatial information may describe the positions of points in 3D space, such as the X, Y, and Z coordinates of the points that make up the mug 900. Note that the spatial information may be available to both the encoder (such as Encoder 202) and the decoder (such as Decoder 220). For example, various techniques (such as K - D tree compression, octree compression, nearest - neighbor point prediction, etc.) may be used to compress and / or encode the spatial information for the mug 900, and the spatial information may be sent to the decoder together with or in addition to the compressed attribute information for the attributes of the points that make up the point cloud (such as the point cloud of the mug 900).

[0189] In some embodiments, a deterministic re - ordering process may be applied to both the encoder side (such as at Encoder 202) and the decoder side (such as at Decoder 220) to organize the points of the point cloud (such as the points representing the mug 900) into a set of levels of detail (LOD). For example, the levels of detail may be generated by a level - of - detail generator 904, which may be included in the prediction / correction evaluator of the encoder (such as the prediction / correction evaluator 206 of Encoder 202 as shown in Figure 2A . In some embodiments, the level - of - detail generator 904 may be a separate component of the encoder (such as Encoder 202). For example, the level - of - detail generator 904 may be a separate component of Encoder 202. Note that in some embodiments, no additional information needs to be included in the bitstream to generate such an LOD structure other than the parameters of the LOD generation algorithm. For example, the parameters that may be included in the bitstream as parameters of the LOD generator algorithm may include:

[0190] i. The maximum number of LODs to be generated, denoted as “N” (e.g., N = 6),

[0191] ii. Initial sampling distance “D0” (e.g., D0 = 64), and

[0192] iii. Sampling distance update factor “f” (e.g., 1 / 2).

[0193] In some embodiments, the parameters N, D0, and f may be provided by a user (such as an engineer configuring the compression process). In some embodiments, the parameters N, D0, and f may be automatically determined by an encoder and / or a decoder using, for example, an optimization program. These parameters may be fixed or adaptive.

[0194] In some embodiments, LOD generation may be performed as follows:

[0195] a. Points of the geometry G (e.g., points of a point cloud organized according to spatial information) (such as points of the mug 900) are marked as unvisited, and a set of visited points V is set to be empty.

[0196] b. Then, the LOD generation process may be iterated. At each iteration j, the level of detail for that refinement level, e.g., LOD(j), may be generated as follows:

[0197] 1. The sampling distance of the current LOD is denoted as D(j) and may be set as follows:

[0198] a. If j = 0, then D(j) = D0.

[0199] b. If j > 0 and j < N, then D(j) = D(j - 1) * f.

[0200] c. If j = N, then D(j) = 0.

[0201] 2. The LOD generation process iterates over all points of G.

[0202] a. In the point evaluation iteration i, the point P(i) is evaluated,

[0203] i. If the point P(i) has been visited, it is ignored and the algorithm jumps to the next iteration (i + 1), e.g., evaluating the next point P(i + 1).

[0204] ii. Otherwise, calculate the distance D(i, V), which is defined as the minimum distance from P(i) over all points in V. Note that V is a list of visited points. If V is empty, the distance D(i, V) is set to 0, which means the distance from point P(i) to the visited points is zero since there are no visited points in the set V. If the shortest distance D(i, V) from point P(i) to any visited point is strictly greater than the parameter D0, then the point is ignored, LoD generates a jump to iteration (i + 1) and evaluates the next point P(i + 1). Otherwise, mark P(i) as a visited point and add the point P(i) to the set V of visited points.

[0205] b. This process can be repeated until all points of the geometry G are traversed.

[0206] 3. The set of points added to V during iteration j describes the refinement level R(j).

[0207] 4. LOD(j) can be obtained by taking the union of all refinement levels R(0), R(1), …, R(j).

[0208] In some embodiments, the above process can be repeated until all LODs are generated or all vertices are visited.

[0209] In some embodiments, the encoder as described above may further include a quantization module (not shown) that quantizes the geometric information included in the "position (x, y, z)" being provided to the geometry encoder 914. Additionally, in some embodiments, the encoder as described above may additionally include a module that removes duplicate points after quantization and before the geometry encoder 914.

[0210] In some embodiments, quantization may also be applied to compress attribute information, such as attribute correction values and / or one or more attribute value starting points. For example, quantization is performed at 910 on the attribute correction values determined by the interpolation-based prediction module 908. Quantization techniques may include uniform quantization, uniform quantization with dead zones, non-uniform / nonlinear quantization, grid quantization, or other suitable quantization techniques.

[0211] Figure 10 An example process for determining points to be included at different refinement layers of a level of detail (LOD) structure is shown according to some embodiments.

[0212] At 1002, an encoder (or decoder) receives or determines a level-of-detail parameter for determining a level-of-detail classification for a point cloud. Simultaneously with, before, or after receiving the level-of-detail parameter, at 1004, the encoder (or decoder) may receive compressed spatial information for the point cloud, and at 1006, the encoder (or decoder) may determine decompressed spatial information for the points of the point cloud. In embodiments that utilize lossy compression techniques for compressing spatial information, at 1006, the compressed spatial information may be compressed at the encoder and may also be decompressed at the encoder to generate a representative sample of the geometric information that will be encountered at the decoder. In some embodiments that utilize lossless compression of spatial information, 1004 and 1106 may be omitted on the encoder side.

[0213] At 1008, a level-of-detail structure generator (which may be on the encoder side or the decoder side) marks all points of the point cloud as "unvisited points".

[0214] At 1010, the level-of-detail structure generator also sets the directory of visited points "V" to be empty.

[0215] At 1012, a sampling distance D(j) is determined for the current refinement level R(j) being evaluated. If the refinement level is the coarsest refinement level, where j = 0, then the sampling distance D(j) is set to be equal to D0, e.g., an initial sampling distance, which is received or determined at 1002. If j is greater than 0 but less than N, then D(j) is equal to D(j - 1) * f. Note that "N" is the total number of levels of detail to be determined. Also note that "f" is a sampling update distance factor, which is set to be less than one (e.g., 1 / 2). Additionally, note that D(j - 1) is the sampling distance used in the previous refinement level. For example, when f is 1 / 2, the sampling distance D(j - 1) is cut in half for subsequent refinement levels, such that D(j) is half the length of D(j - 1). Additionally, note that the level of detail (LOD(j)) is the union of the current refinement level and all lower refinement levels. Thus, the first level of detail (LOD(0)) may include all points included in refinement level R(0). Subsequent levels of detail (LOD(1)) may include all points included in the previous level of detail and additionally all points included in the subsequent refinement level R(1). In this way, points can be added to the previous level of detail for each subsequent level of detail until the level of detail "N" that includes all points of the point cloud is reached.

[0216] To determine the points of the point cloud to be included in the current level of detail being determined, at 1014, a point P(i) to be evaluated is selected, where "i" is one of the current points of the point cloud being evaluated. For example, if the point cloud includes one million points, then the range of "i" can be from 0 to 1,000,000.

[0217] At 1016, it is determined whether the point P(i) currently being evaluated has been marked as an access point. If P(i) is marked as visited, then at 1018, P(i) is ignored, and then the process continues to proceed to evaluate the next point P(i+1), which then becomes the point P(i) currently being evaluated. Then, the process returns to 1014.

[0218] If it is determined at 1016 that the point P(i) has not been marked as an access point, then at 1020, the distance D(i) is calculated for the point P(i), where D(i) is the shortest distance between the point P(i) and any visited points included in the directory V. If there are no points in the directory V, for example, if the directory V is empty, then D(i) is set to zero.

[0219] At 1022, it is determined whether the distance D(i) for the point P(i) is greater than the initial sampling distance D0. If so, then at 1018, the point P(i) is ignored, and the process continues to the next point P(i+1) and returns to 1014.

[0220] If the point P(i) has not been marked as visited and the distance D(i) (i.e., the minimum distance between the point P(i) and a set of points included in V) is less than the initial sampling distance D0, then at 1024, P(i) is marked as visited and added to the set of access points V at the current refinement level R(j).

[0221] At 1026, it is determined whether there is an additional refinement level for the point cloud. For example, if j < N, where N is the LOD parameter that can be transmitted between the encoder and the decoder, then there is a determined additional refinement level. If there are no other determined refinement levels, then the process stops at 1028. If there is a determined additional refinement level, then the process continues to the next refinement level at 1030, and then continues to evaluate the point cloud for the next refinement level at 1012.

[0222] Once the refinement level is determined, the refinement level can be used to generate the LOD structure, where each subsequent LOD level includes all the points of the previous LOD level plus any points determined to be included in the additional refinement level. Since the process of determining the LOD structure is known to the encoder and the decoder, a decoder given the same LOD parameter used at the encoder can reconstruct the same LOD structure at the decoder as the LOD structure generated at the encoder.

[0223] Example Level of Detail Hierarchy

[0224] Figure 11AShows an example LOD according to some embodiments. Note that the LOD generation process may generate a uniform sampling approximation (or level of detail) of the original point cloud, which becomes increasingly refined as more and more points are included. Such features make it particularly suitable for progressive / view-dependent transmission and scalable rendering. For example, 1104 may include more details than 1102, and 1106 may include more details than 1104. Additionally, 1108 may include more details than 1102, 1104, and 1106.

[0225] A hierarchical LOD structure can be used to construct an attribute prediction strategy. For example, in some embodiments, points can be encoded in the same order as they are accessed during the LOD generation phase. The attributes of each point can be predicted by using the K nearest neighbors that have been encoded previously. In some embodiments, "K" is a parameter that can be user-defined or can be determined by using an optimization strategy. "K" can be static or adaptive. In the latter case where "K" is adaptive, additional information describing the parameter can be included in the bitstream.

[0226] In some embodiments, different prediction strategies can be used. For example, one of the following interpolation strategies, a combination of the following interpolation strategies, or the encoder / decoder can adaptively switch between different interpolation strategies can be used. Different interpolation strategies can include interpolation strategies such as: inverse distance interpolation, barycentric interpolation, natural neighbor interpolation, moving least squares interpolation, or other suitable interpolation techniques. For example, interpolation-based prediction can be performed at the interpolation-based prediction module 908 included in the prediction / correction value evaluator of the encoder (such as the prediction / correction value evaluator 206 of encoder 202). Additionally, interpolation-based prediction can be performed at the interpolation-based prediction module 908 included in the prediction evaluator of the decoder (such as the prediction evaluator 224 of decoder 220). In some embodiments, the color space can also be transformed at the color space transformation module 906 before performing the interpolation-based prediction. In some embodiments, the color space transformation module 906 can be included in the encoder (such as encoder 202). In some embodiments, the decoder can also include a module that transforms the transformed color space back to the original color space.

[0227] In some embodiments, quantization can also be applied to the attribute information. For example, quantization can be performed at the quantization module 910. In some embodiments, the encoder (such as encoder 202) can also include the quantization module 910. The quantization techniques employed by the quantization module 910 can include uniform quantization, uniform quantization with dead zones, non-uniform / nonlinear quantization, grid quantization, or other suitable quantization techniques.

[0228] In some embodiments, LOD attribute compression can be used to compress dynamic point clouds as follows:

[0229] a. Assume that FC is the current point cloud frame, and let RF be the reference point cloud.

[0230] b. Assume that M is the motion field that deforms RF into the shape of FC.

[0231] i. M can be calculated on the decoder side, and in this case, the information can be not encoded in the bitstream.

[0232] ii. M can be calculated by the encoder and explicitly encoded in the bitstream.

[0233] 1. M can be encoded by applying the hierarchical compression techniques described herein to the motion vectors associated with each point of RF (e.g., the motion of RF can be considered as an additional attribute).

[0234] 2. M can be encoded as a skeleton / skin-based model with associated local and global transformations.

[0235] 3. M can be encoded as a motion field defined based on an octree structure, which is adaptively refined to accommodate the complexity of the motion field.

[0236] 4. M can be described by using any suitable animation technique, such as keyframe-based animation, deformation techniques, free-form deformation, key-point-based deformation, etc.

[0237] iii. Assume that RF’ is the point cloud obtained after applying the motion field M to RF. Then, not only the "K" nearest neighbor points of FC are considered, but also the "K" nearest neighbor points of RF’ can be used for the attribute prediction strategy.

[0238] In addition, the attribute correction value can be determined based on comparing the interpolation-based prediction value determined at the interpolation-based prediction module 908 with the original uncompressed attribute value. The attribute correction value can be further quantized at the quantization module 910, and the quantized attribute correction value, the encoded spatial information (output from the geometry encoder 902), and any configuration parameters used in the prediction can be encoded at the arithmetic coding module 912. In some embodiments, the arithmetic coding module can use context-adaptive arithmetic coding techniques. Then the compressed point cloud can be provided to a decoder (such as decoder 220), and the decoder can determine a similar level of detail and perform interpolation-based prediction to reconstruct the original point cloud based on the quantized attribute correction value, the encoded spatial information (output from the geometry encoder 902), and the configuration parameters used in the encoder prediction.

[0239] Figure 11BShows an example compressed point cloud file including LOD according to some embodiments. The level of detail attribute information file 1150 includes configuration information 1152, point cloud data 1154, and level of detail point attribute correction values 1156. In some embodiments, the level of detail attribute information file 1150 can be transmitted in batches via multiple data packets. In some embodiments, not all of the parts shown in the level of detail attribute information file 1150 may be included in each data packet for transmitting compressed attribute information. In some embodiments, the level of detail of an attribute information file (such as the level of detail of the attribute information file 1150) can be stored in a storage device (such as a server implementing an encoder or decoder) or other computing devices.

[0240] Figure 12A Shows a method for encoding attribute information of a point cloud using an update operation according to some embodiments.

[0241] At 1202, the point cloud is received by an encoder. The point cloud can be captured by one or more sensors, for example, or can be generated in software, for example.

[0242] At 1204, the spatial or geometric information of the point cloud is encoded as described herein. For example, the spatial information can be encoded using a K-D tree, an octree, a neighborhood prediction strategy, or other suitable techniques for encoding spatial information.

[0243] At 1206, one or more levels of detail are generated as described herein. For example, a similar process as shown in Figure 10 can be used to generate the levels of detail. Note that in some embodiments, the spatial information encoded or compressed at 1204 can be decoded or decompressed to generate a representative decompressed point cloud geometry that the decoder will encounter. Then, this representative decompressed point cloud geometry can be used to generate the LOD structure, as further described in Figure 10

[0244] At 1208, interpolation-based prediction is performed to predict the attribute values of the attributes of the points of the point cloud. At 1210, an attribute correction value is determined based on comparing the predicted attribute value with the original attribute value. For example, in some embodiments, interpolation-based prediction can be performed for each level of detail to determine the predicted attribute values of the points included in the corresponding level of detail. Then, these predicted attribute values can be compared with the attribute values of the original point cloud before compression to determine the attribute correction values for the points of the corresponding level of detail. For example, as shown in Figure 1B Figure 4 to Figure 5 ​The interpolation-based prediction process described in connection with FIG. 8 can be used to determine predicted attribute values for corresponding levels of detail. In some embodiments, attribute correction values can be determined for multiple levels of detail of the LOD structure. For example, a first set of attribute correction values can be determined for points included in a first level of detail, and other sets of attribute correction values can be determined for points included in other levels of detail.

[0245] At 1212, an update operation that affects the attribute correction values predicted at 1210 can optionally be applied. The execution of the update operation is discussed in more detail below in Figure 13A where.

[0246] At 1214, as described herein, the attribute correction values, LOD parameters, encoded spatial information (output from the geometry encoder), and any configuration parameters used in the prediction are encoded.

[0247] In some embodiments, the attribute information encoded at 1214 can include attribute information for multiple or all levels of detail of the point cloud, or can include attribute information for a single level of detail or less than all levels of detail of the point cloud. In some embodiments, the level-of-detail attribute information can be encoded sequentially by the encoder. For example, the encoder can make the first level of detail available before encoding the attribute information for one or more additional levels of detail.

[0248] In some embodiments, the encoder can also encode one or more configuration parameters to be sent to the decoder, such as any of the configuration parameters shown in configuration information 1152 of the compressed attribute information file 1150. For example, in some embodiments, the encoder can encode the number of levels of detail to be encoded for the point cloud. The encoder can also encode a sampling distance update factor, where the sampling distance is used to determine which points will be included in a given level of detail.

[0249] Figure 12B A method for decoding attribute information of a point cloud according to some embodiments is shown.

[0250] At 1252, compressed attribute information for the point cloud is received at the decoder. Additionally, at 1254, spatial information for the point cloud is received at the decoder. In some embodiments, various techniques such as K-D trees, octrees, neighborhood prediction, etc. can be used to compress or encode the spatial information, and at 1254, the decoder can decompress and / or decode the received spatial information.

[0251] At 1256, the decoder determines which level of detail among the number of levels of detail to be decompressed / decoded. The selected level of detail to be decompressed / decoded can be determined based on the viewing mode of the point cloud. For example, a point cloud viewed in a preview mode may require determination of a lower level of detail compared to a point cloud viewed in a full view mode. Additionally, the position of the point cloud in the view being rendered can be used to determine the level of detail to be decompressed / decoded. For example, the point cloud can represent an object such as Figure 9 the coffee cup shown. If the coffee cup is in the foreground of the view being rendered, a higher level of detail can be determined for the coffee cup. However, if the coffee cup is in the background of the view being rendered, a lower level of detail can be determined for the coffee cup. In some embodiments, the level of detail for the point cloud can be determined based on the data budget allocated for the point cloud.

[0252] At 1258, the points included in the first level of detail (or the next level of detail) being determined can be determined as described herein. For the points of the level of detail being evaluated, the attribute values of the points can be predicted based on inverse distance weighted interpolation with the k nearest neighbors for each point being evaluated, where k can be a fixed or adjustable parameter.

[0253] At 1260, in some embodiments, as Figure 12F more detailedly described in, an update operation can be performed on the predicted attribute values.

[0254] At 1262, the attribute correction values included in the compressed attribute information for the point cloud can be decoded for the current level of detail being evaluated, and can be applied at 1258 to correct the predicted attribute values or the updated predicted attribute values determined at 1260.

[0255] At 1264, the corrected attribute values determined at 1262 can be assigned as attributes to the points of the first level of detail (or the current level of detail being evaluated). In some embodiments, the attribute values determined for subsequent levels of detail can be assigned to the points included in the subsequent levels of detail, while the attribute values that have been determined for previous levels of detail are retained by the corresponding points of the previous one or more levels of detail. In some embodiments, new attribute values can be determined for sequential levels of detail.

[0256] In some embodiments, the spatial information received at 1254 can include spatial information for multiple or all levels of detail of the point cloud, or can include spatial information for a single level of detail or less than all levels of detail of the point cloud. In some embodiments, the level of detail attribute information can be received by the decoder sequentially. For example, the decoder can receive the first level of detail and generate the attribute values of the points of the first level of detail before receiving the attribute information for one or more additional levels of detail.

[0257] At 1266, it is determined whether there are additional levels of detail to be decoded. If so, the process returns to 1258 and the decoding of the next level of detail is repeated. If not, the process stops at 1267, but can be resumed at 1256 in response to input that affects the number of levels of detail to be determined (such as changing the view of the point cloud or applying a zoom operation to the point cloud being viewed), as some examples of input that affects the level of detail to be determined.

[0258] In some embodiments, the spatial information may be encoded and decoded via a geometry encoder and an arithmetic encoder, such as the geometry encoder 202 and the arithmetic encoder 212 described above with respect to FIG. 2. In some embodiments, the geometry encoder, such as the geometry encoder 202, may utilize an octree compression technique, and the arithmetic encoder 212 may be a binary arithmetic encoder, as described in more detail below.

[0259] Compared to a multi-symbol codec with an alphabet of 256 symbols (e.g., 8 sub-cubes per cube, and each sub-cube is either occupied or unoccupied 2^8=256), the use of a binary arithmetic encoder as described below reduces the computational complexity of encoding the occupancy symbols of the octree. Additionally, the use of context selection based on the most likely neighbor configuration can reduce the search for neighbor configurations compared to searching all possible neighbor configurations. For example, instead of searching all possible neighborhood configurations, the encoder can keep track of the number of sub-cubes corresponding to the occupancy symbol. Figure 12C There are 10 encoding contexts of the 10 neighborhood configurations 1268, 1270, 1272, 12712, 1276, 1278, 1280, 1282, 1284 and 1286 shown in .

[0260] In some embodiments, an arithmetic encoder (such as arithmetic encoder 212) can use a binary arithmetic codec to encode the occupied symbols of 256 values. Compared with a multi-symbol arithmetic codec, this may be less complicated and more hardware-friendly in terms of implementation. In addition, the arithmetic encoder 212 and / or the geometric encoder 202 can utilize a look-ahead program to calculate 6 neighbors for arithmetic context selection, which may not be as complex as a linear search and may involve a constant number of operations (compared to a linear search that may involve different number operations). In addition, the arithmetic encoder 212 and / or the geometric encoder 202 can utilize a context selection program, which reduces the number of coding contexts. In some embodiments, the binary arithmetic codec, the look-ahead program, and the context selection program can be implemented together or independently.

[0261] Binary Arithmetic Coding

[0262] In some embodiments, to encode spatial information, the occupancy information of each cube is encoded as an 8-bit value that can have a value between 0 - 255. To perform efficient encoding / decoding on such non-binary values, a multi-symbol arithmetic encoder / decoder is typically used, which is computationally complex and less hardware-friendly to implement compared to a binary arithmetic encoder / decoder. However, on the other hand, directly using a conventional binary arithmetic encoder / decoder on such values (e.g., encoding each bit independently) may not be as efficient. However, to efficiently encode non-binary occupancy values with a binary arithmetic encoder, an adaptive lookup table (A-LUT) (tracking N (e.g., 32) of the most frequent occupancy symbols) can be used in conjunction with a cache that tracks the M (e.g., 16) most recently observed occupancy symbols.

[0263] The value of the number M of the last observed distinct occupancy symbols to be tracked and the number N of the most frequent occupancy symbols to be tracked can be user-defined, such as an engineer customizing the encoding technique for a specific application, or can be selected based on an offline statistical analysis of the encoding session. The selection of the M and N values can be based on a trade-off between:

[0264] ● Encoding efficiency,

[0265] ● Computational complexity, and

[0266] ● Memory requirements.

[0267] In some embodiments, the algorithm proceeds as follows:

[0268] ● The adaptive lookup table (A-LUT) is initialized with N symbols provided by the user (e.g., engineer) or computed offline based on statistical information of a similar class of point clouds.

[0269] ● The cache is initialized with M symbols provided by the user (e.g., engineer) or computed offline based on statistical information of a similar class of point clouds.

[0270] ● Each time an occupancy symbol S is encoded, the following steps are applied

[0271] 1. Encode the binary information indicating whether S is in the A-LUT.

[0272] 2. If S is in the A-LUT, encode the index of S in the A-LUT using a binary arithmetic encoder

[0273] ● Suppose (b1, b2, b3, b4, b5) are the five bits of the binary representation of the index of S in the A-LUT. Suppose b1 is the least significant bit, while b5 is the most significant bit.

[0274] ● The index of S can be encoded using three methods as described below, for example by using 31, 9, or 5 adaptive binary arithmetic contexts, as follows

[0275] ○ 31 contexts

[0276] First, the index b5 of S is encoded using the first context (referred to as context 0). When encoding the most significant bit (the first bit to be encoded), there is no information available from the encoding of any other bits, which is why this context is called context zero. Then, when encoding b4 (the second bit to be encoded), two additional contexts can be used, which are called context 1 (if b5 = 0) and context 2 (if b5 = 1). When this method is applied all the way to b1, there are 31 resulting contexts, as shown below, contexts 0 - 30. This method exhaustively uses each bit being encoded to select the adaptive context for encoding the next bit. For example, see Figure 12E .

[0277] ○ 9 contexts

[0278] Remember that the index values of the adaptive lookup table ALUT are assigned based on the frequency of occurrence of symbol S. Therefore, the index value of the most frequently used symbol S in the ALUT will be 0, which means all bits of the index value of the most frequently used symbol S are zero. For example, the smaller the binary value, the higher the frequency of occurrence of the symbol. To encode nine contexts (the most significant bits b4 and b5), if they are 1s, the index value must be relatively large. For example, if b5 = 1, the index value is at least 16 or higher; or if b4 = 1, the index value is at least 8 or higher. Therefore, when encoding 9 contexts, the focus is on the first 7 index entries, such as 1 to 7. For these 7 index entries, adaptive encoding contexts are used. However, for index entries with values greater than 7, the same context is used, such as a static binary encoder. Therefore, if b5 = 1 or b4 = 1, the same context is used to encode the index value. If not, one of the adaptive contexts 1 - 7 is used. Since there is a context 0 for b5, there are 7 adaptive contexts, and an entry strictly greater than 8 has a common context, so there are a total of nine contexts. This simplifies the encoding and reduces the number of contexts to be transmitted compared to using all 31 contexts as described above.

[0279] ○ 5 contexts

[0280] To encode the index value using 5 contexts, determine if b5 = 1. If b5 = 1, encode all bits of the index value from b4 to b1 using a static binary context. If b5 is not equal to 1, encode b4 of the index value and check if b4 is equal to 1 or 0. If b4 = 1, which means the index value is greater than 8, then use the static binary context again to encode bits b3 to b1. Then repeat this reasoning so that if b3 = 1, the static binary context is used to encode bits b2 to b1, and if b2 = 1, the static binary context is used to encode bit 1. However, if bits b5, b4, and b3 are equal to zero, select an adaptive binary context to encode bits 2 and 1 of the index value.

[0281] 3. If S is not in the A-LUT, then

[0282] ● Encode the binary information indicating whether S is in the cache.

[0283] ● If S is in the cache, encode the binary representation of its index using a binary arithmetic encoder

[0284] ○ In some embodiments, encode the binary representation of the index by encoding each bit one by one using a single static binary context. Then, shift the values of these bits up by one, where the least significant bit becomes the next higher significant bit.

[0285] ● Otherwise, if S is not in the cache, encode the binary representation of S using a binary arithmetic encoder

[0286] ○ In some embodiments, encode the binary representation of S using a single adaptive binary context. The value of the known index is between 0 and 255, which means it is encoded in 8 bits. Shift these bits so that the least significant bit becomes the next higher significant bit, and use the same adaptive context to encode all the remaining bits.

[0287] ● Add the symbol S to the cache and evict the earliest symbol in the cache.

[0288] 4. Increment the occurrence count of symbol S in the A-LUT by one.

[0289] 5. Periodically recalculate the list of the N most frequent symbols in the A-LUT

[0290] ● Method 1: If the number of symbols encoded so far reaches a user-defined threshold (e.g., 64 or 128), then recalculate the list of the N most frequent symbols in the A-LUT.

[0291] ● Method 2: Adapt the update cycle to the number of encoded symbols. The idea is to update the probabilities quickly at the beginning and increase the update cycle exponentially as the number of symbols increases:

[0292] ○ Initialize the update cycle _updateCycle to a low number N0 (e.g., 16).

[0293] ○ Whenever the number of symbols reaches the update cycle

[0294] ■ Recalculate the list of the N most frequent symbols in the A-LUT

[0295] ■ Update the update cycle as follows:

[0296] _updateCycle = min(_alpha * _updateCycle, _maxUpdateCycle)

[0297] ■ _alpha (e.g., 5 / 4) and maxUpdateCycle (e.g., 1024) are two user-defined parameters that control the rate of exponential growth and the maximum update cycle value.

[0298] 6. At the start of each level of octree subdivision, the occurrences of all symbols are reset to zero. Set the occurrences of the N most frequent symbols to 1.

[0299] 7. When the number of occurrences of a symbol reaches a user-defined maximum number (e.g., _maxOccurence = 1024), the number of occurrences of all symbols is divided by 2 to keep the number of occurrences within the user-defined range.

[0300] In some embodiments, a circular buffer is used to track elements in the cache. The element to be evicted from the cache corresponds to the position index 0 = (_last++) % CacheSize, where _last is a counter initialized to 0 and incremented each time a symbol is added to the cache. In some embodiments, the cache can also be implemented with a sorted list, which will ensure that the earliest symbol is evicted each time.

[0301] 2. Look-Ahead to Determine Neighbors

[0302] In some embodiments, at each subdivision level of the octree, a cube of the same size is subdivided and the occupancy code for each is encoded.

[0303] ● For subdivision level 0, there may be only a single cube (2 C , 2 C , 2 C ), and there are no neighbors.

[0304] ● For level of detail 1, there can be at most 8 - dimensional cubes in each medium (2 C-1 ,2 C-1 ,2 C-1 ).

[0305] ●…

[0306] ● For level of detail L, there can be at most 8 L - dimensional cubes in each medium (2 C-L ,2 C-L ,2 C-L ).

[0307] In some embodiments, at each level L, a set of non - overlapping look - ahead cubes for each dimension (2H - C+L,2H - C+L,2H - C+L) can be defined, as Figure 12D shown. Note that the look - ahead cubes can accommodate cubes of size 23xH (2C - L,2C - L,2C - L).

[0308] ● At each level L, the cubes contained in each look - ahead cube are encoded without reference to the cubes in other look - ahead cubes.

[0309] ○ During the look - ahead phase, the cubes of the dimensions in the current look - ahead cube (2 C-L ,2 C-L ,2 C-L ) are extracted from the FIFO, and a lookup table depicting whether each (2 C-L ,2 C-L ,2 C-L ) region of the current look - ahead cube is occupied or empty is populated.

[0310] ○ Once the lookup table is populated, the encoding phase of the extracted cubes begins. Here, 6 adjacent occupancy information is obtained by directly retrieving information from the lookup table.

[0311] ○ For the cubes on the boundary of the look - ahead cube, the neighbors outside are assumed to be empty.

[0312] ■ Another alternative can include filling the values of the external neighbors based on extrapolation.

[0313] ○ An effective implementation can be achieved by

[0314] ■ Storing the occupancy information of each group of 8 adjacent (2 C-L ,2 C-L ,2 C-L ) regions on one byte

[0315] ■ Store the occupied bytes in Z-order to maximize memory cache hits

[0316] 3. Context Selection

[0317] In some embodiments, to reduce the number of coding contexts (NC) to a smaller number of contexts (e.g., from 10 to 6), a separate context is assigned to each of the (NC - 1) most likely neighborhood configurations, and the contexts corresponding to the least likely neighborhood configurations share the same one or more contexts. The way to do this is as follows:

[0318] ○ Before starting the coding process, initialize the occurrence counts of 10 neighborhood configurations (e.g., Figure 12C the 10 configurations shown):

[0319] ● Set all 10 occurrence counts to 0

[0320] ● Set the occurrence counts based on offline / online statistics or based on user-provided information.

[0321] ○ At the start of each refinement level of the octree:

[0322] ● Determine the (NC - 1) most likely neighborhood configurations based on the statistics collected during the coding of the previous refinement level.

[0323] ● Compute a lookup table NLUT that maps the indices of the (NC - 1) most likely neighborhood configurations to the numbers 0, 1, …, (NC - 2), and maps the indices of the remaining configurations to NC - 1.

[0324] ● Initialize the occurrence counts of the 10 neighborhood configurations to 0.

[0325] ○ During coding:

[0326] ● Each time a neighborhood configuration is encountered, increment such a configuration by one.

[0327] ● Use the lookup table NLUT[] to determine the context for encoding the current occupancy value based on the neighborhood configuration index.

[0328] Figure 12F An example octree compression technique using a binary arithmetic encoder, cache, and look-ahead table is shown according to some embodiments. For example, Figure 12FAn example of the process described above is shown. At 1288, occupancy symbols for the levels of the octree for the point cloud are determined. At 1290, an adaptive look-ahead table with "N" symbols is initialized. At 1292, a cache with "M" symbols is initialized. At 1294, the symbols for the current octree level are encoded using the techniques described above. At 1296, it is determined whether to encode other octree levels. If so, the process continues at 1288 for the next octree level. If not, the process ends at 1298 and the encoded spatial information of the point cloud is made available, such as being sent to a recipient or being stored.

[0329] Lifting Scheme for Level of Detail Compression and Decompression

[0330] In some embodiments, a lifting scheme can be applied to the point cloud. For example, as described below, the lifting scheme can be applied to irregular points. This is contrary to other types of lifting schemes that can be applied to images with regular points in a plane. In the lifting scheme, for points in the current level of detail, the nearest points in the lower level of detail can be found. These nearest points in the lower level of detail are used to predict the attribute values of the points in the higher level of detail. Conceptually, a chart can be made showing how the points in the lower level of detail are used to determine the attribute values of the points in the higher level of detail. In such a conceptual view, edges can be assigned to the chart between the levels of detail, where there is an edge between each point in the higher level of detail and each point in the lower level of detail, and this edge forms the basis for predicting the attributes of the points in the higher level of detail. As described in more detail below, weights can be assigned to each of these edges, thereby indicating the relative influence. The weight can represent the influence of the attribute value of the point in the lower level of detail on the attribute value of the point in the higher level of detail. Additionally, multiple edges can form a path through the levels of detail, and weights can be assigned to these paths. In some embodiments, the influence of a path can be defined by the sum of the weights of the edges of the path. For example, Equation 1 discussed further below represents such weighting of a path.

[0331] In the lifting scheme, the attribute values of low-influence points can be highly quantized, while the attribute values of high-influence points can be less quantized. In some embodiments, a balance can be achieved between the quality and efficiency of reconstructing the point cloud, where more quantization increases the compression efficiency and less quantization increases the quality. In some embodiments, not all paths may be evaluated. For example, some paths with very little influence may not be evaluated. Additionally, an update operator can smooth the residual differences, such as the attribute correction values, in order to improve the compression efficiency while taking into account the relative influence or importance of the points when smoothing the residual differences.

[0332] Figure 13AShows a direct transformation that can be applied at an encoder to encode attribute information of a point cloud.

[0333] In some embodiments, an encoder may utilize a direct transformation as Figure 13A shown to determine an attribute correction value that is encoded as part of a compressed point cloud. For example, in some embodiments, a direct transformation (such as interpolation-based prediction) may be utilized to determine an attribute value, as described at 1208 in Figure 12A and an update operation may be applied, as described at 1212 in Figure 12A

[0334] In some embodiments, a direct transformation may receive an attribute signal for an attribute associated with a point of a point cloud to be compressed. For example, the attribute may include a color value (such as an RGB color) or other attribute values of a point to be compressed in the point cloud. The geometric structure of a point of the point cloud to be compressed may also be known by the direct transformation that receives the attribute signal. At 1302, the direct transformation may include a splitting operator that splits the attribute signal 1310 into a first (or next) level of detail. For example, for a particular level of detail (such as LOD(N)) that contains a number of X points, a subsample of the attributes of the points (e.g., a sample that contains a number of Y points) may contain attribute values for a number of points that is less than X. In other words, the splitting operator may take as input the attributes associated with a particular level of detail and generate a low-resolution sample 1304 and a high-resolution sample 1306. It should be noted that the LOD structure may be divided into refinement levels, where subsequent levels of refinement include more attributes of points than a base refinement level. A particular level of detail as described below may be obtained by taking the union of all lower levels of detail. For example, the detail level j is obtained by taking the union of all refinement levels R(0), R(1),..., R(j). It should also be noted that, as described above, a compressed point cloud may have a total of N levels of detail, where R(0) is the smallest refinement detail level and R(N) is the highest refinement detail level for the compressed point cloud.

[0335] ​At 1308, a prediction of the attribute value of a point not included in the low-resolution sample 1304 is made based on the points included in the low-resolution sample. For example, based on an inverse distance interpolation prediction technique or any other prediction technique described above. At 1312, the difference between the predicted attribute values of the points left over for the low-resolution sample 1304 is compared with the actual attribute values of the points left over for the low-resolution sample 1304. This comparison determines the difference between the predicted attribute value and the actual attribute value for the corresponding point. These differences (D(N)) are then encoded as attribute correction values for the attributes of the points included in a particular level of detail that were not encoded in the low-resolution sample. For example, for the highest level of detail N, the difference D(N) can be used to adjust / correct the attribute values included in the lower levels of detail. Since at the highest level of detail, the attribute correction values are not used to determine the attribute values of other even higher levels of detail (since for the highest level of detail N, there are no higher levels of detail), an update operation that takes into account the relative importance of these attribute correction values may not be performed. Thus, the difference D(N) can be used to encode the attribute correction values for LOD(N).

[0336] Additionally, direct conversion can be applied to subsequent lower levels of detail, such as LOD(N - 1). However, before applying the direct conversion to subsequent levels of detail, an update operation can be performed to determine the relative importance of the attribute values of the points at the lower level of detail for the attribute values of one or more higher levels of detail. For example, the update operation 1314 can determine the relative importance of the attribute values of the attributes of the points at the lower level of detail included in the higher level of detail, such as for the point attributes included in L(N). Taking into account the relative importance of the corresponding attribute values, the update operator can also smooth the attribute values to improve the compression efficiency of the attribute correction values for subsequent levels of detail, where the smoothing operation is performed such that the modification of the attribute values that have a greater impact on the subsequent levels of detail is less than the modification of the points that have a smaller impact on the subsequent levels of detail. Several methods for performing the update operation are described in more detail below. The lower-resolution sample of the updated level of detail L’(N) is then input into another segmentation operator, and the process is repeated for the subsequent level of detail LOD(N - 1). Note that the attribute signal for the lower level of detail LOD(N - 1) can also be received at the second (or subsequent) segmentation operator.

[0337] Figure 13B An inverse conversion that can be applied at a decoder to decode the attribute information of a point cloud is shown according to some embodiments.

[0338] In some embodiments, the decoder can utilize an inverse conversion process as Figure 13B shown to reconstruct the point cloud from the compressed point cloud. For example, in some embodiments, it can be according to as Figure 13BExecute according to the inverse conversion process described in Figure 12B the prediction described at 1258 in Figure 12B apply the update operator described at 1260 in Figure 12B the attribute correction value described at 1262 in Figure 12B and the details of assigning attributes to points as described at 1264 in

[0339] In some embodiments, the inverse conversion process may receive an updated low-level resolution sample L'(0) for the lowest level of detail of the LOD structure. The inverse conversion process may also receive an attribute correction value for points not included in the updated low-resolution sample L'(0). For example, for a particular LOD, L'(0) may include a subsampling of the points included in the LOD, and prediction techniques may be used to determine the other points of the LOD, such as those that will be included in the high-resolution sample of the LOD. As shown at 1306, an attribute correction value, such as D(0), may be received. At 1318, an update operation may be performed to account for the smoothing of the attribute correction value performed at the encoder. For example, the update operation 1318 may "undo" the update operation performed at 1314, where the update operation performed at 1314 was performed to improve compression efficiency by smoothing the attribute values by considering the relative importance of the attribute values. The update operation may be applied to the updated low-resolution sample L'(0) to generate an "unsmoothed" or unupdated low-resolution sample L(0). The low-resolution sample L(0) may be used by prediction techniques at 1320 to determine the attribute values of points not included in the low-resolution sample. The attribute correction value D(0) may be used to correct the predicted attribute values to determine the attribute values of the points of the high-resolution sample of LOD(0). The low-resolution sample and the high-resolution sample may be combined at the merge operator 1322, and a new updated low-resolution sample for the next level of detail L'(1) may be determined. A similar process may be repeated for the next level of detail LOD(1) as described for LOD(0). In some embodiments, as Figure 13A shown for the encoder and as Figure 13B shown for the decoder may repeat their respective processes for N levels of detail of the point cloud.

[0340] A more detailed example definition of LOD and a method for determining the update operation are described below.

[0341] In some embodiments, the definition of LOD is as follows:

[0342] ● LOD(0) = R(0)

[0343] ● LOD(1) = LOD(0) U R(1)

[0344] ●...

[0345] ●LOD(j) = LOD(j - 1) ∪ R(j)

[0346] ●…

[0347] ●LOD(N + 1) = LOD(N) ∪ R(N) = the entire point cloud

[0348] In some embodiments, assume that A is a set of attributes associated with the point cloud. More precisely, assume that A(P) is a scalar / vector attribute associated with the point P of the point cloud. An example of an attribute is color described by RGB values.

[0349] Assume that L(j) is the set of attributes associated with LOD(j), and H(j) is the set of attributes associated with R(j). Based on the definitions of the level of detail LOD(j), L(j), and H(j), verify the following properties:

[0350] ●L(N + 1) = A and H(N + 1) = {}

[0351] ●L(j) = L(j - 1) ∪ H(j)

[0352] ●L(j) and H(j) are disjoint.

[0353] In some embodiments, a splitting operator (such as splitting operator 1302) takes L(j + 1) as input and generates two outputs: (1) a low-resolution sample L(j) and (2) a high-resolution sample H(j).

[0354] In some embodiments, a merging operator (such as merging operator 1322) takes L(j) and H(j) as input and produces L(j + 1).

[0355] As described in more detail above, a prediction operator can be defined on top of the LOD structure. Assume that (P(i, j))_i are the set points of LOD(j) belonging to R(j) and (Q(i, j))_i, and assume that (A(P(i, j)))_i and (A(Q(i, j)))_i are the attribute values associated with LOD(j) and R(j), respectively.

[0356] In some embodiments, the prediction operator predicts the attribute value by using the representation of the k nearest neighbors of the attribute value in LOD(j - 1) as the attribute value of A(Q(i, j))

[0357]

[0358] where α(P, Q(i, j)) is the interpolation weight. For example, an inverse distance weight interpolation strategy can be used to calculate the interpolation weight.

[0359] The prediction residual, such as the attribute correction value D(Q(i,j)), is defined as follows:

[0360] D(Q(i,j)) = A(Q(i,j)) - Pred(Q(i,j))

[0361] Note that the prediction grading can be described by the orientation graph G defined as follows:

[0362] ● Each point Q in the point cloud corresponds to a vertex V(Q) of the graph G.

[0363] ● If there exist i and j, then two vertices V(P) and V(P) of the graph G are connected by an edge E(P,Q) such that

[0364] ○ Q = Q(i,j) and

[0365] ○

[0366] ● The edge E(Q,P) has a weight α(P,Q(i,j)).

[0367] In such prediction strategies as described above, points with a lower level of detail are more influential because they are more frequently used for prediction.

[0368] Assume that w(P) is the influence weight associated with the point P. w(P) can be defined in various ways.

[0369] ● Method 1

[0370] ○ If there exists a path x = (E(1), E(2), …, E(s)) of the edges connecting two vertices V(P) and V(Q) of G, then they are said to be connected. The weight w(x) of the path x is defined as follows:

[0371]

[0372] ○ Assume that X(P) is the set of paths with P as the destination. w(P) is defined as follows:

[0373] w(P) = 1 + ∑ x∈X(P) (w(x)) 2 [Equation 1]

[0374] ○ The previous definition can be interpreted as follows. Assume that the attribute A(P) is modified by an amount ∈, then all attributes associated with the points connected to P will be perturbed. The sum of the squared errors associated with such perturbations is denoted as SSE(P,∈) and is given by:

[0375] SSE(P,∈) = w(P)∈ 2

[0376] ● Method 2

[0377] ○ As described above, calculating the influence weights can be computationally complex because all paths need to be evaluated. However, since the weight α(E(s)) is usually normalized between 0 and 1, the weight w(x) of a path x decays rapidly with the number of its edges. Therefore, long paths can be ignored without significantly affecting the final influence weights to be calculated.

[0378] ○ Based on the previous property, the definition in [Equation 1] can be modified to consider only paths of limited length, or to discard paths whose known weights are below a user-defined threshold. This threshold can be fixed and known at both the encoder and the decoder, or it can be explicitly signaled or predefined at different stages of the encoding process, for example, once per frame, per LOD, or even after a certain number of signal points.

[0379] ● Method 3

[0380] ○ w(P) can be approximated by the following recursive procedure:

[0381] ● Assume w(P) = 1 for all points

[0382] ● Traverse the points in reverse order of the order defined by the LOD structure

[0383] ● For each point Q(i,j), update the weight of its neighbor as follows

[0384] w(P) ← w(P) + w(Q(i,j),j){α(P,Q(i,j))} γ

[0385] where γ is a parameter that is usually set to 1 or 2.

[0386] ● Method 4

[0387] ○ w(P) can be approximated by the following recursive procedure:

[0388] ● Assume w(P) = 1 for all points

[0389] ● Traverse the points in reverse order of the order defined by the LOD structure

[0390] ● For each point Q(i,j), update the weight of its neighbor as follows

[0391] w(P) ← w(P) + w(Q(i,j),j)f{α(P,Q(i,j))}

[0392] where f(x) is a function whose resulting value is in the range [0,1].

[0393] In some embodiments, an update operator (such as update operator 1314 or 1318) uses the prediction residual D(Q(i,j)) to update the attribute value of LOD(j). The update operator can be defined in different ways, such as:

[0394] ●Method 1

[0395] 1. Assume that Δ(P) is a set of points Q(i,j) such that

[0396] 2. The update operation of P is defined as follows:

[0397] renew

[0398] where γ is a parameter that is usually set to 1 or 2.

[0399] Method 2

[0400] 1. Assume that Δ(P) is a set of points Q(i,j) such that

[0401] 2. The update operation of P is defined as follows:

[0402] renew

[0403] Where g(x) is a function whose result value is in the range [0,1].

[0404] Method 3

[0405] ○ Update(P) is calculated iteratively as follows:

[0406] 1. Initially set Update(P) = 0

[0407] 2. The inverse traversal points in the order defined by the LOD structure

[0408] 3. For each point Q(i,j), calculate its neighbor The associated local updates (u(1),u(2),..,u(k)) are the solutions to the following minimization problem:

[0409]

[0410] 4. Update (P (r)):

[0411] update(P(r))←update(P(r))+u(r)

[0412] ●Method 4

[0413] ○ Calculate Update(P) iteratively as follows:

[0414] 1. Initially set Update(P) = 0

[0415] 2. Traverse the points in the reverse order of the order defined by the LOD structure

[0416] 3. For each point Q(i,j), calculate the local update (u(1), u(2),.., u(k)) associated with it as the solution to the following minimization problem: (u(1), u(2),.., u(k)) = argmin{h(u(1),.., u(k), D(Q(i,j)))}

[0417] where h can be any function.

[0418] 4. Update Update(P(r)):

[0419] Update(P(r)) ← Update(P(r)) + u(r)

[0420] In some embodiments, when using the lifting scheme as described above, a quantization step can be applied to the computed wavelet coefficients. Such a process may introduce noise, and the quality of the reconstructed point cloud may depend on the selected quantization step. Additionally, as described above, perturbing the attributes of points at a lower LOD may have a greater impact on the quality of the reconstructed point cloud than perturbing the attributes of points at a higher LOD.

[0421] In some embodiments, the computed impact weights as described above can be further utilized during the conversion process to guide the quantization process. For example, the coefficient associated with point P can be multiplied by a factor of {w(P)}

[0422] where β is a parameter that is typically set to β = 0.5. After inverse quantization on the decoder side, an inverse scaling process with the same factor is applied. R In some embodiments, the value of the β parameter can be fixed for the entire point cloud and known at both the encoder and decoder, or can be explicitly signaled or predefined at different stages of the encoding process, for example, once per point cloud frame, LOD, or even after a certain number of signal points.

[0423] In some embodiments, a hardware-friendly implementation of the above-described lifting scheme can utilize a fixed-point representation of weights and look-up tables for non-linear operations.

[0424] In some embodiments, a hardware-friendly implementation of the above-described lifting scheme can utilize a fixed-point representation of weights and look-up tables for non-linear operations.

[0425] In some embodiments, the lifting scheme described herein can be used for other applications in addition to compression, such as denoising / filtering, watermarking, segmentation / detection, and various other applications.

[0426] In some embodiments, the decoder can employ the complementary process described above to decode the compressed point cloud compressed using the octree compression technique and binary arithmetic encoder described above.

[0427] In some embodiments, the lifting scheme described above can further implement a bottom-up approach to build the level of detail (LOD). For example, instead of determining the predicted values for the points and then assigning these points to different levels of detail, the predicted values can be determined while determining which points are to be included in which level of detail. Additionally, in some embodiments, the residual values can be determined by comparing the predicted values with the actual values of the original point cloud. This can also be performed while determining which points are to be included in which levels of detail. Additionally, in some embodiments, approximate nearest neighbor search can be used instead of exact nearest neighbor search to accelerate level of detail creation and prediction calculations. In some embodiments, a binary / arithmetic encoder / decoder can be used to compress / decompress the quantized computational wavelet coefficients.

[0428] As described above, the bottom-up approach can build the level of detail (LOD) and simultaneously calculate the predicted attribute values. In some embodiments, such a method can proceed as follows:

[0429] ● Assume (P i ) i=1…N is the set of positions associated with the point cloud points, and assume (M i ) i=1…N is the Morton code associated with (P i ) i=1…N . Assume D0 and ρ are two user-defined parameters that respectively specify the initial sampling distance and the distance ratio between LODs. The Morton code can be used to represent multi-dimensional data in one dimension, where the "Z-order function" is applied to the multi-dimensional data to produce a one-dimensional representation. Note that ρ > 1

[0430] ● First, sort the points in ascending order according to their associated Morton codes. Assume I is the array of point indices sorted according to this process.

[0431] ● The algorithm proceeds iteratively. At each iteration k, the points belonging to LOD k are extracted, and its predictor is built starting from k = 0 until all points are assigned to LODs.

[0432] ● The sampling distance D is initialized to D = D0

[0433] ● For each iteration k, where k = 0…the number of LODs

[0434] ○ Assume that L(k) is the index set of points belonging to the k-th LOD, and O(k) is the set of points belonging to LODs higher than k. L(k) and O(k) are calculated as follows.

[0435] ○ First, initialize O(k) and L(k)

[0436] ■ If k = 0, then L(k) ← {}. Otherwise, L(k) ← L(k-1)

[0437] ■ O(k) ← {}

[0438] ○ Traverse the point indices stored in the array I in order. Each time, select the index i and calculate its distance to the SR1 point that was most recently added to O(k) (e.g., Euclidean distance or other distance). SR1 is a user-defined parameter for controlling the accuracy of the nearest neighbor search. For example, SR1 can be chosen as 8 or 16 or 64, etc. The smaller the value of SR1, the lower the computational complexity and the accuracy of the nearest neighbor search. The parameter SR1 is included in the bitstream. If any SR1 distance is less than D, then append i to the array L(k). Otherwise,

[0439] Append i to the array O(k).

[0440] ■ The parameter SR1 can be adaptively changed based on the LOD or / and the number of points traversed.

[0441] ■ In some embodiments, instead of calculating approximate nearest neighbors, an exact nearest neighbor search technique can be applied.

[0442] ■ In some embodiments, exact and approximate neighbor search methods can be combined. Specifically, depending on the LOD and / or the number of points in I, the method can switch between exact and approximate search methods. Other criteria can include point cloud density, the distance between the current point and the previous point, or any other criterion related to the point cloud distribution.

[0443] ○ Iterate this process until all indices in I have been traversed.

[0444] ○ At this stage, L(k) and O(k) will be calculated and used in the next step to construct predictors associated with the points in L(k).

[0445] ○ More precisely, assume that R(k) = L(k) \ L(k-1) (where \ is the difference operator) is the set of points that need to be added to LOD(k-1) to obtain LOD(k). For each point i in R(k), we want to find the h nearest neighbors of i in O(k) (h is a user-defined parameter for controlling the maximum number of neighbors used for prediction), and calculate the prediction weight associated with i (α j (i)) j=1…h。The algorithm proceeds as follows.

[0446] ○ Initialize the counter j = 0

[0447] ○ For each point i in R(k)

[0448] ■ Assume M i is the Morton code associated with i, and assume M j is the Morton code associated with the j-th element of the array O(k)

[0449] ■ When (M i ≥ M j and j < SizeOf(O(k))), increment the counter j by one (j ← j + 1)

[0450] ■ Calculate the distance from M i to the points associated with the indices of O(k) within the range [j - SR2, j + SR2] of the array, and track the h nearest neighbors (n1, n2,..., nh) and their associated squared distances SR2 is a user-defined parameter used to control the accuracy of the nearest neighbor search. Possible values of SR2 are 8, 16, 32, and 64. The smaller the value of SR2, the lower the computational complexity and the accuracy of the nearest neighbor search. The parameter SR2 is included in the bitstream. The calculation of the prediction weights for attribute prediction can be the same as above.

[0451] ○ The parameter SR2 can be adaptively changed based on the LOD or / and the number of traversed points.

[0452] ○ In some embodiments, instead of calculating approximate nearest neighbors, an exact nearest neighbor search technique can be used.

[0453] ○ In some embodiments, exact and approximate neighbor search methods can be combined. Specifically, depending on the LOD and / or the number of points in I, the method can switch between exact and approximate search methods. Other criteria can include point cloud density, the distance between the current point and the previous point, or any other criterion related to the point cloud distribution.

[0454] ○ If the distance between the current point and the last processed point is less than a threshold, use the neighborhood of the last point as an initial guess and search nearby. The threshold can be adaptively selected based on criteria similar to those above. The threshold can be signaled in the bitstream or be known to both the encoder and the decoder.

[0455] ○ The previous idea can be generalized to n = 1, 2, 3, 4… the last point

[0456] ○Exclude points whose distance is greater than a user-defined threshold. The threshold can be adaptively selected based on criteria similar to those described above. The threshold can be signaled in the bitstream or be known to both the encoder and the decoder.

[0457] ○I ← O(k)

[0458] ○D ← D × ρ

[0459] ○The above methods can be used with any metric (e.g., L2, L1, Lp) or any approximation of these metrics. For example, in some embodiments, distance comparisons can use Euclidean distance comparison approximations such as taxi / Manhattan / L1 approximation or octagon approximation.

[0460] In some embodiments, the lifting scheme can be applied in the context of a hierarchical uncertainty level. In such embodiments, the technique can proceed as follows:

[0461] ● Sort the input points according to the Morton code associated with the coordinates of the input points

[0462] ● Encode / decode the point attributes according to the Morton order

[0463] ● For each point i, find the h nearest neighbors (n1, n2,..., n h )(n j (< i)

[0464] ● Calculate the prediction weights as described above.

[0465] ● Apply the above adaptive scheme to adjust the prediction strategy.

[0466] ● Predict the attributes and entropy encode them as described below.

[0467] Binary Arithmetic Coding of Quantized Lifting Coefficients

[0468] In some embodiments, the lifting scheme coefficients can be non-binary values. In some embodiments, the arithmetic encoder (such as arithmetic encoder 212) of the components of the encoder 202 described above, and using a binary arithmetic codec to encode the occupancy symbols of 256 values can also be used to encode the lifting scheme coefficients. Alternatively, in some embodiments, a similar arithmetic encoder can be used. For example, the technique can proceed as follows: Figure 2B ● One-dimensional attributes

[0469] ● Assume C is the quantization coefficient to be encoded. First, map C to a positive number using a function that maps positive numbers to even numbers and negative numbers to odd numbers.

[0470] ​

[0471] ○Assume M(C) is the mapped value.

[0472] ○Then encode the binary value to indicate whether C is 0

[0473] ○ If C is not zero, two cases are distinguished

[0474] ■ If M(C) is greater than or equal to alphabetSize (for example, Figures 12C to 12F The difference between M(C) and AlphabetSize is encoded using exponential Golomb coding.

[0475] ■ Otherwise, use the above for N Figures 12C to 12F The described method encodes the value of M(C).

[0476] ●Three-dimensional signal

[0477] ○ Assume that C1, C2, C3 are quantized coefficients to be encoded. Assume that K1 and K2 are two indexes of the context used to encode the quantized coefficients C1, C2, and C3.

[0478] ○First, as mentioned above, Figures 12C to 12F As described, C1, C2 and C3 are mapped to positive numbers. Assume that M(C1), M(C2) and M(C3) are the mapped values ​​of C1, C2 and C3.

[0479] ○Encode M(C1).

[0480] ○ In the selection of different contexts based on whether C1 is zero (i.e., the above Figures 12C to 12F M(C2) is encoded when the binary arithmetic context and binarization context are defined.

[0481] ○ Encode M(C3) while choosing different contexts based on the conditions C1 is zero and C2 is zero. If C1 is zero, then we know that the value is at least 16. If the condition C1 is zero, then use binary context K1, if the value is not zero, then decrement the value by 1 (knowing that the value is at least one or more), then check if the value is less than the alphabet size, if so, then encode the value directly. Otherwise, encode the maximum possible value of the alphabet size. The difference between the maximum possible value of the alphabet size and the value of M(C3) will be encoded using exponential Golomb coding.

[0482] Multi-dimensional signals

[0483] ○ The same method described above can be extended to d-dimensional signals. Here, the context for encoding the k-th coefficient depends on the values of the previous coefficients (e.g., the last 0, 1, 2, 3, …, k-1 coefficients).

[0484] ○ The number of previous coefficients to be considered can be adaptively selected depending on any of the criteria described in the previous section for choosing SR1 and SR2.

[0485] The following is a more detailed discussion on how to use the point cloud transfer algorithm to minimize the distortion between the original point cloud and the reconstructed point cloud.

[0486] The attribute transfer problem can be defined as follows:

[0487] a. Assume PC1 = (P1(i)) i∈{1,…,N1} is a point cloud defined by its geometry (i.e., 3D position) (X1(i)) i∈{1,…,N1} and a set of attributes (e.g., RGB color or reflectance) (A(i)) i∈{1,…,N1}

[0488] b. Assume PC2 (P2(j)) j∈{1,…,N2} is a resampled version of PC1, and assume (X2(j)) j∈{1,…,N2} is its geometry.

[0489] c. Then calculate the set of attributes (A2(j)) j∈{1,…,N2} associated with the points of PC2 such that the texture distortion is minimized.

[0490] To solve the texture distortion minimization problem using the attribute transfer algorithm:

[0491] ● Assume P (3→1)(j) ∈ PC1 is the nearest neighbor of P2(j) ∈ PC2 in PC1, and A (2→1)(j) is its attribute value.

[0492] ● Assume P (1→2)(i) ∈ PC2 is the nearest neighbor of P1(i(∈ PC1 in PC2, and A (1→2)(i) is its attribute value.

[0493] ● Assume is the set of points in PC2 that share the point P1(i) ∈ PC1 as their nearest neighbor, and assume (α(j,h)) h∈{1,…,H(j)} is its attribute value

[0494] ● Assume E 2→1 is the asymmetric error calculated from PC2 to PC1:

[0495] ● ​

[0496] ● Assume E 1→2 is the asymmetric error calculated from PC1 to PC2:

[0497] ●

[0498] ● Assume E is the symmetric error that measures the attribute distortion between PC2 and PC1:

[0499] ● E = max(E 2→1 , E 1→2 )

[0500] Then the attribute set (A2(j)) is determined as follows j∈{1,…,N2} :

[0501] a. Initialize E1 ← 0 and E2 ← 0

[0502] b. Traverse all the points of PC2

[0503] 1) For each point P2(j), calculate P (2→1)(j) ∈ PC1 and

[0504] 2) If (E1 > E2 or )

[0505] ● A2(j) = A 2→1 (j)

[0506] 3) Otherwise

[0507] ●

[0508] 4) End condition

[0509] 5) E1 ← E1 + ‖A2(j) - A 2→1 (j)‖ 2

[0510] 6)

[0511] Point Cloud Attribute Transfer Algorithm

[0512] In some embodiments, a point cloud transfer algorithm can be used to minimize the distortion between an original point cloud and a reconstructed version of the original point cloud. The transfer algorithm can be used to evaluate the distortion due to slightly different point positions in the original and reconstructed point clouds. For example, the reconstructed point cloud can have a similar shape to the original point cloud, but can have a.) a different total number of points and / or b.) points that are slightly offset compared to corresponding points in the original point cloud. In some embodiments, the point cloud transfer algorithm can allow for the selection of attribute values for the reconstructed point cloud such that the distortion between the original point cloud and the reconstructed version of the original point cloud is minimized. For example, for the original point cloud, both the position of the points and the attribute values of the points are known. However, for the reconstructed point cloud, the position values may be known (e.g., based on the subsampling process, K-D tree process, or patch image process described above). However, it may still be necessary to determine the attribute values of the reconstructed point cloud. Thus, the point cloud transfer algorithm can be used to minimize the distortion by selecting the attribute values of the reconstructed point cloud that minimize the distortion.

[0513] The distortion from the original point cloud to the reconstructed point cloud can be determined with respect to the selected attribute values. Similarly, the distortion from the reconstructed point cloud to the original point cloud can be determined with respect to the selected attribute values for the reconstructed point cloud. In many cases, these distortions are asymmetric. The point cloud transfer algorithm is initialized with two errors (E21) and (E12), where E21 is the error from the second or reconstructed point cloud to the original or first point cloud, and E12 is the error from the first or original point cloud to the second or reconstructed point cloud. For each point in the second point cloud, it is determined whether the attribute value of the corresponding point in the original point cloud should be assigned to that point, or whether the average attribute value of the nearest neighbors to the corresponding point in the original point cloud should be assigned to that point. The attribute values are selected based on the minimum error.

[0514] Trimming the Search Space for Nearest Neighbor Search in Point Cloud Compression and Decompression

[0515] In various embodiments, the techniques for generating levels of detail (LOD) discussed above can iteratively apply a subsampling process in order to separate the points in the current LOD from the points belonging to the next LOD. In various embodiments, the process can include: checking the distance for each point according to the Euclidean norm (sometimes referred to as the "L2" norm), where the point is represented as a vector to all other points in the point cloud that appear before the current point in the order used to generate the LOD (e.g., Morton order). If any of the distances is higher than a defined threshold, the point can be included in the current LOD. Otherwise, the point can belong to the subsequent LOD. By trimming the search to a limited search range (e.g., a search range of 64 or 128, "SR1") as discussed below in various embodiments, an approximate nearest neighbor search can be performed, which is an order of magnitude faster than other nearest neighbor search techniques that do not utilize an approximation.

[0516] For example, in some embodiments, the search space for nearest neighbor search can be trimmed (or otherwise reduced) by implementing a boundary shape. Consider the following exemplary embodiment where the points of a point cloud are clustered, grouped, or otherwise assigned to point buckets of a fixed size, as Figure 14 shown. Each bucket (such as buckets 1420a, 1420b, 1420c, etc.) can be assigned points determined by a range of space filling curve values 1400 in the point cloud, such as space filling curve ranges 1410a, 1410b, and 1410c assigned to buckets 1420a, 1420b, and 1420c, respectively. Consider the following exemplary embodiment where every 8, 16, or 32 consecutive points in Morton order can be grouped in point bucket 1420. A boundary shape, such as an axis-aligned bounding box 1430, can then be determined and is associated with each bucket (since the bounding box 1430 can be associated with the points included in bucket 1420a). In various embodiments, the bounding box can be calculated incrementally or pre-calculated before starting the subsampling process. Although a Figure 14 bounding box is shown and discussed, other boundary shapes can be implemented in other embodiments, such as a bounding sphere, a bounding capsule, etc.

[0517] The boundary shape can be used to calculate a minimum boundary based on the distance from points outside the point bucket to the points assigned to the point bucket. For example, as Figure 14 indicated, the minimum distance 1440 to a point can be determined relative to the bounding box 1430 (instead of calculating the distance between an external point and the points within bucket 1420a). Figure 15 shows a high-level flowchart for applying a boundary shape to trim the search space for nearest neighbor search according to some embodiments.

[0518] As shown at 1510, in some embodiments, the points of a point cloud can be grouped within different ranges of a space filling curve, where these different ranges include corresponding space filling curve values generated for the points. For example, Morton codes can be generated for the points of the point cloud. Points whose Morton code values fall within the range of Morton code values assigned to group A (e.g., bucket A) can be grouped in group A. Points within Morton code values that fall within the range of Morton code values assigned to group B can be similarly grouped.

[0519] As shown at 1520, in some embodiments, a boundary shape can be determined for the grouped points within different ranges of the space filling curve. For example, a shape that encompasses all the points within the group, such as a box, a sphere, a cube, etc., can be determined. In some embodiments, techniques can be implemented to determine an optimally fitting boundary shape (e.g., a shape that encompasses all the points but covers the minimum amount of space). In some embodiments, the boundary shape can be determined as a preprocessing step before subsampling, or this can be performed iteratively.

[0520] As shown at 1530, in various embodiments, a nearest neighbor search for encoding a point cloud may be performed. For example, as discussed above with respect to FIGS. 1 to Figure 13B as discussed, various nearest neighbor searches may be performed. To determine whether a set of points should be included in the nearest neighbor search, the distance between the point for which the nearest neighbor search is performed and the boundary shape of the set may be determined, as shown at 1540. The determined distance may be compared with a sampling threshold. If the sampling threshold is exceeded, then as shown at 1550, those points in the set whose distance to the corresponding boundary shape exceeds the threshold may be excluded from the nearest neighbor search. In this way, the search space may be reduced by not separately calculating the distances of the points within the exclusion boundary shape. For sets within the boundary shape that do not exceed the threshold, the distances may be determined individually using the points in the set and the point being searched. In some embodiments, those distances within the sampling threshold may be included.

[0521] In some embodiments, the techniques discussed above with respect to Figure 15 and Figure 16 may include further optimization. For example, in some embodiments, the distance metric may be changed from the L2 distance (the L2 distance may be implemented using 3 multiplication operations, which may be expensive if considering HW implementations), alternatively, L1 normalization may be used (e.g., L1 normalization takes the absolute value of a k - dimensional point |x|+|y|+|z| etc., up to k dimensions, and L1 normalization may be implemented using adders, thus providing a cheap hardware implementation). In some embodiments, L1 may be implemented for the initial distance search and a list of potential nearest neighbor candidates may be maintained. Then, a search based on a different distance metric such as L2 normalization (e.g., L2 normalization takes the sum of the squares of a k - dimensional point x 2 +y 2 +z 2 etc., up to k dimensions, and L2 normalization may also optionally take the square root of the sum of the squares) may be applied to refine the nearest neighbor candidates. Note that in some embodiments, the point cloud may be defined in more than three dimensions (e.g., X, Y, and Z), for example, in some embodiments, the fourth dimension may be time etc.

[0522] In at least some embodiments, the techniques discussed above for subsampling to find the nearest neighbors for search may be parallelized by calculating the distances in parallel for each point. In some embodiments, the search range may also be trimmed based on the following two properties of Morton - ordered points, such as by using the smallest quadtree box containing two points to determine whether a point should be considered with respect to the nearest neighbor, and this smallest quadtree box will also contain all the points in Morton order between these two points.

[0523] In various embodiments, search spaces for nearest neighbor searches can be trimmed for predictor generation techniques. For example, predictor generation techniques can attempt to compute, for each point in the current LOD, the k nearest neighbors of the point in a subsequent LOD, and / or the k nearest neighbors of the point with a lower decoder order in the same LOD. In these and other predictor generation scenarios, techniques for trimming search spaces can be implemented to improve search performance.

[0524] For example, in some embodiments, because nearest neighbor searches can be applied to all points in the current LOD, the results for one or more points can be reused for searching other points. In some embodiments, nearest neighbor searches can be performed within a search range (SR2) for even points on a space filling curve (such as a Morton order). For odd points in the space filling curve, the results from the event point distance calculation can be reused, and thus the search space can be reduced. In various embodiments, the selection of points that will benefit from nearest neighbor searches with search range SR2 and points that will reuse search results can be adaptive based on inter-point distance criteria. Some points can also combine the reuse of other point searches with a limited local search with a search range SR3 that is less than SR2.

[0525] Figure 16 An example of the reuse of nearest neighbor search results according to the range of a space filling curve is shown. The space filling curve 1600, which can be a Morton order or other space filling curve, can show different ranges of space filling curve values, such as ranges 1610a, 1610b, 1610c, 1610d, 1610e, etc. As discussed in the above example, the ranges can be alternating odd and even space filling curve values. However, in other embodiments, different ranges (including ranges of different sizes) can be implemented. The new search results generated for the current LOD and the previous search results generated for a previous LOD can identify different ranges of the space filling curve. For example, for the nearest neighbor search 1640, ranges 1610a, 1610c, and 1610e can be identified for reuse, such that the search results (e.g., the distance values that determine which points should be sampled as nearest neighbors) from the nearest neighbor searches in the previous LOD can be reused, as shown at 1620a, 1620c, and 1620e. Some search ranges can be identified for generating new results, such as the space filling curve ranges 1610b and 1610d, which use the search results from the current LOD 1620b and 1620d, respectively.

[0526] Figure 17Shows a high - level flowchart for applying a boundary shape to trim the search space for nearest - neighbor search according to some embodiments. As shown at 1710, according to the various techniques discussed above, as part of subsampling to generate predictors, nearest - neighbor search results for points at a first level of detail (LOD) can be generated. As shown at 1720, according to some embodiments, nearest - neighbor search results can be generated for points within the range identified for the space - filling curve values of the second LOD of the point cloud. As shown at 1730, according to some embodiments, those portions of the nearest - neighbor search results of the points at the first LOD of the point cloud can be selected for points within the range identified for result reuse for the space - filling curve values in the second LOD.

[0527] In some embodiments, various further optimizations can be implemented. For example, as discussed above, different normalization techniques (e.g., L1 norm instead of L2) can be used. In some embodiments, different normalization techniques can be performed to filter search results, such as using L1 normalization for an initial search and maintaining a list of potential nearest - neighbor candidates. Then, an L2 - based search can be utilized to refine the nearest - neighbor candidate list. In some embodiments, according to the techniques discussed above, for nearest - neighbor searches in subsequent LODs, the bounding boxes calculated in the previous section can be reused to trim the search space in the same manner as in the previous section. In some embodiments, for points at the same LOD, new bounding boxes can be calculated as described, and additionally, the bounding boxes can be used to generate search results for those portions identified for the new results.

[0528] In some embodiments, a k - nearest - neighbor graph G can be pre - calculated for points belonging to a subsequent LOD. Then this graph can be used to accelerate the search process. For example, for each point P, a fast search with a limited search range will first be applied to find an approximate nearest - neighbor n0. Then the neighbors of the node in graph G are evaluated to find the k nearest - neighbors of P in graph G. The process of generating graph G can be accelerated by only looking for nearest - neighbors with lower space - filling curve values (e.g., lower Morton order), and making graph G symmetric (e.g., if P is a neighbor of Q, then make Q a neighbor of P). To reduce the memory requirements of G, the number of nearest - neighbor searches in graph G can be adaptively changed based on the point - cloud characteristics.

[0529] All of the above methods can be combined in different ways to achieve different optimizations for reducing the search space. As previously mentioned, in some embodiments, the nearest - neighbor search can be parallelized along different dimensions, such as by performing nearest - neighbor searches for different points in parallel, and / or by searching for neighbors for each point in parallel.

[0530] Further Enhancement of the Trimmed Search Space for Nearest Neighbor Search

[0531] In the context of both the lift / predict scheme and the region adaptive hierarchical transform (RAHT) scheme, the "k" nearest neighbor problem (k-NN) can be formulated and described as follows Figure 18 Let A and B be two sets of points in the discrete metric space R^d, where d is the dimension of the space (e.g., d = 3). For each point in A, it is desired to find its k nearest neighbors in B.

[0532] As discussed above, both A and B can be sorted according to their Morton order. Also as discussed above, an approximate "K" nearest neighbor solution can be calculated as follows:

[0533] ● Assume j is a counter initialized to 0

[0534] ● Assume P(i) is the i-th point of A according to the Morton order

[0535] ● Keep incrementing j until the Morton code of P(i) (denoted as MP(i)) satisfies the following condition

[0536] ● where MQ(j) is the Morton code of the j-th point of B according to the Morton order (denoted as Q(j)).

[0537] ● Apply a finite search in the following subset S(j) of the points of B

[0538] S(j) = {Q(j), Q(j + 1), Q(j - 1), …, Q(j + Δ), Q(j - Δ)}

[0539] Note that when A = B, the same algorithm can be applied, where j = i.

[0540] It should also be noted that the problem can be further restricted by only allowing the "K" nearest neighbors with lower Morton codes. Here, S(j) is defined as follows:

[0541] S(j) = {Q(j), Q(j - 1), …, Q(j - Δ)}

[0542] This method provides a good approximation of the "K" nearest neighbors. However, when a significant jump in terms of the Morton order is observed between adjacent points (see points P and Q in Figure 19 ), the actual nearest neighbors may not be captured.

[0543] Since the determined "K" nearest neighbors are used to predict the attributes of a point and / or determine the level of detail (LOD), improving the accuracy of the "K" nearest neighbor search can improve compression efficiency. For example, a prediction based on more accurate nearest neighbors can produce better prediction results compared to a prediction based on a set of points that inadvertently excludes one or more points that are closer to the point being evaluated than the points included in this set of nearest neighboring points, e.g., the excluded nearest neighboring points. Additionally, if the prediction is more accurate, the residual value, e.g., the attribute correction value, can also be smaller. Since the attribute correction value is encoded in the compressed attribute file, reducing the size of the attribute correction value will also reduce the amount of data that must be encoded and transmitted to the decoder, thereby improving compression efficiency. Further, if the attribute correction value is quantized, a larger quantized attribute correction value may also affect the reconstruction quality of the attributes of the point cloud. Therefore, by using a more accurate set of neighboring points for prediction to improve the prediction accuracy and thus reduce the size of the attribute correction value, the quality of the reconstructed point cloud can also be improved.

[0544] In some embodiments, to further refine the "K" nearest neighbor search, the solution can be calculated as follows:

[0545] ○ Assume j is a counter initialized to 0

[0546] ○ Assume P(i) is the i-th point of A according to the Morton order

[0547] ○ Keep incrementing j until the Morton code of P(i), denoted as MP(i), satisfies the following condition

[0548] MP(i)≥MQ(j),

[0549] where MQ(j) is the Morton code of the j-th point of B, denoted as Q(j).

[0550] ○ Apply a limited search in the following subset S(j) of the points of B

[0551] S(j) = {Q(j), Q(j + 1), Q(j - 1), …, Q(j + Δ), Q(j - Δ)}

[0552] ○ Assume N(i,1), N(i,2), …, N(i,H) are a set of neighbors of P(i) in R d For example, the 6 / 18 / 26 connectivity of P(i) forms the union {P(i)}

[0553] ○ For each point N(i,h), if it exists in B, perform a search

[0554] ○ Alternative 1: Apply a binary search

[0555] ■ Assume MN(i,h) is the Morton code of N(i,h)

[0556] ■ If MN(i,h) ≥ MQ(j - Δ) and MN(i,h) ≤ MQ(j + Δ), do nothing (if it belongs to B, it has already been captured by the initial search).

[0557] ■ If MN(i,h) < MQ(j - Δ), apply binary search to the integer interval [j - Δ - 1 - MQ(j - Δ) + MN(i,h), j - Δ - 1].

[0558] ■ If MN(i,h) > MQ(j + Δ), apply binary search to the integer interval [j + Δ + 1, j + Δ + 1 + MQ(j + Δ) - MN(i,h)].

[0559] ○ Alternative 2: Use a lookup table

[0560] ■ Assume C = [0, …, 2 c -1] × [0, …, 2 c -1] × [0, …, 2 c -1] is the boundary cube of B (i.e., )

[0561] ■ First, assume any point X in the mapped C is

[0562] ● True (or -1), if the point does not belong to B

[0563] ● False (or the index of X in B), otherwise.

[0564] ■ Use the LUT to check if MN(i,h) belongs to B. If MN(i,h) belongs to B, add MN(i,h) to S(j).

[0565] ■ Note 1: The LUT can be implemented as a table or using any hierarchical structure (e.g., octree, kd - tree, hierarchical grid, …) to save memory by taking advantage of the sparse nature of the point cloud.

[0566] ■ Note 2: The LUT can store multiple indices per location to handle duplicate points.

[0567] ■ Note 3: Allocating the LUT that holds C can be memory - expensive. To reduce such requirements, the boundary cube C can be divided into cubes of size 2 eThe smaller sub - cuboid {E(a,b,c)}. Since the points are traversed in Morton order, all points in a sub - cuboid are traversed in order before switching to the next sub - cuboid. Therefore, only a LUT capable of storing a single sub - cuboid is required. Whenever P(i) enters a new sub - cuboid, the LUT is initialized. Only the points of B that are in this sub - cuboid are added to the LUT. For points on the boundary of the sub - cuboid, two strategies are possible:

[0568] ●Ignore neighborhood relationships across sub - cuboid boundaries

[0569] ●Use binary search (i.e., alternative 1) to determine if its neighbor exists.

[0570] ○Alternative 3: Combine alternative 1 and alternative 2

[0571] Note that the connectivity of neighbors, such as the connectivity of 6 adjacent voxels, can be similar to Figure 12C the connectivity 1286 shown. Although not shown, similar connectivities for more voxels (such as 18 or 26) can be used. Additionally, as Figure 12B shown, the cuboid can be partitioned and further divided into sub - cuboids. In some embodiments, the voxel in which the point for which the nearest neighbor search is being performed can also be searched to see if there is another adjacent point in the same voxel. In this case, when the voxel in which the point being evaluated resides is further included in the set of voxels to be considered, the connectivity can be 7, 19, 27, etc.

[0572] Also, as can be seen above, the binary search used to determine if the Morton code of a point in a point cloud matches the Morton code of an adjacent voxel can exclude the Morton codes used in the initial search based on Morton codes. This is because the points with those Morton codes have already been identified by the first search, so if they have already been identified as one of the nearest neighbor points, there is no need to search for them again in adjacent voxels.

[0573] Also, in some embodiments, to further refine the "K" nearest neighbor search, the solution can be calculated as follows:

[0574] ○The search can be conducted in a manner opposite to that just described above. For example, first apply a search based on adjacent voxels (e.g., neighbors N(i,1), N(i,2), …, N(i,H)). Then, further refine the search by applying a trimmed Morton order search.

[0575] ○Apply a k - NN search based on N(i,1), N(i,2), …, N(i,H)

[0576] ○No neighbor is found

[0577] ■ Search in S(j) = {Q(j), Q(j+1), Q(j-1), …, Q(j+Δ), Q(j-Δ)}

[0578] ○ Find a neighbor

[0579] ■ Assume j * is the index of the nearest neighbor found in B

[0580] ■ In S(j * ) = {Q(j * ), Q(j * +1), Q(j * -1), …, Q(j * +Δ), Q(j * -Δ)} perform a search

[0581] ○ Find k′ < k neighbors. Assume is the index of the nearest neighbor found in B

[0582] ■ Method 1

[0583] ● In perform a search

[0584] ■ Method 2

[0585] ● Assume is the index of the nearest neighbor found in B

[0586] ● In perform a search

[0587] Additionally, in some embodiments, when searching for the "K" nearest neighbors, boosting schemes and / or region adaptive hierarchical transform (RAHT) schemes such as Figures 13A to 13B described may be further utilized. For example, the boosting / prediction scheme ensures that at each LODl, the distance between every two points is higher than a predefined threshold d(l). The RAHT scheme merges all points sharing the same Morton code with an offset of l bits at each level. If the distance sequence is constrained as {d(1), d(2), … d(L)}, as follows, the same properties as RAHT can be achieved:

[0588] ○

[0589] ○ d(l+1) = 2 × d(l)

[0590] The example described below focuses on the boosting / prediction scheme. However, a similar approach can also be applied to the RAHT scheme.

[0591] Based on the above constrained distance model, all points with known LOD l = 2 have distances higher than √3×2^(n0+1). Thus, if the coordinates of the points of B are divided by 2^(n0+1) (e.g., shifting all coordinates by (n0+1) bits), all points will still have different coordinates while significantly shrinking the size of the LUT and / or the range of the binary search described in the previous section.

[0592] Exemplary Applications for Point Cloud Compression and Decompression

[0593] Figure 20 Shows a compressed point cloud being used in a 3D remote display application according to some embodiments.

[0594] In some embodiments, sensors (such as sensor 102), encoders (such as encoder 104 or encoder 202), and decoders (such as decoder 116 or decoder 220) can be used to transmit point clouds in 3D remote display applications. For example, at 2002, a sensor (such as sensor 102) can capture a 3D image, and at 2004, the sensor or a processor associated with the sensor can perform 3D reconstruction based on the sensed data to generate a point cloud.

[0595] At 2006, an encoder (such as encoder 104 or 202) can compress the point cloud, and at 2008, the encoder or a post-processor can package the compressed point cloud and transmit the compressed point cloud via network 2010. At 2012, the data packet can be received at a target location including a decoder (such as decoder 116 or decoder 220). At 2014, the decoder can decompress the point cloud, and at 2016, the decompressed point cloud can be rendered. In some embodiments, the 3D remote display application can transmit point cloud data in real time such that the display at 2016 can represent the image being observed at 2002. For example, at 2016, a camera in a canyon can allow a remote user to experience walking through a virtual canyon.

[0596] Figure 21 Shows a compressed point cloud being used in a virtual reality (VR) or augmented reality (AR) application according to some embodiments.

[0597] In some embodiments, the point cloud may be generated in software (e.g., as opposed to being captured by a sensor). For example, at 2102, virtual reality or augmented reality content is generated. The virtual reality or augmented reality content may include point cloud data and non-point cloud data. For example, as one example, non-point cloud characters may traverse a terrain represented by the point cloud. At 2104, the point cloud data may be compressed, and at 2106, the compressed point cloud data and non-point cloud data may be packaged and transmitted via network 2108. For example, the virtual reality or augmented reality content generated at 2102 may be generated at a remote server and transmitted via network 2108 to a VR or AR content consumer. At 2110, the data packet may be received and synchronized at the device of the VR or AR consumer. At 2112, a decoder operating at the device of the VR or AR consumer may decompress the compressed point cloud, and the point cloud and non-point cloud data may be rendered in real time, for example, in a head-mounted display of the device of the VR or AR consumer. In some embodiments, the point cloud data may be generated, compressed, decompressed, and rendered in response to the VR or AR consumer manipulating the head-mounted display to look in different directions.

[0598] In some embodiments, point cloud compression as described herein may be used in a variety of other applications such as geographic information systems, live sports events, museum displays, autonomous navigation, and the like.

[0599] Exemplary Computer System

[0600] Figure 22 An exemplary computer system 2200 is shown in accordance with some embodiments, which may implement an encoder or decoder or any other component described herein (e.g., any component described above with reference to FIGS. 1 to Figure 21 any of the components described). The computer system 2200 may be configured to execute any or all of the embodiments described above. In different embodiments, the computer system 2200 may be any of a variety of types of devices, including but not limited to: personal computer systems, desktop computers, laptop computers, notebooks, tablets, all-in-one computers, slate computers or netbook computers, mainframe computer systems, handheld computers, workstations, network computers, cameras, set-top boxes, mobile devices, consumer devices, video game controllers, handheld video game devices, application servers, storage devices, televisions, video recording devices, peripherals (such as switches, modems, routers), or generally any type of computing or electronic device.

[0601] Various embodiments of the point cloud encoder or decoder described herein may be executed on one or more computer systems 2200, which may interact with a variety of other devices. Note that, according to various embodiments, the above with respect to FIGS. 1 to Figure 21Any of the components, acts, or functionality described can be implemented on one or more computers of a computer system 2200 configured to Figure 22 In an illustrated embodiment, the computer system 2200 includes one or more processors 2210 coupled to a system memory 2220 via an input / output (I / O) interface 2230. The computer system 2200 also includes a network interface 2240 coupled to the I / O interface 2230, and one or more input / output devices 2250, such as a cursor control device 2260, a keyboard 2270, and a display 2280. In some cases, it is contemplated that an embodiment can be implemented using a single instance of the computer system 2200, while in other embodiments, multiple such systems or multiple nodes that make up the computer system 2200 can be configured to host different portions or instances of the embodiment. For example, in one embodiment, some elements can be implemented by one or more nodes of the computer system 2200 that are different from those nodes implementing other elements.

[0602] In various embodiments, the computer system 2200 can be a single-processor system including one processor 2210 or a multi-processor system including several processors 2210 (e.g., two, four, eight, or another suitable number). The processor 2210 can be any suitable processor capable of executing instructions. For example, in various embodiments, the processor 2210 can be a general-purpose or embedded processor implementing any one of a variety of instruction set architectures (ISAs), such as the x86, PowerPC, SPARC, or MIPS ISA, or any other suitable ISA. In a multi-processor system, each of the processors 2210 typically, but not necessarily, implements the same ISA.

[0603] The system memory 2220 can be configured to store point cloud compression or point cloud decompression program instructions 2222 and / or sensor data accessible by the processor 2210. In various embodiments, the system memory 2220 can be implemented using any suitable memory technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash-type memory, or any other type of memory. In the illustrated embodiment, the program instructions 2222 can be configured to implement an image sensor control application incorporating any of the functionality described above. In some embodiments, the program instructions and / or data can be received, sent, or stored on a different type of computer-accessible medium or a similar medium separate from the system memory 2220 or the computer system 2200. Although the computer system 2200 is described as implementing the functionality of the functional blocks of the preceding figures, any functionality described herein can be implemented by such a computer system.

[0604] In one embodiment, the I / O interface 2230 may be configured to coordinate I / O communications between the processor 2210, the system memory 2220, and any peripheral devices (including the network interface 2240 or other peripheral interfaces such as the input / output device 2250) in the device. In some embodiments, the I / O interface 2230 may perform any necessary protocol, timing, or other data conversions to convert data signals from one component (e.g., the system memory 2220) into a format suitable for use by another component (e.g., the processor 2210). In some embodiments, the I / O interface 2230 may include support for devices attached via various types of peripheral buses, such as variants of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard. In some embodiments, the functionality of the I / O interface 2230 may, for example, be divided into two or more separate components, such as a north bridge and a south bridge. Additionally, in some embodiments, some or all of the functionality of the I / O interface 2230 (such as the interface to the system memory 2220) may be incorporated directly into the processor 2210.

[0605] The network interface 2240 may be configured to allow the exchange of data between the computer system 2200 and other devices attached to the network 2285 (e.g., carrier or proxy devices) or between nodes of the computer system 2200. In various embodiments, the network 2285 may include one or more networks, including but not limited to a local area network (LAN) (e.g., Ethernet or enterprise network), a wide area network (WAN) (e.g., the Internet), a wireless data network, some other electronic data network, or some combination thereof. In various embodiments, the network interface 2240 may support communication via a wired or wireless general data network, such as any suitable type of Ethernet network; communication via a telecommunications / telephone network, such as an analog voice network or a digital fiber optic communication network; communication via a storage area network, such as a Fibre Channel SAN, or communication via any other suitable type of network and / or protocol.

[0606] In some embodiments, the input / output device 2250 may include one or more display terminals, keyboards, keypads, touchpads, scanning devices, voice or optical recognition devices, or any other devices suitable for inputting or accessing data by one or more computer systems 2200. Multiple input / output devices 2250 may be present in the computer system 2200 or may be distributed across various nodes of the computer system 2200. In some embodiments, similar input / output devices may be separate from the computer system 2200 and may interact with one or more nodes of the computer system 2200 via a wired or wireless connection (such as via the network interface 2240).

[0607] As Figure 22As shown, the memory 2220 may include program instructions 2222 that may be executable by the processor to implement any of the elements or actions described above. In one embodiment, the program instructions may execute the method described above. In other embodiments, different elements and data may be included. Note that the data may include any data or information described above.

[0608] Those skilled in the art will appreciate that the computer system 2200 is merely illustrative and is not intended to limit the scope of the embodiments. Specifically, the computer system and devices may include any combination of hardware or software that can perform the indicated functions, including computers, network devices, Internet devices, PDAs, wireless telephones, pagers, etc. The computer system 2200 may also be connected to other devices not shown, or alternatively may operate as a stand-alone system. Additionally, the functions provided by the illustrated components may in some embodiments be combined in fewer components or distributed in additional components. Similarly, in some embodiments, the functions of some of the illustrated components may not be provided, and / or other additional functions may be available.

[0609] Those skilled in the art will also recognize that although various items are shown as being stored in memory or on a storage device during use, for purposes of memory management and data integrity, these items or portions thereof may be transferred between memory and other storage devices. Alternatively, in other embodiments, some or all of these software components may be executed in the memory of another device and communicate with the illustrated computer system via inter-computer communication. Some or all of the system components or data structures may also be stored (e.g., as instructions or structured data) on a computer-accessible medium or portable article for reading by a suitable drive, examples of which are described above. In some embodiments, instructions stored on a computer-accessible medium separate from the computer system 2200 may be transmitted to the computer system 2200 via a transmission medium or signal (such as an electrical, electromagnetic, or digital signal transmitted via a communication medium such as a network and / or wireless link). Various embodiments may also include receiving, sending, or storing instructions and / or data implemented in accordance with the above description on a computer-accessible medium. Generally speaking, a computer-accessible medium may include a non-transitory computer-readable storage medium or memory medium, such as a magnetic or optical medium, e.g., a disk or DVD / CD-ROM, a volatile or non-volatile medium, such as RAM (e.g., SDRAM, DDR, RDRAM, SRAM, etc.), ROM, etc. In some embodiments, a computer-accessible medium may include a transmission medium or signal, such as an electrical, electromagnetic, or digital signal transmitted via a communication medium such as a network and / or wireless link.

[0610] In various embodiments, the methods described herein may be implemented in software, hardware, or a combination thereof. Additionally, the order of the blocks of the methods may be varied, and various elements may be added, reordered, combined, omitted, modified, etc. Various modifications and changes will be apparent to those of ordinary skill in the art who have benefited from this disclosure. The various embodiments described herein are intended to be illustrative and not restrictive. Many variations, modifications, additions, and improvements are possible. Accordingly, multiple examples may be provided for components described herein as a single example. The boundaries between various components, operations, and data repositories are to some extent arbitrary, and particular operations are illustrated in the context of specific exemplary configurations. Other allocations of functionality are anticipated and may fall within the scope of the appended claims. Finally, the structures and functions presented as discrete components in exemplary configurations may be implemented as a combined structure or component. These and other variations, modifications, additions, and improvements may fall within the scope of the embodiments as defined in the following claims.

Claims

1. One or more non-transitory computer-readable storage media storing program instructions that, when executed on or across one or more computing devices, cause the one or more computing devices to: Group points of a point cloud into one or more groups, where the points are grouped based on one or more space-filling curve value ranges, and where points of the point cloud having space-filling curve values within a given space-filling curve value range of the one or more space-filling curve value ranges are grouped into the same group of the one or more groups as other points of the point cloud having space-filling curve values within the given space-filling curve value range; For a respective group of the one or more groups of grouped points, determine a boundary shape that bounds the points included in the respective group; And Perform an adjacent point search, where, when performing the adjacent point search, the program instructions cause the one or more computing devices to: For the one or more groups of grouped points, determine a respective distance between a point of the point cloud for which the adjacent point search is being performed and the boundary shape of the one or more groups of grouped points; And Exclude from the adjacent point search those points included in a respective group of the one or more groups of grouped points for which the determined distance to the boundary shape of the respective group exceeds a distance threshold.

2. The one or more non-transitory computer-readable storage media according to claim 1, wherein, To perform the adjacent point search, the program instructions cause the one or more computing devices to: Determine a respective distance between the point and points included in a respective group of the one or more groups that have not been excluded from the adjacent point search; And Compare the respective distance between the point and the points of those groups that have not been excluded from the adjacent point search with a threshold distance to identify one or more points of the point cloud as adjacent points of the point.

3. The one or more non-transitory computer-readable storage media of claim 2, wherein the program instructions cause the one or more computing devices to further: Determine one or more levels of detail (LOD) of the point cloud, where the one or more levels of detail include a subset of the points of the point cloud, and where, to determine the one or more levels of detail (LOD), the program instructions cause the one or more computing devices to: Select a first point or one or more other points of the point cloud to be included in a given level of detail; Determine adjacent points within the threshold distance for the selected points; and Avoid including adjacent points of the selected points included in the given level of detail in the given level of detail.

4. One or more non-transitory computer-readable storage media according to claim 3, wherein the program instructions cause the one or more computing devices to: determine a first level of detail among the one or more levels of detail using the grouped points and the corresponding boundary shapes, and as part of performing an adjacent point search, repeatedly use the determined grouped points and the corresponding boundary shapes to determine one or more additional levels of detail among the one or more levels of detail.

5. One or more non-transitory computer-readable storage media according to claim 1, wherein the program instructions further cause the one or more computing devices to: predict an attribute value for a corresponding point among the points of the point cloud, wherein to predict the attribute value for a given point, the program instructions cause the one or more computing devices to: determine a corresponding distance between the given point and adjacent points not excluded from the adjacent point search; and predict the attribute value for the given point based on the corresponding attribute values of the adjacent points not excluded from the adjacent point search and the corresponding distances determined for the adjacent points not excluded from the adjacent point search.

6. One or more non-transitory computer-readable storage media according to claim 5, wherein the program instructions further cause the one or more computing devices to: determine a corresponding attribute correction value based on a difference between the predicted attribute value for a corresponding point among the points of the point cloud and a known attribute value for the points of the point cloud; and encode the determined attribute correction value in a compressed bitstream for the point cloud.

7. One or more non-transitory computer-readable storage media according to claim 5, wherein the program instructions further cause the one or more computing devices to: receive a compressed bitstream for the point cloud, the compressed bitstream including attribute correction values for the points of the point cloud; and apply the attribute correction values to the predicted attribute values for the points of the point cloud, wherein the predicted attribute values are predicted based on the attribute values and distances to adjacent points not excluded from the adjacent point search.

8. One or more non-transitory computer-readable storage media according to claim 5, wherein the program instructions cause the one or more computing devices to: use a distance calculated using the L-1 norm to determine the corresponding distance between the given point and adjacent points not excluded from the adjacent point search.

9. One or more non-transitory computer-readable storage media according to claim 5, wherein the program instructions cause the one or more computing devices to: use a distance calculated using the L-1 norm to determine the corresponding distance between the given point and adjacent points not excluded from the adjacent point search, and use the L-2 norm to refine a set of determined adjacent points to determine the corresponding distance, wherein a distance calculated using the L-2 norm is used to further evaluate the points determined to be adjacent points using the L-1 norm.

10. One or more non-transitory computer-readable storage media according to claim 9, wherein the distance calculated using the L-1 norm or the L-2 norm is determined in K dimensions, where K is three or greater.

11. One or more non-transitory computer-readable storage media according to claim 1, wherein the boundary shape is a cube or a rectangular prism having width, height, and depth dimensions parallel to the coordinate axes of the coordinates of the points defining the point cloud.

12. One or more non-transitory computer-readable storage media according to claim 11, wherein at least some of the boundary shapes include smaller boundary shapes corresponding to a subgroup of points whose space-filling curve values in the point cloud are within a subrange of the space-filling curve value range for the at least some boundary shapes.

13. One or more non-transitory computer-readable storage media according to claim 12, wherein the program instructions further cause the one or more computing devices to: For points not excluded as neighboring points, based on the distance between the point and the corresponding boundary shape among the at least some boundary shapes: For one or more of the subgroups, determine the corresponding distance between the point in the point cloud for which the neighboring point search is being performed and the smaller boundary shape for the one or more subgroups; and Exclude from the neighboring point search those points included in the corresponding subgroup among the one or more subgroups for which the determined distance to the smaller boundary shape for the corresponding subgroup exceeds the distance threshold.

14. A device, the device comprising: A memory storing program instructions; And One or more processors configured to execute the program instructions to: Group points of a point cloud into one or more groups, wherein the points are grouped based on one or more space-filling curve value ranges, and wherein points in the point cloud whose space-filling curve values are within a given space-filling curve value range among the one or more space-filling curve value ranges are grouped into the same group among the one or more groups as other points in the point cloud whose space-filling curve values are within the given space-filling curve value range; For a corresponding one of the one or more groups of grouped points, determine a boundary shape defining the points included in the corresponding group; And Perform a neighboring point search, wherein during the performance of the neighboring point search, the program instructions cause the one or more computing devices to: For the one or more groups of grouped points, determine the corresponding distance between the point in the point cloud for which the neighboring point search is being performed and the boundary shape for the one or more groups of grouped points; And Exclude from the neighboring point search those points included in the corresponding group among the one or more groups of grouped points for which the determined distance to the boundary shape for the corresponding group exceeds a distance threshold.

15. The apparatus according to claim 14, wherein the program instructions, when executed by the one or more processors, further cause the one or more processors to: Predict an attribute value for the point based on attribute values determined for neighboring points of the point for which the attribute value is being predicted for the point cloud.

16. The apparatus according to claim 15, wherein the program instructions, when executed by the one or more processors, further cause the one or more processors to: Receive a bitstream, the bitstream including: Spatial information for points of the point cloud; And Compressed attribute information for the points of the point cloud; And Apply the compressed attribute information included in the bitstream to adjust the predicted attribute value for the points of the point cloud to determine a reconstructed attribute value for the points of the point cloud.

17. The apparatus according to claim 16, wherein the program instructions, when executed by the one or more processors, further cause the one or more processors to: Determine one or more levels of detail (LOD) of the point cloud, wherein the one or more levels of detail include a subset of the points of the point cloud, and wherein, to determine the one or more levels of detail (LOD), the program instructions cause the one or more processors to: Select a first point or one or more other points of the point cloud to be included in a given level of detail; Determine neighboring points within the distance threshold for the selected points; and Avoid including neighboring points of the selected points included in the given level of detail in the given level of detail, wherein the predicted attribute value and the applying the compressed attribute information included in the bitstream to adjust the predicted attribute value are performed for a limited number of points of the point cloud included in the given level of detail being reconstructed in the level of detail.

18. One or more non-transitory computer-readable storage media storing program instructions that, when executed on one or more computing devices or across the one or more computing devices, cause the one or more computing devices to: Group points within different ranges of a space-filling curve based on corresponding space-filling curve values generated for points of a point cloud; Perform a nearest neighbor search for points of the point cloud, wherein, when performing the nearest neighbor search, the program instructions cause the one or more computing devices to: Exclude from the nearest neighbor search those points in one or more groups that have a distance value from the point for which the nearest neighbor search is being performed to one or more corresponding boundary shapes of the one or more groups that exceeds a sampling threshold; Evaluate space-filling curve values of neighboring voxels from the point for which the nearest neighbor search is being performed to determine whether any of the space-filling curve values generated for the points of the point cloud fall within one of the neighboring voxels; And Include in the result of the nearest neighbor search: A set of nearest neighboring points to the point for which the nearest neighbor search is performed, the set of nearest neighboring points being determined based on a search that excludes those points in the set whose distance values to the corresponding boundary shape exceed the sampling threshold; and One or more nearest neighboring points to the point for which the nearest neighbor search is performed, if the one or more nearest neighboring points are found in one of the neighboring voxels, wherein the one or more nearest neighboring points are excluded from the nearest neighboring points determined based on the search of the points grouped according to the space filling curve values.

19. One or more non-transitory computer-readable storage media according to claim 18, wherein the nearest neighbor search is performed as part of a prediction process performed by an encoder.

20. One or more non-transitory computer-readable storage media according to claim 18, wherein the nearest neighbor search is performed as part of a prediction process performed by a decoder.

21. One or more non-transitory computer-readable storage media according to claim 18, wherein the result of the nearest neighbor search, including one or more points that are excluded based on the distance values to the boundary shape of the grouped points but are found in one of the neighboring voxels, is used to predict an attribute value for a given point of the point cloud at an encoder or a decoder.

Citation Information

Patent Citations

  • Point cloud attribute compression method based on hierarchical division

    CN108632621A

  • Point cloud geometry compression

    US20190075320A1