Method and system for assessing point cloud

The PQM metric addresses the limitations of existing point cloud quality assessment methods by incorporating resolution, accuracy, coverage, and artifact-score sub-metrics, providing a comprehensive evaluation of point cloud quality for improved 3D perception applications.

WO2025255559A1PCT designated stage Publication Date: 2025-12-11THE RES FOUNDATION FOR THE STATE UNIV OF NEW YORK

Patent Information

Application Number
PCT/US2025/032796
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-06
Filing Date
2025-06-06
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing point cloud quality assessment techniques fail to capture the complexity of point clouds, particularly in terms of coverage, local variations in density, and the presence of artifacts, leading to inaccurate navigation and measurement in applications like autonomous vehicles and robotics.

Method used

A comprehensive point quality metric (PQM) is introduced, comprising resolution, accuracy, coverage, and artifact-score sub-metrics to evaluate point cloud quality, providing a normalized value between 0 and 1, with customizable weights for different applications.

Benefits of technology

PQM offers a thorough assessment of point cloud quality, enabling better debugging and optimization of 3D perception applications by capturing various aspects of point cloud quality, including completeness, accuracy, and artifact presence, thus improving system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025032796_11122025_PF_FP_ABST
    Figure US2025032796_11122025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method of assessing a candidate point cloud is provided. The method includes receiving a source point cloud representing a ground truth and the candidate point cloud. Each of the candidate point cloud and the source point cloud is divided into a plurality of cells of equal size. Two or more sub-metrics are calculated. For example, two or more of a resolution sub-metric, an accuracy sub-metric, a coverage sub-metric, and an artifact sub-metric may be calculated. A point quality metric based on the sub-metrics (e.g., the coverage sub-metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub- metric). The point quality' metric is output as a normalized value.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 011520.01958 METHOD AND SYSTEM FOR ASSESSING POINT CLOUD Cross-Reference to Related Applications

[0001] This application claims priority to U.S. Provisional Application No.63 / 657,108, filed on June 6, 2024, now pending, the disclosure of which is incorporated herein by reference Statement Regarding Federally Sponsored Research

[0002] This invention was made with government support under grant number 1846320 awarded by National Science Foundation. The government has certain rights in the invention. Field of the Disclosure

[0003] The present disclosure relates generally to point cloud processing and quality assessment. More specifically, the disclosure relates to computer-implemented methods and systems for evaluating the quality of point clouds, such as, for example, those generated by mapping systems, by comparing them against ground truth reference data using multiple quantitative metrics. Background of the Disclosure

[0004] Point clouds are three-dimensional data structures made up of a set of data points in space, typically representing the external surface of objects or environments. Point clouds are widely used in various applications including autonomous vehicles, robotics, augmented reality, virtual reality, 3D mapping, surveying, and architectural modeling. These point clouds are commonly generated by mapping systems such as Light Detection and Ranging (LIDAR) sensors, stereo cameras, structured light scanners, and photogrammetry systems.

[0005] The quality of such point clouds is critical for the success of downstream applications. Poor quality can lead to inaccurate navigation decisions in autonomous vehicles, imprecise measurements in surveying applications, poor object recognition in robotics. However, assessing point cloud quality presents significant challenges. Traditional point cloud quality assessment techniques often focus on single metrics. Such approaches fail to capture the varied aspects of point cloud quality.Attorney Docket No.: 011520.01958

[0006] There exists a need for an automated method for assessing point cloud quality that accounts for the complexity of point clouds and provides actionable metrics for system optimization and performance evaluation. Brief Summary of the Disclosure

[0007] Advancements in sensors, algorithms, and compute hardware has made 3D perception feasible in real-time. Current methods to compare and evaluate quality of a 3D model such as Chamfer, Hausdorff, and Earth-mover’s distance are uni-dimensional and have limitations; including inability to capture coverage, local variations in density and error, and are significantly affected by outliers. The present disclosure provides an evaluation framework for point clouds that may include four metrics: resolution (^^^) to quantify ability to distinguish between the individual parts in the point cloud, accuracy (^^^) to measure registration error, coverage (^^^) to evaluate portion of missing data, and artifact-score (^^௧) to characterize the presence of artifacts. Through detailed analysis, we demonstrate the complementary nature of each of these dimensions, and the improvement they provide compared to uni-dimensional measures highlighted above. Further, we demonstrate the utility of ^^^^^^^^^^3^^ by comparing our metric with the uni-dimensional metrics for two 3D perception applications (SLAM and point cloud completion). The presently-disclosed techniques advance our ability to reason between point clouds and helps better debug 3D perception applications by providing richer evaluation of their performance. Description of the Drawings

[0008] For a fuller understanding of the nature and objects of the disclosure, reference should be made to the following detailed description taken in conjunction with the accompanying drawings.

[0009] Figure 1. A chart illustrating a method according to an embodiment of the present disclosure.

[0010] Figure 2. A diagram of a system according to another embodiment of the present disclosure.

[0011] Figure 3: Point Quality Metric (PQM) is the only metric that correctly identifies FAST-LIO2 (W. Xu, Y. Cai, D. He, J. Lin, and F. Zhang, “FAST-LIO2: Fast Direct LiDAR-Attorney Docket No.: 011520.01958 Inertial Odometry,” IEEE Transactions on Robotics, pp.1–21, 2022. Conference Name: IEEE Transactions on Robotics) as having the highest quality among all candidates despite chamfer distance (CD) and hausdorff distance (HD) indicating LeGO-LOAM (T. Shan and B. Englot, “LeGO-LOAM: Lightweight and Ground-Optimized Lidar Odometry and Mapping on Variable Terrain,” in 2018 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS), (Madrid), pp.4758–4765, IEEE, Oct.2018) and Puma (I. Vizzo, X. Chen, N. Chebrolu, J. Behley, and C. Stachniss, “Poisson Surface Reconstruction for LiDAR Odometry and Mapping,” in 2021 IEEE International Conference on Robotics and Automation (ICRA), (Xi’an, China), pp. 5624–5630, IEEE, May 2021) have the highest quality, respectively, in a visual comparison.

[0012] Figure 4: Left: Ground truth from HILTI SLAM Dataset (L. Zhang, M. Helmberger, L. F. T. Fu, D. Wisth, M. Camurri, D. Scaramuzza, and M. Fallon, “Hilti-Oxford Dataset: A Millimetre-Accurate Benchmark for Simultaneous Localization and Mapping,” 2022. Publisher: arXiv Version Number: 1) using Z+F Imager 5016; Right: FAST-LIO2 using Hesai Pandar XT-32.

[0013] Figure 5: Valid points – Candidate points within ^^ of source point are marked valid while others are marked invalid (artifacts).

[0014] Figure 6: Simulation Worlds used for evaluation. Mai City (left), Village (center), Warehouse (right)

[0015] Figure 7: Ablation study. (a) Original unmodified point cloud; (b) Arbitrary clusters added in error; (c) Clusters of points removed to test completeness; (d) Random points added with gaussian noise (50%); and (e) Points down-sampled to 20%.

[0016] Figure 8: Qualitative evaluation of SLAM systems. The figure shows, from top to bottom, dense maps generated using PUMA, FAST-LIO, LeGo-LOAM, and ground truth. Worlds from left to right are Mai-City, Warehouse, Village, and Exp04 sequence from the HILTI dataset.

[0017] Figure 9: Ablation Study. (a), (b), (c), and (d) depict the impact of adding artifacts, removing points, adding noise, and down-sampling on the proposed quality metrics, individual sub-metrics, CD and HD. The presently disclosed method provides further insight into map quality.Attorney Docket No.: 011520.01958

[0018] Figure 10: Degradations on a sample point cloud (evaluated against original); Left to right: point cloud is, down-sampled, noise is added, cropped and artifacts added. Zoomed-in view below degraded point cloud, and the corresponding Empir3D metric reflects the degradation. (^^^,^^^,^^^,^^௧) is resolution, accuracy, coverage, and artifact score respectively and ^^^and ^^^are Chamfer and Hausdorff distances.

[0019] Figure 11 demonstrates cells of size ^^, green-labeled cells contain both ground truth (green (light gray)) and candidate (blue (dark gray)) points making them covered, red- labeled cells only contain candidate points making them artifacts and gray-labeled with onlyground truth points showing missing coverage (or un-covered). Top row 3rd cell shows ^^ ^ ^^ asthe distance considered to compute accuracy as described herein.

[0020] Figure 12. Ablation Study on street block dataset; Left to right: Down-sampled, Noise Added, Cropped Simulated Artifacts.

[0021] Figure 13. Simulation dataset; Point cloud built using FAST-LIO2 (Top) and LeGO-LOAM (Bottom). Zoomed in view for qualitative assessment.

[0022] Figure 14. Real-world evaluation of Dense SLAM – Point clouds map generated using FAST-LIO2 (Spot robot + Ouster OS-1128 LiDAR) on the left, and ground truth on the right (robotic total-station).

[0023] Figure 15. Evaluation on Davis dataset, zoomed-in view shows variations in detail for different SLAM methods. Top to Bottom: FAST-LIO2, Ground Truth, LeGO-LOAM. Zoomed in view of staircase on the right for qualitative assessment.

[0024] Figure 16. Comparing different SLAM methods.

[0025] Figure 17. Point clouds generated using three completion networks; left to right: Input (partial cloud), ECG, TOP-NET, PCN and ground truth. Qualitative results are corroborated by quantitative evaluation with Empir3D shown in Table VI.

[0026] Figure 18. Poses and point cloud for the Davis dataset

[0027] Figure 19. Warehouse environment in simulation. Robot third-person view (left) and sensor output (right)Attorney Docket No.: 011520.01958

[0028] Figure 20. Evaluation of completed point clouds created when given a partial point cloud.

[0029] Figure 21. Plot shows runtime (Z-axis) of Empir3D, ^^^and ^^^on point clouds of varying resolution (Y-axis) and region sizes (X-axis). Candidate Map: Warehouse dataset with point cloud generated using FAST-LIO2 (Number of Points at 100% = 67,690,672).

[0030] Figure 22. Graphs indicating performance and system load.

[0031] Figure 23. Plot shows change in Empir3D metrics with increase in region size (^^). Resolution and accuracy are affected by averaging while coverage and artifact-score are not.

[0032] Figure 24. Additional warehouse point clouds.

[0033] Figure 25. Top: Anomaly (change) detection using Empir3D. Empir3D allows real-time anomaly detection at > 5 FPS on an Ouster OS-1128. Figure shows anomaly detection on sample dataset, a box is moved and anomalies are highlighted in purple. The anomalies measure out to 0.021 which indicates 2.1% of the scene has changed. Detailed Description of the Disclosure

[0034] With reference to Figure 1, in a first aspect, the present disclosure may be embodied as a computer-implemented method 100 of assessing a candidate point cloud. The method 100 includes receiving 103 a source point cloud and the candidate point cloud (i.e., receiving both point clouds—source point cloud and candidate point cloud). For example, the source point cloud may represent a ground truth.

[0035] Each of the source point cloud and the candidate point cloud are divided 106 into a plurality of cells of equal size. The cell size may be a configurable parameter (e.g., user- selectable, etc.) that may affect the granularity of the quality assessment. In some embodiments, the method 100 may further include dividing 130 each of the source point cloud and the candidate point cloud into a plurality of regions for computational efficiency.

[0036] A number of quality sub-metrics are calculated (“sub-metrics” are also referred to herein as “metrics”). The method 100 includes calculating two or more of a resolution sub- metric, accuracy sub-metric, coverage sub-metric (sometimes referred to herein as “completeness”), and an artifact sub-metric.Attorney Docket No.: 011520.01958

[0037] The method 100 may include calculating 109 a resolution sub-metric. The resolution sub-metric may quantify the spatial sampling quality of the candidate point cloud relative to the source point cloud. For example, in some embodiments, for each cell, the system calculates the average distance between neighboring points within that cell for both the source and candidate point clouds. The resolution sub-metric for each cell may then be calculated as the ratio of the average inter-point distance in the applicable cell of the source point cloud to the average inter-point distance in the applicable cell of the candidate point cloud. In another example, the resolution may be determined based on density (points per volume) of each cell.

[0038] In embodiments wherein the point clouds are divided into regions, the resolution sub-metric may be calculated by computing the average distance between points within each region of the source point cloud and computing the average distance between points within each corresponding region of the candidate point cloud. The resulting average distances across the regions of source point cloud may be averaged, and the resulting average distances across the regions of candidate point cloud may be averaged. A ratio of the average distance in the source point cloud to the average distance in the candidate point cloud is calculated.

[0039] The method 100 may include calculating 112 an accuracy sub-metric. The accuracy sub-metric may measure the geometric precision of the candidate point cloud by evaluating how closely candidate points match the positions of corresponding points in the source point cloud.

[0040] In an example, for each point in the candidate point cloud within a given cell, the system identifies the nearest neighbor point in the corresponding cell of the source point cloud. The distance between each candidate point and its nearest neighbor in the source point cloud is calculated. Points in the candidate point cloud that are farther than the pre-determined distance threshold from any point in the source point cloud may be excluded from the accuracy calculation. The predetermined distance threshold may be a tunable parameter that may be set based on, for example, the expected accuracy requirements of a specific application.

[0041] The method 100 may include calculating 115 a coverage sub-metric (sometimes referred to herein as completeness). The coverage sub-metric may indicate a level of overlap between the candidate point cloud and the source point cloud. In an example, for each cell, the system determines whether the cell is occupied by points from the source point cloud, theAttorney Docket No.: 011520.01958 candidate point cloud, both point clouds, or neither point cloud. A cell may be considered occupied if it contains at least one point from the respective point cloud.

[0042] The coverage sub-metric may calculated as the ratio of the number of cells occupied by points from both the source point cloud and the candidate point cloud, to the number of cells occupied by points from the source point cloud.

[0043] The method 100 may include calculating 118 an artifact sub-metric. The artifact sub-metric may indicate a proportion of anomalous points in the candidate point cloud (e.g., points in the candidate point cloud that do not correspond to features present in the source point cloud, etc.) The artifact sub-metric may be calculated as the ratio of the number of cells occupied by points from the candidate point cloud but not occupied by points from the source point cloud to the total number of cells occupied by points from the candidate point cloud.

[0044] A point quality metric (an overall quality metric) is calculated 121 based on the coverage sub-metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub- metric across all cells (in cases where less than all four sub-metrics are calculated, the point quality metric is based only those sub-metrics calculated). For example, the point quality metric may be the weighted sum of the sub-metrics. The weights allow users to prioritize different quality aspects based on application requirements. For example, navigation applications might prioritize coverage and artifact metrics, while measurement applications might prioritize accuracy and resolution metrics. In some embodiments, certain sub-metrics may be calculated by subtracting from 1.0. For example, the artifact sub-metric may be subtracted from 1.0 because lower values indicate better quality from this sub-metric. Depending on the actual method of calculating each sub-metric, the value(s) may be subtracted from 1.0.

[0045] The method 100 includes outputting 124 the point quality metric as a normalized value. For example, the point quality metric may be a normalized value between 0 and 1, wherein a value of 1 indicates the highest quality and a value of 0 indicates the lowest quality.

[0046] In embodiments wherein the source point cloud and the candidate point cloud are divided, one or more of the resolution sub-metric and the accuracy sub-metric may be calculated for each region. The values may be averaged or otherwise synthesized into an overall value for a particular sub-metric.Attorney Docket No.: 011520.01958

[0047] Referring to Figure 2, in another aspect, the present disclosure may be embodied as a system 10 for assessing a candidate point cloud. The system 10 may include a memory 30 for storing the candidate point cloud and a source point cloud. The system 10 includes a processor 20 in electronic communication with the memory. The processor 20 is configured (e.g., programmed) to perform any of the methods disclosed herein. For example, the processor may be configured to: divide each of the candidate point cloud and the source point cloud into a plurality of cells of equal size; calculate a resolution sub-metric; calculate an accuracy sub- metric; calculate a coverage sub-metric; calculate an artifact sub-metric; and calculate a point quality metric based on the coverage sub-metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub-metric across all cells; and output the point quality metric as a normalized value.

[0048] In some embodiments, the processor is a parallel processor configured to calculate one or more of the coverage sub-metric, the artifact sub-metric, the accuracy sub- metric, and the resolution sub-metric in parallel across the plurality of cells using a multi- threaded implementation. For example, in embodiments wherein the source point cloud and the candidate point cloud are divided into regions, the parallel processor may calculate sub-metrics for each region in parallel with one or more other region sub-metric calculations. Multi-threaded implementations can distribute cell calculations across multiple processor cores or threads. GPU- based implementations can leverage the massively parallel architecture of graphics processors to simultaneously process hundreds or thousands of cells. The parallel processing capability enables real-time quality assessment for applications such as live mapping system monitoring or autonomous vehicle operation.

[0049] In another aspect, the present disclosure may be embodied as a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method for assessing the quality of a point cloud. The method may be any of the methods disclosed herein. ADDITIONAL DISCUSSION – First Example

[0050] I. INTRODUCTION

[0051] Dense maps play a crucial role in numerous applications including but not limited to autonomous driving, search and rescue, service robotics, and augmented reality. Among theAttorney Docket No.: 011520.01958 different ways to build dense maps, dense Simultaneous Localization and Mapping (SLAM) is particularly interesting due to its ability to produce high-fidelity maps as point clouds that are used for tasks such as localization, re-localization, place recognition, and cross-robot localization but also a need for realtime execution requiring various tradeoffs in map quality. The point clouds produced by these algorithms have been proposed for use in advanced monitoring, sophisticated manipulation, augmented reality, and fine-grained control.

[0052] LiDAR-based SLAM methods are increasingly popular thanks to advances in technology. The affordability and improved accuracy of LiDAR sensors now allow for the real- time creation of high-quality, dense point clouds. Figure 4 provides a comparison between point clouds generated using an engineering-grade LiDAR, specifically the Z+F Imager 5016 from the HILTI SLAM dataset (on the left), and an inexpensive Hesai Pandar XT-32 using FAST-LIO2, an online direct registration-based SLAM method (on the right). It is clear from the visual comparison that FAST-LIO2 produces point clouds that are nearly as high-quality as those generated by the engineering-grade LiDAR, which is generally very expensive.

[0053] The evaluation of point clouds poses a significant challenge due to the complexity of capturing all aspects of map- ping accuracy. While some measures, such as the Absolute Trajectory Error (ATE), can assess the difference between expected and measured translation and rotation, they do not fully account for map quality and are only indirect measures. Given a reference map, it is important to identify how much of the reference map is captured by a mapping method, how close the created map is to the reference, whether the method created anamolies that do not exist in the reference (we call these artifacts) and if the density of the resultant point cloud is similar to the reference or sparser.

[0054] Popular methods for comparing point clouds and meshes such as Chamfer distance (CD), Hausdorff distance (HD), and Earth Mover’s distance (EMD), have limitations. CD is insensitive to point density and significantly influenced by outliers. Therefore, it serves as a poor performance metric to characterize point cloud completeness or map artifacts. On the other hand, while EMD can detect changes in density, the requirement for a one-to-one correspondence between compared maps is usually too strict and can lead to ignoring local fine- grained structural details. Additionally, EMD is significantly more computationally expensive than CD, which can limit its practical applications. Overall, neither CD nor EMD is ideally suitable for evaluating the quality of generated shapes, as they may fail to capture coverage orAttorney Docket No.: 011520.01958 completeness, structural information, and local variations in error. Therefore, an ideal evaluation method should be efficient and accurately reflect the presence of artifacts or missing data while considering all factors affecting the quality of point clouds for the above applications. A good metric should have the following: ^ The metric should capture coverage and completeness of the point cloud, as well as structural information, to provide a comprehensive evaluation of quality. ^ The metric should be computationally efficient and handle large datasets to be useful for practical applications ^ The metric should accurately penalize artifacts while rewarding higher density and resolution.

[0055] To address these challenges, the present disclosure provides a novel point quality evaluation metric (PQM) that provides a comprehensive and thorough assessment of point cloud quality. PQM comprises four sub-metrics, each evaluating a different aspect of point clouds’ quality: ^ Completeness (Coverage): Measures the proportion of missing data in a point cloud map. It is advantageous for applications such as autonomous driving and robotics, where having a complete point cloud is essential for ensuring safety. ^ Artifact Score: Measures the proportion of non-existent artifacts added in error. It is useful in detecting the impact of artifacts on visual fidelity, especially in augmented reality and virtual environment creation. ^ Accuracy: Measures how close the points are to their true positions. It is advantageous for infrastructure inspection and manufacturing quality control, where registration accuracy plays a crucial role. ^ Resolution: Measures the density of the point cloud map. It is an indicator of how detailed the map is and can enhance fine-grained manipulation and object recognition precision.Attorney Docket No.: 011520.01958

[0056] PQM provides a comprehensive evaluation of point cloud quality by addressing various aspects of Li-DAR maps, making it a valuable tool for several applications. The following disclosure provides: ^ PQM for evaluating point cloud maps. ^ Provide an efficient multi-threaded implementation. ^ Evaluate the metric in simulation over three maps and 3 SLAM systems ^ Perform an ablation study on the effect of mapping errors

[0057] II. RELATED WORK

[0058] SLAM systems can produce point cloud maps with varying levels of density and fidelity using different sensors. Visual SLAM systems that use, for example, monocular cameras, stereo cameras, and RGBD cameras to produce dense or sparse point clouds have been proposed. Recent advances in sensor technology, efficient libraries, and faster computing have enabled realtime LiDAR mapping. LiDAR SLAM systems continue to improve, both in localization and mapping performance.

[0059] There is a growing interest in dense 3D mapping using LiDAR SLAM. LOAM, LeGO-LOAM, LIOSAM, LVI-SAM, LINS, and FAST-LIO2 generate relatively dense point clouds using online localization and mapping. Methods like Puma and SHINE output meshes by performing offline mapping and localization either solely with sequential LiDAR scans or with additional odometry information.

[0060] SLAM systems are generally evaluated for their localization and re-localization performance with the Absolute Trajectory Error (ATE) with changes in environmental factors such as illumination. Although ATE is a good measure of a SLAM system’s localization performance, it is a poor measure of map quality. In some cases, ATE can be used to evaluate the overall structure of the map, not density and completeness. For example, ORB-SLAM is known for good localization and tracking performance even though it produces sparse point cloud maps. In other cases, trajectory error may not be a sufficient metric to evaluate mapping performance. A previous WiFi-based distributed mapping system cannot be evaluated with the ATE because a ground truth trajectory is hard to obtain in a distributed mapping scenario. Thus, knownAttorney Docket No.: 011520.01958 landmark (April tag) positions were used to evaluate their system indicating a need for a metric to evaluate the map quality directly.

[0061] Reconstruction error calculated over point clouds is a direct approach to evaluating map accuracy. Chamfer distance (CD), Hausdorff distance (HD), and Earth Mover’s distance (EMD) are popular distance metrics used in computing reconstruction error in point clouds. However, there are limitations in using distance-based reconstruction error as a performance metric to measure map quality.

[0062] For example, CD (Eq.1) is computed as the sum of distances in two point clouds, usually referred to as source and candidate. For each point in the source, the distance to its nearest neighbor in the candidate point cloud is computed and vice versa. The sum of distances over both point clouds is the CD. It is fast to compute, and it can capture the overall similarity between two point clouds. However, it does not account for the local variations and structural information in the point clouds, which can be important in some applications. Secondly, it is insensitive to density distribution. Finally, it is significantly influenced by outliers.

[0063] As distance between two points in the source and candidate point clouds. This means that for each point in one point cloud, the distance to the farthest point in the other point cloud is calculated, and the maximum of all such distances is the HD. It captures the similarity between two point clouds, including their overall arrangement. However, it is computationally expensive to compute and is not as efficient as CD.

[0064] Dense point clouds generated by some SLAM systems like FAST-LIO2 rival that of survey and engineering grade LiDARs as seen in (Figure 4). This means they can be used for applications that need high-resolution point clouds such as GIS analysis, infrastructure inspection, 3D reconstruction, and object detection to name a few. With hardware and algorithmic advancements that produce such detailed point clouds, there is a need for a way to measure the difference in the quality of these point clouds. Popular metrics like CD, HD, andAttorney Docket No.: 011520.01958 EMD have difficulty in capturing coverage, completeness, structural information, local variations in error, and are computationally expensive. Additionally, these methods do not account for artifacts or missing data and do not evaluate the components of quality discretely. Although other methods exist, they focus on specific applications like visual quality and point cloud generation acting as a loss function for neural network training. One method provides a way to measure the accuracy and completeness of meshes generated by multi-view stereo reconstruction but doesn’t not account for resolution and artifacts. Another method complains about the lack of ground truth to evaluate point clouds. We address this by using simulated datasets where ground truth from the simulation environment is available in the form of meshes. We then sample these meshes to acquire ground truth point clouds.

[0065] Therefore, the present disclosure provides the Point Quality Metric (PQM) which addresses some limitations of existing metrics, namely: (i) capturing a notion of map coverage, (ii) penalizing non-existent artifacts, (iii) measuring accuracy and, (iv) rewarding higher density and resolution. The presently disclosed technique provides a framework for the comprehensive evaluation of two point clouds based on their completeness, accuracy, artifacts, and resolutions.

[0066] III. METHOD

[0067] This section describes the presently disclosed metric PQM and the evaluation framework including an ablation study to independently measure the effectiveness of each sub- metric. As mentioned above, point clouds generated by LiDAR SLAM methods, although dense, can be inaccurate and incomplete due to the path taken by the robot and registration errors. Further, these point clouds can contain artifacts (anomalies or points not present in ground truth) that degrade the overall quality.

[0068] A. Point Quality Metric (PQM)

[0069] We denote the source (ground truth) point cloud by ^^ ൌ ^^^, which we refer to as^^^^^^^. Similarly, we denote the candidate point cloud by ^^ ൌ ^^^, referred as ^^^^^^^, where ^^^ and^^^ are in ^^ௗ and ^^ ൌ 1, ... ,^^. An objective is to measure the difference in quality between thepoint cloud ^^^^^^^and the candidate point cloud ^^^^^^^. Quality, as defined in Sec.III, is a weighted combination of the four sub-metrics: completeness, artifact score, accuracy, and resolution. Each sub-metric contributes to the overall quality of the point cloud, and evaluating them independently enables us to assess the effect of each sub-metric on the overall quality.Attorney Docket No.: 011520.01958

[0070] ^^^ொெdenotes the overall quality given by eq.7 while ^^^, ^^௧, ^^^and ^^^denote the individual sub-metrics completeness (sometimes referred to as coverage), artifact score, accuracy, and resolution respectively.

[0071] The point clouds are divided into smaller portions of equal size or “cells,” with each cell having a size of ^^. This enables comparison in parallel and provides insights into the quality of different areas within the point cloud. The point cloud is split into ^^ such cells, andsub-metrics are computed for each cell. Cells are denoted as ^^^^^^^^^ೕ ∈ ^^^^^^^ and ^^^^^^^^^ೕ ∈ ^^^^^^^,where ^^ ൌ 1, ... ,^^. PQM may be normalized between 0 and 1, where 1 represents the bestquality and 0 represents the worst quality. In contrast, geometric distance metrics such as CD, HD, and EMD are typically calculated such that a score of 0 represents a perfect match, and any value greater than 0 represents a degree of mismatch.

[0072] 1) Resolution: Resolution per cell may be considered as the ratio of the density (pts / volume) of ^^^^^^^^^, a cell in ^^^^^^^to density (pts / volume) of ^^^^^^^^^, a cell in ^^^^^^^given in Eq.3. Overall resolution (^^^) is the mean of ^^^over ^^ cells. Resolution determines the level of detail in the point cloud. Low resolution can cause loss of texture and smaller objects making the point cloud unusable for applications that require high fidelity and detail. ^^ ^^ ^^^^^ ൌ ^ ಳ^^^ (3)

[0073] 2) Accuracy:between every point in ^^^^^^^^^, a cell in ^^^^^^^to the nearest neighbor in ^^^^^^^^^, a cell in ^^^^^^^given distance is less than threshold ^^, to the product of the number of points in ^^ and ^^, given in(eq. 4). The normalization may be performed over (|^^| ൈ ^^) as this is the maximum distancepossible if all points in ^^^^^^^^^are valid (i.e., have neighbors within ^^ distance in ^^^^^^^^^). Overall accuracy (^^^) is the mean of ^^^over all ^^ cells. 1 ^^^(4)where,Attorney Docket No.: 011520.01958 ^^ ^^,^^ ൌ ^^∈ m^^in^^‖^^ െ ^^‖ଶ , if min ‖^^ െ ^^‖ଶ ^ ^^^ ^ ಲ ^∈^^^^ಲ0 , otherwise

[0074] the ratio of valid points (Figure 5) of ^^^^^^^^^(i.e., points within ^^ distance of ^^^^^^^^^) to the total points in ^^^^^^^^^given by (eq.5). This provides a measure of how complete a given cell is compared to the ground truth and can be used to estimate missing areas in the candidate point cloud. Overall completeness is given by ^^^, the mean of all (^^^) over ^^ cells. ฬ^^^^ ∈ ^^^^^^^^^:^ m∈^^in ‖^^ െ ^^‖ଶ ^ ^^ൠฬ^^ ^^ಲ^ ^(5)

[0075] 4)not in ^^^^^^^^^. These are generated due to reflections, distortion, or misregistration of points. Artifact score is the ratio of valid points of ^^^^^^^^^(i.e., points within ^^ distance of ^^^^^^^^^) to the total points in ^^^^^^^^^given by (eq.6). Similar to III-A.3, the overall artifact score is given by (^^௧). ฬ^^^^ ∈ ^^^^^^^^^:^∈ m^^in^^ ‖ ‖ಲ ^^ െ ^^ ଶ ^ ^^ൠฬ^ (6)

[0076] Overall Map Quality:

[0077] PQM may be computed as the mean of the weighted sum of ^^^, ^^^, ^^^, and ^^௧overall ^^ cells. The weights ^^^^,^^^,^^^,^^௧^ correspond to each of these sub-metrics, respectively. For all experiments in thiswe equally weight each submetric(^^^^ ,^^^,^^^ ,^^௧^ ൌ 0.25) to ensure that they contribute equally to the overall quality score.However, these weights can be adjusted to meet the specific requirements of a given application. ே 1 ^^ ^^^ ^ ^ ^ ^ ^ ^ ^^^(1)Attorney Docket No.: 011520.01958

[0078] B. Controlled Ablation Study

[0079] To evaluate PQM, we performed an ablation study using a prototype point cloud (Stanford Bunny) model. The study involved applying various degradation to ^^^^^^^, the source model, and using PQM, CD, and HD to evaluate the quality at each step.

[0080] 1) Artifacts: We added points to ^^^^^^^to simulate artifacts (^^^^^^^is a copy of ^^^^^^^with added artifacts), which can be caused by sensor noise or registration errors. A set percent (^^) of points were added to each cell from a uniformly sampled sphere artifact 1 / 10ththe size of the cell, placed at the center of the cell. This also affects resolution since the number of points in ^^^^^^^increases.

[0081] 2) Completeness: A patch of points was removed from ^^^^^^^to simulate incompleteness, which can be caused due to inconsistent mapping, down-sampling, and / or sensor noise. A set percent (^^) of nearest neighbors of a randomly selected point were removed per cell.

[0082] 3) Accuracy: Gaussian noise was added to ^^^^^^^to simulate loss of accuracy. Gaussian noise was applied to the candidate point cloud with 0 ^^ and a finite ^^ value, where ^^ is the variance applied to the points in a random normal direction.

[0083] 4) Resolution: Uniform down-sampling was applied to ^^^^^^^to simulate a reduction in resolution and reduce the complexity of the point cloud while preserving its overall structure. This was achieved by sampling every ^^thpoint in the current cell, where ^^ is the control parameter.

[0084] Figure 7 shows the bunny model with 50% degradation (for artifacts, completeness, and accuracy) and a sampling rate of 5 (for resolution) where ^^^^^^^(green (or light gray)) is the source point cloud and ^^^^^^^(red (or dark gray)) are candidates after degradation.

[0085] C. Map Evaluation Framework

[0086] To further study PQM, we collected LiDAR scans in purpose-built simulation environments (Figure 6) and built point cloud maps using several popular LiDAR SLAM systems. The simulation environments were built using the Gazebo simulator and mesh modelsAttorney Docket No.: 011520.01958 of the worlds. We simulated an Ouster OS1-128 LiDAR on a Clearpath Husky platform to collect the LiDAR scans. The simulation worlds included the outdoor world provided by the Mai-City dataset and two other worlds called “Village” and “Warehouse” as seen in Figure 6. The environments were designed to represent city blocks and warehouses, complete with buildings, trees, storage pallets, and other elements commonly found in the real world. The use of simulation allows for more controlled testing conditions, ground truth maps, and the generation of visual data from multiple views for testing and evaluation.

[0087] The candidate point clouds were generated using Lego-LOAM, FAST-LIO2, and Puma. LeGO-LOAM stands for ground-optimized LiDAR odometry and mapping. The system outputs a dense point cloud and real-time odometry. FAST-LIO2 employs a tightly-coupled LiDAR+IMU (inertial measurement unit (IMU)) method to generate dense point clouds in real time. This is achieved through the direct registration of scans with minimal downsampling facilitated by the use of an iKDtree for fast point-wise and block-wise operations. Puma employs a unique approach to mesh generation by performing frame-to-mesh registration and Poisson Surface Reconstruction to generate a mesh. While this process is not real-time, the resulting meshes are lightweight and highly representative of the real world.

[0088] Finally, the point clouds generated by these methods (candidates) in the simulated worlds (Figure 8) and their respective ground truths (source) were used to evaluate PQM. We also measured CD and HD between the candidate and source point clouds. The next section shows the comparative results.

[0089] IV. EXPERIMENTAL EVALUATION

[0090] In the study, PQM was implemented using Open3D and PDAL libraries and ran entirely on a desktop workstation with 6 core / 12 thread CPU and 32 GB of Memory. We also provided a CUDA accelerated implementation with PyTorch, which can be advantageous for large point clouds.

[0091] In some tests, experimental evaluation was performed on simulation datasets including, but not limited to, the HILTI dataset, and the Stanford bunny. Candidate point clouds were evaluated against the source point cloud sampled from meshes, which, as mentioned earlier, were used in the simulation worlds. Meshes were sampled into point clouds where the number of points was the maximum of all point clouds generated by the three SLAM methods.Attorney Docket No.: 011520.01958

[0092] A. Ablation Study

[0093] To evaluate the performance of PQM in isolation we degraded the source point cloud as described in Section III-B. Experiments were performed for a range of degradation (0% to 90% with an increment of 5% and sampling rate of 1 to 19) with varying cell sizes i.e., 0.05, 0.04, 0.03, 0.02 (meter). Figure 9 shows results with cell size 0.05 (dividing the model into 4×4×3 cells). ^^ was maintained as half the average distance between points in ^^^^^^^. ^^ is a tuneable parameter and can be chosen based on application. All trials were evaluated withweights ^^^^ ,^^௧,^^^,^^^^ ൌ 0.25, and ^^ ൌ 0.0002.

[0094] 1) Artifact Score: Table I and Figure 9 show adding artifacts proportionally decreased ^^௧, while ^^^,^^^were unchanged. ^^^also shows change due to the spherical artifact as mentioned in Section III-B.1.

[0095] 2) Completeness: Table I and Figure 9 show completeness ^^^decreased as points were removed from ^^^^^^^, while ^^^,^^௧were unchanged. ^^^also shows change due to the decrease in total points as mentioned in Section III-B.2.

[0096] 3) Accuracy: Table I and Figure 9 show decrease in ^^^and in-turn in ^^^ொெ. ^^ forgaussian noise was constrained to ^^ ൌ 0.0002 to ensure ^^^ and ^^௧ were not affected. Thenegligible change in other sub-metrics can be due to the overflow of boundary points. This demonstrates the efficacy of using PQM where deviations in accuracy might be small.

[0097] 4) Resolution: Table I shows a decrease in resolution with an increase in sampling rate. This is unlike IV-A.2 as removing a continuous patch of points affected completeness more than resolution, while down-sampling degraded ^^^and ^^^almost equally, if not exactly as seen in Figure 9.

[0098] 5) Chamfer and Hausdorff distance: Similarly, Figure 9, shows the change in CD and HD for each degradation. CD (light) shows a gradual increase, which signifies an increasing degradation in quality, but it fails to give insights about the type of quality that is affected. Further, its value is unbounded and hence cannot justify normalized quality value for a reasonable comparison. In the case of HD (dark), values do not show any correlation with the applied degradation.Attorney Docket No.: 011520.01958

[0099] Results of these experiments show that PQM is effective in detecting and quantifying changes in map quality due to different types and levels of degradation. The scores generated by PQM correlate with the expected change in resulting value for each targeted sub- metric.

[0100] Overall, the experiment demonstrates the effectiveness of the PQM metric in evaluating the quality of point clouds and detecting changes in map quality due to different types and levels of degradation and hence can act as a tool to detect point cloud suitability for specific applications.

[0101] B. Evaluation on SLAM systems

[0102] The performance of PQM on large, realistic point cloud maps generated using the SLAM systems mentioned above was evaluated exhaustively in three simulated worlds and the exp04 sequence of the HILTI dataset. The results of this evaluation are presented in Figure 8, which shows the qualitative differences between the maps generated by the SLAM systems and the ground truth. Table II shows the sub-metrics for each map. Each sub-metric evaluates a certain aspect of quality as defined in Sec. III. PQM was computed as the weighted sum of these sub-metrics, where the weights (Eq.7) were set by users based on application. For all evaluations, the weights were set to 0.25, making the contribution of each sub-metric equal foroverall quality. We also set ^^ ൌ 0.1 and cell size ^^ ൌ 10 for all tests in Table II. More suitablevalues were not explored for all maps due to long computing times. A study of the effect of ^^ and ^^ is shown in Table III.Attorney Docket No.: 011520.01958

[0103] Visua y, s appa e a e - ge e a es e ghest-quality maps compared to other methods in terms of resolution, completeness, and accuracy. However, the highlighted CD and HD distance values in Table II do not reflect this observation. We can see that CD and HD distances only correctly identified the best map 1 / 4 and 0 / 4 times, respectively. On the other hand, PQM consistently identified point clouds generated by FAST-LIO2 as the ones with the highest overall quality. The sub-metrics are also indicative of PQM’s performance, with only ^^^being incorrect in 2 instances (HILTI and Warehouse). This is likely due to thethreshold of ^^ ൌ 0.1 used. However, a lower ^^ value correctly identified ^^^, as seen inTable III. Highlighted values show that as ^^ was reduced (^^ ൌ 0.01), quality was calculatedcorrectly 2 / 3 times, which is much greater.

[0104] V. DISCUSSION

[0105] Need for reference point clouds: For the scope of the present method, we assume that candidate maps are compared against ground truth point clouds or meshes, or similar source point clouds. While acquiring the ground truth maps can be difficult in real-worldAttorney Docket No.: 011520.01958 scenarios, we make the following observations: (i) To effectively quantify the mapping performance of a SLAM system a ground truth or reference is advantageous; (ii) Ground truth can be generated using an engineering grade LiDAR which can generate centimeter-accurate if not millimeter-accurate point clouds (Figure 4); and (iii) Alternatively, we can evaluate in simulation where ground truth is readily available. Recent advancements in high fidelity simulators make this a lucrative option. Our evaluation framework used three simulation scenarios as shown in Figure 4 (Mai City, Village, Warehouse).

[0106] Customizability: While we provide one holistic PQM metric, the intent is to expose the various dimensions of importance as highlighted by the sub-metrics—completeness, artifact score, accuracy, and resolution. The user may customize the weights and the various parameters to better suit their application. This will allow better quantification of map quality. On those lines, the introduction of cell size ^^ and threshold ^^ enables the user to tune PQM to a particular use case. The smaller the ^^ the lower the tolerance is for accuracy, completeness, and artifact score. The cell size helps reduce the computational complexity for large point clouds, metrics can be computed in parallel, and as cell size is decreased the resolution for local variations increases. ADDITIONAL DISCUSSION – Second Example

[0107] 1. Introduction

[0108] 3D perception is crucial for various applications including autonomous driving (Caesar et al.2020; Maddern et al.2017), infrastructure inspection (Jun Wang et al. 2014; Valença et al.2017), augmented reality (Mahmood, Han, and Lee 2020; Chen et al.2019), mobile manipulation (Rusu et al.2009; Seita et al.2023), 3D reconstruction (Yi et al.2017; Brostow et al.2008), object detection (C. Xu et al.2020a; Y. Zhou and Tuzel 2017), and GIS applications (Xie, Tian, and Zhu 2019; Westoby et al.2012). Each of these applications uses a 3D point cloud as input. Such 3D point clouds can be produced by various methods such as dense SLAM (W. Xu et al.2022a; Rosinol et al.2020; Mathieu Labbé and Michaud 2019), structure-from-motion (Photogrammetry) (Smith, Carrivick, and Quincey 2015; Westoby et al. 2012; Schonberger and Frahm 2016), survey grade scanners (“Robotic Total Stations — Leica- Geosystems.com”) and generative / learning-based methods (Yuan et al.2018; Tchapmi et al. 2019; Pan 2020; Zhao et al.2021; H. Zhou et al.2022; Fan, Su, and Guibas 2017;).Attorney Docket No.: 011520.01958

[0109] A logical question to be asked is how good a constructed 3D point cloud is in comparison to the real-world scene it represents and / or for the intended application.3D point clouds are typically evaluated using distance-based similarity metrics comparing the constructed 3D point cloud with a reference. Such metrics quantify similarity of the unordered set of points in the constructed point cloud with the ones in the reference. Prevalent methods are Chamfer distance (^^^), Hausdorff distance (^^^) (Huttenlocher and Kedem 1990), and earth-mover’s distance (^^^^) (Fredman and Tarjan 1987). Chamfer distance is the sum of squared distances between closest point pairs in two shapes. Earth-mover’s distance is the sum of distances between closest point pairs where pairing is bijective and Hausdorff distance is the greatest of distances between closest point pairs. Though popular, each of these measures are uni- dimensional and have their own limitations in comparing two point clouds. ^^^and ^^^have limited sensitivity to point density and are significantly influenced by outliers (T. Wu et al. 2021). While ^^^^can detect changes in density, the bijectivity requirement can lead to ignoring local fine-grained structural details. Also, ^^^^is significantly more computationally expensive, which can limit its practicality.

[0110] Some other relevant methods (Javaheri et al.2022; T. Wu et al.2021; Seitz et al. 2006; Sinha and Fleuret 2023) focus on specific applications like visual quality and point cloud generation acting as a loss function for neural network training. (Seitz et al.2006) provides a way to measure the accuracy and coverage of meshes generated by multi-view stereo reconstruction.

[0111] Evaluating 3D point clouds is challenging and based on the present disclosure, is benefited by quantification of multiple factors for a comprehensive comparison. Below is a list of factors which may be beneficial: ^ It should reward a constructed point cloud if its density is high (resolution), it matches points from the reference (accuracy), and it is able to capture most of the reference points spatially (coverage) ^ It should penalize the constructed point cloud if it has points in areas where the reference doesn’t (artifacts) ^ It should be computationally efficient to process large point cloudsAttorney Docket No.: 011520.01958

[0112] To address these issues, the present disclosure provides ^^^^^^^^^^3^^, an Evaluation Methodology for PoIntcloud Reasoning in 3D. ^^^^^^^^^^3^^ comprises of four metrics, each evaluating a specific aspect of point cloud quality:

[0113] Resolution: Resolution is the ability to resolve areas in a point cloud. It is an indicator of how detailed the point cloud is.

[0114] Accuracy: Measures how close the points are to their true positions.

[0115] Coverage: Measures the areas of the reference that the constructed point cloud covers. Larger the score, larger the overlap between the two point clouds.

[0116] Artifact Score: Measures the proportion of anomalous points (artifacts) added in error in the constructed point cloud.

[0117] To demonstrate these metrics visually, Figure 10 provides a comparison between point clouds with some variations (down-sampling, adding noise, removing some portions, and adding artifacts). We show ^^^^^^^^^^3^^ metrics and two other popular point cloud comparison metrics - Chamfer distance (^^^) (Achlioptas et al.2018) and Hausdorff distance (^^^) (Huttenlocher and Kedem 1990).

[0118] Design of the ^^^^^^^^^^3^^ metrics took multiple iterations striving to maximize two aspects—comprehensive evaluation of point cloud quality while providing independence across metrics to limit the number of metrics we used to capture the comparison. As a result, ^^^^^^^^^^3^^ provides a detailed evaluation of point cloud quality by addressing various aspects of point clouds, making it a valuable tool for several applications. The contributions of this paper are as follows: ^ ^^^^^^^^^^3^^, a comprehensive framework for evaluating constructed point clouds with a reference. ^ We evaluate the framework on a real-world, simulated and generated point clouds demonstrating the utility of each metric for real-world applications. ^ Perform an ablation study to show changes in metrics caused by a change to point clouds.Attorney Docket No.: 011520.01958

[0119] 2. Related Work

[0120] New-age sensors like high-resolution cameras, LiDARs, and RADARs can produce rich, dense point clouds. Visual SLAM systems that use monocular cameras (Mur-Artal and Tardós 2017), stereo cameras (M. Labbé and Michaud 2014), and RGB-D cameras (F. Endres et al.2014) to produce dense (F. Endres et al.2014; M. Labbé and Michaud 2014) or sparse (Mur-Artal and Tardós 2017) point clouds have been proposed. Photogrammetry tools like (Schönberger et al.2016; Schönberger and Frahm 2016; Moulon et al.2016) have enabled quick and easy access to reconstructing 3D environments for applications in game development, augmented reality and geospatial surveying.

[0121] Point clouds from 3D reconstruction

[0122] Recent advances in sensor technology, efficient libraries (Agarwal, Mierle, and Team 2022; Rusu and Cousins 2011), Neural-network architectures (Qi et al.2017a) and faster computing have enabled real-time dense mapping. Subsequently, Vision and LiDAR based SLAM systems (J. Zhang and Singh 2014; Rosinol et al.2020) have followed suit, especially in dense mapping performance. (J. Zhang and Singh 2014; Shan and Englot 2018; Shan, Englot, et al.2020; Shan et al.2021; Qin et al.2020; W. Xu et al.2022b; He et al.2023) generate relatively dense point clouds using online localization and mapping. (Vizzo et al.2021; X. Zhong et al. 2023) output meshes by performing offline mapping and localization either solely with sequential LiDAR scans or with additional position information.

[0123] SLAM methods are generally evaluated for their localization and re-localization performance with the Absolute Trajectory Error (ATE) as seen in (Bujanca et al.2021; Sturm et al.2012; Felix Endres et al.2012) with changes in environmental factors such as illumination. Although ATE is a good measure of a SLAM system’s localization performance, it is a poor measure of map quality. In some cases, ATE can be used to evaluate the overall structure of the map, not density and completeness. For example, ORB-SLAM (Mur-Artal and Tardós 2017) is known for good localization and tracking performance, even though it produces sparse point cloud maps. In (Adhivarahan and Dantu 2019), the authors use a WiFi-based distributed mapping system which cannot be evaluated with the ATE since a ground truth trajectory is hard to obtain in a distributed mapping scenario. Thus, the authors use known landmark (AprilTag (John Wang and Olson 2016)) positions to evaluate their system, indicating a need for a metric to evaluate the map quality directly. (Felix Endres et al.2012; Sturm et al.2012)Attorney Docket No.: 011520.01958 highlights the lack of ground truth to evaluate point clouds; we address this by using simulated datasets as well as capturing ground truth using poses measured using a robotic total station and stitching corresponding LiDAR scans.

[0124] Point clouds in Learning

[0125] In addition to dense mapping and 3D reconstruction, contemporary learning methods and networks have enabled applications like object detection (Qi et al.2019; Y. Zhou and Tuzel 2018; Shi, Wang, and Li 2019; B. Yang, Luo, and Urtasun 2018; Lang et al.2019), segmentation (Qi et al.2017b; Zhao et al.2021; Lai et al.2022; C. Xu et al.2020b) point cloud completion (Yuan et al.2018; Tchapmi et al.2019; Pan 2020; Pan et al.2021; Huang et al.2020) , super-resolution (Dinesh, Cheung, and Bajić 2019; Guan et al.2020; Ledig et al.2017; Shan, Wang, et al.2020; H. Wu, Zhang, and Huang 2019; Zyrianov, Zhu, and Wang 2022) , image-to- point cloud generation(Fan, Su, and Guibas 2017), image-to-mesh generation (Wang_2018_ECCV), denoising and compression. These learning methods are usually evaluated using popular distance-based metrics and in some cases introduce non-standard metrics to capture subjective perceptual quality, which makes bench-marking a challenging task and highlights the need for a better metric.

[0126] Point clouds in Multimedia and AR / VR

[0127] Applications such as augmented and virtual reality, social media avatars, and game development have greatly benefited from the recent rise in point cloud acquisition and processing techniques (Bonatto et al.2016). These multimedia and AR / VR methods are generally evaluated for their perceptual quality and compression losses. (Q. Yang, Ma, et al. 2022; Q. Yang et al.2021; Q. Yang, Liu, et al.2022; Liu et al.2022; Diniz, Garcia Freitas, and Farias 2021; Meynet et al.2020; Viola, Subramanyam, and Cesar 2020) propose no-reference and full-reference ways to evaluate the perceptual quality of point clouds with a focus on visual fidelity and predicting subjective quality. While these methods do a good job at evaluating the perceptual quality of the point clouds, they do not account for all the geometric aspects of quality.Attorney Docket No.: 011520.01958

[0128] Popular distance-based metrics for point cloud comparison

[0129] While distance-based metrics can be convenient and relatively fast to compute, they are uni-dimensional and only output a single measure of similarity which cannot account of all aspects of quality. They also fail to reward and penalize candidate point clouds based on specific dimensions of quality making them difficult to be used as a reasonable feedback signal for improvement in learning methods and SLAM systems.

[0130] Chamfer Distance (^^^^)

[0131] ^^^is computed as the sum of distances in two point clouds, usually referred to as source and candidate. For each point in the source, the distance to its nearest neighbor in the candidate point cloud is computed and vice versa. The sum of distances over both point clouds is the ^^^. It is fast to compute, and it can capture the overall similarity between two point clouds. However, it does not account for the local variations and structural information in the point clouds, which can be important in some applications. Secondly, it is insensitive to density distribution and significantly influenced by outliers. While being influenced by outliers can provide insights into point cloud similarity, it can cause over-penalization where the candidate point cloud is unjustly penalized even if it has good coverage and accuracy. ^^^^^^,^^^ ൌ ^ m^∈i^n ∥ ^^ െ ^^ ∥ଶ^ ^ m^∈i^n ∥ ^^ െ ^^ ∥ଶ (1)

[0132]

[0133] ^^^is calculated as the maximum distance between two points in the source and candidate point clouds. This means that for each point in one point cloud, the distance to the farthest point in the other point cloud is calculated, and the maximum of all such distances is the ^^^. It captures the similarity between two point clouds, including their overall arrangement. However, it fails to capture any local variations and density of the point clouds. ^^^^^^,^^^ ൌ max ൬ sup^^^^^^, ^^^,  sup^^^^^^,^^^ ^ (2)Attorney Docket No.: 011520.01958

[0134] Earth-mover’s Distance (^^^^^^)

[0135] ^^^^solves the optimal-transport problem, also known as the assignment or correspondence problem, by finding a bijective mapping between the two point sets. It is known that the optimal bijection is unique and is invariant to infinitesimal movement (Fan, Su, and Guibas 2016). While this makes ^^^^one of the most precise measures of distance-basedsimilarity, its ∼ ^^^^^ଶlog^^^ (Fredman and Tarjan 1987) complexity makes it impractical forlarge point clouds. Additionally, the bijectivity requirement is not realistic when point clouds are in the 10^points range. ^^^^^^^,^^^ ൌథ m:^i→n^ ^∥^^ െ ^^^^^^ ∥ଶ (3)where, ^^:^^ → ^^ is a

[0136] 3. ^^^^^^^^^^^^^^ Framework

[0137] The presently disclosed ^^^^^^^^^^3^^ framework provides a multi-dimensional comparison between two point clouds. This could be a ground truth point cloud and a captured orgenerated point cloud. We denote the source (ground truth) point cloud by ^^ ൌ ^^^, which werefer to as ^^^^^^^. Similarly, we denote the candidate point cloud by ^^ ൌ ^^^, referred as ^^^^^^^,where ^^^ and ^^^ are in ^^ଷ and ^^ ൌ 1, … ,^^. An objective is to measure the difference in qualitybetween the source (^^^^^^^) and the candidate (^^^^^^^).

[0138] The advantageous feature of a high-quality point cloud is to represent the inherent continuous structure of a 3D environment or an object as best as possible. This is a challenging task because most sensors produce discrete outputs. Generative networks and mapping algorithms therefore either produce results in the form of point clouds or use interpolation to produce continuous meshes. The discrete nature of sensors, the error in their measurements, and the ability of the algorithm to handle these errors lead to variations in point cloud quality. Errors may also occur due to faulty depth estimation, low sample size that affects interpolation accuracy, poor generalization of neural-networks, and misplaced points due to errors in pose information when using SLAM.

[0139] We therefore define quality as a composition of four metrics: resolution, accuracy, coverage, and artifact-score. Each metric contributes to the overall quality of the pointAttorney Docket No.: 011520.01958 cloud and evaluating them independently enables us to assess the effect of each metric on the overall quality. ^^^, ^^^, ^^^and ^^௧denote the individual sub-metrics resolution, accuracy, coverage, artifact-score respectively.

[0140] Region Splitting: To efficiently evaluate the point clouds, each point cloud may be divided into smaller regions of equal size (^^). This enables parallel computation and provides insights into the values of the metrics of different areas within the point cloud. Point clouds may be split into ^^ such regions, and metrics are computed for each. Regions are denoted as ^^^^^^^ೕ∈^^^^^^^ and ^^^^^^^ೕ ∈ ^^^^^^^, where ^^ ൌ 1 …^^.

[0141] Independent of regions, we define cells as volumes of size ^^, where ^^ is a hyper-parameter set by the user based on expected precision. Let set ^^ ൌ ^^^^|^^ ൌ 1,2. … ^^^ be a set ofall cells such that the total volume occupied by ^^ is equal to the total volume occupied by ^^^^^^^and ^^^^^^^. Further, let ^^^ ⊆ ^^ and ^^^ ⊆ ^^ where ^^^ and ^^^ are sets of cells occupied by points of^^^^^^^and ^^^^^^^respectively.

[0142] Empire3D Metrics: Metrics may be normalized between 0 and 1, representing the lowest and highest values of quality respectively. In contrast, geometric distance metrics such as ^^^, ^^^, and ^^^^are typically calculated such that a score of 0 represents a perfect match, and any value greater than 0 represents a degree of mismatch.

[0143] 3.1 Resolution^^^^^^

[0144] Resolution (per region) may be defined as the ratio of the average distance between points of ^^^^^^^to the average distance between points of ^^^^^^^given by ^^^. Overall resolution is the mean of ^^^over ^^ regions. Resolution determines the level of detail in the point cloud. Low resolution can cause loss of texture and smaller objects making the point cloud unusable for applications that require high fidelity and detail. ே 1^^^‾^^^^^ ൫^^^^^^^^൯ ^^ ^^ ^ (4)where,Attorney Docket No.: 011520.01958 ∑௫^∈^ min ∥ ^^ െ ^^ ∥^ ^^ ൌ ^௫ ∈^^ ^ ଶ^^^‾^^^^ ^ ^ ೕ ^ ; ^^ ് ^^

[0145] 3.2

[0146] Error may be defined as the ratio of the sum of distances between every point in ^^^^^^^to the nearest neighbor in ^^^^^^^given distance is less than threshold ^^ (shown in Figure 11), to the product of the number of points in ^^^^^^^and ^^, given by ^^^, here ^^ is theprecision set by the user. The normalization may be performed over (|^^^^^^^| ൈ ^^) as this is themaximum distance possible if all points in ^^^^^^^are valid (i.e., have neighbors within ^^ distance in ^^^^^^^). Overall accuracy (^^^) is the mean of ^^^over ^^ regions. ே 1 æ 1 ö (5) where,min∥ ^^ െ ^^ ∥ଶ , if  min ∥ ^^ െ ^^ ∥ଶ^ ^^^^^^^,^^^ ൌ ^^∈^^^ಲ ^∈^^^ಲ

[0147] In some embodiments, accuracy is computed on points that are not artifacts, which means any point not within set precision ^^ is considered an artifact, and accuracy is not penalized for the same. This ensures that for a given change in the candidate, per point the change only contributes to either of the metrics.

[0148] 3.3 Coverage ^^^^^^

[0149] Coverage is the ratio of number of cells occupied by points of ^^^^^^^and ^^^^^^^(shown in Figure 11) to the number of cells occupied by points of ^^^^^^^. Coverage may be computed per region as well as the whole point cloud separately, which may provide insights about local coverage (per cell) in addition to overall coverage. Overall coverage, unlike accuracy and resolution, is not an average of coverage per region over ^^ regions, rather it is computed on the entire volume bounded by the two point clouds.Attorney Docket No.: 011520.01958 |^^ ∩ |^^^ ൌ ^ ^ ^^^|^^^ (6) ^|

[0150] Coverage isand resolution which are computed as functions of point-point distance. This provides independence from accuracy and resolution sub-metrics, i.e., change in density of points or addition of Gaussian noise (within set precision ^^) has little to no effect on coverage, this is further explored in Section 4, below.

[0151] 3.4 Artifact Score ^^^^^^

[0152] Artifacts may be defined as the points in ^^^^^^^but not in ^^^^^^^(shown in Figure 11). These are generated due to reflections, distortion, or incorrect registration of points. Artifact score quantifies the lack of artifacts, i.e., the score is high if the candidate has low artifacts. Artifacts may be calculated as the ratio of number of cells occupied by points of ^^^^^^^but not occupied by points of ^^^^^^^to the number of cells occupied by points of ^^^^^^^. Someembodiments may use “artifact score” which may be defined as 1 െ ^^^^^^^^^^^^^^^^^^. This is similar tocoverage in the way that it is computed as a function of the volume occupied as compared to point-point distance. ^^|^^ \^௧ ൌ 1 െ ^ ^ ^^|^ (7)

[0153] 3.5 Relationship

[0154] A challenge in identifying these metrics can be to provide that each of them are independent and that, together, they cover all aspects of map quality. Here, we analyze the independence of these metrics.

[0155] Accuracy and Artifact-score: If points in ^^^^^^^drift away from points in ^^^^^^^, ^^^and ^^௧can both see change based on the ^^ value set. If a point moves more than ^^, it is counted as an artifact, affecting ^^௧but if it moves within the cell bounded by ^^ it only affects ^^^as defined above. If noise is introduced to ^^^^^^^, both ^^^and ^^௧can see change as some points may move out of the cells and some may move within.Attorney Docket No.: 011520.01958

[0156] Resolution and Accuracy: Since ^^^and ^^^are both distance-based metrics as shown above, they may appear to perform a similar role in quality measurement. However, ^^^is measured with distances between points in the same point cloud, whereas ^^^is measured with distances between points from different point clouds. In other words, ^^^changes when points within ^^^^^^^drift away from each other or when points within ^^^^^^^drift away from each other, whereas ^^^changes when points from ^^^^^^^drift away from points in ^^^^^^^.

[0157] Resolution and Coverage: A change in ^^^at lower values affects ^^^. When distances between points increase beyond ^^, ^^^decreases with ^^^since there are gaps between points in the map. We note that this is consistent with the definition of ^^^. We also note that not all reductions in ^^^will translate to a change in resolution. For example, a high-resolution map of a building floor with a room missing will be measured with lower coverage with no effect on ^^^.

[0158] Further, in section 4.1, we perturb the point cloud in various ways and show that ^^^^^^^^^^3^^ captures these perturbations in at least one of the metrics demonstrating comprehensiveness in quantifying map quality.

[0159] 4. Evaluation

[0160] We demonstrated the applicability of the presently disclosed framework with extensive experimental evaluation. We performed an ablation study using a custom dataset which demonstrated ^^^^^^^^^^3^^’s metrics’ utility, independence and consistency in evaluating aspects of point clouds quality. Next, ^^^^^^^^^^3^^ was evaluated on two applications: dense SLAM and learning-based point cloud completion using simulation and real-world experiments. This demonstrated the broad applicability of ^^^^^^^^^^3^^ for a broad class of 3D perception applications and improvement on other distance-based metrics.

[0161] 4.1 Ablation Study

[0162] The ablation study was performed using a prototype point cloud representing a simulated city-block bounded in a (40×40×10m) region containing approximately 1.28 million points (Figure 12). We used the simulation for accurate ground truth so we can study each metric of ^^^^^^^^^^3^^ in detail. The study involved applying various degradations to ^^^^^^^(source model), to produce a degraded model which is referred to as, ^^^^^^^in each case, and usingAttorney Docket No.: 011520.01958 ^^^^^^^^^^3^^, ^^^and ^^^to evaluate the quality at each step ^^^^^is not considered as heavy imbalance in candidate and reference point clouds prevents effective bijectivity / correspondence and ^^^^fails to provide any valuable insights into quality). Table IV shows results of this experiment. TABLE IV. Ablation Study Results | ^^ ൌ 0.1

[0163] . .

[0164] We down-sampled the source point clouds to simulate the reduction in resolution while preserving its overall structure. For each resolution ablation, we halved the total number of points. Uniform sampling was used to ensure points were removed consistently. When the resolution was reduced, the resolution metric ^^^also decreased. Since the points were uniformly sampled, they created no artifacts in this study, which is reflected in the consistency of the ^^௧and ^^^values. An observation is that ^^^changes with change in resolution. This is tied to the set precision. If a lower precision is set, the effect of the change in resolution is less on ^^^.

[0165] 4.1.2 Accuracy

[0166] To simulate loss of accuracy, Gaussian noise ^^^0,  ^^ଶ^ was applied to each axisof each point where ^^ଶis the variance applied to the points in a random normal direction. This resulted in a significant change to each metric due to potentially shifting points into other regions, causing a loss of coverage, resolution, and an increase in artifacts. When a noise with^^ ൌ 0.01 was applied, ^^^ showed a value of 0.8417. If ^^ was increased to 0.2, the ^^^ scoreincreased. This can be explained by how ^^^, ^^௧and ^^^are defined, adding Gaussian noise moves points out of cells into other cells affecting ^^^and ^^௧. Similarly, it also changes the distances between points as noise doesn’t translate all points uniformly affecting ^^^.Attorney Docket No.: 011520.01958

[0167] 4.1.3 Coverage

[0168] Spatial coverage was reduced by cropping the point cloud along a particular axis to simulate a lack of coverage. Table IV shows two ablated point clouds cropped to 40% the original in both the X and Y axes. Results show that this is reflected in ^^^as expected, due to its formulation as the ratio of the number of un-cropped points to the number of points in the original.

[0169] 4.1.4 Artifacts

[0170] Artifacts were simulated by shifting the source point cloud resulting in points leaving their respective cells, thereby inducing artifacts. When a shift was applied in the X-axis, the resulting artifact score ^^௧dropped down. However, if the precision (^^) was increased, the artifact score subsequently increased. This is because more points were considered valid when the ^^ value was increased. This same trend was observed when the point cloud was shifted in both X and Y with a larger value. This invariably affected coverage which was expected given points were non-uniformly distributed in the candidate. This provides some insight into the effect of setting the precision, a higher precision (lower ^^) will lead to a higher artifact score.

[0171] The ablation study shows that each our perturbations affect one of the ^^^^^^^^^^3^^ metrics but not others. For a real-world application, such comparisons provide hints on the benefits of using one method to construct 3D point clouds in comparison to another. Further, the ablation study shows that when precision is increased, the corresponding metric decreases noticeably. Precision is intended to be set based on expected quality and application and a veryhigh precision i.e., ^^ → 0 indicates a smaller expected margin of error. For example, whenconsidering a robot manipulation application, the precision can be set based on the size of the objects being manipulated.

[0172] Overall, the experiment demonstrates the effectiveness of ^^^^^^^^^^3^^ in evaluating the quality of point clouds and detecting changes in quality due to different types of degradation. ^^^^^^^^^^3^^ metrics scale in a proportional manner with change to the point clouds and ^^ provides control in quality assessment. This is in contrast to ^^^and ^^^’s behaviour where the change in these metrics cannot be effectively explained based on the degradation performed.

[0173] 4.2 Evaluation on dense SLAMAttorney Docket No.: 011520.01958

[0174] To further evaluate ^^^^^^^^^^3^^ we constructed point clouds using dense SLAM methods. First, collect LiDAR scans in a custom simulation environment and build point cloud maps using several popular LiDAR SLAM systems. The simulation environments were built in Gazebo (Koenig and Howard 2004) with mesh models of the worlds and a simulated Ouster OS1-128 LiDAR mounted on a Clearpath Husky (“Husky UGV - Outdoor Field Research Robot by Clearpath — Clearpathrobotics.com”; “GitHub - Husky / Husky_simulator: Simulator Packages for the Clearpath”). Worlds include Mai City (Vizzo et al.2021) and another named Warehouse (Figure 13). These are designed to represent real-world environments with elements commonly found in the real world. Ground truth meshes were sampled to obtain ground truth point cloud since ^^^^^^^^^^3^^ compares point clouds. We also matched the number of points from the maximum of the candidate point clouds to keep the comparison fair.

[0175] Next, we tested ^^^^^^^^^^3^^ on real-world data where LiDAR scans were collected using an Ouster OS1-128 LiDAR and pose using a Leica Geosystems TS15 Robotic Total Station (“Robotic Total Stations — Leica-Geosystems.com”). These scans were stitched using the captured poses and ICP (Rusinkiewicz and Levoy 2001; “GitHub - CloudCompare / CloudCompare: CloudCompare Main Repository — Github.com”) to generate a ground truth map with mm precision in poses. The dataset is named Davis for ease of reference (Figure 14, Figure 15). For the candidate point cloud, we captured LiDAR and IMU data using a Boston Dynamics Spot equipped with an Ouster OS1-128 LiDAR with a built-in IMU by walking it in the same building and use a visual SLAM method to build corresponding point cloud map. We also tested ^^^^^^^^^^3^^ on (L. Zhang et al.2022) which contains a ground truth point cloud generated using an engineering-grade LiDAR. We studied these four datasets (two simulation, two real-world) with three candidate SLAM methods totalling 12 point clouds with their corresponding ground truth. Candidate point clouds were generated using LeGO-LOAM (Shan and Englot 2018), FAST-LIO2 (W. Xu et al.2022a), and SHINE (X. Zhong et al.2023). The first two output a dense point cloud and real-time odometry while the last one employs a unique approach to mesh generation by employing hierarchical implicit neural representations to generate a mesh. (Note: The output mesh was sampled into a point cloud similar to thesimulation ground truth, this led to resolution metric ^^^ ൌ 1 as the average distance betweenpoints is the same as the ground truth.)

[0176] These were used to evaluate ^^^^^^^^^^3^^ metrics as well as ^^^and ^^^(^^^^was not considered as heavy imbalance in candidate and reference point clouds prevents effectiveAttorney Docket No.: 011520.01958 bijectivity / correspondence and ^^^^fails to provide any insights into quality). The results of this evaluation are presented in Table V. Figures 13 and15 show the qualitative differences between the point clouds for the Davis and Warehouse datasets (Note: Evaluation was performed on the entire point cloud and not only the zoom-in view). Each metric evaluates a certain aspectof quality as described in Section 3. We set precision ^^ ൌ 0.5 and ^^ based on point cloud size forthe tests in Table V. Other values were explored but not presented in the interest of space as results are consistent with the definition. TABLE V. Evaluation on SLAM maps | ^^ ൌ 0.5

[0177] sua y, s apparen a e po n c ou map genera e w T-LIO2 has significantly higher detail than the one built using LeGO-LOAM. Subsequently, FAST-LIO2 receives higher ^^^, ^^^and ^^^but a lower ^^௧while ^^^and ^^^identify LeGO-LOAM as the one nearest to ground truth. In this example, ^^^and ^^^do not provide any insight into the point clouds’ quality, i.e., the extreme imbalance between the point densities of the two candidates. It is noted that ^^^^^^^^^^3^^ metrics may not always agree with our perception of quality due to the multi-dimensional nature of quality assessment. However, they are fundamentally true to their definition which is consistent in Rଷand accurate based on (^^). We see this in point cloud maps built using LeGO-LOAM; although sparse, they are accurate. Their sparse nature contributes to the lack of artifacts which is reflected in ^^௧but negatively affects ^^^and ^^^.

[0178] Figure 16 shows different views of the Warehouse dataset, including magnified views for qualitative analysis. These images clearly show that SHINE and FAST-LIO2 significantly outperform LeGO-LOAM. ^^^^^^^^^^3^^ values corroborate with qualitative results byAttorney Docket No.: 011520.01958 identifying SHINE’s point cloud as the one with highest ^^^, ^^^, and ^^^whereas ^^^identifies FAST-LIO2. Of note, SHINE has a lower artifact-score (compared to FAST-LIO2) indicating the presence of artifacts, observed in Figure 16 where the pillars appear distorted due to artifacts. Further, ^^^^^^^^^^3^^ also accurately quantifies the lower resolution in LeGO-LOAM compared to others.

[0179] Constructing ground truth point clouds:

[0180] Acquiring ground truth for point cloud evaluation in a real-world setting is a challenging task. Although CAD and BIM files can be used, they’re generally outdated and seldom represent the current environment. In outdoor settings, topographical maps and digital elevation models (DEMs) are sparse and lack rich 3D information. A common practice to build dense 3D models is the use of survey grade LiDAR scanners (Arayici 2007; Mandlburger et al. 2023), these use ground truth poses (from RTK GPS or Total-Stations) along with LiDAR scans to stitch a high-resolution 3D representation of the environment.

[0181] We used a similar approach to build ground truth point clouds for the experiments. First, we captured ground truth poses using a Robotic Total-Station (Kizil and Tisor 2011) which is essentially a theodolite with an integrated distance meter that can measure distances and angles. This enables extremely precise pose estimation (millimeter-level) and is widely used in infrastructure and geospatial surveying (Chekole 2014; Tarvo Mill and Liias 2013). Second, we captured LiDAR scans at the exact locations where poses were measured. This was achieved by placing a LiDAR mounted on a custom tripod, the tripod was equipped with a nadir pointed laser that projected a cross-hair onto carefully placed markers on the ground to align scans. For the Davis dataset, over 400 poses and corresponding scans were captured with this method. Finally, the scans were stitched into a dense point cloud of the environment with the recorded poses providing initial alignment. Iterative closest point (ICP) was used for fine registration.

[0182] Simulation environment: The simulator built to study ^^^^^^^^^^3^^’s performance supports various environments; these are mesh files of objects and structures and can be built using popular tools like Blender (Community 2018). The simulator supports three variants of the Ouster OS series LiDARs, namely the OS-0, OS-1, and OS-2 sensors. These allow a maximum vertical field-of-view of 90, 45, and 22.5 degrees respectively. Each of the sensors can beAttorney Docket No.: 011520.01958 configured in three resolutions - 128, 64 and 32 channels. This totals nine distinct LiDARs in simulation. We also add noise to simulate realistic outputs which can be tuned to match the sensor. Additionally, the simulator outputs IMU data and ground truth odometry for use in SLAM methods that need IMU or need external pose information.

[0183] 4.3 Evaluation of Learning-based Point Cloud Completion

[0184] A recent approach to point cloud generation is point cloud completion. We analyze three networks that output a completed point cloud when given a partial point cloud. For this study, we consider PCN (Yuan et al.2018), TopNet (Tchapmi et al.2019) and ECG (Pan 2020) point cloud completion models and the MVP dataset (Pan et al.2021) for their evaluation (see Figure 20). The resulting completed point clouds are evaluated against ground truth using ^^^^^^^^^^3^^, ^^^and ^^^. Table VI shows the quantitative results of this experiment while Figure 17 shows the resulting point clouds. On visual inspection of the point clouds, it is evident that the point clouds generated using ECG and PCN exhibit the highest quality, and this is corroborated by ^^^^^^^^^^3^^’s metrics. Unlike the SLAM experiments, these findings are corroborated by ^^^and ^^^. This reinforces our hypothesis regarding ^^^and ^^^’s limitations; both these metrics are able to identify the highest quality point clouds when the size of the pointcloud is small and where the density is roughly uniform but fail to do so in the SLAM study due to large size and unevenness of the point clouds. Learning-based methods are rapidly becoming the dominant way to generate point clouds including point cloud completion (Yuan et al.2018; Tchapmi et al. 2019; Pan 2020), image-based 3d reconstruction (Fan, Su, and Guibas 2017), and image-to-mesh generation, etc., and we believe that ^^^^^^^^^^3^^ is the right framework to compare and evaluate their outputs. TABLE VI: Evaluation on Point cloud completion || ^^ ൌ 0.3Attorney Docket No.: 011520.01958

[0185] 4.4 Compute Performance

[0186] For the demonstration, ^^^^^^^^^^3^^ was implemented using Open3D (Q.-Y. Zhou, Park, and Koltun 2018), PDAL (Contributors 2020), Scikit-Learn (Pedregosa et al.2011), NumPy (Harris et al.2020) and PyTorch (Paszke et al.2019) and was intended for dense pointclouds (^ 10^ points). To handle such large point clouds, we provided a fast multi-threadedimplementation that computes the regions in parallel, as well as a GPU-accelerated implementation capable of utilizing a GPU if one exists. In addition to that, ^^^^^^^^^^3^^ computes in ^^^^^log^^^ which is similar if not faster than popular methods while being multi-dimensional (Figure 21).

[0187] This experiment was conducted to compare an implementation of ^^^^^^^^^^3^^ with Chamfer and Hausdorff distances for 4 resolutions of the point cloud generated using FAST- LIO2 on the Warehouse dataset. Different resolution maps were obtained by down-sampling the original point cloud. As expected, reducing the number of points decreased the computation time for all. However, at different region sizes, we were at least twice as fast as ^^^and ^^^.

[0188] ^^^^^^^^^^3^^ compute performance was evaluated by calculating quality metrics for a sample point cloud. We computed ^^^^^^^^^^3^^ for a wide range of region sizes while ^^^and ^^^were calculated on the entire point clouds. Region sizes of 1, 2, 5, 10, 15, 20, 30, 50, 100 meters were chosen and run-time along with CPU and memory utilization was measured. Further, we repeated the experiments by limiting number of CPU cores for multi-threading providing some insights on computing on resource constrained hardware.

[0189] Using ^^^and ^^^as baselines for reference, we observed that their CPU utilization was low and correspondingly the time taken to complete the evaluation was high. In contrast, this experimental implementation of ^^^^^^^^^^3^^, traded CPU utilization for execution time. Figure 22 illustrates this trade-off along with the effect of choosing different region sizes for parallelization on a benchmark point cloud. We observe that for smaller region size configuration, the CPU utilization is higher. As the region size is increased, the utilization dips as the region size grows closer to the size of the point cloud. Eventually, ^^^^^^^^^^3^^’s parallel implementation will reach a CPU utilization parity with ^^^and ^^^where entire point clouds are loaded on a single thread.Attorney Docket No.: 011520.01958

[0190] On the other hand, we observe a trend in the execution time that might appear unusual at first glance. The time taken by ^^^^^^^^^^3^^ starts to decrease with increase in region size but rises again for even higher values. We note that the reasons for this effect are two-fold. With much smaller region sizes, the individual region comparisons are quick. The major reason for this is that there are fewer points to consider for the closest point match. But, there are many regions to compute leading to a longer time. As we increase the region size, the number of regions to compare reduces giving ^^^^^^^^^^3^^ a performance boost. At the same time, the time taken for the closest point search for individual regions increases. Finally, as the region size grows close to the full map, the time taken for closest point matches dominates leading to longer execution times. Further, we note that as an added benefit of using lower region sizes is the lower memory usage.

[0191] Since ^^^^^^^^^^3^^ metrics may be averaged, a choice of region size will affect the actual values of some metrics. While coverage and artifact-score are unaffected by region size, accuracy and resolution are affected by the averaging. However, this should not affect comparisons between different point clouds as long as the same region size is used for the comparisons.

[0192] 5. Applications and Use-cases

[0193] To illustrate ^^^^^^^^^^3^^’s usage, consider the problem of selecting a suitable LiDAR sensor for dense mapping. Given Ouster OS-0 and OS-1 LiDARs, with vertical FOV of 90 and 45 degrees respectively, a range of 35m and 90m respectively and a vertical resolution of 128 channels, we simulate both sensors in our simulator and map the Warehouse environment using FAST-LIO2. Figure 24 shows the point clouds and corresponding ^^^^^^^^^^3^^ metrics. FAST-LIO2 with OS-0 generated ^^^^^^ைௌି^and OS-1 generated ^^^^^^ைௌି^.

[0194] ^^^^^^ைௌି^has higher coverage at 90.5% compared to 89.36% in ^^^^^^ைௌି^, this is primarily due to the larger V-FOV of the OS-0 LiDAR. ^^^^^^ைௌି^on the other hand, has higher resolution at 97.86% compared to 95.02% in ^^^^^^ைௌି^, this can be attributed to the spreading out of the 128 available vertical channels over different FOVs that produce different densities. Results indicate that the OS-1 is a better choice in this case due to the higher resolution in the maps despite the expanded coverage of OS-0.Attorney Docket No.: 011520.01958

[0195] This style of assessment can help pick sensors and algorithms based on expected coverage and resolution. Higher coverage and resolution generally result in improved SLAM performance, both in localization and mapping. This is because an increase in these metrics indicates an increase the number of features detected in the point cloud. To demonstrate this further, we compute ISS features (Y. Zhong 2009) on both ^^^^^^ைௌି^and ^^^^^^ைௌି^. ^^^^^^ைௌି^contains 1,531,103 features while ^^^^^^ைௌି^contains 1,509,107, this confirms the hypothesis that an increase of 1.14% in coverage results in 1.45% increase in the number of points detected.

[0196] Application: Anomaly Detection

[0197] Empir3D may also be used for anomaly detection. An objective in anomaly detection is to measure and localize changes in a scene using a reference point cloud map. Traditional distance measures can achieve this but are often slow due to the reasons outlined in section 2. By leveraging Empir3D, we can accelerate this process, as Empir3D is computationally efficient and enables real-time performance.

[0198] To validate this, we conducted experiments using an Ouster OS-1128 Channel LiDAR mounted on a Boston Dynamics Spot robot. The robot operated in a controlled environment where objects could be moved, added, or removed. Figure 25 demonstrates results from one such experiment. In this scenario, a large cardboard box was relocated, and Empir3D’s anomaly detection output is highlighted (annotated and outlined in boxes). The detection identifies two regions of anomalies: new points (artifacts) are visible in the area where the box was originally placed, and the box’s new location shows coverage. The terms ‘coverage’ and ‘artifacts’ are interchangeable, depending on whether the initial LiDAR frame is considered the reference or the candidate. This interchangeability does not impact usability or performance.

[0199] Furthermore, we can quantify the detected change. The change illustrated in Figure 25 measures to 0.021, indicating that 2.1% of the scene has changed. We achieved real- time performance at 5 frames per second with a LiDAR output of 5.2 million points per second. Performance can be further enhanced by limiting the field of view (FOV) to the region of interest (ROI).

[0200] 6. DiscussionAttorney Docket No.: 011520.01958

[0201] Multi-Dimensional Evaluation: Evaluations in Section 4 help demonstrate the utility of ^^^^^^^^^^3^^’s multi-dimensional approach to quality assessment. We emphasize that the objective of ^^^^^^^^^^3^^ is not to categorically align with any specific qualitative assessment, but rather to illustrate and quantify various aspects of point quality. Such a comprehensive assessment provides valuable feedback to point cloud construction methods and allows the developers to improve them. It also allows specific applications to identify quantifiable metrics in point cloud quality and how they correspond to application accuracy (object recognition, for example). The comprehensive assessment is clearly articulated in the SLAM and point cloud completion experiments, where ^^^and ^^^only provide a single number while ^^^^^^^^^^3^^ is able to quantify each aspect of quality.

[0202] The assessments further highlight certain trends in the behavior of ^^^and ^^^. Both these distance measures perform relatively well when the point clouds are small and densities are consistent. This is demonstrated in the evaluation of point cloud completion networks where they identify ECG and PCN’s outputs as highest quality which is in agreement with ^^^^^^^^^^3^^ metrics and qualitative results. This behavior does not hold when point clouds have high-imbalance and / or are large in size. In the SLAM experiments, in some datasets they identify LeGO-LOAM’s point clouds as the ones with the highest quality which contradicts qualitative results. Overall, ^^^and ^^^fail to provide any real insight into the point clouds’ quality.

[0203] Applications: Although ^^^^^^^^^^3^^ is demonstrated on dense point clouds for convenience in this disclosure, applications exist in various domains (and the scope of the disclosure should be interpreted to include) such as:

[0204] Optimizing point cloud construction: Algorithms such as Visual SLAM intend to recreate 3D structure for navigation, manipulation etc. ^^^^^^^^^^3^^ provides better insight into the algorithm performance, thereby better informing the developer of how it could be used for the end application.

[0205] Improving learning on point clouds: Chamfer loss is a popular loss function used for 3D deep learning tasks. The various ^^^^^^^^^^3^^ metrics can provide a way to learn in a structured manner for applications such as depth completion, point cloud generation, etc.Attorney Docket No.: 011520.01958

[0206] Sensor characterization: ^^^^^^^^^^3^^ can be used to quantify how well a sensor or suite of sensors are able to see all obstacles in a scene. Such characterization could be useful for a sensor suite on an autonomous car, for example, to identify potential blind spots.

[0207] Need for reference point clouds: As ^^^^^^^^^^3^^ and other methods evaluated in this paper are full-reference similarity measures, source point clouds (e.g., ground truth point clouds are necessary for evaluation. Ground truth point clouds can be generated using better sensors (e.g., an engineering-grade LiDARs or Total-Station). Alternatively, we can evaluate methods in a simulation where ground truth is readily available. However, most point cloud construction methods need to be evaluated using a reference—and ^^^^^^^^^^3^^ requires the same.

[0208] Although the present disclosure has been described with respect to one or more particular embodiments, it will be understood that other embodiments of the present disclosure may be made without departing from the spirit and scope of the present disclosure.

Claims

Attorney Docket No.: 011520.01958 What is claimed is:

1. A computer-implemented method of assessing a candidate point cloud, comprising: receiving a source point cloud representing a ground truth and the candidate point cloud; dividing each of the candidate point cloud and the source point cloud into a plurality of cells of equal size; calculating a resolution sub-metric; calculating an accuracy sub-metric; calculating a coverage sub-metric to indicate a level of overlap between the candidate point cloud and the source point cloud; calculating an artifact sub-metric to indicate a proportion of anomalous points in the candidate point cloud; and calculating a point quality metric based on the coverage sub-metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub-metric; and outputting the point quality metric as a normalized value.

2. The method of claim 1, wherein the resolution sub-metric is based on a ratio of an average distance between points within each cell of the source point cloud to an average distance between points of each corresponding cell of the candidate point cloud.

3. The method of claim 1, wherein the accuracy sub-metric is based on a ratio of a sum of distances between every point in each cell of the candidate point cloud to a nearest neighbor in each corresponding cell of the source point cloud given distance is less than a predetermined distance threshold, to a product of a number of points in each cell of the candidate point cloud and the predetermined distance threshold.

4. The method of claim 3, wherein the predetermined distance threshold is a tunable parameter selected based on an application-specific requirement for accuracy, coverage, or artifact tolerance.

5. The method of claim 1, wherein the coverage sub-metric is based on a ratio of a number of cells occupied by points of the source point cloud and the candidate point cloud to a number of cells occupied by points of the source point cloud.Attorney Docket No.: 011520.01958 6. The method of claim 1, wherein the artifact sub-metric is based on a ratio of a number of cells occupied by points of the candidate point cloud but not occupied by points of the source point cloud to a number of cells occupied by points of the candidate point cloud.

7. The method of claim 1, wherein the size of each cell of the plurality of cells is a tunable parameter configured to balance computational efficiency and resolution of local variations in the candidate point cloud.

8. The method of claim 1, wherein the point quality metric is calculated as a weighted sum, and wherein the weights for the coverage sub-metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub-metric are adjustable to prioritize specific sub-metrics based on an application.

9. The method of claim 1, further comprising calculating one or more of the coverage sub- metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub-metric in parallel across the plurality of cells using a multi-threaded implementation.

10. The method of claim 1, further comprising dividing each of the source point cloud and the candidate point cloud into a plurality of regions for computational efficiency, and where one or more of the resolution sub-metric and the accuracy sub-metric are calculated for each region.

11. A system for assessing a candidate point cloud, comprising: a memory for storing the candidate point cloud and a source point cloud; and a processor in electronic communication with the memory, the processor configured to: divide each of the candidate point cloud and the source point cloud into a plurality of cells of equal size; calculate a resolution sub-metric; calculate an accuracy sub-metric; calculate a coverage sub-metric to indicate a level of overlap between the candidate point cloud and the source point cloud; calculate an artifact sub-metric to indicate a proportion of anomalous points in the candidate point cloud; and calculate a point quality metric based on the coverage sub-metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub-metric; and output the point quality metric as a normalized value.Attorney Docket No.: 011520.01958 12. The system of claim 11, wherein the processor calculates the resolution sub-metric based on a ratio of an average distance between points within each cell of the source point cloud to an average distance between points of each corresponding cell of the candidate point cloud.

13. The system of claim 11, wherein the processor calculates the accuracy sub-metric based on a ratio of a sum of distances between every point in each cell of the candidate point cloud to a nearest neighbor in each corresponding cell of the source point cloud given distance is less than a predetermined distance threshold, to the product of a number of points in each cell of the candidate point cloud and the predetermined distance threshold.

14. The system of claim 13, wherein the predetermined distance threshold is a tunable parameter selected based on an application-specific requirement for accuracy, coverage, or artifact tolerance.

15. The system of claim 11, wherein the processor calculates the coverage sub-metric based on a ratio of a number of cells occupied by points of the source point cloud and the candidate point cloud to a number of cells occupied by points of the source point cloud.

16. The system of claim 11, wherein the processor calculates the artifact sub-metric based on a ratio of a number of cells occupied by points of the candidate point cloud but not occupied by points of the source point cloud to a number of cells occupied by points of the candidate point cloud.

17. The system of claim 11, wherein the size of each cell of the plurality of cells is a tunable parameter configured to balance computational efficiency and resolution of local variations in the candidate point cloud.

18. The system of claim 11, wherein the processor calculates the point quality metric as a weighted sum, and wherein the weights for the coverage sub-metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub-metric are adjustable to prioritize specific sub- metrics based on an application.

19. The system of claim 11, wherein the processor is a parallel processor configured to calculate one or more of the coverage sub-metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub-metric in parallel across the plurality of cells using a multi-threaded implementation.Attorney Docket No.: 011520.01958 20. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method for assessing the quality of a point cloud, the method comprising: receiving a source point cloud and a candidate point cloud; each of the candidate point cloud and the source point cloud into a plurality of cells of equal size; computing a coverage sub-metric, an artifact sub-metric, an accuracy sub-metric, and a resolution sub-metric for each cell of the plurality of cells; calculating a point quality metric based on the coverage sub-metric, the artifact sub-metric, the accuracy sub-metric, and the resolution sub-metric; and outputting the point quality metric as a normalized value.

Citation Information

Patent Citations

  • Estimation of density distortion metric for processing of point cloud geometry

    US20240020880A1

Cited By

  • Method for detecting and evaluating mechanical properties of 3D printing-based ABS material test piece

    CN122474228A