System and method for generating a multi-resolution voxel space

JP2025518696A5Pending Publication Date: 2026-06-03ZOOX INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ZOOX INC
Filing Date
2023-05-24
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Existing systems face challenges in efficiently generating and aligning multi-resolution voxel spaces, particularly in environments with sparse data or large initial scan errors, which can consume significant processing resources and time.

Method used

The system aligns coarser voxel resolutions first and iteratively adds finer resolutions, using eigenvalue weighting and quality metrics to optimize the alignment process, thereby reducing computational resources and time required for convergence.

Benefits of technology

This approach enables faster and more efficient convergence of multi-resolution voxel spaces, allowing for more accurate environmental mapping and improved operational efficiency for autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In this application, a method for representing a scene or a map is described based on statistical data of captured environmental data. In some cases, the data (covariance data, average data, or the like) may be stored as a multi-resolution voxel space including a plurality of semantic layers. In some cases, individual semantic layers may include a number of voxel grids having different resolutions. A number of multi-resolution voxel spaces may be integrated or aligned to generate a combined scene based on voxel covariance detected at one or more resolutions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to systems and methods for generating a multi - resolution voxel space.

Background Art

[0002] [Cross - Reference to Related Applications] This application claims priority to U.S. Application No. 17 / 804,744, filed May 31, 2022, entitled "Systems and Methods for Generating a Multi - Resolution Voxel Space", the entire disclosure of which is incorporated herein by reference.

[0003] [Background] In an environment, data can be captured and represented as a map of the environment. Often, such a map can be used by a vehicle navigating within the environment, but the map can be used for various purposes. In some cases, the environment can be represented as a two - dimensional map, while in other cases, the environment can be represented as a three - dimensional map. Further, the surfaces within the environment are often represented using a plurality of polygons or triangles.

Brief Description of the Drawings

[0004] For a detailed description, reference is made to the accompanying drawings. In the figures, the left - most digit of the reference number identifies the figure in which the reference number first appears. The use of the same reference number in different figures indicates similar or identical components or features.

[0005]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0006] The techniques described herein are directed to generating an alignment between map data composed of multi - resolution voxel spaces. In some examples, such multi - resolution voxel spaces can be composed of voxels that store statistical information regarding associated measurements, where the associated measurements include, but are not limited to, the spatial mean, covariance, and weighting of a point distribution of data representing a physical environment. This map data can be composed of multiple voxel grids (e.g., "volumetric pixels" or a discretized volumetric representation including voxels) or voxel layers that represent a physical environment at various resolutions or physical distances. For example, each voxel layer can represent a physical environment at a resolution that is a multiple (e.g., 2 times, etc.) of that in the proceeding layer. That is, the voxels in the first layer can represent a first volume (e.g., 10 cm × 10 cm × 10 cm, etc.), while the voxels in the second layer can represent a second volume (e.g., 20 cm × 20 cm × 20 cm, etc.).

[0007] Data associated with voxels in a multi-resolution voxel space can be represented as a plurality of covariance ellipsoids. This covariance ellipsoid representation can be generated based on the calculated mean and covariance values of the data points associated with individual voxels. For example, each of the ellipsoids can have a shape determined by one or more eigenvectors associated with the covariance matrix of the voxel (such as three eigenvectors associated with the X, Y, and Z measurements of the associated data points, etc.). In some cases, voxel data can be associated with semantic information such as classification information and / or segmentation information, and data associated with a certain classification can be associated with a specific multi-resolution voxel space associated with that certain classification. In this example, the covariance semantic layer of each voxel can be composed of data points associated with a specific semantic class (such as trees, vehicles, buildings, etc.) as covariance ellipsoids.

[0008] In some cases, map data represented by a multi-resolution voxel space can be generated from data points representing a physical environment, such as the output of a light detection and ranging (lidar) system. For example, the system may receive lidar data represented as a plurality of lidar points, or a point cloud. The system may assign the lidar point cloud to voxels of a voxel grid having multiple resolutions (e.g., finer resolutions, coarser resolutions, etc.) based at least in part on a local reference frame of a vehicle (such as a system that captures lidar points), or otherwise associate them. The system may then integrate or otherwise combine groups of voxels (or data associated with the voxel groups) of one or more resolutions substantially simultaneously to generate a final multi-resolution voxel space. As a non-limiting example, measurements may first be associated with a voxel grid of the finest resolution, and other groups of layers (having coarser resolutions) may be calculated based on the finer resolution and, for example, integrated. In some examples, such association and calculation may be performed substantially simultaneously (e.g., within the technical tolerance). In one specific example, groups of voxels in the neighborhood are integrated by taking a weighted sum of the individual Gaussian distributions of each voxel in the finer resolution grid.

[0009] In some implementations, the system may utilize a multi-resolution voxel space to generate an alignment between voxel spaces in order to assist in mapping the physical environment and generating a scene, as well as to assist in determining the position of a vehicle within the map or scene. For example, once a multi-resolution voxel space (e.g., a target multi-resolution voxel space, etc.) is generated for a particular scan or dataset representing the physical environment (e.g., determined during driving), the system may determine an alignment between the generated multi-resolution voxel space and a reference multi-resolution voxel space representing the scene. In some cases, the alignment may be generated by finding a correspondence between voxels at each resolution of the reference multi-resolution voxel space and the target multi-resolution voxel space. For example, the system may search among a group of voxels within a threshold distance or threshold number of voxels for each voxel at a particular resolution in the target multi-resolution voxel space, including the average target point at the corresponding particular resolution of the reference multi-resolution voxel space for the occupied voxel. In an example including a semantic layer, the system may search in the vicinity of a group of voxels including the average target point at the corresponding particular resolution of the corresponding semantic layer in the reference multi-resolution voxel space for each voxel at a particular resolution of each semantic layer in the target multi-resolution voxel space.

[0010] Among the identified group of voxels in the reference multi-resolution voxel space, the system may select a voxel having a center of gravity close to the voxel in the target multi-resolution voxel space. Thereafter, the system may, in at least some instances, determine a residual (or error, etc.) for each of such matching voxels, which may be at least partially based on such matching normal vectors, and subsequently perform an optimization on all such residuals. This optimization may minimize the distance between the centers of gravity and / or the average pairs of such voxels. In this way, an alignment or error between two voxels may be determined.

[0011] Even if each layer can be integrated substantially simultaneously during alignment, a coarser resolution (e.g., a resolution corresponding to a larger voxel group) can result in matching earlier than a finer resolution. In this way, matching at a coarser resolution can help bring the two multiresolution voxel spaces into closer alignment, enabling the finer resolution to start matching and complete the alignment process. In some cases, by integrating the captured sensor data into a multiresolution voxel space representing the environment, a vehicle group can initialize or identify its position within the environment with higher accuracy and / or more quickly than a system that uses conventional map data composed of polygons and / or meshes. Further, by storing voxel groups in a multiresolution voxel space, the data can be stored in a more easily indexable / searchable manner, thereby improving processing speed and throughput. For example, if a coarse resolution is acceptable for a practical task, that coarse layer can be loaded into memory, thereby reducing the amount of data accessed and processed for a desired operation.

[0012] In some examples, the system can start alignment or pre-align the target multiresolution voxel space and the reference multiresolution voxel space using position data, location data, and / or the like associated with the autonomous vehicle at the time the data used to generate the target multiresolution voxel space was generated. For example, the position data and / or location data can include global positioning system (GPS) data (or other satellite-based position data), odometry data, inertial measurement unit (IMU) data, prior known positions and / or degrees of freedom determined based on the alignment of a prior multiresolution voxel space to a reference, and the like.

[0013] In some cases, depending on the seed data (e.g., lidar data or point cloud data, etc.) used to generate two multi-resolution voxel spaces, it may be difficult for the system to generate an alignment, and / or it may require a large amount of processing resources and / or time for the alignment between the voxel spaces to converge. For example, when the voxel groups are sparse or far apart from each other, two voxels (e.g., one from each voxel space) may hardly overlap, and when aligning the voxels based on covariance, a large amount of time and resources may be consumed. In other examples, when the initial scan error is large (e.g., exceeding 10 m), in addition to consuming a large amount of resources and / or time, it may be difficult for the voxel space to converge at all. In such cases, the systems and methods described herein can assist and / or improve the convergence rate and speed, and reduce the consumption of processing resources or computing resources associated with generating an alignment between multi-resolution voxel spaces.

[0014] In one embodiment, the system may align or converge a coarser voxel resolution before starting the alignment of a finer resolution. For example, the system may enable voxel groups larger than a predetermined threshold to be aligned during the initial stages of the process. For example, in one specific example, the system may enable voxels with a resolution of 25 meters or more to start the alignment in the initial stage. Thereafter, the system may iteratively add the next finer resolution to the convergence process, for example, when the system determines an error less than the error threshold. In some cases, the error threshold for the next finer resolution may be such that the average error of the voxel group is less than half of the current finest resolution (e.g., if the current finest resolution is 25 meters, the system may add the next finer resolution when the error is 12.5 meters or less). In other examples, the error threshold for the next finer resolution may be such that the average error of the voxel group is less than or equal to one quarter of the size of the current finest resolution, or something similar.

[0015] In this example, the system may continue to iterate the alignment stage until each resolution is added. In some cases, the system may iterate for a predetermined number of iterations with all resolutions (e.g., 1 iteration, 2 iterations, 3 iterations, 5 iterations, or the like) after all resolutions are added, or until the average error is below a final error threshold, or until the change in the sum of voxel residuals is below a change threshold.

[0016] In another example, also during alignment, the system may evaluate, weight, or otherwise score the eigenvalues of a group of voxels and utilize the weighting to select eigenvectors and / or voxels for use in the alignment. For example, as described above, a voxel may have a set of three or more eigenvalues (providing the episodic shape of the voxel). The system may then determine the score or weight of an eigenvector or voxel, at least in part based on the size of the eigenvalue. In some implementations, the system may evaluate individual eigenvalues against one or more predetermined heuristics to determine the weight. For example, the system may then determine whether one or more of the eigenvalues for a voxel are below one or more thresholds when generating the score. In some cases, one or more of the thresholds may be relative to the size of the voxel or the size of the associated resolution. In other embodiments, the system may utilize one or more machine learning models to evaluate and / or score individual eigenvalues. The system may then select a group of voxels associated with a higher score in the alignment process (e.g., based on those eigenvalues or other characteristics described herein, such as the number of points, quality metrics, etc.).

[0017] In some examples, a multi - resolution voxel space may include multiple voxel layers for a given resolution, and individual layers may be associated with various semantic classes. For example, a voxel layer at a first resolution may be associated with buildings, a voxel layer at a second resolution may be associated with ground surfaces, and / or a voxel layer at a third resolution may be associated with vegetation. In this example, the system may also adapt or regress a quality or confidence metric for individual voxels within the resolution using the semantic classes assigned to the voxels. For example, the system may generate and / or train a noise model for evaluating the quality of a group of voxels based at least in part on the eigenvalues of the voxels, the resolution of the voxels, the number of points associated with the integrated voxels, and the semantic class of the integrated voxels. The system may then select a group of voxels for use in an alignment process using the quality or confidence metric. For example, the system may determine a scale factor correction for the multi - resolution voxel space by randomly sampling (such as by Monte Carlo method) a group of voxels based at least in part on the quality or confidence metric. In some cases, the scale factor correction may be utilized as a metric to determine the overall quality of the resulting alignment.

[0018] In some cases, by aligning coarser resolutions before finer resolutions, using weighting on the eigenvalues of the group of voxels to determine eigenvectors for use in alignment, and using the quality or confidence metric of the group of voxels for use in alignment, the system can not only converge the multi - resolution voxel space in a more time - efficient manner using fewer processing resources, but also generate a high - quality multi - resolution voxel space that better represents the corresponding physical environment.

[0019] As described herein, the systems and methods enable the process for alignment between multi-resolution voxel spaces to converge in a more efficient manner. For example, the systems and methods enable convergence between multi-resolution voxel spaces in fewer time periods while consuming fewer resources than conventional systems. In some cases, by shortening the time period associated with convergence, an autonomous vehicle can make operational decisions, including safety-related decisions, in a more timely manner. Further, the systems and methods described herein can result in a more accurate alignment between a target multi-resolution voxel space and a reference multi-resolution voxel space, thereby enabling an autonomous vehicle to perform operations using a more accurate perception of the environment surrounding the vehicle, and generally resulting in a safer operation of such a system.

[0020] FIG. 1 is an exemplary process flow diagram 100 illustrating an exemplary data flow of a system configured to align data representing a physical environment with a scene, as described herein. In the illustrated example, the system can be configured to store a scene as a multi-resolution voxel space, as well as data representing the environment. As described above, a multi-resolution voxel space may have a plurality of semantic layers, where each semantic layer represents voxels as covariance ellipsoids at different resolutions.

[0021] In one particular example, a sensor system 102, such as a lidar, radar, sonar, infrared, camera, or other imaging device, can capture data representing the physical environment surrounding the system. In some cases, the captured data can be a plurality of data points 104, such as a point cloud generated from the output of a lidar scan. In this example, the data point cloud 104 can be received by a multi-resolution voxel space component 106.

[0022] The multi-resolution voxel space component 106 may be configured to generate a target multi-resolution voxel space from the data point cloud 104. In some cases, the multi-resolution voxel space component 106 may process data points via a classification technique and / or a segmentation technique. For example, the multi-resolution voxel space component 106 may assign a type or class to the data points using one or more neural networks (e.g., a deep neural network, a convolutional neural network, etc.), a regression technique, etc., and identify and classify the data point cloud 104 using semantic labels. In some cases, the semantic labels may include classes or entity types such as vehicles, pedestrians, cyclists, animals, buildings, trees, road surfaces, curbs, sidewalks, unknowns, etc. In additional and / or alternative examples, the semantic labels may include one or more characteristics associated with the data point 104. For example, these characteristics may include, but are not limited to, an X-position (global position and / or local position), a Y-position (global position and / or local position), a Z-position (global position and / or local position), an orientation (e.g., roll, pitch, yaw, etc.), an entity type (e.g., classification, etc.).

[0023] In some examples, the step of generating the target multi-resolution voxel space 108 may include associating data associated with static objects (e.g., buildings, trees, leaves, etc.) with the target multi-resolution voxel space 108 while filtering out data associated with dynamic objects (e.g., representing pedestrians, vehicles, etc.). In an alternative implementation, the data point cloud 104 may be output by a recognition pipeline or component with semantic labels added. For example, the data point cloud 104 may be received as part of a sparse object state representation output by a recognition component, details of which are described in U.S. Application No. 16 / 549,694, which is hereby incorporated by reference in its entirety.

[0024] In the present example, the multi-resolution voxel space component 106 can assign the semantically labeled data point cloud 104 to the semantic layer of the target multi-resolution voxel space 108 having corresponding semantic labels (e.g., trees, buildings, pedestrians, and the like). For example, the multi-resolution voxel space component 106 can project the data point cloud 104 into a common reference frame and assign it to an appropriate point cloud associated with the corresponding semantic class. Then, for each point cloud, the multi-resolution voxel space component 106 can assign each data point 104 to a voxel of the voxel grid of the finest resolution of each semantic layer (e.g., a basic voxel grid, etc.). In some specific examples, the multi-resolution voxel space can be a single layer storing a plurality of statistical values including each semantic class of the voxel group.

[0025] Once each of the data point clouds 104 regarding the corresponding cloud is assigned to a voxel, the multi-resolution voxel space generation component 106 can calculate spatial statistics (e.g., spatial mean, covariance, weighting, and / or the number of data points 104 assigned to the voxel, etc.) regarding each voxel of the grid of the finest resolution of the semantic layer. Once the basic voxel grid or the voxel grid of the finest resolution of the semantic layer is completed, the multi-resolution voxel space generation component 106 can iteratively or recursively generate each of the voxel grids of the next coarser resolution for each of the semantic layers. For example, examples of the processes associated with the generation of the voxel space are described in U.S. Patent No. 11,288,861 and U.S. Application No. 16 / 722,771, which are hereby incorporated by reference in their entirety for all purposes.

[0026] Once the target multi-resolution voxel space 108 is generated from the data points 104, the target multi-resolution voxel space 108 can be aligned with a reference multi-resolution voxel space 110 (e.g., a pre-generated multi-resolution voxel space representing a shared scene or physical environment, etc.). For example, in the illustrated example, the multi-resolution voxel space alignment component 112 can generate an alignment 114 between the newly generated target multi-resolution voxel space 108 and the reference multi-resolution voxel space 110, for example, to assist in the localization, object tracking, and / or navigation of an autonomous vehicle with respect to a physical environment. In some examples, to generate an alignment 110 between the target multi-resolution voxel space 108 and the reference multi-resolution voxel space 110, the multi-resolution voxel space alignment component 112 can first select one or more coarse resolutions (e.g., a resolution exceeding a size threshold, etc.) for the individual semantic layers and initiate the determination of an alignment or offset between the voxels of the target multi-resolution voxel space 108 and the reference multi-resolution voxel space 110. In some examples, the multi-resolution voxel space alignment component 112 can utilize odometry, position data, orientation data, trajectory data, or the like to determine an initial alignment for initiating an alignment between the voxel groups of the target multi-resolution voxel space 108 and the reference multi-resolution voxel space 110.

[0027] Next, the multi-resolution voxel space alignment component 112 can iteratively add the next finer resolution to the convergence process, for example, when the system determines an error below an error threshold. In some cases, the error threshold for the next finer resolution can be such that the average error of the voxel group (e.g., the distance between centroids, or other such as described herein) is less than or equal to half the size of the voxels of the current resolution (e.g., the width of the individual voxels used for alignment). In other examples, the error threshold for the next finer resolution can be such that the average error of the voxels is less than or equal to one quarter of the size of the current finest resolution, or the like.

[0028] In some examples, the multi - resolution voxel space alignment component 112 may continue to iterate through the alignment stages until each resolution is added / used. In some cases, the multi - resolution voxel space alignment component 112 may, after all resolutions are added, use all resolutions for a predetermined number of iterations (e.g., 1 iteration, 2 iterations, 3 iterations, 5 iterations, or the like), or until the average error is below a final error threshold, or until the change in the sum of voxel residuals is below a change threshold.

[0029] In some cases, the multi-resolution voxel space alignment component 112 may also evaluate, weight, or otherwise score the eigenvalues of individual voxels during alignment, and use the weights to select eigenvectors and / or voxels for use in alignment 114. For example, as described above, a voxel may have a set of three or more eigenvalues, and the multi-resolution voxel space component 106 may determine the score or weight of an eigenvector or voxel based at least in part on any one of one or more of the eigenvalues. In the current example, it may happen that the smaller the eigenvalue, the higher the score. In some implementations, the system may evaluate individual eigenvalues against one or more predetermined heuristics or thresholds to determine the weights. For example, the multi-resolution voxel space alignment component 112 may determine whether one or more of the eigenvalues for a voxel are below one or more thresholds (such as size thresholds, error thresholds, or the like) when generating a score. In some cases, one or more of the thresholds may be relative to the size of the voxel or the size of the associated resolution. In other implementations, the multi-resolution voxel space alignment component 112 may utilize one or more machine learning models to evaluate and / or score individual eigenvalues. Thereafter, the multi-resolution voxel space alignment component 112 may select a group of voxels associated with higher scores (e.g., based on the number of points, eigenvalues, eigenvectors, etc.) for use in the alignment process. In some examples, the system may select the eigenvalue having the minimum value.

[0030] In some implementations, the multi-resolution voxel space alignment component 112 may also adapt or regress a quality or confidence metric for individual voxel groups using the semantic classes assigned to the voxels. For example, the multi-resolution voxel space alignment component 112 may generate and / or train a noise model for evaluating the quality of the voxel groups to be aligned, based at least in part on the eigenvalues of the voxels, the voxel resolution, the number of points associated with the voxels, and the semantic class of the voxels. The multi-resolution voxel space alignment component 112 may then select or weight the voxel groups for use in the alignment process using the quality or confidence metric. In various examples, a number of multi-resolution voxel spaces may be generated for each such classification, and the processes described herein may be performed for any number of different classifications where the resulting transformations between the voxel spaces are combined. In at least some such examples, such combinations may be based, for example, on the covariance associated with the different classifications.

[0031] Figures 2 through 4 are flow diagrams illustrating exemplary processes associated with the generation of a multi-resolution voxel space as described herein. The process is illustrated as a collection of blocks in a logical flow diagram, and these blocks represent sequences of operations, some or all of which may be implemented in hardware, software, or a combination thereof. In a software context, the blocks represent computer-executable instructions stored on one or more computer-readable media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular abstract data types.

[0032] The order in which the operations are described should not be construed as limiting. Any number of the described blocks can be combined in any order and / or in parallel to implement that process, or an alternative process, and not all of the blocks need to be executed. For purposes of explanation, the processes herein are described with reference to the frameworks, architectures, and environments described in the examples herein, but those processes may be implemented in a variety of other frameworks, architectures, or environments.

[0033] FIG. 2 is another flowchart illustrating an exemplary process 200 associated with the generation of a multi - resolution voxel space, as described herein. As described above, the system can generate an alignment between a target multi - resolution voxel space representing a physical environment and a reference multi - resolution voxel space representing the same physical environment. In some cases, the convergence between a group of voxels in the target multi - resolution voxel space and a group of voxels in the reference multi - resolution voxel space (e.g., generated from one or more pre - scans of the physical environment) can be time - consuming and resource - intensive. As described below, process 200 reduces the time and resources required to achieve convergence or a desired alignment between the two multi - resolution voxel spaces.

[0034] In step 202, the system can receive a first multi - resolution voxel space and a second multi - resolution voxel space. For example, the system can receive two or more misaligned multi - resolution voxel spaces, such as a target and a reference. For example, the two or more multi - resolution voxel spaces can be generated from data captured as part of multiple spins in a lidar system, data from two or more vehicles that supplement data representing the same physical environment as described above.

[0035] In step 204, the system may determine a resulting voxel and a residual representing the resulting voxel, based at least in part on a first voxel associated with a first multiresolution voxel space and a second voxel associated with a second multiresolution voxel space. For example, for a pair of associated voxels (e.g., having the closest centroid), the system may determine various metrics corresponding to a combination (e.g., mean, number of points, covariance, eigenvalues, eigenvectors, etc.) of all associated point clouds, a vector or distance between the mean of the first voxel and the mean of the second voxel, and the like. In at least some examples, the residual associated with the alignment of the first voxel and the second voxel may be determined as the dot product between a vector between the mean of the first voxel and the mean of the second voxel and an individual eigenvalue (such as the smallest eigenvalue). In other words, this residual may be equal to the dot product between the unit vector of the eigenvector of the combined voxels and the result of subtracting the second mean from the first mean.

[0036] In step 206, the system may determine a quality metric for the resulting voxels based at least in part on the voxel resolution, eigenvalues, number of points within the voxel, semantic class, and one or more models. For example, a first voxel layer at a certain resolution may be associated with a building, a second voxel layer at the same resolution may be associated with the ground surface, and / or a third voxel layer at the same resolution may be associated with vegetation. In this example, the system may adapt or regress a matching quality or confidence metric for individual voxels within the voxel layer using the semantic class assigned to the resulting voxels. For example, the system may input, for each voxel, the voxel resolution of the resulting voxel, the eigenvalues of the resulting voxel, the number of points within the resulting voxel, and the semantic class of the resulting voxel into a machine learning model or other pre-trained model that will output the quality metric. For example, the system may generate and / or train a noise model for evaluating the matching quality of voxels based at least in part on the eigenvalues, resolution, number of points associated with the voxel, and semantic class of the voxel. In other cases, the system may utilize a predetermined function whose result or output is the quality metric. In these cases as well, this function may also receive as input the voxel resolution of the resulting voxel, the eigenvalues of the resulting voxel, the number of points within the resulting voxel, and the semantic class of the resulting voxel.

[0037] In step 208, the system may generate an alignment between the first multi-resolution voxel space and the second multi-resolution voxel space based at least in part on the quality metric and the residuals. For example, the system may scale or weight pairs of voxels (e.g., residuals, etc.) based at least in part on the quality metric. In some examples, the system may compare the quality metric to an expected error and utilize the residuals of pairs of voxels having a quality metric below the expected error to generate the alignment. In some examples, this residual may be scaled based at least in part on the quality metric. In such examples, those pairs associated with a lower quality metric will be less relied upon when compared to other pairs having a higher quality metric. In various examples, such a quality metric may be a numerical value from 0 to 1.

[0038] In step 210, the system may determine whether the voxel groups have converged. For example, the system may iterate until the convergence is below a distance threshold, the number of iterations, steps, or levels has been executed, and / or the change in the residuals between a previous iteration and the current iteration is below a residual threshold. As one illustrative example, convergence may be achieved when the error represented by the vector between the averages of pairs of voxels in the target multi-resolution voxel space and the reference multi-resolution voxel space is less than an error threshold. If the voxels have not converged, process 200 may return to step 204. Otherwise, once the voxels have converged, process 200 may proceed to step 212.

[0039] In step 212, the system may apply a scale factor correction to the alignment. For example, the system may determine the scale factor correction to the alignment by randomly sampling (such as by the Monte Carlo method) the residuals of the voxel groups. The system may then apply the scale factor correction to one or more alignments and, in step 214, output the alignment.

[0040] FIG. 3 is an exemplary flow diagram illustrating an exemplary process 200 associated with the alignment of a multi - resolution voxel space, as described herein. As described above, the system can generate an alignment between a target multi - resolution voxel space and a reference multi - resolution voxel space that represents a physical environment. In some cases, the convergence between a group of voxels in the target multi - resolution voxel space and a group of voxels in the reference multi - resolution voxel space (e.g., generated from one or more pre - scans of the physical environment) may take time and resources. As described below, process 200 reduces the time and resources required to achieve convergence, or a desired alignment, between two multi - resolution voxel spaces.

[0041] In step 302, the system can receive the resolution of the multi - resolution voxel space. For example, the system can evaluate or select eigenvalues of groups of voxels for use in the alignment process for each semantic layer and each resolution. As described above, an individual group of voxels can include three eigenvectors that represent the corresponding data associated with a particular voxel or provide a physical shape. These eigenvectors can be used to determine whether a voxel is suitable or preferable for use in the alignment process.

[0042] In step 304, the system determines at least one voxel of a resolution having at least one eigenvalue below a threshold, where the threshold may be related to the size of the voxel (e.g., in any dimension). For example, for each eigenvalue larger than the size threshold, the system may determine that the data in the direction represented by the eigenvalue should not be trusted. However, instead of discarding the entire voxel, the system may determine usability based on each individual direction (e.g., eigenvalue, etc.). Thus, if a voxel has three eigenvalues larger than the size threshold, the entire voxel may be ignored for the alignment process and / or the current iteration of the alignment process. In some cases, the size threshold may be based on the resolution associated with the voxel, a predetermined value, the classification of the associated point cloud, or the like.

[0043] In step 306, the system may associate a weight with each individual voxel to assist in the selection of a group of voxels and / or a group of eigenvalues that will provide usable data for generating an alignment. For example, the system may assign a weight from 0 to 1 where values above a first size threshold may become zero and values below a second size threshold may be assigned a weight of 1. Values between the first size threshold and the second size threshold will be assigned weights from 0 to 1 such as 0.3, 0.5, 0.7, and the like.

[0044] In step 308, the system may align at least one voxel to a second voxel based at least in part on eigenvalues and / or weights. For example, the system may select at least one voxel based on weights. However, the system may use only two of the three eigenvalues of the voxel as input to the alignment process based on the individual weights of each eigenvalue. For example, if a voxel has two eigenvalues equal to 1 and a single eigenvalue with a weight of 0, the system may ignore the single eigenvalue when determining the alignment. In other examples, the system may use one or three of the eigenvalue groups as input to the alignment process. In this way, the voxel groups may be integrated or aligned using higher quality eigenvectors and / or eigenvalues.

[0045] FIG. 4 is an exemplary flow diagram illustrating an exemplary process 400 associated with the generation of a multi - resolution voxel space, as described herein. As described above, the system may generate an alignment between two multi - resolution voxel spaces representing a physical environment. In some cases, the convergence of the alignment between a voxel group of a target multi - resolution voxel space (e.g., generated from a current scan of the physical environment) and a voxel group of a reference multi - resolution voxel space (e.g., a previously generated map of the physical environment, etc.) may take time and resources. As will be described below, process 400 reduces the time and resources required to achieve the convergence of the alignment between two multi - resolution voxel spaces (such as the target multi - resolution voxel space and the reference multi - resolution voxel space).

[0046] In step 402, the system may receive a target multi-resolution voxel space and a reference multi-resolution voxel space. As described herein, the target multi-resolution voxel space may be generated from data captured by an autonomous vehicle operating in a physical environment. For example, the autonomous vehicle may use various sensors associated with the vehicle to capture and / or generate data representing the physical environment. In some cases, the sensor data may include image data, lidar data, point cloud data, environmental data, radar data, sonar data, infrared data, and the like. The system may generate semantic point cloud data from the data representing the physical environment. For example, semantic classification may be associated with various points and classes are separated. For example, the system may segment and classify the data representing the physical environment. For example, the system may utilize one or more machine learning models to segment and classify the data representing the physical environment. In some cases, the segmented and classified data may be stored or organized in a semantic layer (e.g., each layer includes data corresponding to the assigned class). For example, in one specific example, the assignment of semantic classes to a data point cloud is described in U.S. Application No. 15 / 820,245, which is hereby incorporated by reference in its entirety. In some examples, the system may generate a voxel covariance grid for each semantic class at multiple resolutions for at least some (including all) of the semantic classes. For example, for each semantic layer, the system may generate a voxel group at one or more resolutions (e.g., each resolution may have a voxel group of different physical sizes). In some cases, the size of the resolution may be based on a square, such as 25 centimeters, 1 meter, 16 meters, 25 meters, or the like. The reference multi-resolution voxel space may be pre-generated from a prior scan of the physical environment or captured data and provided as a map of the physical environment that can be used by the autonomous vehicle in decision-making and processes.For example, in one specific example of generating a multi-resolution voxel space, it is described in U.S. Application No. 17 / 446,344, which is hereby incorporated by reference in its entirety.

[0047] In step 404, the system may determine a set of resolutions that exceed a resolution threshold. For example, the system may initiate an alignment process by restricting the set of resolutions at which voxel groups will be processed to one or more coarse resolutions. For example, the system may restrict the set of resolutions to a resolution of 25 meters or 16 meters, or a similar value or greater. In some examples, the set of resolutions may be a single coarsest resolution.

[0048] In step 406, the system may align the voxel groups of the set of resolutions to update the alignment between the target multi-resolution voxel space and the reference multi-resolution voxel space. For example, the system may enable aligning the voxels of the set of resolutions, at least in part, based on voxel covariance as described above with respect to FIGS. 2 and 3. In this manner, since only the coarse resolution is used to generate an alignment within a predetermined error threshold, the system and process 400 enable improved efficiency associated with generating the alignment and improved (e.g., faster while consuming fewer resources) convergence with respect to finer resolutions.

[0049] As a non-limiting illustrative example, the system may generate an alignment from pairs of voxels (e.g., corresponding voxel groups of the target multi-resolution voxel space and the reference multi-resolution voxel space, etc.). In some cases, the alignment may be generated by determining the matching residuals between two voxels of the same semantic class (e.g., one from each space, etc.). For example, the system may determine the average of a first voxel (a voxel of the target space) and a second voxel (a voxel of the reference space) and determine a vector between the average of the first voxel and the average of the second voxel.

[0050] Also, the system generates voxels obtained as a result (such as through statistical analysis or the sum of two voxels), and based at least in part on the average of the voxels obtained as a result and a vector, the system can determine eigenvalues for pairs of voxels. Then, the system selects either the smallest eigenvalue or each eigenvalue below a threshold (since the longer the eigenvalue, the greater the error and the lower the accuracy of the representation information), and can generate a scalar value or a residual using the selected eigenvalue and a vector (such as an inner product). In this example, the residual can represent the error between a first voxel and a second voxel in the direction of the eigenvalue / eigenvector. In this example, the residual may be calculated for each of the three eigenvalues, and as a result, three residuals representing the error in three directions are generated, and it should be understood that these residuals can be used to align pairs of voxels, thereby assisting in the generation of an alignment between a target multi-resolution voxel space and a reference multi-resolution voxel space.

[0051] Next, the system can compare the residual to an expected error determined as the output of a predetermined function that utilizes the number of points within a voxel group, the semantic class of the voxel group, the resolution of the voxel group, and the eigenvalues, as described above with respect to FIG. 3. In this example, if the residual is below the expected error, the system can determine rotations and translations (such as through a least squares operation) between a first voxel and a second voxel that are utilized to update the alignment.

[0052] In step 408, the system may determine whether the average residual of the voxel group associated with the set of resolutions is less than or equal to the error threshold. In some cases, the error threshold may be half the size of the finest resolution within the set of resolutions. For example, if the finest resolution of the set of resolutions is 25 meters, the system may repeatedly execute the alignment step until the average residual is less than 12.5 meters. In other examples, the error threshold for the next finer resolution may be one-fourth of the size of the current finest resolution, or an average residual less than or equal to a similar value.

[0053] If the average residual of the voxel group is not less than or equal to the error threshold, process 400 returns to step 406 and performs another iteration to improve the alignment by continuing to align the voxel group of the set of resolutions. However, if the average residual of the voxel group is less than or equal to the error threshold, process 400 proceeds to step 410.

[0054] In step 410, the system may determine whether an even finer resolution is available. If no further resolutions are available, the system may output the alignment between the target multi-resolution voxel space and the reference multi-resolution voxel space in step 412. In some examples, it should be understood that process 400 may pre-execute one or more iterations of step 406 (such as until the final error threshold is reached or exceeded and / or until a predetermined number of iterations is exhausted) once the finest resolution is added to the set of resolutions.

[0055] Otherwise, process 400 proceeds to step 414. In step 414, the system may add the next finer resolution to the set of resolutions and update the error threshold. For example, the error threshold may be reduced to a value proportional to the size of the next finer resolution (e.g., half the size of the next finer resolution). Once the next finer resolution is added to the set of resolutions and the error threshold is updated, process 400 returns to step 410 and the system may re-align the set of resolution voxels as described above.

[0056] FIG. 5 is a block diagram showing an exemplary system 500 for implementing a multi-resolution voxel space alignment system as described herein. In this embodiment, system 500 may be an autonomous vehicle 502 that includes a vehicle computing device 504, one or more sensor systems 506, one or more communication connections 508, and one or more drive systems 510.

[0057] The vehicle computing device 504 may include one or more processors 512 (or processing resources) and a computer-readable medium 514 communicatively coupled to the one or more processors 512. In the illustrated example, vehicle 502 is an autonomous vehicle, but vehicle 502 can be any other type of vehicle or any other system (e.g., a robotic system, a camera-enabled smartphone, etc.). In the illustrated example, the computer-readable medium 514 of the vehicle computing device 504 stores a multi-resolution voxel space component 516, a planning component 518, and a prediction component 520, as well as other components 522 associated with the autonomous vehicle. Additionally, the computer-readable medium 514 may store sensor data 524 and a multi-resolution voxel space 526. It should be understood that in some implementations, the systems and data stored on the computer-readable medium may be additionally or alternatively accessible from vehicle 502 (e.g., stored on another computer-readable medium remote from vehicle 502 or otherwise accessible from that computer-readable medium).

[0058] As described above, the multi-resolution voxel space generation component 516 may generate a multi-resolution voxel space, and the multi-resolution voxel space component 516 may output an alignment between two or more multi-resolution voxel spaces, as described above.

[0059] In some implementations, the prediction component 520 may be configured to estimate current properties or states and / or predict future properties or states of an object (e.g., a vehicle, a pedestrian, an animal, etc.), such as pose, speed, trajectory, velocity, yaw, yaw rate, roll, roll rate, pitch, pitch rate, position, acceleration, or other properties, based at least in part on the multi-resolution voxel space 526 output by the multi-resolution voxel space component 516.

[0060] In addition, vehicle 502 can include one or more communication connections 508 that enable communication between vehicle 502 and one or more other local or remote computing devices. For example, communication connection 508 can facilitate communication with other local computing devices on vehicle 502 and / or drive system 510. Also, communication connection 508 can enable vehicle 502 to communicate with other nearby computing devices (e.g., other nearby vehicles, traffic signals, etc.). Further, communication connection 508 can enable vehicle 502 to communicate with remotely located remote control computing devices or other remote services.

[0061] Communication connection 508 can include a physical interface and / or a logical interface for connecting vehicle computing device 504 to another computing device (e.g., computing device 530, etc.) and / or a network such as network 528. For example, communication connection 508 can enable Wi-Fi-based communication via a frequency defined by the IEEE802.11 standard, short-range radio frequencies such as Bluetooth (registered trademark), cellular phone communication (e.g., 2G, 3G, 4G, 4G LTE, 5G, etc.), or any suitable wired or wireless communication protocol that enables each computing device to interface with other computing devices. In some examples, communication connection 508 of vehicle 502 can transmit or send multi-resolution voxel space 526 to computing device 530.

[0062] In at least one example, the sensor system 506 can include lidar sensors, radar sensors, ultrasonic transducers, sonar sensors, position sensors (e.g., GPS, compass, etc.), inertial sensors (e.g., inertial measurement unit (IMU), accelerometer, magnetometer, gyroscope, etc.), cameras (e.g., RGB, IR, intensity, depth, time-of-flight, etc.), microphones, wheel encoders, environmental sensors (e.g., temperature sensors, humidity sensors, light sensors, pressure sensors, etc.), and one or more time-of-flight (ToF) sensors, etc. The sensor system 506 can include multiple instances for each of these sensors or other types of sensors. For example, the lidar sensors can include individual lidar sensors disposed at the corners, front, rear, sides, and / or top of the vehicle 502. As another example, the camera sensors can include multiple cameras disposed at various positions with respect to the exterior and / or interior of the vehicle 502. The sensor system 506 can provide inputs to the vehicle computing device 504. Additionally, or alternatively, the sensor system 506 can transmit sensor data to one or more computing devices 530 at a specific frequency, after a predetermined time period has elapsed, in near real-time, etc., via one or more networks 528.

[0063] In at least one example, vehicle 502 can include one or more drive systems 510. In some examples, vehicle 502 can have a single drive system 510. In at least one example, if vehicle 502 has multiple drive systems 510, the individual drive systems 510 can be disposed at opposing ends of vehicle 502 (e.g., the front and rear, etc.). In at least one example, drive system 510 can include one or more sensor-systems 506 as described above and can detect the situation around drive system 510 and / or vehicle 502. By way of non-limiting example, sensor-system 506 can include one or more wheel encoders (e.g., rotary encoders, etc.) for sensing the rotation of the wheels of the drive system, inertial sensors- (e.g., inertial measurement units, accelerometers, gyroscopes, magnetometers, etc.) for measuring the orientation and acceleration of the drive system, cameras or other image sensors-, ultrasonic sensors- for acoustically detecting objects around the drive system, lidar sensors-, radar sensors-, etc. Some sensors- such as wheel encoders can be specific to drive system 510. In some cases, sensor-system 506 on drive system 510 can overlap or complement the corresponding systems of vehicle 502.

[0064] In at least one example, the components described herein can process sensor data 524 and can transmit their respective outputs to one or more computing devices 530 via one or more networks 528, as described above. In at least one example, the components described herein can transmit their respective outputs to one or more computing devices 530 at a particular frequency, after a predetermined time period has elapsed, in near real-time, etc.

[0065] In some examples, vehicle 502 can send sensor data to one or more computing devices 530 via network 528. In some examples, vehicle 502 can send raw sensor data 524 or processed multi - resolution voxel space 526 to computing device 530. In other examples, vehicle 502 can send processed sensor data 524 and / or a representation of the sensor data (e.g., an object recognition track, etc.) to computing device 530. In some examples, vehicle 502 can send sensor data 524 to computing device 530 at a specific frequency, after a predetermined time period has elapsed, in near real - time, and so on. In some cases, vehicle 502 can send sensor data (raw or processed) to computing device 530.

[0066] Computing system 530 can include a processor 532, a multi - resolution voxel space component 536, and a computer - readable medium 534 that stores other components 538, sensor data 540, and multi - resolution voxel space 542 received from vehicle 502. In some examples, the multi - resolution voxel space component 536 can be configured to generate a multi - resolution voxel space 542 or align multi - resolution voxel spaces 542 generated from data captured by multiple vehicles 502 to form a more complete scene of various physical environments and / or connect various scenes together as an extended physical environment of signals. In some cases, the multi - resolution voxel space component 536 can be configured to generate one or more models from sensor data 524 that can be used for machine learning and / or future code testing.

[0067] The processor 512 of the vehicle 502 and the processor 532 of the computing device 530 can be any suitable processors capable of processing data and executing instructions for performing operations as described herein. By way of non-limiting example, processors 512 and 532 can include one or more central processing units (CPUs), graphics processing units (GPUs), or any other device or portion of a device that processes electronic data and can transform that electronic data into other electronic data that can be stored in registers and / or a computer-readable medium. In some examples, integrated circuits (e.g., ASICs, etc.), gate arrays (e.g., FPGAs, etc.), and other hardware devices can also be considered processors insofar as they are configured to implement encoded instructions.

[0068] The computer-readable media 514 and 534 are examples of non-transitory computer-readable media. The computer-readable media 514 and 534 can store an operating system and one or more software applications, instructions, programs, and / or data for implementing the functions attributable to the methods and various systems described herein. In various implementations, the computer-readable media can be implemented using any suitable computer-readable media technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash-type memory, or any other type of computer-readable media capable of storing information. The architectures, systems, and individual elements described herein can include many other logical, programmatic, and physical components, and those shown in the accompanying figures are merely examples relevant to the description herein.

[0069] As can be understood, the components described herein are described as being divided for illustrative purposes. However, the operations performed by the various components can be combined or performed in any other component.

[0070] Although FIG. 5 is illustrated as a distributed system, it should be noted that, in an alternative, the components of vehicle 502 can be associated with computing device 530 and / or the components of computing device 530 can be associated with vehicle 502. That is, vehicle 502 can perform one or more of the functions associated with computing device 530 and vice versa.

[0071] FIG. 6 is an illustration 600 showing exemplary resolutions of a multi - resolution voxel space 602 compared to a representation 604 of captured data, as described herein. In this example, the multi - resolution voxel space 602 includes a number of layers or resolutions generally indicated by 602(A) through 602(C) and semantic layers generally indicated by 606(A) through 606(C). For example, in this example, the voxel group of layer 606(A) corresponds to leaves and is represented as a shaded voxel group with a dark outline, the voxel group of layer 606(B) corresponds to the ground surface and is represented as a non - shaded voxel group with a bright outline, and the voxel group of layer 606(C) corresponds to buildings and stationary objects and is represented as a non - shaded voxel group with a dark outline. As shown, both the multi - resolution voxel space 602 and the representation 604 correspond to physical locations or spaces in the real world.

[0072] [Article Example] A. The system comprises: a system. That is, one or more processors. And one or more non-transitory computer-readable media storing instructions executable by the one or more processors. When these instructions are executed, they cause the system to perform operations including: Receiving a target multi-resolution voxel space, where the target multi-resolution voxel space represents a discrete volumetric portion of the physical environment and includes a first plurality of voxels defined by a first resolution and a second resolution, and the first resolution is coarser than the second resolution. Receiving a reference multi-resolution voxel space, where the reference multi-resolution voxel space represents a discrete volumetric portion of the physical environment and includes a second plurality of voxels defined by the first resolution and the second resolution. Determining a resulting voxel as a first result based at least in part on a first voxel of the first resolution of the target multi-resolution voxel space and a second voxel of the first resolution of the reference multi-resolution voxel space. Determining a first quality metric of the resulting voxel as a first result based at least in part on the number of points associated with the resulting voxel as a first result, the semantic class associated with the resulting voxel as a first result, the eigenvalue associated with the resulting voxel as a first result, or one or more of the first resolutions. Determining a first residual based at least in part on the resulting voxel as a first result. Determining an alignment between the target voxel space and the reference multi-resolution voxel space based at least in part on the first quality metric and the first residual. And performing an operation of the autonomous vehicle based at least in part on the alignment.

[0073] The system of item B.A, wherein the operation further includes the following operations in response to determining that the average residual of the alignment is below an error threshold. That is, determining a resulting voxel as a second result based at least in part on a third voxel at a second resolution of the target multi-resolution voxel space and a fourth voxel at a second resolution of the reference multi-resolution voxel space. Determining a second quality metric of the resulting voxel as a second result based at least in part on the number of points associated with the resulting voxel as a second result, the semantic class associated with the resulting voxel as a second result, the eigenvalue associated with the resulting voxel as a second result, or one or more of the second resolutions. Determining a second residual based at least in part on the resulting voxel as a second result. Determining a final alignment based at least in part on the second quality metric and the second residual. And, performing an operation of an autonomous vehicle based at least in part on the final alignment.

[0074] The system of item C.A, wherein the first voxel and the second voxel are determined based on having a minimum distance between their centers of gravity.

[0075] The system of item D.A, wherein determining the alignment further includes scaling the first residual by a first quality metric.

[0076] The system of item E.A, wherein determining the first residual includes the following. That is, determining a set of data associated with the first voxel and the second voxel. Determining an eigenvector of the set of data. Determining a vector between the average of the data associated with the first voxel and the average of the data associated with the second voxel. And, determining the inner product between the eigenvector and that vector as the residual.

[0077] A system of the F.A item, wherein the first voxel includes a first eigenvalue, a second eigenvalue, and a third eigenvalue, and its operation further includes the following. That is, determining that at least one of the first eigenvalue, the second eigenvalue, or the third eigenvalue is less than or equal to a size threshold. Determining a first weight based at least in part on the size of the first eigenvalue. Determining a second weight based at least in part on the size of the second eigenvalue. Determining a third weight based at least in part on the size of the third eigenvalue. And, before determining the resulting voxel as the first result, applying the first weight to the first eigenvalue, applying the second weight to the second eigenvalue, and applying the third weight to the third eigenvalue.

[0078] G. One or more non-transitory computer-readable media that, when executed, store instructions that cause one or more processors to perform operations including the following. That is, determining a quality metric of a voxel associated with a first multi-resolution voxel space and a second multi-resolution voxel space based at least in part on one or more of the number of points associated with the voxel, the semantic class associated with the voxel, the eigenvalue associated with the voxel, or the resolution associated with the voxel. Determining a residual associated with the voxel. And, determining an alignment between the first multi-resolution voxel space and the second multi-resolution voxel space based at least in part on the quality metric and the residual.

[0079] H. A non-transitory computer-readable medium of item G, which is as follows. The first multi-resolution voxel space represents a discrete volume portion of the physical environment and includes a first plurality of voxels defined by a first resolution and a second resolution, and the first resolution is coarser than the second resolution. And, the second multi-resolution voxel space represents a discrete volume portion of the physical environment and includes a second plurality of voxels defined by the first resolution and the second resolution, and its resolution is the first resolution.

[0080] A non-transitory computer-readable medium of item I.H, wherein the voxel is a first voxel, associated with a first resolution, and the operation further includes the following. That is, in response to determining that the average residual of the alignment is below a threshold, the number of points associated with a second voxel, the semantic class associated with the second voxel, the eigenvalue associated with the second voxel, or one or more of the resolutions associated with the voxel and the second voxel of the second resolution, determining a second quality metric in the second voxel associated with the first multi-resolution voxel space and the second multi-resolution voxel space. Determining a second residual associated with the second voxel. And generating an updated alignment based at least in part on the second quality metric, the second residual, and the alignment.

[0081] A non-transitory computer-readable medium of item J.I, including one or more of controlling a vehicle based at least in part on the updated alignment or creating a map based at least in part on the updated alignment.

[0082] A non-transitory computer-readable medium of item K.G, wherein the voxel is a static combination of first data associated with a voxel of a first multi-resolution voxel space and second data associated with a voxel of a second multi-resolution voxel space.

[0083] A non-transitory computer-readable medium of item L.K, wherein determining the residual further includes determining the inner product of the eigenvector of the voxel and a vector indicating the separation between the voxel of the first multi-resolution voxel space and the voxel of the second multi-resolution voxel space.

[0084] A non-transitory computer-readable medium of item M.G, wherein the voxel is a first voxel, and determining a first residual includes the following. That is, determining a set of data associated with the first voxel and a second voxel. Determining an eigenvector of the set of data. Determining a vector between the average of the data associated with the first voxel and the average of the data associated with the second voxel. And determining, as the residual, the inner product between the eigenvector and that vector.

[0085] A method, including the following. That is, determining a quality metric of voxels associated with a first multi-resolution voxel space and a second multi-resolution voxel space, at least partially based on one or more of the number of points associated with the voxels, the semantic class associated with the voxels, the eigenvalues associated with the voxels, or the resolution associated with the voxels. Determining a residual associated with the voxels. And determining an alignment between the first multi-resolution voxel space and the second multi-resolution voxel space, at least partially based on the quality metric and the residual.

[0086] A method of item O.N, which is as follows. The first multi-resolution voxel space represents a discrete volume portion of a physical environment and includes a first plurality of voxels defined by a first resolution and a second resolution, and the first resolution is coarser than the second resolution. And the second multi-resolution voxel space represents a discrete volume portion of a physical environment and includes a second plurality of voxels defined by the first resolution and the second resolution, and its resolution is the first resolution.

[0087] A method of P.N terms, which is as follows. The voxel is the first voxel, associated with the first resolution, and the method further comprises the following. That is, in response to determining that the average residual of the alignment is below a threshold, the number of points associated with the second voxel, the semantic class associated with the second voxel, the eigenvalue associated with the second voxel, or one or more of the resolutions associated with the voxel and the second voxel of the second resolution, determining a second quality metric of the second voxel associated with the first multi-resolution voxel space and the second multi-resolution voxel space. Determining a second residual associated with the second voxel. And generating an updated alignment based at least in part on the second quality metric, the second residual, and the alignment.

[0088] A method of Q.P terms, further comprising outputting the updated alignment in response to determining that the updated alignment reaches a convergence threshold or exceeds a convergence value.

[0089] A method of R.P terms, wherein the voxel is a static combination of first data associated with a voxel in the first multi-resolution voxel space and second data associated with a voxel in the second multi-resolution voxel space.

[0090] A method of S.R terms, wherein determining the residual further includes determining the inner product of the unit vector of the voxel and the value obtained by subtracting the average of the voxels in the second multi-resolution voxel space from the average of the voxels in the first multi-resolution voxel space.

[0091] A method of T.N terms, wherein the operation further includes performing an operation of an autonomous vehicle based at least in part on the alignment.

[0092] The exemplary clauses described above are described with respect to one particular implementation, but in the context of this specification, it should be understood that the content of the exemplary clauses can also be implemented via methods, devices, systems, computer-readable media, and / or other implementations. Further, any of Examples A through T may be implemented alone or in combination with one or more of the other Examples A through T.

[0093] [Conclusion] As can be understood, the components described in this specification are described as being divided for illustrative purposes. However, the operations performed by the various components can be combined or performed in any other component. It should also be understood that the components or steps described with respect to one example or implementation can be used in conjunction with the components or steps of other examples. For example, the components and instructions of FIG. 5 can utilize the processes and flows of FIGS. 1 through 4.

[0094] Although one or more examples of the techniques described in this specification are described, various changes, additions, substitutions, and equivalents thereof are included within the scope of the techniques described herein.

[0095] In the description of the examples, reference is made to the accompanying drawings which form a part of this specification, and the accompanying drawings illustrate certain examples of the claimed subject matter by way of illustration. It should be understood that other examples may be used and that changes or modifications such as structural changes may be made. Such examples, changes, or modifications are not necessarily departing from the scope of the claimed subject matter. The steps in this specification can be presented in a certain order, but in some cases, the order can be changed, whereby at various times or in a different order, certain inputs can be provided without changing the functionality of the described system and method. Also, the disclosed procedures can be executed in a different order. Further, the various calculations described in this specification need not be executed in the disclosed order, and other examples using alternative orders of calculation can be readily implemented. Not only can the order be changed, but in some cases, those calculations can also be broken down into sub-calculations having the same result.

Claims

1. It is a method, A step of determining a quality metric for a voxel associated with a first multi-resolution voxel space and a second multi-resolution voxel space, based at least partially on one or more of the number of points associated with a voxel, the semantic class associated with the voxel, the eigenvalue associated with the voxel, or the resolution associated with the voxel. A step of determining the residual associated with the voxel, A step of determining the alignment between the first multi-resolution voxel space and the second multi-resolution voxel space, based at least in part on the quality metric and the residuals. A method that includes [a certain feature].

2. The first multi-resolution voxel space represents a discrete volume portion of the physical environment and includes a first plurality of voxels defined by a first resolution and a second resolution, wherein the first resolution is coarser than the second resolution. The second multi-resolution voxel space represents a discrete volume portion of the physical environment and includes a second plurality of voxels defined by the first resolution and the second resolution. The resolution is the first resolution, The method according to claim 1.

3. The voxel is a first voxel, associated with the first resolution, and the method is In response to determining that the mean residual of the alignment is below a threshold, A step of determining a second quality metric for the second voxel associated with the first multi-resolution voxel space and the second multi-resolution voxel space, based at least partially on one or more of the number of points associated with the second voxel, the semantic class associated with the second voxel, the eigenvalues ​​associated with the second voxel, or the resolution associated with the voxel and the second resolution of the second voxel; A step of determining a second residual associated with the second voxel, A step of generating an updated alignment based at least partially on the second quality metric, the second residual, and the alignment. The method according to claim 2, further comprising:

4. A step of controlling the vehicle based at least in part on the updated alignment, or A step of creating a map based at least partially on the updated alignment. The method according to claim 3, further comprising one or more of the above.

5. The method according to claim 1, wherein the voxel is a statistical combination of first data associated with a voxel in the first multi-resolution voxel space and second data associated with a voxel in the second multi-resolution voxel space.

6. The method according to claim 5, wherein the step of determining the residual further includes determining the inner product of the eigenvalue of the voxel and a vector indicating the separation between the voxel in the first multi-resolution voxel space and the voxel in the second multi-resolution voxel space.

7. The voxel is the first voxel, and the step of determining the residual is, The steps include determining the dataset associated with the first voxel and the second voxel, The steps include determining the eigenvalues ​​of the aforementioned dataset, A step of determining a vector between the mean of the data associated with the first voxel and the mean of the data associated with the second voxel, The residual is determined by the step of determining the inner product between the eigenvalue and the vector. The method according to claim 3, including the method described in claim 3.

8. A computer program product comprising encoded instructions that, when executed on a computer, implements the method according to any one of claims 1 to 7.

9. It is a system, One or more processors, When executed, one or more of the processors will Determining the quality metrics of the voxels associated with a first multi-resolution voxel space and a second multi-resolution voxel space based at least partially on one or more of the number of points associated with the voxels, the semantic class associated with the voxels, the eigenvalues ​​associated with the voxels, or the resolution associated with the voxels, Determining the residual associated with the voxel, Determining the alignment between the first multi-resolution voxel space and the second multi-resolution voxel space based at least in part on the quality metric and the residuals. One or more non-temporary computer-readable media that store instructions for performing an operation including the above, A system equipped with these features.

10. The first multi-resolution voxel space represents a discrete volume portion of the physical environment and includes a first plurality of voxels defined by a first resolution and a second resolution, wherein the first resolution is coarser than the second resolution. The second multi-resolution voxel space represents a discrete volume portion of the physical environment and includes a second plurality of voxels defined by the first resolution and the second resolution. The resolution is the first resolution, The system according to claim 9.

11. The voxel is a first voxel, and is associated with the first resolution. The aforementioned operation is, In response to determining that the mean residual of the alignment is below a threshold, Determining a second quality metric for the second voxel associated with the first multi-resolution voxel space and the second multi-resolution voxel space, based at least partially on one or more of the number of points associated with the second voxel, the semantic class associated with the second voxel, the eigenvalues ​​associated with the second voxel, or the resolution associated with the voxel and the second resolution of the second voxel, Determining the second residual associated with the second voxel, To generate an updated alignment based at least partially on the second quality metric, the second residual, and the alignment. Further including, The system according to claim 10.

12. The system according to claim 11, further comprising outputting the updated alignment in response to determining whether the updated alignment reaches or exceeds a convergence threshold.

13. The system according to claim 12, wherein the voxel is a statistical combination of first data associated with a voxel in the first multi-resolution voxel space and second data associated with a voxel in the second multi-resolution voxel space.

14. The system according to claim 9, further comprising the step of determining the residual by determining the dot product of the unit vector of the voxel and the value obtained by subtracting the average of the voxel in the second multi-resolution voxel space from the average of the voxel in the first multi-resolution voxel space.

15. The system according to any one of claims 9 to 14, wherein the operation further comprises performing an operation of the autonomous vehicle based at least in part on the alignment.