Point diagram registration using semantic information
By introducing semantic information and the ICP algorithm into the object detection and tracking system, the accuracy and speed problems of point map registration in the existing technology are solved, and more efficient object localization and navigation support is achieved.
Patent Information
- Application Number
- CN202380095980.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-24
- Publication Date
- 2025-11-07
AI Technical Summary
Existing object detection and tracking systems fail to effectively utilize semantic information for point map registration, making it difficult to balance detection speed and accuracy in computationally intensive environments.
By using semantic information for point map registration, utilizing machine learning models such as deep neural networks to detect object features, and combining the ICP algorithm for alignment transformation between point maps, a more accurate 3D environment model is generated.
It improves the accuracy and speed of object detection and tracking systems, especially in scenarios with limited computing resources, providing more accurate object localization and navigation decision support.
Smart Images

Figure CN120917482A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Aspects of the disclosure generally relate to object detection and tracking. In some implementations, systems and techniques for performing reference map registration using semantic information are described. BACKGROUND
[0002] Object detection and tracking can be used to identify objects (e.g., from a digital image or from a video frame of a video clip) and track the objects over time. Object detection and tracking can be used in different fields, including transportation, video analytics, security systems, robotics, aviation, etc. In some fields, tracking an object can determine the positioning of other objects (e.g., target objects) in an environment, such that the tracking object can navigate the environment accurately. To make accurate motion and trajectory planning decisions, the tracking object can also have the ability to estimate various target object properties such as pose (e.g., including position and orientation) and size. SUMMARY
[0003] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary has the sole purpose to present certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0004] Systems and techniques for performing point map registration using semantic information are disclosed. According to at least one example, a method for performing point map registration (using a preamble) is provided. The method includes obtaining an image comprising a scene; determining a first point map representing the scene based on the image, wherein the first point map comprises a first plurality of three-dimensional (3D) points representing the scene; determining a first plurality of point groupings, wherein each point grouping of the first plurality of point groupings is determined based on proximity of pairs of points of the first plurality of 3D points, wherein each point grouping comprises two or more 3D points from the first plurality of 3D points; obtaining a second point map representing the scene, wherein the second point map comprises a second plurality of 3D points representing the scene; obtaining a second plurality of point groupings, wherein each point grouping of the second plurality of point groupings comprises a plurality of 3D points from the second plurality of 3D points; determining a correspondence between each point grouping of the first plurality of point groupings and each point grouping of the second plurality of point groupings; and determining an alignment transformation for aligning the first plurality of 3D points and the second plurality of 3D points based on the correspondence between each point grouping of the first plurality of point groupings and each point grouping of the second plurality of point groupings.
[0005] In another example, an apparatus for performing point graph registration is provided, the apparatus comprising at least one memory and at least one processor coupled to the at least one memory. The at least one processor is configured to: obtain an image comprising a scene; determine a first point graph representing the scene based on the image, wherein the first point graph comprises a first plurality of 3D points representing the scene; determine a first plurality of point groupings, wherein each point grouping of the first plurality of point groupings is determined based on a proximity of pairs of points of the first plurality of 3D points, wherein each point grouping comprises two or more 3D points from the first plurality of 3D points; obtain a second point graph representing the scene, wherein the second point graph comprises a second plurality of 3D points representing the scene; obtain a second plurality of point groupings, wherein each point grouping of the second plurality of point groupings comprises a plurality of 3D points from the second plurality of 3D points; determine a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings; and determine an alignment transform for aligning the first plurality of 3D points and the second plurality of 3D points based on the correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings.
[0006] In another example, a non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: obtain an image comprising a scene; determine a first point graph representing the scene based on the image, wherein the first point graph comprises a first plurality of 3D points representing the scene; determine a first plurality of point groupings, wherein each point grouping of the first plurality of point groupings is determined based on a proximity of pairs of points of the first plurality of 3D points, wherein each point grouping comprises two or more 3D points from the first plurality of 3D points; obtain a second point graph representing the scene, wherein the second point graph comprises a second plurality of 3D points representing the scene; obtain a second plurality of point groupings, wherein each point grouping of the second plurality of point groupings comprises a plurality of 3D points from the second plurality of 3D points; determine a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings; and determine an alignment transform for aligning the first plurality of 3D points and the second plurality of 3D points based on the correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings.
[0007] In another example, an apparatus for performing point graph registration is provided. The apparatus comprises: means for obtaining an image comprising a scene; determining a first point graph representing the scene based on the image, wherein the first point graph comprises a first plurality of 3D points representing the scene; means for determining a first plurality of point groupings, wherein each point grouping of the first plurality of point groupings is determined based on proximity of pairs of points of the first plurality of 3D points, wherein each point grouping comprises two or more 3D points from the first plurality of 3D points; means for obtaining a second point graph representing the scene, wherein the second point graph comprises a second plurality of 3D points representing the scene; means for obtaining a second plurality of point groupings, wherein each point grouping of the second plurality of point groupings comprises a plurality of 3D points from the second plurality of 3D points; means for determining a correspondence between each point grouping of the first plurality of point groupings and each point grouping of the second plurality of point groupings; and determining an alignment transformation for aligning the first plurality of 3D points and the second plurality of 3D points based on the correspondence between each point grouping of the first plurality of point groupings and each point grouping of the second plurality of point groupings.
[0008] In some aspects, one or more of the apparatuses described herein is, is part of, or includes a vehicle or a computing device or system of a vehicle, a mobile device (e.g., a mobile telephone or so-called “smart phone” or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other device. In some aspects, an apparatus includes one or more cameras for capturing one or more images. In some aspects, the apparatus includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the apparatus can include one or more sensors. In some cases, the one or more sensors can be used to determine a location and / or pose of the apparatus, a state of the apparatus, and / or for other purposes.
[0009] Other objects and advantages associated with the aspects disclosed herein will be apparent to those skilled in the art based on the accompanying drawings and detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The accompanying drawings are presented to aid in the description of various aspects of the disclosure and are provided solely for illustration of various aspects.
[0011] Figure 1 is an image illustrating a plurality of vehicles traveling on a road in accordance with some examples;
[0012] Figure 2A is a diagram illustrating an example point graph of a road in accordance with some examples;
[0013] Figure 2B and Figure 2C is a diagram illustrating example point graph registration according to some examples;
[0014] Figure 3 is a block diagram illustrating an example of a point graph registration system according to some examples;
[0015] Figure 4A and Figure 4B is a diagram illustrating an example point graph according to some examples;
[0016] Figure 5A , Figure 5B and Figure 5C is a diagram illustrating an example dynamic time warping calculation according to some examples;
[0017] Figure 6 is a diagram illustrating point grouping and correspondence between individual points in two point graphs according to some examples;
[0018] Figure 7 is a diagram illustrating an example registration between point graphs according to some examples;
[0019] Figure 8 is a flowchart diagram illustrating an example of a process for performing object detection and tracking using the techniques described herein according to some examples;
[0020] Figure 9 is a block diagram illustrating an example of a deep neural network according to some examples;
[0021] Figure 10 is a diagram illustrating an example of a Cifar-10 neural network according to some examples;
[0022] Figures 11A-11C is a diagram illustrating an example of a single-shot object detector according to some examples;
[0023] Figures 12A-12C is a diagram illustrating an example of a You Only Look Once (YOLO) detector according to some examples; and
[0024] Figure 13 is a block diagram of an example computing device that can be used to implement some aspects of the techniques described herein according to some examples. DETAILED DESCRIPTION
[0025] For illustrative purposes, certain aspects of the present disclosure are provided below. Alternate aspects can be devised without departing from the scope of the present disclosure. Additionally, well-known elements of the disclosure, related to those components that are known, not specifically described in detail in some instances to not obscure the related details and keep the discussion relevant for the aspects of the disclosure. Some aspects described herein can be applied independently of one another, and some of them can be combined in different ways. In the following description, for purposes of explanation, specific details are set forth to provide a thorough understanding of the aspects of the application. However, it will be apparent to one skilled in the art that various aspects can be practiced without these specific details. The drawings and descriptions are not intended to be limiting.
[0026] The following description provides for example aspects only and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the following description of the example aspects will provide those skilled in the art with an enabling description of how the example aspects can be implemented. It should be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.
[0027] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term “aspects of the disclosure” does not require that all aspects of the disclosure include the discussed feature, advantage or mode of operation.
[0028] Object detection can be used to detect or identify objects in an image or frame. Object tracking can be performed to track detected objects over time. For example, an image of an object can be acquired, and object detection can be performed on the image to detect one or more objects in the image. In some cases, a three-dimensional (3D) point map can include a 3D representation of an object in an image.
[0029] Object detection and tracking can be used in driving systems, video analytics, security systems, robotic systems, aerial systems, extended reality (XR) systems (e.g., augmented reality (AR) systems, virtual reality (VR) systems, mixed reality (MR) systems, etc.), and the like. In such systems, an object (referred to as a tracking object) that tracks other objects (referred to as target objects) in an environment can determine the position and / or size of the other objects. Determining the position and / or size of target objects in an environment allows the tracking object to accurately navigate the environment by making intelligent motion planning and / or trajectory planning decisions.
[0030] In some cases, machine learning models (e.g., deep neural networks) can be used to perform object detection and localization in some cases. Machine learning based object detection can be computationally intensive, can be difficult to implement in cases where detection speed is a high priority, and present other challenges. For example, machine learning based object detection can be computationally intensive because they are typically run on entire images and run at various scales (implicitly or explicitly) to capture target objects (e.g., target vehicles) at different distances from a tracked object (e.g., a tracked or autonomous vehicle). Examples of multiple scales that can be considered by a neural network based object detector are shown and described below with respect to Figures 11A-11C and Figures 12A-12C Examples of multiple scales that can be considered by a neural network based object detector are shown and described below.
[0031] In some cases, an object detection and tracking system can track (e.g., using an object tracker) a location of a tracked object (e.g., a tracked vehicle) relative to a reference map. Object tracking can be performed across multiple consecutive images (or frames), such as images (or frames) received by a tracked object, e.g., captured by image capture devices such as cameras, light detection and ranging (LiDAR) sensors, and / or radar sensors of the tracked object. Object tracking can also be performed using data (e.g., images or frames) from multiple different sensors. For example, object tracking can be performed by a tracking system that analyzes data from both LiDAR sensors and image capture devices. In some cases, a two-dimensional (2D) representation of an object captured by a sensor such as an image capture device, LiDAR, and / or radar can be converted to a 3D representation of an environment surrounding the tracked object. In some cases, the 3D representation of the environment can be a point map. However, existing systems do not consider semantic information of the environment.
[0032] Systems, apparatuses, processes (methods), and computer-readable media (collectively referred to as “systems and techniques”) for performing road registration using semantic information are described herein. For example, the systems and techniques provide solutions for improving registration between a reference point map (e.g., from an accurate road map) and a point map (also referred to herein as a point cloud) generated based on captured images (e.g., by a tracking and detection system). As used herein, the term registration refers to a process of determining a spatial transformation (e.g., a rotation and / or a translation) that aligns two (or, in some cases, more) point maps. The systems and techniques described herein can be applied to any scenario, such as scenarios that require fast and / or accurate detection, scenarios where computational resources are limited, etc.
[0033] In some aspects, a detection and tracking system that tracks an object (e.g., a vehicle) can receive or obtain images that include a target object (e.g., a road). Some detection and tracking systems can also generate a 3D model of the environment surrounding the tracked object, including a 3D point map of the road (or other surface). In some cases, features of the road such as lane markings (e.g., lane lines) can be represented as individual points in the 3D point map of the environment surrounding the tracked object. A point map registration system can perform registration between a point map generated based on images captured by the object tracking and detection system and a representation of a reference point map. In some cases, accurate registration can provide a better understanding of the positioning of the tracked object relative to the road (e.g., lane positioning). In some implementations, the systems and techniques described herein can be used for accurate real-time vehicle positioning tracking.
[0034] In some cases, semantic information can be used to improve registration between point maps. For example, in cases where a wheel of a vehicle is driving on a road (or on another surface), a registration system can utilize semantic information about the road to improve registration. For example, points in a point map that correspond to road markings (e.g., lane lines) can be grouped together in groups that correspond to individual line segments. For example, points in a point map can be grouped together as lines based on the distance between adjacent points. In some cases, the correspondence between line segments can be included in generating a registration between a sensed point map generated based on a captured image of a road and a reference point map of the road.
[0035] Examples are described herein using a vehicle as an illustrative example of a tracked object and using a road as an illustrative example of a target object. However, one of ordinary skill will appreciate that the systems and related techniques described herein can be included in and performed by any other system or device for detecting and / or tracking any type of object in one or more images. Examples of other systems that can perform the techniques described herein or that can include components for performing the techniques described herein include robotic systems, extended reality (XR) systems (e.g., augmented reality (AR) systems, virtual reality (VR) systems, mixed reality (MR) systems, etc.), video analytics, security systems, aerial systems, etc. Examples of other types of objects that can be detected include people or pedestrians, infrastructure (e.g., roads, signs, etc.), etc. In one illustrative example, a vehicle that tracks can perform one or more of the techniques described herein to detect a pedestrian or infrastructure object (e.g., a road sign) in one or more images.
[0036] The systems and techniques described herein provide advantages over existing object detection and tracking systems. For example, the systems and techniques can be used to update a point graph (e.g., for a sensed road network) so that the point graph has accurate registration with a reference point graph and accordingly provides more accurate references for object localization (e.g., real-time object localization, such as real-time vehicle localization).
[0037] Various aspects of the application will be described with respect to the drawings. Figure 1 is an image 100 that illustrates an environment that includes a number of vehicles traveling on a roadway. The vehicles include a tracking vehicle 102 (as an example of a tracking object), a target vehicle 104, a target vehicle 106, and a target vehicle 108 (e.g., as examples of tracking objects). The tracking vehicle 102 can track the target vehicles 104, 106, and 108 and / or lane lines 111 in order to navigate the environment. For example, the tracking vehicle 102 can determine a localization of the tracking vehicle 102 to determine when to slow down, speed up, change lanes, and / or perform some other function. Although the tracking vehicle 102 is referred to as the tracking vehicle 102 and the vehicles 104, 106, and 108 are referred to as target vehicles, if the vehicles 104, 106, and 108 are tracking other vehicles, they can also be referred to as tracking vehicles, in which case the other vehicles become target vehicles. Figure 1 , although the tracking vehicle 102 is referred to as the tracking vehicle 102 and the vehicles 104, 106, and 108 are referred to as target vehicles, if the vehicles 104, 106, and 108 are tracking other vehicles, they can also be referred to as tracking vehicles, in which case the other vehicles become target vehicles.
[0038] Figure 2A is an example of an illustration that illustrates information that can be included in a point graph 200 and / or generated using points in a point graph 200. Figure 2A The example in illustrates three lanes of a highway from a top perspective (or “bird’s eye view”), including a left lane 222, a middle lane 224, and a right lane 226. Each lane is shown as having a center line and two boundary lines, with the middle lane 224 sharing a boundary line with the left lane 222 and the right lane 226. A vehicle 220 is shown in the middle lane 224. One or more cameras on the vehicle can capture images of the environment around the vehicle, as described herein. In some examples, the point graph 200 can include points (or path points) that represent lane lines. For example, each line can be defined by a number of points. In some cases, the roadway represented by the point graph 200 can be represented as a plane in a three-dimensional representation of the tracking vehicle 220 and its surrounding environment.
[0039] In some aspects, point map 200 may include multiple map points corresponding to one or more reference locations in 3D space. In one example, using an autonomous vehicle as an illustrative example of an object, the points of point map 200 define stationary physical reference locations associated with a road, such as road lanes and / or other data. For example, point map 200 may represent a lane on a road as a connected set of points. Line segments are defined between two map points, where multiple line segments define different lines of the lane (e.g., lane boundary lines and center lines). Line segments may form piecewise linear curves defined using map points. For example, the connected set of points (or segments) may represent the center lines and boundary lines of lanes on a road, which allows the autonomous vehicle to determine its location on the road and the location of a target object on the road. In some cases, point map 200 may represent a road (or a local area of a road) as a plane in 3D space (e.g., a road plane or ground plane).
[0040] In some cases, reference locations in the environment can be included in a reference point map (e.g., Figure 3 In the reference point diagram 307). In some cases, the reference point diagram (e.g., Figure 3 The reference point map (307) may be referred to as a high-resolution (HD) map. In some cases, different reference point maps may be maintained for different regions of the world (e.g., reference point map for New York City, reference point map for San Francisco, reference point map for New Orleans, etc.). In some examples, different reference point maps may be included in separate data files (e.g., Geo-JavaScript Object Notation (GeoJSON) files, ShapeFiles, comma-separated value (CSV) files, and / or other files).
[0041] In some implementations, point map registration can be used to determine the correspondence between point map 200 and a reference point map. In one illustrative example, point map registration can determine the correspondence between portions of point map 200 and the reference point map corresponding to the same location (e.g., a specific location on a highway). In one illustrative example, the Iterative Closest Point (ICP) algorithm can be used to perform registration. The ICP algorithm may include a data association step and a transformation step. In one illustrative example, data association may be based on each point in point map 200 and its correspondence to a reference point map (e.g., ...). Figure 3The nearest point between point map 200 and the reference point map 307 is determined. In some cases, the initial alignment between point map 200 and the reference point map can be approximated based on one or more assumptions (e.g., orientation relative to a reference direction). In some cases, the transformation step can be performed by determining the centroid of each point map and performing a transformation to align the centroid of point map 200 with the centroid of the reference point map. In each iteration of the ICP algorithm, a new data association can be performed on the point maps transformed from previous iterations. In some cases, after determining the new data association, a new transformation can be determined to align the point maps based on the new data association. In some cases, the ICP algorithm can be repeated until the maximum number of iterations is reached or until the data alignment and transformation are close to convergence.
[0042] Figure 2B and Figure 2C An example point map registration is shown, illustrating how the ICP algorithm is used to determine the correspondence between points in a point map. Figure 2B In the figure, the thin lines correspond to points in reference point Figure 229. As shown, the thick lines correspond to points traced by the object being tracked (e.g., Figure 1 The points captured by the tracking vehicle 102) are shown in Figure 231. Figure 2B In the examples, the solid lines in point diagrams 229 and 231 are composed of points that are very close to each other (see, for example, see...). Figure 4B ).exist Figure 2B In the example, the relative positions of point plots 229 and 231 can correspond to the initial estimate of the correspondence between point plots. Figure 2C An example of a transformed point graph 233 is illustrated, generated by applying a transformation (e.g., a transformation matrix) determined by the point-to-point correspondences, which are determined by the ICP algorithm applied to point graph 231. Figure 2C In the illustrated example, point plot 233 is shown as having a larger misalignment relative to reference point plot 229 than the initial estimate shown in point plot 231. In some cases, the increased misalignment may be due to mismatched points 235 present in one point plot but not in the other. Figure 2B and Figure 2C In the example, the registration technique applies minimal semantic information about the data in each point map 229, 231. In some cases, semantic information (e.g., the grouping of points in point maps 229, 231 corresponds to line segments) can be used to improve registration.
[0043] Figure 3is a block diagram illustrating an example of a point map registration system 300 for performing registration between a point map generated from a captured image and a reference point map. The point map registration system 300 can be included in a tracking object that tracks one or more target objects. In some cases, the point map registration system 300 can be included in and / or in communication with an object detection and tracking system. As described above, a tracking object refers to an object that tracks one or more other objects, which are referred to as target objects. In one illustrative example, the point map registration system 300 can include, be included in, and / or be in communication with an autonomous driving system included in an autonomous vehicle (as an example of a tracking object). In another illustrative example, the point map registration system 300 can include, be included in, and / or be in communication with an autonomous navigation system included in a robotic device or system. While examples are described herein using autonomous driving systems, autonomous navigation systems, and / or autonomous vehicles for illustrative purposes, one of ordinary skill will appreciate that the point map registration system 300 and related techniques described herein can be included in and performed by any other system or device for determining registration between point maps.
[0044] The point map registration system 300 can be used to estimate a position of a tracking object relative to a reference point map using point detection from radar, using radar images, a combination thereof, and / or using other information. In one illustrative example, the point map registration system 300 can use wheel keypoint detection from a camera and corresponding vehicle type classification, point detection from radar, and optionally object detection from an imaging radar to estimate a positioning and / or size of a target vehicle detected on a road. As described in more detail below, the point map registration system 300 can apply any combination of one or more of a camera-based object type likelihood filter, a target positioning estimation technique for object (e.g., vehicle or other object) size estimation (e.g., based on observed wheel keypoint locations), a radar-based length estimation technique, and / or imaging radar-based object detection, and can implement an estimation model to track a best estimate of a size (e.g., length and / or other size dimension) of an object using measurements provided by map-based size determination, radar-based size estimation, and / or imaging radar detection.
[0045] The dot map registration system 300 includes various components, including one or more cameras 302, a feature detection engine 304, a line generation engine 306, a correspondence engine 308, and a transformation engine 310. The components of the dot map registration system 300 may include software, hardware, or both. For example, in some embodiments, the components of the dot map registration system 300 may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or the components of the dot map registration system may include computer software, firmware, or any combination thereof for performing the various operations described herein and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of a computing device implementing the dot map registration system 300.
[0046] Although the dot plot registration system 300 is shown as including certain components, those skilled in the art will understand that the dot plot registration system 300 may include components such as... Figure 3 The components shown may have more or fewer components. For example, the dot map registration system 300 may include one or more input devices and one or more output devices (not shown), or it may be part of a computing device or object that includes one or more input devices and one or more output devices. In some specific embodiments, the dot map registration system 300 may also include or may be a computing device that includes the following (e.g., Figure 3 A portion of the computing system 1300 includes: one or more memory devices (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices); one or more processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices) that communicate with and / or are electrically connected to one or more memory devices; one or more wireless interfaces for performing wireless communication (e.g., including one or more transceivers and baseband processors for each wireless interface); one or more wired interfaces (e.g., serial interfaces such as Universal Serial Bus (USB) inputs, lighting connectors, and / or other wired interfaces) and / or other components for performing communication via one or more hardwired connections.
[0047] As described above, the point graph registration system 300 can be implemented by and / or included in a computing device or other object. In some cases, multiple computing devices can be used to implement the point graph registration system 300. For example, a computing device used to implement the point graph registration system 300 can include a computer or multiple computers that are part of a device or object, such as a vehicle, a robotic device, a surveillance system, and / or any other computing device or object that has the resource capabilities to perform the techniques described herein. In some implementations, the point graph registration system 300 can be integrated with (e.g., integrated into software, added as one or more plug-ins, included as one or more library functions, or otherwise integrated with) one or more software applications, such as an autonomous driving or navigation software application or suite of software applications. The one or more software applications can be installed on the computing device or object that implements the point graph registration system 300.
[0048] The one or more cameras 302 of the point graph registration system 300 can capture one or more images 303. In some cases, the one or more cameras 302 can include multiple cameras. For example, an autonomous vehicle that includes the point graph registration system 300 can have one or more cameras located at the front of the vehicle, one or more cameras located at the back of the vehicle, one or more cameras located on each side of the vehicle, and / or other cameras. In another example, a robotic device that includes the point graph registration system 300 can include multiple cameras located on various parts of the robotic device. In another example, an aerial device that includes the point graph registration system 300 can include multiple cameras located on different parts of the aerial device.
[0049] The one or more images 303 can include still images or video frames. The one or more images 303 each contain an image of a scene. Figure 3 An example of an image 305 is shown in FIG. 3. The image 305 illustrates an example of an image captured by a camera of a tracking vehicle, including multiple target vehicles. When an image frame is captured, the image frame can be part of one or more video sequences. In some cases, the images captured by the one or more cameras 302 can be stored in a storage device (not shown), and the one or more images 303 can be retrieved or otherwise obtained from the storage device. In some implementations, the reference point graph 307 can be obtained from the storage device. In some examples, the reference point graph 307 can correspond to an HD road graph, a map of a geographic region, and / or the like. The one or more images 303 can be raster images composed of pixels (or voxels) optionally with depth maps, vector images composed of vectors or polygons, or a combination thereof. The images 303 can include one or more two-dimensional representations of a scene along one or more planes (e.g., a plane in the horizontal or x-direction and a plane in the vertical or y-direction), or one or more three-dimensional representations of a scene.
[0050] The feature detection engine 304 can obtain and process one or more images 303 to detect and / or track one or more objects in the one or more images 303. The feature detection engine 304 can output the objects as detected and tracked objects. The feature detection engine 304 can determine a classification (referred to as a class) or category of each object detected in an image. For example, the feature detection engine 304 can determine a lane line classification for the lane line 311. In some cases, the feature detection engine 304 can output multiple classes for a detected object, as well as a confidence score indicating a confidence that the object belongs to each class (e.g., a confidence score that the object is a lane line is 0.85, a confidence score that the object is a center divider is 0.14, and a confidence score that the object is a motorcycle is 0.01).
[0051] In some cases, the feature detection engine 304 can output the localization of objects detected in the sensed point map 331. For example, the feature detection engine 304 can detect and output points (e.g., spatial localizations in a 3D coordinate system) corresponding to the lane line 311. In some cases, one or more classifications determined by the feature detection engine 304 can be used to identify a particular object (e.g., the lane line 311) for additional processing (e.g., registration). In some cases, the sensed point map 331 can be incomplete (e.g., due to occlusions of objects in the scene). In some examples, missing information in the sensed point map 331 can be supplemented by the reference point map 307 by registering the sensed point map 331 to the reference point map 307 (e.g., occluded portions of the environment can be present in the reference road map 307).
[0052] Any suitable object detection and / or classification technique can be performed by the feature detection engine 304. In some cases, the feature detection engine 304 can use a machine learning based object detector, such as using one or more neural networks. For example, a deep learning based object detector can be used to detect and classify objects in the one or more images 303. In one illustrative example, a Cifar-10 neural network based detector can be used to perform object classification to classify objects. In some cases, the Cifar-10 detector can be trained to classify only certain objects, such as only lane lines. Additional details of the Cifar-10 detector are described below with respect to FIGS. 4A-4B. Figure 10 Additional details of the Cifar-10 detector are described below.
[0053] Another illustrative example of a deep learning-based detector is a single shot multibox detector (SSD) that includes a neural network and can be applied to multiple object classes. A feature of the SSD model is the use of multi-scale convolutional bounding box outputs attached to multiple feature maps at the top of the neural network. Such a representation allows the SSD to effectively model different bounding box shapes. It has been shown that given the same VGG-16 base architecture, the SSD outperforms its state-of-the-art object detector counterpart in both accuracy and speed. The SSD deep learning detector is described in more detail in K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," CoRR, abs / 1309.1556, 79014, which is hereby incorporated by reference in its entirety for all purposes. Additional details of the SSD detector are described below with respect to Figures 11A-11C Additional details of the YOLO detector are described below.
[0054] Another illustrative example of a deep learning-based detector that can be used to detect and classify objects in one or more images 303 includes a You Only Look Once (YOLO) detector. The YOLO detector processes images at 40 fps - 90 fps with an mAP of 78.6% (based on VOC 2007) when running on a Titan X. The YOLO deep learning detector is described in more detail in J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, "You only look once: Unified, real-time object detection," arXiv preprint arXiv: 1506.02640, 2015, which is hereby incorporated by reference in its entirety for all purposes. Additional details of the YOLO detector are described below with respect to Figures 12A-12C Although the Cifar-10, SSD, and YOLO detectors are provided as illustrative examples of deep learning-based object detectors, one of ordinary skill in the art will appreciate that any other suitable object detection and classification can also be performed by the feature detection engine 304.
[0055] Figure 4A and Figure 4B is an illustration of an example reference point map 407 (e.g., the reference point map 307 of Figure 3 ) and an example sensed point map 431 (e.g., the sensed point map 331 of Figure 3 ) generated based on one or more images. In Figure 4AIn the illustration 400, the points of the reference point map 407 have a thin line appearance, and the points of the sensed point map 431 have a thick line appearance. Figure 4B A subset 450 of the points included in the reference point map 407 and the sensed point map 431 is illustrated, where each point has a stronger visual effect. In the illustrated example, the subset 450 includes the points of the reference point map 407 and the sensed point map 431 that are within a distance threshold of each other. Figure 4B In the illustrated example, the reference point map 407 includes small ovals corresponding to the points in the reference point map 407. Similarly, the sensed point map 431 depicts larger ovals representing the points in the sensed point map 431. Figure 4B The depictions are provided for understanding the content of the point maps 407, 431. Although the points of the point maps 407, 431 are illustrated as ovals in the Figure 4B illustrations, it should be understood that each oval can represent a value, such as a 3D localization value for each point.
[0056] Returning to Figure 3 , the line generation engine 306 can group the points in a point map (e.g., the sensed point map 331, the reference point map 307) into point groups. As an illustrative example, the groups generated by the line generation engine 306 semantically represent lines. In some cases, other groupings based on different semantic representations can be used without departing from the scope of the present disclosure.
[0057] As illustrated in Figure 3 , the line generation engine 306 can obtain the sensed point map 331 output from the feature detection engine 304. In some cases, the line generation engine 306 can also obtain the reference point map 307. In some examples, a group can be generated by determining whether a distance between a particular point and a nearest neighboring point of the particular point is below a distance threshold. In one illustrative example, points can be grouped together into a point grouping by determining whether the Euclidean distance between a point and another neighboring point is less than 0.15 meters (m). In some cases, the line generation engine 306 can perform the grouping for each point in a given point map and generate a plurality of point group representations corresponding to lines in the point map (e.g., the reference point map 307, the sensed point map 331). As used herein, a point grouping representing a line is also referred to as a line representation of a point map. In Figure 3 the illustrated example, the line generation engine 306 can output a sensed line representation 309 of the sensed point map 331 and a reference line representation 313 of the reference point map 307. In some implementations (not shown), the reference point map 307 can itself include a line representation, and the point map registration system 300 can utilize the line representation of the reference point map 307 stored in the storage device. In some cases, the stored line representation of the reference point map 307 can be used in place of and / or in conjunction with the reference line representation 313 of the reference point map 307 in the line generation engine 306.
[0058] In some cases, the line generation engine 306 can generate groupings including curve segments as long as the threshold criteria for grouping points into point groupings are met. In some cases, different and / or additional criteria can be used during the generation of the line representations 309, 313. For example, in some cases, a linearity threshold can be used to distinguish intersecting lines into separate groupings within 309, 313. In another illustrative example, the spatial positioning of the lines generated by the line generation engine 306 can be compared to a road plane (e.g., as described with respect to FIG. 3). Other criteria can also be used to distinguish lane lines within the line representations 309, 313 without departing from the scope of the present disclosure. Figure 2A
[0059] Each of the line representations 309, 313 can include a different number of groupings (e.g., a different number of lines). For example, the sensed line representation 309 of the sensed point map 331 can include M elements, where M is an integer. Equation [1] below illustrates a sensed line representation 309 including M elements:
[0060]
[0061] Similarly, the reference line representation 313 of the reference point map 307 can include N elements, where N is an integer. Equation [2] below illustrates a reference line representation 313 including N elements:
[0062]
[0063] In some cases, the sensed line representation 309 can include more point groupings than the reference line representation 313 (e.g., M > N). In some cases, the reference line representation 313 can include more point groupings than the sensed line representation 309 (e.g., M < N). In some examples, the line representations 309, 313 can include an equal number of point groupings (e.g., M = N).
[0064] Referring to Figure 4B , the example reference lines 413 can provide examples of lines determined by the line generation engine 306 and included in the reference line representation 313. Similarly, the example sensed lines 409 can provide examples of lines determined by the line generation engine 306 and included in the sensed line representation 309.
[0065] Returning to Figure 3 In one illustrative example, the correspondence engine 308 can obtain the M line representations included in the sensed line representation 309 and the N line representations included in the reference line representation 313 from the line generation engine 306. In some implementations, the correspondence engine 308 can determine a correspondence between the lines in the sensed line representation 309 and the lines in the reference line representation 313.
[0066] In some cases, the Dynamic Time Warping (DTW) algorithm can be used to determine the correspondence between individual lines in the sensed line representation 309 (e.g., grouping of individual points) and their corresponding individual lines in the reference line representation 313. In some aspects, the DTW algorithm can be used to compare sequential data from different datasets, where the ordering between the datasets is consistent. In some cases, the ordering may correspond to time. In an exemplary example, the DTW algorithm can be used to compare audio to determine whether audio data corresponds to the same word. In this case, the audio data may be stored in a time series (e.g., sequentially sampled data captured by a microphone). While time can provide a convenient ordering relationship for certain types of data, the DTW algorithm can also be used for other types of data where the data in the datasets being compared have a comparable ordering.
[0067] In some examples, such as with lane markings on a road, the order of sequential data can be based on the mode of transport (e.g., Figure 1 The direction of travel of the tracking vehicle 102 is determined. For example, the individual points of each line in line representations 309 and 313 can be ordered sequentially based on their positioning along an axis. For example, the axis used to sequentially order the points of the line could be an axis parallel to the direction of travel of the tracking vehicle, an axis parallel to the optical axis of the camera capturing the road image, etc. The grouping of individual points in sensing line representation 309 can be identified by index i, where each point group... The line including each point P can be represented as shown in the following formula [3]:
[0068]
[0069] In some cases, the reference line representation 313 of the reference point map 307 may also be ordered relative to the same axis orientation used to sort the points in the sensing line representation 309. In some cases, the reference point map may include information about orientation (e.g., relative to the four fundamental directions, geographic coordinates, etc.). The individual point groups of the reference line representation 313 may be identified by index j, where each point group... The line including each point P can be represented as shown in the following formula [4]:
[0070]
[0071] Figure 5A and Figure 5B An example comparison of the DTW algorithm is shown for two data sequences. Figure 5A Data graph 500 illustrates data sequence A and data sequence B. Figure 5A In the example, the horizontal axis represents an index, which indicates the sequential order of the data points. In some cases, the horizontal axis may represent time. In some cases, the index may correspond to a sequence of events. Figure 3correspondence engine 308 generates. Figure 5A The vertical axis illustrates data values, which are not shown with units. In determining a point graph (e.g., Figure 3 The line representations (e.g., Figure 3 In one illustrative example of a correspondence between line representations 309, 313 of a point graph 307, a sensed point graph 331, values can correspond to the positioning of points of a line segment, the distance of a point of a line segment from an origin position (e.g., for capturing Figure 3 In some cases, a DTW algorithm can compute a DTW "distance" between two data sequences.
[0072] Referring to Figure 5B , a DTW distance computation matrix 550 is shown for computing a DTW distance between data sequence A and data sequence B. As shown, according to an index order, the values 551 of sequence A are listed alongside each row of matrix 550, starting with the first value "1" at the bottom row of matrix 550 and ending with the final value "3" at the top row of matrix 550. Similarly, the values 552 of sequence B are listed below each column of matrix 550, starting with the first value "1" of the sequence at the leftmost column of matrix 550 and ending with the final value "3" at the rightmost column of the matrix. As shown, each cell of matrix 550 includes a distance value computed for a pair of points in the two sequences A and B. The value in each cell of matrix 550 can be determined according to the following equation [1] based on the components of the values in the corresponding row / column of the series:
[0073] D(i,j) = |A i -B j | + min(D[i - 1,j - 1], D[i - 1,j], D[i,j - 1]) [5]
[0074] where i is an index of sequence A and j is an index of sequence B, D(i,j) is a distance corresponding to a pair of values A i and B j , |A i -B j | is a modulus of A i and B j , and the min() function takes the minimum of the three adjacent distance measurements.
[0075] Applying equation [5] to cell 554 yields D(6,7) = |2-4| + min(9,16,12) = 2 + 9 = 11. In the example of Figure 5B the final DTW distance 556 computed between sequence A and sequence B equals 13. The cells included in the distance computation for cell 554 function are contained within the dashed line 555.
[0076] In addition to determining line-to-line correspondences, the DTW algorithm can be used to determine point-to-point correspondences between points of the compared sequences. Figure 5C An example 570 of point-to-point correspondences of sequence A and sequence B as discussed with respect to Figure 5A and Figure 5B is illustrated. In the example of Figure 5C , point-to-point correspondences are illustrated by dark, thick outline cells. The point-to-point correspondences can be determined based on the minimum DTW distance between points of sequence A and sequence B. Figure 5C Correspondences 560 (e.g., dashed lines) are illustrated that illustrate the correspondences identified with dark, thick outline cells of Figure 5B In some cases, the mapping between points of two sequences can not be 1 : 1. For example, the point of sequence A with index = 2 can correspond to points of sequence B with index = 2, 3, and 4. In some cases, the DTW algorithm can determine point-to-point correspondences of points within each of the compared sequences (e.g., for M x N computed DTW distances).
[0077] In some cases, for the purposes of registration, point-to-point correspondences between point maps (e.g., reference point map 307, sensed point map 331) of Figure 3 may be determined by first determining line-to-line correspondences. For example, line-to-line correspondences can be determined based on the mutual minimum distance between lines, DTW distances, or any combination thereof. In some cases, point-to-point correspondences can be generated during the computation of line-to-line correspondences. In one illustrative example, point-to-point correspondences can be determined based on DTW distances (e.g., based on thick outline cells of Figure 5B In some cases, point-to-point correspondences can be computed using other suitable techniques that are constrained by the line-to-line correspondences. In some cases, correspondence engine 308 can output point-to-point correspondences to transformation engine 310.
[0078] Figure 6 Examples of line-to-line and point-to-point correspondences that can be determined by correspondence engine 308 of Figure 3 are illustrated. In the illustration of Figure 6 , reference point map 607 can correspond to reference point map 407 of Figure 4B , sensed point map 631 can correspond to sensed point map 431 of Figure 4B , sensed lines 609 can correspond to sensed lines 409 of Figure 4B , and reference lines 613 can correspond to reference lines 413 of Figure 4B . Line correspondences between sensed lines 609 and reference lines 613 are illustrated as thick dashed lines 640. Multiple point-to-point correspondences 660 are illustrated by dashed lines. Point-to-point correspondences 660 can correspond toFigure 5C The point-to-point correspondence 660 is shown proximate to the distal ends of the sensing line 609 and the reference line 613 for illustrative purposes. However, it should be understood that the point-to-point correspondence for each point in each line can be determined as described above with respect to the correspondence engine 308 of FIG. 3. Figure 3
[0079] Returning to FIG. 3, Figure 3 The transform engine 310 can determine a transform for transforming the sensing point cloud 331 to align with the reference point cloud 307. As described above, the transform engine 310 can obtain the point-to-point correspondence determined by the correspondence engine 308. In one illustrative example, the transform engine 310 can utilize any suitable computation to determine a transform to align the sensing point cloud 331 and the reference point cloud 307. For example, a least squares technique can be used to generate a transform between the sensing point cloud 331 and the reference point cloud 307. In some cases, the least squares technique can include a transform that minimizes the least squares distance error between the sensing point cloud 331 and the reference point cloud 307. In one illustrative example, a least squares technique described in K. S. Arun, T. S. Huang, and S. D. Blostein, "Least-Squares Fitting of Two 3-D Point Sets," IEEE Transactions on pattern analysis and machine intelligence, September 1998, which is hereby incorporated by reference in its entirety for all purposes, can be used to generate a transform matrix.
[0080] For example, the transform of a point X S in the sensing point cloud 331 is described with respect to the following equation [6]:
[0081] X T = X S *T = X S *R + t [6]
[0082] where X T is the translation of the point X S , T is a translation matrix that includes a rotation parameter R and a translation parameter t. An example of a rotation parameter is shown in the following equation [7]:
[0083]
[0084] In the above formula, α is yaw (horizontal rotation), β is pitch (vertical rotation), and γ is roll (left-right rotation). Pitch, roll, and yaw can be conceptualized as yaw being a horizontal rotation relative to the ground (e.g., from left to right relative to the horizontal axis), pitch being a vertical rotation relative to the ground (e.g., up and down relative to the horizontal axis), and roll being a left-right rotation relative to the horizon (e.g., left and right relative to the horizontal axis). The translation vector t can be represented as shown in formula (6):
[0085]
[0086] Where X t Y t and Z t These are the translation components in the X, Y, and Z coordinates of Euclidean geometry.
[0087] In some cases, the translated sensing point map 370 can be generated by the transformation engine 310. In some examples, the translated sensing point map 370 can be provided to the correspondence engine 308, and the correspondence between 370 and the reference point map 307 can be determined. In some examples, the determination of the correspondence by the correspondence engine 308 and the determination of the translation by the transformation engine 310 can be iterative until convergence is reached or the maximum number of iterations is reached.
[0088] Figure 7 Examples are given of those that can be generated by Figure 3 The example point map registration system 300 performs point map registration 700. Figure 7 In the illustrated example, based on images (e.g., Figure 3 Sensing point map 731 (represented by thick lines) generated from sensing point map 331) is based on the sensing point map 731 generated from the sensing point map 331. Figure 3 The dot map registration system 300 determines the transformation. In the illustrated example, the sensed dot map 731 is aligned with the reference dot map 707 (represented by thin lines). As shown, it can be relative to... Figure 2C The alignment shown (e.g., based on the ICP algorithm) improves the sensing point map 731 and the reference point map 707 (e.g., Figure 3 Alignment of the reference point map 307. As described above, by utilizing the semantic information that groups of points closely adjacent to each other in point maps 707 and 731 correspond to lines (e.g., lane lines), the correspondence between point maps 707 and 731 can be improved. For example, sensing the length difference between lines in point map 731 and reference point map 707 can result in a larger DTW distance value, thereby utilizing the line length as additional semantic information. Additionally, as shown, the mismatch point 735 included in reference point map 707 can be aligned by the point map registration system 300 (e.g., by...). Figure 3correspondence engine 308) determines that there are no corresponding points in the sensed point graph 731 (e.g., the line formed by grouping the mismatched points 735 does not match any line in the sensed point graph 731). Thus, the mismatched points 735 can be excluded from the transformation computation of the transformation engine 310.
[0089] Accordingly, based on the above, it should be appreciated that the point graph registration systems and techniques described herein can improve point graph registration. For example, a correspondence between point groupings representing lines (e.g., lane lines) in a point graph can be determined. In contrast, a pure point-to-point algorithm (e.g., an ICP algorithm) can ignore semantic information that particular points of a point graph correspond to a lane line. In some cases, due to mismatched points between a reference point graph (e.g., the reference point graph 307) and a sensed point graph (e.g., the sensed point graph 331) determined based on images captured by one or more cameras (e.g., the one or more cameras 302), an alignment resulting from point-to-point registration can be relatively coarse and / or error prone. In some cases, the systems and techniques described herein can utilize directional information to sequentially order points included in a point grouping (e.g., points of a particular lane line segment) with respect to an axis. In some cases, a DTW algorithm can be used to determine a correspondence between sequentially ordered point groupings in a reference point graph and a point graph generated based on captured images (also referred to herein as a sensed point graph). In some cases, by determining line-to-line correspondences in addition to point-to-point correspondences (e.g., based on determined line-to-line correspondences), point-to-point correspondences can be improved. Based on the improved correspondences, a transformation can be determined by the systems and techniques that provide an improved alignment between a reference point graph and a point graph generated based on captured images. Figure 3 Figure 3 Accordingly, based on the above, it should be appreciated that the point graph registration systems and techniques described herein can improve point graph registration. For example, a correspondence between point groupings representing lines (e.g., lane lines) in a point graph can be determined. In contrast, a pure point-to-point algorithm (e.g., an ICP algorithm) can ignore semantic information that particular points of a point graph correspond to a lane line. In some cases, due to mismatched points between a reference point graph (e.g., the reference point graph 307) and a sensed point graph (e.g., the sensed point graph 331) determined based on images captured by one or more cameras (e.g., the one or more cameras 302), an alignment resulting from point-to-point registration can be relatively coarse and / or error prone. In some cases, the systems and techniques described herein can utilize directional information to sequentially order points included in a point grouping (e.g., points of a particular lane line segment) with respect to an axis. In some cases, a DTW algorithm can be used to determine a correspondence between sequentially ordered point groupings in a reference point graph and a point graph generated based on captured images (also referred to herein as a sensed point graph). In some cases, by determining line-to-line correspondences in addition to point-to-point correspondences (e.g., based on determined line-to-line correspondences), point-to-point correspondences can be improved. Based on the improved correspondences, a transformation can be determined by the systems and techniques that provide an improved alignment between a reference point graph and a point graph generated based on captured images. Figure 3
[0090] In some examples, the alignment between the reference point graph and the point graph generated based on the captured images can be used to determine a position of the tracked object. In one illustrative example, the position can be used to determine at least one or more of a lane position, a direction of travel, a lane position of one or more target objects, etc. of the tracked object. In some implementations, the position can be such that the autonomous navigation system can perform one or more navigation operations based on the position. For example, the autonomous navigation can generate a notification (e.g., generate a visual, audio, and / or haptic notification, generate and / or send a message, etc.), perform a motion plan (e.g., adjust a speed, avoid a collision, change a lane of travel, etc.), a navigation plan (e.g., adjust a route of travel), and / or any other suitable operation based on the position.
[0091] Figure 8 is a flowchart illustrating an example of a process 800 for performing point graph registration in accordance with some aspects of the disclosed technology. In some implementations, the process 800 can include obtaining (e.g., by the alignment engine 306), at block 802, a reference point graph (e.g., the reference point graph 307) and a sensed point graph (e.g., the sensed point graph 331) based on images captured by one or more cameras (e.g., the one or more cameras 302). Figure 3 one or more cameras 302) includes an image of a scene (e.g., Figure 3 one or more images 303).
[0092] At block 804, the process 800 can include determining (e.g., by the feature detection engine 304 of the system 400) a first point graph representing the scene based on the image (e.g., Figure 3 and Figure 4A and Figure 4B 431) of the system 400. In some cases, the first point graph includes a first plurality of 3D points representing the scene.
[0093] At block 806, the process 800 can include determining (e.g., by the line generation engine 306 of the system 400) a first plurality of point groupings (e.g., Figure 3 and Figure 4A and Figure 4B sensing lines 409) of the system 400. In some examples, each point grouping of the first plurality of point groupings is determined based on a proximity of pairs of points of the first plurality of 3D points. In some aspects, each point grouping includes two or more 3D points from the first plurality of 3D points. In some examples, individual point groupings of the first plurality of point groupings correspond to lines in the image.
[0094] In some examples, the process 800 includes, for a first point of the first plurality of 3D points, determining that a first Euclidean distance between the first point of the first plurality of 3D points and a second point of the first plurality of 3D points is less than a threshold distance. In some cases, the second point of the first plurality of 3D points is adjacent to the first point of the first plurality of 3D points, and based on determining that the first Euclidean distance between the first point of the first plurality of 3D points and the second point of the first plurality of 3D points is less than the threshold distance, grouping the first point and the second point into a first point grouping of the first plurality of point groupings. In some examples, the process 800 includes, for the second point of the first plurality of 3D points, determining that a second Euclidean distance between the second point of the first plurality of 3D points and a third point of the first plurality of 3D points is greater than the threshold distance, and based on determining that the second Euclidean distance between the second point of the first plurality of 3D points and the third point of the first plurality of 3D points is greater than the threshold distance, excluding the third point of the first plurality of 3D points from the first point grouping of the first plurality of point groupings.
[0095] At block 808, the process 800 can include obtaining a second point graph representing the scene (e.g., Figure 2B and Figure 2C reference point graph 229, Figure 4A and Figure 4Bof the reference point map 407). In some examples, the second point map includes a second plurality of 3D points representing the scene. In some cases, the reference point map includes a reference point map. In some cases, the reference point map includes an HD reference point map. In some aspects, the reference point map is obtained from a reference map service.
[0096] At block 810, the process 800 can include obtaining a second plurality of point groupings (e.g., by the point grouping engine 306 of the system 300 Figure 4A and Figure 4B of the reference line 413). In some cases, each point grouping of the second plurality of point groupings includes a plurality of 3D points from the second plurality of 3D points. In some examples, the process 800 includes determining, based on determining a correspondence between a first point grouping of the first plurality of point groupings and a second point grouping of the second plurality of point groupings, a correspondence between respective points of the first point grouping of the first plurality of point groupings and respective points of the second point grouping of the second plurality of point groupings.
[0097] In some examples, the process 800 includes applying an ordinal ordering to points included in each respective point grouping of the first plurality of point groupings based on respective locations of each pixel included in the respective point grouping along an axis, and applying an ordinal ordering to points included in each point grouping of the second plurality of point groupings based on respective locations of each pixel included in the respective point grouping along the axis. In some cases, the axis corresponds to at least one or more of a direction of motion of the tracked object or an optical axis associated with the image.
[0098] At block 812, the process 800 can include determining a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings (e.g., by the correspondence engine 308 of the system 300 Figure 3 In some cases, determining a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings includes determining that at least one or more of the respective point groupings of the first plurality of point groupings or the respective point groupings of the second plurality of point groupings do not have a corresponding respective point grouping.
[0099] In some examples, the process 800 includes determining a dynamic time warping distance between each point grouping of the first plurality of point groupings and each point grouping of the second plurality of point groupings, and determining, based on the dynamic time warping distance, a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings.
[0100] At block 814, the process 800 can include determining, based on the correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings, a correspondence between respective points of the first point grouping of the first plurality of point groupings and respective points of the second point grouping of the second plurality of point groupings (e.g., by the correspondence engine 308 of the system 300 Figure 3the alignment transform (e.g., by the alignment engine 310 of the transformation engine 310) to align the first plurality of 3D points and the second plurality of 3D points.
[0101] In some examples, the process 800 includes applying the alignment transform to one of the second plurality of 3D points or the first plurality of 3D points representing the scene to generate a transformed first plurality of 3D points, and determining a correspondence between the transformed first plurality of 3D points and the other of the first plurality of 3D points or the second plurality of 3D points.
[0102] In some examples, the process 800 includes aligning the first plurality of 3D points and the second plurality of 3D points based on the alignment transform, and determining the position of the tracked object relative to the second point map based on aligning the first plurality of 3D points and the second plurality of 3D points. In some examples, the second point map includes a reference point map. In some examples, the process 800 includes performing a navigation operation based on the position.
[0103] In some examples, the processes described herein (e.g., the process 800 and / or other processes described herein) can be performed by a computing device or apparatus (e.g., a vehicle computer system). In one example, the process 800 can be performed by the point map registration system 300 shown. In another example, the process 800 can be performed by a computing device having the computing system 1300 shown. For example, a vehicle having the computing architecture shown can include components of the point map registration system 300 shown and can implement operations of the process 800 shown. Figure 3 In some examples, the processes described herein (e.g., the process 800 and / or other processes described herein) can be performed by a computing device or apparatus (e.g., a vehicle computer system). In one example, the process 800 can be performed by the point map registration system 300 shown. In another example, the process 800 can be performed by a computing device having the computing system 1300 shown. For example, a vehicle having the computing architecture shown can include components of the point map registration system 300 shown and can implement operations of the process 800 shown. Figure 13 In some examples, the processes described herein (e.g., the process 800 and / or other processes described herein) can be performed by a computing device or apparatus (e.g., a vehicle computer system). In one example, the process 800 can be performed by the point map registration system 300 shown. In another example, the process 800 can be performed by a computing device having the computing system 1300 shown. For example, a vehicle having the computing architecture shown can include components of the point map registration system 300 shown and can implement operations of the process 800 shown. Figure 13 In some examples, the processes described herein (e.g., the process 800 and / or other processes described herein) can be performed by a computing device or apparatus (e.g., a vehicle computer system). In one example, the process 800 can be performed by the point map registration system 300 shown. In another example, the process 800 can be performed by a computing device having the computing system 1300 shown. For example, a vehicle having the computing architecture shown can include components of the point map registration system 300 shown and can implement operations of the process 800 shown. Figure 3 In some examples, the processes described herein (e.g., the process 800 and / or other processes described herein) can be performed by a computing device or apparatus (e.g., a vehicle computer system). In one example, the process 800 can be performed by the point map registration system 300 shown. In another example, the process 800 can be performed by a computing device having the computing system 1300 shown. For example, a vehicle having the computing architecture shown can include components of the point map registration system 300 shown and can implement operations of the process 800 shown. Figure 8 In some examples, the processes described herein (e.g., the process 800 and / or other processes described herein) can be performed by a computing device or apparatus (e.g., a vehicle computer system). In one example, the process 800 can be performed by the point map registration system 300 shown. In another example, the process 800 can be performed by a computing device having the computing system 1300 shown. For example, a vehicle having the computing architecture shown can include components of the point map registration system 300 shown and can implement operations of the process 800 shown.
[0104] The process 800 is illustrated as a logical flow diagram, the operation of which represents a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored, for example, in a memory, that, when executed by a processor, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0105] Additionally, process 800 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that executes jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0106] As mentioned above, object detection and tracking systems can use machine learning-based object detectors (e.g., based on deep neural networks) to perform object detection. Figure 9 This is an exemplary example of a deep neural network 900, which can be used to analyze objects containing target objects (such as those located in...). Figure 3 The image of lane lines 311 in image 303 is used to perform object detection, as discussed above. The deep neural network 900 includes an input layer 920 configured to take in input data, such as a preprocessed (scaled) subimage containing the target object to be detected. In one exemplary example, input layer 920 may include data representing pixels of an input image or video frame. The neural network 900 includes multiple hidden layers 922a, 922b through 922n. Hidden layers 922a, 922b through 922n comprise “n” hidden layers, where “n” is an integer greater than or equal to one. Multiple hidden layers can be made to include as many layers as needed for a given application. The neural network 900 also includes an output layer 924 that provides the output produced by the processing performed by hidden layers 922a, 922b through 922n. In one exemplary example, output layer 924 may provide a classification of objects in an image or input video frame. The classification may include a category identifying the type of object (e.g., person, dog, cat, or other object).
[0107] Neural network 900 is a multi-layered neural network composed of interconnected nodes. Each node can represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains information while processing it. In some cases, neural network 900 may include a feedforward network, in which case there are no feedback connections where the network's output is fed back into itself. In some cases, neural network 900 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read.
[0108] Information can be exchanged between nodes through node-to-node interconnections between layers. Nodes of the input layer 920 can activate a set of nodes in the first hidden layer 922a. For example, as shown, each of the input nodes of the input layer 920 is connected to each of the nodes of the first hidden layer 922a. The nodes of the hidden layers 922a, 922b through 922n can transform information by applying an activation function to the information of each input node. Information derived from the transformation can then be passed to and can activate the nodes of the next hidden layer 922b, which can perform their own designated functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable function. The outputs of the hidden layer 922b can then activate the nodes of the next hidden layer, and so on. The outputs of the last hidden layer 922n can activate one or more nodes of the output layer 924, at which the output is provided. In some cases, although the nodes (e.g., nodes 926) in the neural network 900 are shown as having multiple output lines, the nodes have a single output, and all lines shown as outputting from the nodes represent the same output value.
[0109] In some cases, each node or interconnection between nodes can have a weight, which is a set of parameters derived from training of the neural network 900. Once the neural network 900 is trained, it can be referred to as a trained neural network, which can be used to classify one or more objects. For example, an interconnection between nodes can represent a piece of information learned about the nodes of the interconnection. The interconnection can have a tunable digital weight that can be tuned (e.g., based on a training data set), allowing the neural network 900 to adapt to inputs and be able to learn as more and more data is processed.
[0110] The neural network 900 is pre-trained to process features from data in the input layer 920 using different hidden layers 922a, 922b through 922n in order to provide an output through the output layer 924. In an example in which the neural network 900 is used to identify objects (e.g., lane lines) in an image, the neural network 900 can be trained using training data that includes both images and labels. For example, training images can be input into the network, where each training image has a label indicating the class of one or more objects in each image (essentially, indicating to the network what the objects are and what features they have). In one illustrative example, the training image can include an image of the number 2, in which case the label for the image can be [0 0 1 0 0 0 0 0 0].
[0111] In some cases, the neural network 900 can adjust the weights of the nodes using a training process known as backpropagation. Backpropagation can include forward pass, loss function, backward pass, and weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. For each training image set, the process can repeat for a number of iterations until the neural network 900 is trained well enough such that the weights of the layers are accurately tuned.
[0112] For the example of identifying objects in an image, the forward pass can include passing a training image through the neural network 900. Initially, the weights are randomized before training the neural network 900. The image can include, for example, an array of numbers representing the pixels of the image. Each number in the array can include a value from 0 to 255 that describes the intensity of the pixel at that location in the array. In one example, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or luminance and two chrominance components, among others).
[0113] For the first training iteration of the neural network 900, the output will likely include values that do not favor any particular class due to the random selection of weights at initialization. For example, if the output is a vector with probabilities of the object including different classes, the probability value for each of the different classes can be equal or at least very similar (e.g., for ten possible classes, each class can have a probability value of 0.1). With the initial weights, the neural network 900 is unable to determine low-level features and, thus, cannot make an accurate determination of what the classification of the object can be. A loss function can be used to analyze the error in the output. Any suitable loss function definition can be used. One example of a loss function includes mean squared error (MSE). MSE is defined as which computes the sum of one-half of the square of the actual answer minus the predicted (output) answer. The loss can be set equal to the value of E total .
[0114] For the first training image, the loss (or error) will be high because the actual value will be very different from the predicted output. The goal of the training is to minimize the amount of loss such that the predicted output is the same as the training label. The neural network 900 can perform the backward pass by determining which inputs (weights) contribute the most to the loss of the network and can adjust the weights such that the loss is reduced and eventually minimized.
[0115] The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a particular layer) can be computed to determine the weights that contribute the most to the loss of the network. After the derivative is computed, the weight update can be performed by updating all of the weights of the filter. For example, the weights can be updated such that they change in the opposite direction of the gradient. The weight update can be represented as where w represents a weight, w i represents an initial weight, and η represents a learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, and a lower value indicates smaller weight updates.
[0116] The neural network 900 can include any suitable deep network. One example includes a convolutional neural network (CNN) that includes an input layer and an output layer with a plurality of hidden layers in between. The hidden layers of the CNN include a series of convolutional layers, nonlinear layers, pooling layers (for down-sampling), and fully connected layers. The neural network 900 can include any other deep network other than a CNN, such as an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), etc.
[0117] Figure 10 is a diagram illustrating an example of a Cifar-10 neural network 1000. In some cases, the Cifar-10 neural network can be trained to classify only certain objects, such as lane lines. As shown, the Cifar-10 neural network 1000 includes various convolutional layers (Conv1 layer 1002, Conv2 / Relu2 layer 1008, and Conv3 / Relu3 layer 1014), a number of pooling layers (Pool1 / Relul layer 1004, Pool2 layer 1010, and Pool3 layer 1016), and a rectified linear unit layer mixed in. Normalization layers Norml 1006 and Norm2 1012 are also provided. The last layer is an ip1 layer 1018.
[0118] Another deep learning-based detector that can be used to detect or classify objects in an image includes an SSD detector, which is a fast single shot object detector that can be applied to multiple object species or classes. Traditionally, SSD models are designed to use multi-scale convolutional bounding box outputs attached to multiple feature maps at the top of a neural network. Such a representation allows the SSD to efficiently model different box shapes, such as when the size of an object is unknown in a given image. However, using the systems and techniques described herein, sub-image extraction and sub-image width and / or height scaling can allow an object detection and tracking system to avoid having to work with different box shapes. Instead, the object detection model of the detection and tracking system can perform object detection on the scaled image in order to detect the orientation and / or location of an object (e.g., a target vehicle) in the image.
[0119] Figures 11A-11C is a diagram illustrating an example of a single shot object detector that models different box shapes. Figure 11A includes an image, and Figure 11B and Figure 11Cincluding an illustration of how an example SSD detector (with a VGG deep network base model) operates. For example, the SSD matches objects to default boxes of different aspect ratios (as shown by the dashed rectangles in Figure 11B and Figure 11C Each element of a feature map has multiple default boxes associated with it. Any default box that has an intersection-over-union with an actual box that exceeds a threshold (e.g., 0.4, 0.5, 0.6, or other suitable threshold) is considered a match for an object. For example, two of the boxes in the 8x8 box (box 1102 and box 1104 in Figure 11B match the cat, and one of the boxes in the 4x4 box (box 1106 in Figure 11C matches the dog. The SSD has multiple feature maps, where each feature map is responsible for objects of different scales, allowing it to identify objects over a large range of scales. For example, Figure 11B the boxes in the 8x8 feature map are smaller than the boxes in the 4x4 feature map of Figure 11C In one illustrative example, the SSD detector can have a total of six feature maps.
[0120] For each default box in each cell, the SSD neural network outputs a probability vector of length c, where c is the number of classes, indicating the probability that the box contains an object of each class. In some cases, a background class is included indicating that there is no object in the box. The SSD network also outputs (for each default box in each cell) an offset vector with four entries containing the predicted offsets needed to match the default box to the bounding box of the underlying object. The vector is given in the format (cx, cy, w, h), where cx indicates the center x, cy indicates the center y, w indicates the width offset, and h indicates the height offset. The vector is only meaningful if the default box actually contains an object. For the image shown in Figure 11A all of the probability labels will indicate the background class except for the three matching boxes (two for the cat and one for the dog).
[0121] As described above, using the systems and techniques described herein, the number of scales is reduced to a scaled sub-image on which an object detection model can perform object detection to detect the position of an object (e.g., a target vehicle).
[0122] Another deep learning-based detector that can be used by an object detection model to detect or classify objects in an image includes a You Only Look Once (YOLO) detector, which is an alternative to the SSD object detection system. Figures 12A-12C is an illustration of an example of a You Only Look Once (YOLO) detector according to some examples. In particular, Figure 12A includes an image, andFigure 12B and Figure 12C A diagram illustrating how a YOLO detector operates. A YOLO detector can apply a single neural network to a full image. As shown, the YOLO network divides the image into multiple regions and predicts bounding boxes and probabilities for each region. These bounding boxes are weighted by the predicted probabilities. For example, as shown, the YOLO detector divides the image into a 13x13 grid of cells. Each cell is responsible for predicting five bounding boxes. A confidence score is provided that indicates the certainty that the predicted bounding box actually encloses an object. This score does not include a classification of the object that can be in the box, but indicates whether the shape of the box is appropriate. The predicted bounding boxes are shown in Figure 12A Figure 12B
[0123] Each cell also predicts a class for each bounding box. For example, a probability distribution over all possible classes is provided. Any number of classes can be detected, such as a bicycle, a dog, a cat, a person, a car, or other suitable object classes. The confidence score and class prediction for a bounding box are combined into a final score that indicates the probability that the bounding box contains a particular type of object. For example, the gray box with a thick border on the left side of the image in Figure 12B Figure 12C Figure 12C
[0124] In some cases, a computing device or apparatus can include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of processes described herein. In some examples, a computing device can include a display, one or more network interfaces configured to communicate and / or receive data, any combination thereof, and / or other components. The one or more network interfaces can be configured to communicate and / or receive wired and / or wireless data, including data according to 3G, 4G, 5G, and / or other cellular standards, data according to WiFi (802.1 lx) standards, data according to Bluetooth TM standard data, data according to Internet Protocol (IP) standards, and / or other types of data.
[0125] Components of computing devices can be implemented in circuitry. For example, components can include or be implemented using electronic circuitry or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or can include or be implemented using computer software, firmware, or any combination thereof for performing various operations described herein and / or can be implemented using computer software, firmware, or any combination thereof for performing various operations described herein.
[0126] Figure 13 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. Specifically, Figure 13 An example of a computing system 1300, which can be, for example, any computing device making up an internal computing system, a remote computing system, a camera, or any component thereof, is illustrated in which components of the system communicate with each other using connections 1305. Connections 1305 can be physical connections that use a bus, or direct connections into a processor 1310, such as in a chipset architecture. Connections 1305 can also be virtual, networking, or logical connections.
[0127] In some aspects, the computing system 1300 is a distributed system in which the functionality described in this disclosure can be distributed within one data center, multiple data centers, a peer-to-peer network, and the like. In some aspects, one or more of the described system components represent a number of such components, each performing a part or all of the function of the described component. In some aspects, the components can be physical or virtual devices.
[0128] The example system 1300 includes at least one processing unit (CPU or processor) 1310 and connections 1305 that couple various system components including the system memory to the processor 1310, such as read-only memory (ROM) 1320 and random access memory (RAM) 1325. The computing system 1300 can include a cache of the processor 1310 in the direct connection with, in close proximity to, or integrated as part of the processor 1310, a high-speed cache 1312.
[0129] The processor 1310 can include any general purpose processor and a hardware service or software service (such as the services 1332, 1334, and 1336 stored in the memory device 1330 configured to control the processor 1310), as well as a specific purpose processor where software instructions are incorporated into the actual processor design. The processor 1310 can essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, and cache, etc. Multiple cores can be symmetric or asymmetric.
[0130] To enable user interaction, the computing system 1300 includes an input device 1345, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech and the like. The computing system 1300 can also include output devices 1335, which can be one or more of a number of output mechanisms known to those of skill in the art. In some instances, multi-modal systems can enable a user to provide multiple types of input to communicate with the computing system 1300. The computing system 1300 can include communication interface 1340, which can generally govern and manage the user input and system output.
[0131] The communication interface can receive and / or send wired or wireless communications using a wired and / or wireless transceiver, including utilizing audio jacks / plugs, microphone jacks / plugs, universal serial bus (USB) ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, dedicated wired ports / plugs, wireless signal transmissions, low power (BLE) wireless signal transmissions, wireless signal transmissions, radio frequency identification (RFID) wireless signal transmissions, near field communication (NFC) wireless signal transmissions, dedicated short-range communication (DSRC) wireless signal transmissions, 802.11 Wi-Fi wireless signal transmissions, wireless local area network (WLAN) signal transmissions, visible light communication (VLC), worldwide interoperability for microwave access (WiMAX), infrared (IR) communication wireless signal transmissions, public switched telephone network (PSTN) signal transmissions, integrated services digital network (ISDN) signal transmissions, 3G / 4G / 5G / LTE cellular data network wireless signal transmissions, ad hoc network signal transmissions, radio wave signal transmissions, microwave signal transmissions, infrared signal transmissions, visible light signal transmissions, ultraviolet light signal transmissions, wireless signal transmissions along the electromagnetic spectrum, or some combination thereof.
[0132] The communication interface 1340 can also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining a location of the computing system 1300 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States Global Positioning System (GPS), the Russian Global Navigation Satellite System (GLONASS), the Chinese BeiDou Navigation Satellite System (BDS), and the European Galileo GNSS. There is no restriction on the operating system used with the present architectures, and, therefore, the foundational aspects of the subject matter are readily applicable to other operating systems or platforms.
[0133] The storage device 1330 can be a nonvolatile and / or non-transitory and / or computer-readable memory device and can be a hard disk or other types of computer readable media which can store data that can be accessed by a computer, such as magnetic cassettes, flash memory cards, solid-state memory devices, digital versatile disks, cassette tapes, floppy disks, flexible disks, hard disks, magnetic tape, magnetic strip / stripes, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, compact disc read-only memory (CD-ROM) optical discs, rewritable compact discs (CDs) optical discs, digital video disc (DVD) optical discs, Blu-ray disc (BDD) optical discs, holographic optical discs, another optical medium, a secure digital (SD) card, micro secure digital (microSD) card, Memory Stick® card, Smart Card chip, EMV chip, Subscriber Identity Module (SIM) card, mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EEPROM (FLASHEPROM), cache memory (LI / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or combinations thereof.
[0134] The storage device 1330 can include software services, servers, services, or the like that, when code defining such software is executed by the processor 1310, cause the system to perform a function. In some aspects, a hardware service that performs a particular function can include the software components stored in a computer-readable medium that are necessary to perform the function coupled with the necessary hardware components, such as a processor 1310, the connection 1305, an output device 1335, and the like. The term computer-readable medium includes, but is not limited to, portable or fixed storage devices, optical storage devices, and various other mediums capable of storing, containing or carrying instruction and / or data. A computer-readable medium can include a non-transitory medium in which data can be stored and which does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections.
[0135] Examples of non-transitory media can include, but are not limited to, magnetic-based storage such as hard disks or tape; optical-based storage such as compact discs (CDs) or digital versatile discs (DVDs); flash memory, memory or memory devices; and the like. A computer-readable medium can have stored thereon code and / or machine-executable instructions that can represent a procedure, function, subprogram, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by the passage of information between components of the code segments or the hardware circuit. Information can be passed between components of a code segment or components of a hardware circuit via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0136] In the description above, specific details are provided to provide a thorough understanding of the aspects and examples provided herein. However, a person of ordinary skill in the art will recognize that the application is not limited to the specific details described herein. Thus, although illustrative aspects of the application have been described in detail herein, it is understood that the inventive concepts can be embodied in other ways and that the appended claims are not limited to the details described herein, except insofar as the existing technology limits the application. Various features and aspects of the applications described above can be used individually or jointly. Further, aspects can be utilized in any number of environments and applications beyond the scope of the description, falling within the broader spirit and scope of the application. As such, the application should be regarded as illustrative rather than restrictive. Methods are described in a particular, sequential order. However, it should be appreciated that in other alternatives, the methods can be performed in an order different from that described. The various aspects described herein can be implemented in software and / or hardware, such as within a suitably-programmed general purpose computer, microprocessor, microcontrolled, or other computing system or device.
[0137] For the sake of clarity, in some instances the techniques can be presented in terms of schemes that include particular devices, device components, steps or routines embodied in software or combinations of software and hardware. Additional components, other than those shown and / or described herein, can be utilized. For example, circuits, systems, networks, processes and other components can be shown as components in block diagram form to avoid obscuring the aspects being presented in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques have not been shown in detail to avoid obscuring aspects of the aspects.
[0138] Also, those skilled in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0139] Various aspects can be described herein in terms of processes or methods being performed by functional building blocks, such as modules, circuits, steps, or routines. Although described as processes or methods, the processes or methods can be performed by hardware, software, or a combination of hardware and software. The described processes or methods can be embodied in computer-readable instructions or other computer-readable media, which can be executed by a processor. When the processes or methods are embodied in software, the software can be executed by a processor, such as a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the general purpose processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0140] The processes and methods described above can be implemented using stored computer-executable instructions or computer-executable instructions accessed from a computer-readable medium. Such instructions can comprise, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessed over a network. Computer-executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used by the described examples, and / or information created during performance of the described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, etc.
[0141] In some aspects, computer-readable storage devices, media, and memories can include cables or wireless signals containing bitstreams and the like. However, where mentioned, non-transitory computer-readable storage media expressly excludes media such as power supply, carrier waves, electromagnetic waves, and signals per se.
[0142] Those skilled in the art will understand that information and signals can be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that can be referenced throughout the above description can be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof consistent with the particular application, as would be understood by one of ordinary skill in the art.
[0143] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein can be implemented or performed with a hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof and can be embodied in any of a number of various forms. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program products) for performing the necessary tasks can be stored in a computer-readable or machine-readable medium. A processor(s) can execute the necessary tasks. Examples of the various shapes include: a laptop device, a smart phone, a mobile phone, a tablet device, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, and the like. The functionality described herein can also be embodied in peripheral devices or in interposers. By way of further example, such functionality can also be implemented on circuit boards among different chips or different processes executing on a single device.
[0144] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described by the present disclosure.
[0145] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses such as application specific integrated circuits (ASICs), or other devices having multiple uses such as peripheral devices (e.g., hard drives or similar). Any features described as modules or components can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms and / or operations described above. The computer-readable data storage medium can form part of a computer program product, which can include packaging materials. The computer-readable medium can comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques can be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0146] The program code can be executed by a processor, which can include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor; but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein can refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.
[0147] Those of ordinary skill in the art will appreciate that the less than (“<”) and greater than (“>”) symbols or terms used herein can be replaced with less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of this description.
[0148] Where components are described as being "configured to" perform certain operations, such configuration can be accomplished, for example, by designing electronic circuitry or other hardware to perform the operation, by programming programmable electronic circuitry (e.g., microprocessors or other suitable electronic circuits) to perform the operation, or any combination thereof.
[0149] The phrase "coupled to" means that any component is directly or indirectly physically connected to another component, and / or that any component is directly or indirectly in communication with another component (e.g., connected to another component through a wired or wireless connection and / or other suitable communication interface).
[0150] Claim language reciting "at least one of a set" and / or "one or more of a set"indicates that one member of the set or multiple members of the set satisfy the claim. For example, claim language expounding "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language expounding "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. Language reciting "at least one of a set" and / or "one or more of a set" does not limit the set to the items listed. For example, claim language expounding "at least one of A and B" or "at least one of A or B" can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.
[0151] Exemplary aspects of the present disclosure include the following:
[0152] Aspect 1. An apparatus for tracking objects performing point graph registration, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: obtain an image comprising a scene; determine a first point graph representing the scene based on the image, wherein the first point graph comprises a first plurality of 3D points representing the scene; determine a first plurality of point groupings, wherein each point grouping of the first plurality of point groupings is determined based on proximity of pairs of points of the first plurality of 3D points, wherein each point grouping comprises two or more 3D points from the first plurality of 3D points; obtain a second point graph representing the scene, wherein the second point graph comprises a second plurality of 3D points representing the scene; obtain a second plurality of point groupings, wherein each point grouping of the second plurality of point groupings comprises a plurality of 3D points from the second plurality of 3D points; determine a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings; and determine an alignment transform for aligning the first plurality of 3D points and the second plurality of 3D points based on the correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings.
[0153] Aspect 2. The apparatus of aspect 1, wherein determining a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings comprises determining that at least one or more of the respective point groupings of the first plurality of point groupings or the respective point groupings of the second plurality of point groupings do not have a corresponding respective point grouping.
[0154] Aspect 3. The apparatus of any one of aspects 1-2, wherein the at least one processor is configured to: apply the alignment transform to one of the second plurality of 3D points or the first plurality of 3D points representing the scene to generate a transformed first plurality of 3D points; and determine a correspondence between the transformed first plurality of 3D points and the other of the first plurality of 3D points or the second plurality of 3D points.
[0155] Aspect 4. The apparatus of any one of aspects 1 through 3, wherein, to determine the first plurality of point groupings, the at least one processor is configured to: determine, for a first point of the first plurality of 3D points, that a first Euclidean distance between the first point of the first plurality of 3D points and a second point of the first plurality of 3D points is less than a threshold distance, wherein the second point of the first plurality of 3D points is adjacent to the first point of the first plurality of 3D points; and based on determining that the first Euclidean distance between the first point of the first plurality of 3D points and the second point of the first plurality of 3D points is less than the threshold distance, group the first point and the second point into a first point grouping of the first plurality of point groupings.
[0156] Aspect 5. The apparatus of any one of aspects 1 through 4, wherein, to determine the first plurality of point groupings, the at least one processor is configured to: determine, for the second point of the first plurality of 3D points, that a second Euclidean distance between the second point of the first plurality of 3D points and a third point of the first plurality of 3D points is greater than the threshold distance; and based on determining that the second Euclidean distance between the second point of the first plurality of 3D points and the third point of the first plurality of 3D points is greater than the threshold distance, exclude the third point of the first plurality of 3D points from the first point grouping of the first plurality of point groupings.
[0157] Aspect 6. The apparatus of any one of aspects 1 through 5, wherein the at least one processor is configured to determine, based on determining a correspondence between a first point grouping of the first plurality of point groupings and a second point grouping of the second plurality of point groupings, a correspondence between respective points of the first point grouping of the first plurality of point groupings and respective points of the second point grouping of the second plurality of point groupings.
[0158] Aspect 7. The apparatus of any one of aspects 1 through 6, wherein, to determine the correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings, the at least one processor is configured to: apply a sequential ordering to points included in each respective point grouping of the first plurality of point groupings based on respective locations of each pixel included in the respective point grouping along an axis; and apply the sequential ordering to points included in each point grouping of the second plurality of point groupings based on respective locations of each pixel included in the respective point grouping along the axis.
[0159] Aspect 8. The apparatus of any one of aspects 1-7, wherein the axis corresponds to at least one or more of a direction of motion of the tracked object or an optical axis associated with the image.
[0160] Aspect 9. The apparatus of any one of aspects 1-8, wherein, to determine the correspondence between individual point groupings of the first plurality of point groupings and individual point groupings of the second plurality of point groupings, the at least one processor is configured to: determine a dynamic time warping distance between each point grouping of the first plurality of point groupings and each point grouping of the second plurality of point groupings; and determine the correspondence between individual point groupings of the first plurality of point groupings and individual point groupings of the second plurality of point groupings based on the determined dynamic time warping distances.
[0161] Aspect 10. The apparatus of any one of aspects 1-9, wherein individual point groupings of the first plurality of point groupings correspond to lines in the image.
[0162] Aspect 11. The apparatus of any one of aspects 1-10, wherein the tracked object is a tracked vehicle.
[0163] Aspect 12. The apparatus of any one of aspects 1-11, wherein the apparatus of the tracked object is a computing system of the tracked vehicle.
[0164] Aspect 13. The apparatus of any one of aspects 1-12, wherein the at least one processor is configured to: align the first plurality of 3D points and the second plurality of 3D points based on the alignment transform; and determine a position of the tracked object relative to a second point map based on aligning the first plurality of 3D points and the second plurality of 3D points, wherein the second point map comprises a reference point map.
[0165] Aspect 14. The apparatus of any one of aspects 1-13, wherein the at least one processor is configured to: perform a navigation operation based on the position.
[0166] Aspect 15. The apparatus of any one of aspects 1-14, wherein the reference point map comprises a HD reference point map.
[0167] Aspect 16. The apparatus of any one of aspects 1-15, wherein the reference point map is obtained from a reference map service.
[0168] Aspect 17. A method for performing point graph registration, the method comprising: obtaining an image comprising a scene; determining a first point graph representing the scene based on the image, wherein the first point graph comprises a first plurality of 3D points representing the scene; determining a first plurality of point groupings, wherein each point grouping of the first plurality of point groupings is determined based on proximity of pairs of points of the first plurality of 3D points, wherein each point grouping comprises two or more 3D points from the first plurality of 3D points; obtaining a second point graph representing the scene, wherein the second point graph comprises a second plurality of 3D points representing the scene; obtaining a second plurality of point groupings, wherein each point grouping of the second plurality of point groupings comprises a plurality of 3D points from the second plurality of 3D points; determining a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings; and determining an alignment transform for aligning the first plurality of 3D points and the second plurality of 3D points based on the correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings.
[0169] Aspect 18. The method of aspect 17, wherein determining a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings comprises determining that at least one or more of the respective point groupings of the first plurality of point groupings or the respective point groupings of the second plurality of point groupings do not have a corresponding respective point grouping.
[0170] Aspect 19. The method of any one of aspects 17-18, further comprising: applying the alignment transform to one of the second plurality of 3D points or the first plurality of 3D points representing the scene to generate a transformed first plurality of 3D points; and determining a correspondence between the transformed first plurality of 3D points and the other of the first plurality of 3D points or the second plurality of 3D points.
[0171] Aspect 20. The method of any one of aspects 17-19, further comprising: for a first point of the first plurality of 3D points, determining that a first Euclidean distance between the first point of the first plurality of 3D points and a second point of the first plurality of 3D points is less than a threshold distance, wherein the second point of the first plurality of 3D points is adjacent to the first point of the first plurality of 3D points; and based on determining that the first Euclidean distance between the first point of the first plurality of 3D points and the second point of the first plurality of 3D points is less than the threshold distance, grouping the first point and the second point into a first point grouping of the first plurality of point groupings.
[0172] Aspect 21. The method of any one of aspects 17 to 20, further comprising: determining, for the second point of the first plurality of 3D points, that a second Euclidean distance between the second point of the first plurality of 3D points and a third point of the first plurality of 3D points is greater than the threshold distance; and based on determining that the second Euclidean distance between the second point of the first plurality of 3D points and the third point of the first plurality of 3D points is greater than the threshold distance, excluding the third point of the first plurality of 3D points from the first point group of the first plurality of point groups.
[0173] Aspect 22. The method of any one of aspects 17 to 21, further comprising: based on determining the correspondence between the first point group of the first plurality of point groups and the second point group of the second plurality of point groups, determining a correspondence between respective points of the first point group of the first plurality of point groups and respective points of the second point group of the second plurality of point groups.
[0174] Aspect 23. The method of any one of aspects 17 to 22, further comprising: based on respective locations of each pixel included in each respective point group of the first plurality of point groups along an axis, applying an ordinal ordering to points included in each point group of the first plurality of point groups; and based on respective locations of each pixel included in each respective point group of the second plurality of point groups along the axis, applying the ordinal ordering to points included in each point group of the second plurality of point groups.
[0175] Aspect 24. The method of any one of aspects 17 to 23, wherein the axis corresponds to at least one or more of a direction of motion of a tracked object or an optical axis associated with the image.
[0176] Aspect 25. The method of any one of aspects 17 to 24, further comprising: determining a dynamic time warping distance between each point group of the first plurality of point groups and each point group of the second plurality of point groups; and based on the determined dynamic time warping distances, determining a correspondence between respective point groups of the first plurality of point groups and respective point groups of the second plurality of point groups.
[0177] Aspect 26. The method of any one of aspects 17 to 25, wherein respective point groups of the first plurality of point groups correspond to lines in the image.
[0178] Aspect 27. The method of any one of aspects 17 to 26, further comprising: aligning the first plurality of 3D points and the second plurality of 3D points based on the alignment transform; and determining a position of a tracked object relative to the second point map based on aligning the first plurality of 3D points and the second plurality of 3D points, wherein the second point map comprises a reference point map.
[0179] Aspect 28. The method of any one of aspects 17 to 27, further comprising performing a navigation operation based on the position.
[0180] Aspect 29. The method of any one of aspects 17 to 28, wherein the reference point map comprises a HD reference point map.
[0181] Aspect 30. The method of any one of aspects 17 to 29, wherein the reference point map is obtained from a reference map service.
[0182] Aspect 31. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform any of the operations of aspects 1 to 30.
[0183] Aspect 32. An apparatus comprising means for performing any of the operations of aspects 1 to 30.
Claims
1. An apparatus for tracking objects that performs point graph registration, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: obtain an image comprising a scene; determine a first point graph representing the scene based on the image, wherein the first point graph comprises a first plurality of three-dimensional (3D) points representing the scene; determine a first plurality of point groupings, wherein each point grouping of the first plurality of point groupings is determined based on proximity of pairs of points of the first plurality of 3D points, wherein each point grouping comprises two or more 3D points from the first plurality of 3D points; obtain a second point graph representing the scene, wherein the second point graph comprises a second plurality of 3D points representing the scene; obtain a second plurality of point groupings, wherein each point grouping of the second plurality of point groupings comprises a plurality of 3D points from the second plurality of 3D points; determine a correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings; and based on the correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings, determine an alignment transform for aligning the first plurality of 3D points and the second plurality of 3D points.
2. The apparatus of claim 1, wherein determining the correspondence between respective point groupings of the first plurality of point groupings and respective point groupings of the second plurality of point groupings comprises determining that at least one or more of the respective point groupings of the first plurality of point groupings or the respective point groupings of the second plurality of point groupings do not have a corresponding respective point grouping.
3. The apparatus of claim 1, wherein the at least one processor is configured to: apply the transform to one of the second plurality of 3D points or the first plurality of 3D points representing the scene to generate a transformed first plurality of 3D points; and determine a correspondence between the transformed first plurality of 3D points and the other of the first plurality of 3D points or the second plurality of 3D points. To determine the first plurality of point groupings, the at least one processor is configured to:
4. The apparatus of claim 1, wherein, for a first point of the first plurality of 3D points, determine that a first Euclidean distance between the first point of the first plurality of 3D points and a second point of the first plurality of 3D points is less than a threshold distance, wherein the second point of the first plurality of 3D points is adjacent to the first point of the first plurality of 3D points; and based on determining that the distance between the first point of the first plurality of 3D points and the second point of the first plurality of 3D points that is adjacent to the first point is less than the threshold distance, group the first point and the second point into a first point grouping of the first plurality of point groupings. To determine the first plurality of point groupings, the at least one processor is configured to: for a first point of the first plurality of 3D points, determine that a first Euclidean distance between the first point of the first plurality of 3D points and a second point of the first plurality of 3D points is less than a threshold distance, wherein the second point of the first plurality of 3D points is adjacent to the first point of the first plurality of 3D points; and 5. The apparatus of claim 4, wherein, based on determining that the distance between the first point of the first plurality of 3D points and the second point of the first plurality of 3D points that is adjacent to the first point is less than the threshold distance, group the first point and the second point into a first point grouping of the first plurality of point groupings. determining, for the second point of the first plurality of 3D points, that a second Euclidean distance between the second point of the first plurality of 3D points and a third point of the first plurality of 3D points is greater than the threshold distance; and based on determining that the second Euclidean distance between the second point of the first plurality of 3D points and the third point of the first plurality of 3D points is greater than the threshold distance, excluding the third point of the first plurality of 3D points from the first point group of the first plurality of point groups.
6. The apparatus of claim 1, wherein the at least one processor is configured to determine, based on determining a correspondence between a first point group of the first plurality of point groups and a second point group of the second plurality of point groups, a correspondence between respective points of the first point group of the first plurality of point groups and respective points of the second point group of the second plurality of point groups.
7. The apparatus of claim 1, wherein, To determine the correspondence between respective point groups of the first plurality of point groups and respective point groups of the second plurality of point groups, the at least one processor is configured to: apply a sequential ordering to points included in each point group of the first plurality of point groups based on respective locations of each pixel included in each respective point group of the first plurality of point groups along an axis; and apply the sequential ordering to points included in each point group of the second plurality of point groups based on respective locations of each pixel included in each respective point group of the second plurality of point groups along the axis.
8. The apparatus of claim 7, wherein the axis corresponds to at least one or more of a direction of motion of the tracked object or an optical axis associated with the image.
9. The apparatus of claim 1, wherein, To determine the correspondence between respective point groups of the first plurality of point groups and respective point groups of the second plurality of point groups, the at least one processor is configured to: determine a dynamic time warping distance between each point group of the first plurality of point groups and each point group of the second plurality of point groups; and determine the correspondence between respective point groups of the first plurality of point groups and respective point groups of the second plurality of point groups based on the dynamic time warping distance.
10. The apparatus of claim 1, wherein respective point groups of the first plurality of point groups correspond to lines in the image.
11. The apparatus of claim 1, wherein the tracked object is a tracked vehicle.
12. The apparatus of claim 11, wherein the apparatus of the tracked object is a computing system of the tracked vehicle.
13. The apparatus of claim 1, wherein the at least one processor is configured to: align the first plurality of 3D points and the second plurality of 3D points based on the alignment transform; and determine a position of the tracked object relative to a second point map based on aligning the first plurality of 3D points and the second plurality of 3D points, wherein the second point map comprises a reference point map.
14. The apparatus of claim 13, wherein the at least one processor is configured to perform a navigation operation based on the positioning.
15. The apparatus of claim 13, wherein the reference point map comprises a high definition (HD) reference point map.
16. The apparatus of claim 15, wherein the reference point map is obtained from a reference map service.
17. A method for performing point map registration, the method comprising: obtaining an image comprising a scene; determining a first point map representing the scene based on the image, wherein the first point map comprises a first plurality of 3D points representing the scene; determining a first plurality of point groupings, wherein each point grouping of the first plurality of point groupings is determined based on proximity of pairs of points of the first plurality of 3D points, wherein each point grouping comprises two or more 3D points from the first plurality of 3D points; obtaining a second point map representing the scene, wherein the second point map comprises a second plurality of 3D points representing the scene; obtaining a second plurality of point groupings, wherein each point grouping of the second plurality of point groupings comprises a plurality of 3D points from the second plurality of 3D points; determining a correspondence between each point grouping of the first plurality of point groupings and each point grouping of the second plurality of point groupings; and based on the correspondence between each point grouping of the first plurality of point groupings and each point grouping of the second plurality of point groupings, determining an alignment transform for aligning the first plurality of 3D points and the second plurality of 3D points.
18. The method of claim 17, wherein determining a correspondence between each point grouping of the first plurality of point groupings and each point grouping of the second plurality of point groupings comprises determining that at least one or more of the each point grouping of the first plurality of point groupings or the each point grouping of the second plurality of point groupings does not have a corresponding each point grouping.
19. The method of claim 17, further comprising: applying the alignment transform to one of the second plurality of 3D points or the first plurality of 3D points representing the scene to generate a transformed first plurality of 3D points; and determining a correspondence between the transformed first plurality of 3D points and the other of the first plurality of 3D points or the second plurality of 3D points.
20. The method of claim 17, further comprising: for a first point of the first plurality of 3D points, determining that a first Euclidean distance between the first point of the first plurality of 3D points and a second point of the first plurality of 3D points is less than a threshold distance, wherein the second point of the first plurality of 3D points is adjacent to the first point of the first plurality of 3D points; and based on determining that the first Euclidean distance between the first point of the first plurality of 3D points and the second point of the first plurality of 3D points is less than the threshold distance, grouping the first point and the second point into a first point grouping of the first plurality of point groupings.
21. The method of claim 20, further comprising: determining, for the second point of the first plurality of 3D points, that a second Euclidean distance between the second point of the first plurality of 3D points and a third point of the first plurality of 3D points is greater than the threshold distance; and based on determining that the second Euclidean distance between the second point of the first plurality of 3D points and the third point of the first plurality of 3D points is greater than the threshold distance, excluding the third point of the first plurality of 3D points from the first point group of the first plurality of point groups.
22. The method of claim 17, further comprising: based on determining the correspondence between the first point group of the first plurality of point groups and the second point group of the second plurality of point groups, determining a correspondence between each point of the first point group of the first plurality of point groups and each point of the second point group of the second plurality of point groups.
23. The method of claim 17, further comprising: applying a sequential ordering to points included in each respective point group of the first plurality of point groups based on respective locations of each pixel included in each respective point group of the first plurality of point groups along an axis; and applying the sequential ordering to points included in each point group of the second plurality of point groups based on respective locations of each pixel included in each respective point group of the second plurality of point groups along the axis.
24. The method of claim 23, wherein the axis corresponds to at least one or more of a direction of motion of a tracked object or an optical axis associated with the image.
25. The method of claim 17, further comprising: determining a dynamic time warping distance between each point group of the first plurality of point groups and each point group of the second plurality of point groups; and based on the dynamic time warping distance, determining a correspondence between each point group of the first plurality of point groups and each point group of the second plurality of point groups.
26. The method of claim 17, wherein each point group of the first plurality of point groups corresponds to a line in the image.
27. The method of claim 17, further comprising: aligning the first plurality of 3D points and the second plurality of 3D points based on the alignment transform; and based on aligning the first plurality of 3D points and the second plurality of 3D points, determining a position of a tracked object relative to the second point map, wherein the second point map comprises a reference point map.
28. The method of claim 27, further comprising performing a navigation operation based on the position.
29. The method of claim 27, wherein the reference point map comprises an HD reference point map.
30. The method of claim 27, wherein the reference point map is obtained from a reference map service.