Map creation and localization for autonomous driving application
By utilizing consumer-grade sensors and crowdsourcing data from multiple vehicles, the method addresses the challenges of high-cost sensor requirements and limited HD map availability, achieving accurate and cost-effective HD map creation and localization for autonomous driving.
Patent Information
- Application Number
- JP2025029758
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-08-31
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-10
AI Technical Summary
Conventional high-definition (HD) map generation for autonomous driving is hindered by the high cost of sophisticated sensors, limited availability of HD maps in remote areas, and the need for frequent updates due to changing road conditions.
A method for map creation and localization using consumer-grade sensors, where data is crowdsourced from multiple vehicles and processed to minimize bandwidth and memory requirements, enabling the generation of fused HD maps that can be used for localization with high accuracy.
This approach allows for the creation of accurate and reliable HD maps in a cost-effective manner, enabling precise localization even in areas with limited sensor data, and facilitates frequent updates to account for changing road conditions.
Smart Images

Figure 2025087751000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to map creation and localization for autonomous driving applications.
Background Art
[0002] Mapping and localization are extremely important processes for autonomous driving functionality. High definition (HD) maps, sensor perception, or a combination thereof are often used to localize a vehicle with respect to an HD map in order to make planning and control decisions. Typically, conventional HD maps are generated using a survey vehicle equipped with sophisticated and very accurate sensors. However, these sensors are prohibitively expensive to implement in consumer vehicles. In a normal deployment, a survey vehicle has the ability to generate an HD map suitable for localization after a single drive. Unfortunately, due to the lack of survey vehicles with such high-cost sensors, the availability of HD maps for any given area can be concentrated around large cities or bases where the survey vehicles operate. For more remote areas, areas with less human traffic, and areas further away from the bases of these activities, the data used to generate an HD map may be collected from only a single drive - i.e., if the data is even available at all. As a result, a single drive can result in low-quality or otherwise unsuitable sensor data for this purpose - for example, due to occlusion, moving objects, adverse weather effects, debris, construction artifacts, temporary hardware failures, and / or other problems that can degrade the quality of the collected sensor data - in which case the HD map generated from the sensor data may not be safe or reliable for use in localization.
[0003] To correct these quality concerns, another survey vehicle may be required to perform another drive at the location where the quality has been impaired. However, the identification, deployment, sensor - data generation, and map - update process may require a long period by a systematic data - collection and map - creation process, and can render the HD map unusable until the update is performed. If road conditions or layout change frequently or drastically over time - e.g., due to construction - there may be no mechanism to identify these changes, and even if identified, there is no way to generate updated data without deploying another survey vehicle, so the problem worsens.
[0004] In addition, consumer vehicles may not be equipped with the same high - quality, high - cost sensors, so localization to the HD map - even when available - may not be possible using a number of sensor modalities - e.g., cameras, LiDAR, RADAR, etc. - because the quality and type of data do not match the data used to generate the HD map. As a result, localization relies only on global navigation satellite system (GNSS) data and, in some specific situations, still achieves only a few meters of accuracy and / or has a variable accuracy - even for the most expensive and accurate sensor models. The few - meter inaccuracy can place the vehicle in a lane different from the current driving lane or on one side of the road other than the one currently being traveled. Thus, the use of conventional solutions for generating HD maps can result in inaccurate maps that cause significant obstacles to the achievement of safe or reliable highly autonomous vehicles (e.g., levels 3, 4, and 5) when impaired by that inaccurate localization. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0005] Embodiments of the present disclosure relate to a method for map creation and localization for autonomous driving applications. Specifically, embodiments of the present disclosure include data generation, map creation using the generated data, and an end-to-end system for localization to a created map that can be used with consumer sensors common in commercially available vehicles. For example, during the data generation process, a data collection vehicle and / or a consumer vehicle using consumer quality sensors can be used to generate sensor data. The resulting data can correspond to a map stream - sensor data, perceptual output from a deep neural network (DNN), and / or relative trajectory (e.g., rotation and translation) data corresponding to any number of drives by any number of vehicles. As such, in contrast to the systematic data collection efforts of conventional systems, the present system can crowdsource data generation using a large number of vehicles and a large number of drives. To reduce the bandwidth and memory requirements of the system, the data from the map stream is minimized (e.g., by removing moving objects with a filter, performing LiDAR plane slicing or LiDAR point reduction, converting perceptual or camera-based output to 3D position information, running a campaign for only specific data types, etc.) and / or compressed (e.g., using delta compression techniques). As a result of the map stream data being generated using consumer-grade sensors, the sensor data - after being converted to a map format for localization - can be used directly for localization without relying solely on GNSS data. Further, since relative trajectory information corresponding to each drive is tracked, this information can be used to generate individual road segments (e.g., road segments of sizes such as 25 meters, 50 meters, etc.) to which localization can be performed, thereby enabling localization accuracy within the centimeter range.
[0006] During map creation, the map stream can be used to generate map data representing data generated over multiple drives - and ultimately, a fused HD map. Additionally, when a new map stream is generated, these additional drives can be merged, combined, or integrated with existing map stream data and used to further enhance the robustness of the HD map. For example, each map stream can be converted into a respective map, and any number of drive segments from any number of maps (or, corresponding map streams) can be used to generate a fused HD map representation of a particular drive segment. Pairs of drive segments can be geometrically registered with respect to each other to determine pose links representing the rotation and translation between the poses (or, frames) of the pair of drives. A frame graph representing the pose links can be split into road segments - for example, road segments used for relative localization - and the poses corresponding to each road segment can be optimized. The resulting finalized poses within each segment can be used to fuse various sensor data and / or perception outputs for generating the final fused HD map. As a result, and since the map data corresponds to consumer quality sensors, sensor data and / or perception results (e.g., landmark positions) from the HD map can be used directly for localization (e.g., by comparing current real-time sensor data and / or perception with corresponding map information) in addition to using GNSS data in the embodiments.
[0007] For example, when localizing to a fused HD map, individual localization results can be generated based on a comparison of the sensor data and / or perceptual output from a certain sensor modality with the map data corresponding to the same sensor modality. For example, a cost space can be sampled in each frame using data corresponding to the sensor modality, an aggregated cost space can be generated using a plurality of individual cost spaces, and filtering (e.g., using a Kalman filter) can be used to finalize with respect to the localization results of a particular sensor modality. This process can be repeated for any number of sensor modalities - for example, LiDAR, RADAR, camera, etc. - and the results can be fused together to determine the final fused localization result for the current frame. The fused localization result can then be advanced to the next frame and used to determine the fused localization for the next frame, and so on. As a result of the HD map including individual road segments for localization and the results of each road segment having a corresponding global position, a global localization result can also be realized when the vehicle localizes in a local or relative coordinate system corresponding to the road segment.
[0008] The present system and method for map creation and localization for autonomous driving applications will be described in detail below with reference to the accompanying drawings.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 5E
Figure 5F
Figure 6A
Figure 6B
Figure 6C
Figure 6D
Figure 6E
Figure 6F
Figure 6G
Figure 6H
Figure 6I
Figure 7
Figure 8A
Figure 8B
Figure 9A
Figure 9B
Figure 9C
Figure 10A
Figure 10B
Figure 10C
Figure 11A
Figure 11B
Figure 12A
Figure 12B
Figure 12C
Figure 12D
Figure 13
Figure 14
Figure 15A
Figure 15B
Figure 15C
Figure 15D
Figure 16
Figure 17
DETAILED DESCRIPTION OF THE INVENTION
[0010] Systems and methods related to map creation and localization for autonomous driving applications are disclosed. This disclosure may be described with respect to an exemplary autonomous vehicle 1500 (an example of which is described herein with respect to FIGS. 15A-15D and is alternatively referred to herein as "vehicle 1500" or "ego vehicle 1500"), but this is not intended to be limiting. For example, the systems and methods described herein may be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more advanced driver assistance systems (ADAS)), robots, warehouse vehicles, off-road vehicles, flying vessels, boats, and / or other types of vehicles. Additionally, this disclosure may be described with respect to autonomous driving, but this is not intended to be limiting. For example, the systems and methods described herein may be used in robotics (e.g., mapping and localization for robotics), aviation systems (e.g., mapping and localization for drones or other aircraft), marine systems (e.g., mapping and localization for boats), simulation environments (e.g., for mapping and localization of virtual vehicles in a virtual simulation environment), and / or, for example, in other technical fields for data generation and curation, map creation, and / or localization.
[0011] Referring to FIG. 1, FIG. 1 is a data flow diagram of a process 100 of a map creation and localization system according to some embodiments of the present disclosure. It should be understood that this and other configurations described herein are merely described as examples. Other configurations and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be completely omitted. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components and in any suitable combination and location. The various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor that executes instructions stored in a memory.
[0012] In some embodiments, vehicle 1500 may include similar components, features, and / or functionality of vehicle 1500 described herein with respect to FIGS. 15A - 15D. Additionally, map creation 106 may, in an embodiment, be executed at a data center and may be executed using components, features, and / or functionality similar to those described herein with respect to exemplary computing device 1600 and / or exemplary data center 1700. In some embodiments, the entire end - to - end process 100 of FIG. 1 may be executed within a single vehicle 1500. Although only a single vehicle 1500 is shown in FIG. 1, this is not intended to be limiting. For example, any number of vehicles 1500 may be used to generate sensor data 102 for map stream generation 104, and any number of (different) vehicles 1500 may be used to generate sensor data 102 for localization 110. Additionally, in addition to sensor configurations and / or other vehicle attributes, the type, model, year, and / or type of vehicle may be the same, similar, and / or different for each vehicle 1500 used in the map stream generation 104, map creation 106, and / or localization 110 processes.
[0013] Process 100 may include operations for map stream generation 104, map creation 106, and localization 110. For example, process 100 may be executed as part of an end-to-end system that relies on a map stream generated using sensor data 102 from any of several vehicles 1500 over any number of drives, map creation 106 that uses received map stream data from vehicle 1500, and localization 110 to a map (e.g., a high-definition (HD) map) generated using the map creation 106 process. The map represented by map data 108 may, in some non-limiting embodiments, include data collected from different sensor modalities or for individual sensors of an individual sensor modality or maps generated by such data. For example, map data 108 may represent a first map (or, map layer) corresponding to image-based localization, a second map (or, map layer) corresponding to LiDAR-based localization, a third map (or, map layer) corresponding to RADAR-based localization, and so on. In some embodiments, for example, image-based localization may be performed using, for example, a first map (or, map layer) corresponding to a forward-facing camera and a second map (corresponding or not corresponding to the same map layer) corresponding to a rear-facing camera. As such, depending on the sensor configuration of vehicle 1500 that receives map data 108 and performs its localization 110, only the necessary portions of map data 108 may be transmitted to vehicle 1500. As a non-limiting example, if a first vehicle 1500 includes cameras and RADAR sensors but does not include a LiDAR sensor, the camera map layer and the RADAR map layer may be transmitted to vehicle 1500, and the LiDAR map layer may not be transmitted. As a result, since the LiDAR map layer may not be transmitted and / or stored in vehicle 1500, the memory usage in vehicle 1500 is reduced and the bandwidth is maintained.
[0014] In addition, when vehicle 1500 uses map data 108 for localization 110 to navigate the environment, a health check 112 can be performed to ensure that the map is up-to-date or accurate, taking into account changes in road conditions, road structures, construction, and / or the like. As such, if a portion or segment of the map is determined to be of less than the desired quality - for example, as a result of difficulty localizing to the map - map stream generation 104 can be executed and used to update the map corresponding to the specific portion of the segment via map creation 106. As a result, the end-to-end system can be used not only to generate maps for localization 110, but also to ensure that the maps are kept up-to-date for accurate localization over time. Map stream generation 104, map creation 106, and the localization 116 process are described in further detail herein, respectively.
[0015] Map stream generation To generate a map stream, any number of vehicles 1500 - for example, consumer vehicles, data collection vehicles, combinations thereof - may perform any number of drives. For example, each vehicle 1500 may drive on various road segments from locations in a town, city, state, country, continent, and / or around the world, and during the drive, any number of sensors - for example, LiDAR sensor 1564, RADAR sensor 1560, cameras 1568, 1570, 1572, 1574, 1598, etc., inertial measurement unit (IMU) sensor 1566, ultrasonic sensor 1562, microphone 1596, speed sensor 1544, steering sensor 1540, global navigation satellite system (GNSS) sensor 1558, etc. - may be used to generate sensor data 102. Each individual vehicle 1500 can generate sensor data 102 and may use the sensor data 102 for map stream generation 104 corresponding to a particular drive of the vehicle 1500. Map streams generated to correspond to different drives from a single vehicle 1500 and drives from any number of other vehicles 1500 may be used in map creation 106, as described in more detail herein.
[0016] As a result of multiple map streams being used to generate map data 108 for any particular road segment, the individual map streams from each drive need not be highly accurate or highly faithful as in conventional systems. For example, conventional systems use survey vehicles equipped with sensor types that are prohibitively expensive and thus not preferable for installation in consumer vehicles (e.g., because the cost of the vehicle would increase significantly). However, the sensors of these survey vehicles can produce sensor data that is reliable enough even after a single drive. However, the negative aspect is that if a particular drive includes many dynamic or transient factors, such as traffic, construction artifacts, debris, obstructions, adverse weather effects, temporary hardware failures, or other sources of concerns about sensor data quality, and is not limited to these, a single drive may not produce data suitable for generating an accurate map for localization. In addition, since road conditions change and the number of available survey vehicles is small, the map may not be updated quickly - for example, the map may not be updated until another survey vehicle traverses the same route. In contrast, in the system of the present disclosure, by leveraging consumer vehicles with low-cost mass-market sensors, any number of map streams from any number of drives can be used to generate map data 108 more quickly and more frequently. As a result, individual map streams from drives in the presence of obstructions or other quality concerns can be relied on less, and map streams from higher-quality sensor data can be relied on more. In addition, when road structure, layout, conditions, surroundings, and / or other information change, a health check 112 can be performed more quickly - for example, in real-time or substantially in real-time - to update map data 108. The result of this process is a more cloud-sourced approach to map stream generation rather than a systematic data collection effort of conventional approaches.
[0017] Referring now to FIG. 2, FIG. 2 is a data flow diagram of a map stream generation process 104 according to some embodiments of the present disclosure. For example, the process 104 may correspond to generating a map stream from a single drive by the vehicle 1500. This process 104 may be repeated by any number of vehicles 1500 over any number of drives. Sensor data 102, as described herein, may correspond to sensor data 102 from any number of different sensor modalities and / or any number of sensors of a single modality. For example, the sensor data 102 may be the sensor types described herein with respect to the vehicle 1500 of FIGS. 15A - 15D - for example, GNSS sensor 1558 (e.g., a global navigation satellite system sensor), RADAR sensor 1560, ultrasonic sensor 1562, LiDAR sensor 1564, ultrasonic sensor, IMU sensor 1566 (e.g., an accelerometer, gyroscope, magnetic compass, magnetometer, etc.), microphone 1576, stereo camera 1568, wide view camera 1570 (e.g., a fish eye camera), infrared camera 1572, surround camera 1574 (e.g., a 360 - degree camera), long - range and / or mid - range camera 1578, speed sensor 1544 (e.g., for measuring the speed of the vehicle 1500), and / or any other sensor type. In some embodiments, the sensor data 102 may be included directly in the map stream 210, for example, with or without compression using a data compressor 208. For example, for LiDAR data and / or RADAR data, the detections represented by the sensor data 102 may be in a three - dimensional (3D) coordinate space (e.g., world space), and the LiDAR points and / or RADAR points (or detections) may be used directly to generate a LiDAR map layer and / or a RADAR map layer, respectively. In some embodiments, the sensor data 102 may be converted from a two - dimensional (2D) coordinate space (e.g., image space) to a 3D coordinate space - for example, using a data converter 206 - and then included in the map stream 210 (e.g., after compression using a data compressor 208 in an embodiment).
[0018] In some embodiments, LiDAR slicing can be performed on LiDAR data - for example, on a point cloud generated using raw LiDAR data - to slice the LiDAR data into different altitude ranges. The LiDAR data can be sliced into any number of altitude ranges. In some embodiments, the LiDAR altitude ranges can be defined relative to the origin or rig of the vehicle 1500. As a non-limiting example, the LiDAR data can be sliced into a ground slice (e.g., 5 meters to 300 meters relative to the origin of the vehicle 1500), a giraffe plane slice (e.g., 2.5 meters to 5 meters relative to the origin of the vehicle 1500), and / or a ground plane slice (e.g., -2.5 meters to 0.5 meters relative to the origin of the vehicle 1500). When LiDAR slicing is performed, different slices can be stored as separate LiDAR layers within the map stream 210 and / or used to generate separate LiDAR map layers during map creation 106 for localization 110. In some embodiments, data corresponding to certain slices can be filtered out or removed so that less data is encoded into the map stream 210 and less data is sent to the cloud for map creation 106. For example, detections far from the ground plane may not be usable, accurate, and / or clear enough for localization, so the ground slice may be less valuable than the giraffe plane slice or the ground plane slice. In such an example, the ground slice (e.g., 5 meters to 300 meters) can be filtered out.
[0019] Sensor data 102 can include data that can be used to track the absolute position of vehicle 1500 and / or the local or relative position of vehicle 1500 - for example, data generated by IMU sensor 1566, GNSS sensor 1558, speed sensor 1544, camera sensor, LiDAR sensor 1564, and / or other sensor types. For example, in each frame of sensor data 102 generated from any number of sensor modalities, a global position - for example, using GNSS sensor 1558 - can be recorded for that frame of sensor data 102. This global position - for example, in the WGS84 reference system - can be used to generally position vehicle 1500 within the global coordinate system. However, when using an HD map for autonomous driving operations, world-scale localization accuracy is not as valuable as localization accuracy within a given road segment. For example, when driving on Highway 101 in Santa Clara, California, the position of vehicle 1500 relative to Interstate 95 in Boston, Massachusetts, is not as critical as the position of the vehicle several meters ahead on Highway 101, such as 50 meters, 100 meters, 1.609 kilometers (1 mile), etc. In addition, GNSS sensors - even the highest quality ones - can be inaccurate within a range exceeding 5 or 10 meters, and thus, global localization can result in low accuracy and / or accuracy that is subject to change due to satellite relative positioning. This can still be the case - for example, global localization can be off by 5 or 10 meters - but relative localization to the current road segment can be accurate within 5 or 10 centimeters. When driving autonomously, to ensure safety, relative local layout or localization to that road segment achieves higher accuracy and precision than global-only approaches.As such, as will be described in more detail herein, when performing localization - the system of the present disclosure can use GNSS coordinates to determine which road segment the vehicle 1500 is currently traveling on, and then can more accurately localize the vehicle 1500 using the local or relative coordinate system of the determined road segment (e.g., in embodiments, without requiring a high - cost GNSS sensor type that is not practical for consumer vehicle implementations). As such, after the vehicle 1500 is localized to a given road segment, the vehicle 1500 can localize itself from road segment to road segment as it travels, so GNSS coordinates may not be required for accurate and precise localization. In some non - limiting embodiments, as described herein, each road segment may be 25 meters, 50 meters, 80 meters, 100 meters, and / or another distance.
[0020] To generate the data of the map stream 210 that can be used to generate an accurate and precise local or relative localization, sensor data 102, such as an IMU sensor 1566, a GNSS sensor 1558, a speed sensor 1544, a wheel sensor (e.g., counting the imprints of the wheels of the vehicle 1500), a perception sensor (e.g., camera, LiDAR, RADAR, etc.), and / or other sensor types can be used to track the movement (e.g., rotation and translation) of the vehicle 1500 at each frame or time step. The trajectory or ego-motion of the vehicle 1500 can be used to generate the trajectory layer of the map stream 210. As a non-limiting example, the movement of the perception sensor can be tracked to determine the corresponding movement of the vehicle 1500 - for example, referred to as visual odometry. The IMU sensor 1566 can be used to track the rotation or pose of the vehicle 1500, and the speed sensor 1544 and / or the wheel sensor can be used to track the distance the vehicle 1500 has traveled. As such, in the first frame, the first pose (e.g., angles along the x, y, and z axes) and the first position (e.g., (x, y, z)) of the rig or origin of the vehicle 1500 can be determined using the sensor data 102. In the second frame, the second pose and the second position of the rig or origin of the vehicle 1500 (e.g., relative to the first position) can be determined using the sensor data, etc. As a result, the trajectory can be generated at points (corresponding to frames), where each point can encode information corresponding to the relative position of the vehicle 1500 with respect to the previous point. Additionally, the sensor data 102 and / or the output 204 of the DNN 202 captured at each of these points or frames can be associated with the points or frames.As such, when creating an HD map (e.g., the map data 108 of FIG. 1), the sensor data 102 and / or output 204 of the DNN 202 can have a known position relative to the origin or rig of the vehicle 1500, and since the origin or rig of the vehicle 1500 can have a corresponding position on the global coordinate system, the sensor data 102 and / or output 204 of the DNN 202 can also have a position on the global coordinate system.
[0021] By using the relative movement of the vehicle 1500 from frame to frame, accuracy can be maintained even when the vehicle is in a tunnel, in a city, and / or in another environment where GNSS signals may be weak or lost. However, even when using relative movement and adding vectors for each frame or time step (e.g., representing the translation and rotation between frames), the relative position may drift after a period or distance has elapsed. As a result, relative movement can be reset or readjusted when drift is detected at predefined intervals and / or based on some other criteria. For example, there may be anchor points within a global coordinate system that can have known positions, and the anchor points can be used to readjust or reset relative movement in a frame.
[0022] In some embodiments, sensor data 102 is applied to one or more deep neural networks (DNNs) 202 that are trained to compute various different outputs 204. Prior to application or input to the DNN 202, the sensor data 102 may undergo preprocessing, for example, to transform, crop, enlarge, reduce, zoom in, rotate, and / or otherwise modify the sensor data 102. For example, if the sensor data 102 corresponds to camera image data, the image data may be adjusted by cropping, reducing, enlarging, flipping, rotating, and / or otherwise, to the appropriate input format of each DNN 202. In some embodiments, the sensor data 102 may include image data representing an image, image data representing a video (e.g., a snapshot of a video), and / or sensor data representing a representation of the sensor's field of perception (e.g., a depth map of a LIDAR sensor, a value graph of an ultrasonic sensor, etc.). For example, any type of image data format, e.g., and without limitation, a compressed image in a JPEG (Joint Photographic Experts Group) or luminance / chrominance (YUV) format, a compressed image as a frame resulting from a compressed video format such as H.264 / Advanced Video Coding (AVC) or H.265 / High Efficiency Video Coding (HEVC), a raw image derived from an RCCB (Red Clear Blue), RCCC (Red Clear), or other type of imaging sensor, and / or other formats, may be used. Additionally, in some instances, the sensor data 102 may be used without preprocessing (e.g., in a raw or captured format), while in other instances, the sensor data 102 may undergo preprocessing (e.g., using a sensor data preprocessor (not shown), such as noise balancing, demosaicing, scaling, cropping, expanding, white balancing, tone curve adjustment, etc.).
[0023] When the sensor data 102 corresponds to LiDAR data, for example, raw LiDAR data can be adjusted by accumulation, ego-motion correction, and / or other methods, and / or can be converted to another representation, such as a 3D point cloud representation (e.g., from a top-down view, sensor perspective view, etc.), a 2D projected image representation (e.g., a LiDAR range image), and / or another representation. Similarly, for RADAR and / or other sensor modalities, the sensor data 102 can be converted to a suitable representation for input to their respective DNNs 202. In some embodiments, the DNN 202 can process two or more different sensor data inputs - from any number of sensor modalities - to generate the output 204. As such, herein, the sensor data 102 can refer to unprocessed sensor data, preprocessed sensor data, or a combination thereof.
[0024] Examples are described herein with respect to the use of the DNN 202, but this is not intended to be limiting. For example, and without limitation, the DNN 202 can be any type of machine learning model or algorithm, such as linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k-nearest neighbor (Knn), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptron, long / short-term memory / LSTM, Hopfield, Boltzmann, deep belief, deconvolutional, adversarial generation, liquid state machine, etc.), region of interest detection algorithms, machine learning models using computer vision algorithms, and / or other types of algorithms or machine learning models.
[0025] As an example, DNN 202 can process sensor data 102 to produce detections of lane dividing lines, road boundaries, signs, poles, trees, static objects, vehicles and / or other dynamic objects, waiting states, intersections, distances, depths, dimensions of objects, and the like. For example, the detections can correspond to a location (e.g., in 2D image space, in 3D space, etc.), geometry, pose, semantic information, and / or other information related to the detection. As such, for a lane boundary line, the position and / or type of the lane boundary line (e.g., dashed line, solid line, yellow, white, crosswalk, bicycle lane, etc.) can be detected by DNN 202 that processes sensor data 102. Regarding a sign, the position or other waiting state information and / or its type (e.g., give way road, stop, crosswalk, traffic signal, give way warning light, construction, speed limit, exit, etc.) can be detected using DNN 202. For a detected vehicle, motorcyclist, and / or other dynamic actor or road user, the position and / or type of the dynamic actor can be identified and / or tracked and / or used to determine a waiting state within the scene (e.g., when a vehicle moves in a particular way with respect to an intersection, e.g., by coming to a stop, its corresponding intersection or waiting state can be detected as an intersection with a stop sign or traffic signal).
[0026] The output 204 of the DNN 202 can, in embodiments, undergo post - processing, for example, by converting the raw output into a useful output - for example, if the raw output corresponds to the confidence of each point (e.g., in LiDAR, RADAR, etc.) or pixel (e.g., of a camera image) where the point or pixel corresponds to a particular object type, the post - processing can be performed to determine each point or pixel corresponding to a single instance of the object type. This post - processing can include temporal filtering, weighting, outlier removal (e.g., removing pixels or points determined to be outliers), upscaling (e.g., the output can be predicted at a lower resolution than the input sensor data instance and the output can be scaled back up to the input resolution), downscaling, curve fitting, and / or other post - processing techniques. The output 204 - after post - processing, in embodiments - can be in a 2D coordinate space (e.g., image space, LiDAR range image space, etc.) and / or in a 3D coordinate system. If the output is in a 2D coordinate space and / or in a 3D coordinate space other than the 3D world space, in embodiments, the data converter 206 can convert the output 204 into the 3D world space.
[0027] In some non-limiting examples, DNN 202 and / or output 204 may be similar to those described in U.S. Patent Application No. 16 / 286,329, filed on February 26, 2019; U.S. Patent Application No. 16 / 355,328, filed on March 15, 2019; U.S. Patent Application No. 16 / 356,439, filed on March 18, 2019; U.S. Patent Application No. 16 / 385,921, filed on April 16, 2019; U.S. Patent Application No. 16 / 535,440, filed on August 8, 2019; U.S. Patent Application No. 16 / 728,595, filed on December 27, 2019; U.S. Patent Application No. 16 / 728,598, filed on December 27, 2019; U.S. Patent Application No. 16 / 813,306, filed on March 9, 2020; U.S. Patent Application No. 16 / 848,102, filed on April 14, 2020; U.S. Patent Application No. 16 / 814,351, filed on March 10, 2020; U.S. Patent Application No. 16 / 911,007, filed on June 24, 2020; and / or U.S. Patent Application No. 16 / 514,230, filed on July 17, 2019, each of which is hereby incorporated by reference in its entirety.
[0028] In an embodiment, the data converter 206 can convert all of the output 204 and / or the sensor data 102 into a 3D world space coordinate system having the rig of the vehicle 1500 as the origin (e.g., (0, 0, 0)). The origin of the vehicle 1500 may be on the vehicle 1500, along the axles of the vehicle 1500, and / or at any position on the vehicle or at the foremost or rearmost point relative to the vehicle 1500. In some non-limiting embodiments, the origin may correspond to the center of the rear axle of the vehicle 1500. For example, at a given frame or time step, the sensor data 102 can be generated. As a non-limiting example, a first subset of the sensor data 102 can be generated in the 3D world space related to the origin and can be directly used - for example, after compression by the data compressor 208 - to generate the map stream 210. A second subset of the sensor data 102 can be generated in the 3D world space but may not be related to the origin of the vehicle 1500. As such, the data converter 206 can convert the sensor data 102 - for example, using the unique and / or incidental parameters of each sensor - so that the 3D world space position of the sensor data 102 is related to the origin of the vehicle 1500. A third subset of the sensor data 102 can be generated in the 2D space. The data converter 206 can convert this sensor data 102 - for example, using the unique and / or incidental parameters of each sensor - so that the 2D space position of the sensor data 102 is within the 3D space and is related to the origin of the vehicle 1500.
[0029] In an embodiment in which DNN 202 is implemented, the output 204 - for example, before or after post-processing - can be generated in a 2D space and / or a 3D space (associated with or not associated with the origin). Similar to the description herein regarding directly converting the position from sensor data 102 in a 3D world space with respect to the origin of the vehicle 1500, the 2D and / or 3D output 204 (not associated with the origin) can be converted by the data converter 206 into a 3D space associated with the origin of the vehicle 1500. As such, and as described herein, the origin of the vehicle 1500 has a known relative position with respect to the current section of the road or within a series of map stream frames, and since the current road section has a relative position within a global coordinate system (e.g., WGS84 reference system), the position of the sensor data 102 and / or the output 204 from the DNN 202 can also have a relative position with respect to the current road segment and the global coordinate system.
[0030] Regarding the detected road boundaries and lane boundaries - for example, detected using a DNN 202 that processes sensor data 102 - a landmark filter may be executed to stitch and / or smooth the detected road boundaries and / or lane boundaries. For example, the 3D positions of the lane boundaries and / or road boundaries may include gaps in detection, may include noise, and / or may not be accurate, precise, or artifact-free in other ways that are optimal or desirable. As a result, landmark filter processing may be executed to stitch together detections within and / or across frames such that a virtual continuous lane divider and road boundaries are generated. These continuous lane dividers and / or road boundaries may be similar to a lane graph used to define some lanes on a driving surface, the positions of the lanes, and / or the positions of the road boundaries. In some embodiments, smoothing may be performed with a generated solid line to more accurately reflect known geometric information of the lane boundaries and road boundaries. For example, if the detection of a lane boundary in a frame is distorted with respect to previous and / or subsequent detections, the distorted portion of the lane boundary may be smoothed to more accurately conform to the known pattern of the lane boundary. As such, the encoded information within map stream 210 may correspond to these continuous lane boundaries and / or road boundaries in addition to, or alternatively to, encoding each detection into map stream 210.
[0031] In some embodiments, sensor data 102 and output 204 may be generated always and for each frame, and all of the data may be sent to a map-making cloud or server as map stream 210. However, in some embodiments, sensor data 102 and / or output 204 may not be generated for each frame, not all of the data may be sent in map stream 210, or a combination thereof. For example, a map stream campaign may be implemented that identifies what type and amount of data to collect, where to collect the data, how often to collect the data, and / or other information. The campaign may fill in gaps, provide additional data to increase accuracy, update the map when a change in the road is detected (e.g., via health check 112), and / or for other reasons, enable targeted or selective generation of certain types of and / or at certain locations of data. As an example, a map stream campaign may be executed that identifies vehicle 1500 at a particular location and instructs vehicle 1500 to generate (or prioritize the generation of) a certain data type - e.g., LiDAR data and RADAR data starting at a location and spanning some distance - to reduce computation (e.g., by vehicle 1500 during map creation 106 when generating map stream 210 and when processing map stream data) and bandwidth (e.g., for sending map stream 210 to the cloud). In such an example, several drives through a particular section of road may have encountered many obstructions, or vehicle 1500 that executed the drive may not have had a particular sensor modality. As such, the map stream campaign may instruct vehicle 1500 to collect data corresponding to previously obstructed data and / or generate data for the missing modality.As another example, a map stream campaign may be generated to more accurately identify lane dividers, signs, traffic lights, and / or other information, and the instructions to vehicle 1500 may be to generate sensor data 102 and / or output 204 that can be used to generate or update the HD map with this information. As such, the DNN 202 that calculates information about lane dividers, signs, traffic lights, etc. can be executed using each sensor data type, and the output 204 - after post-processing, data conversion, compression, etc. - can be sent to the cloud for map creation via the map stream 210.
[0032] In some embodiments, the map stream campaign may be part of a map health check 112 - for example, after the HD map has been generated and used for localization. For example, if a discrepancy is detected between the HD map represented by the map data 108 and the current sensor data or DNN detections, the health checker may trigger vehicle 1500 to generate and / or upload new map stream data for that location. For example, in some embodiments, vehicle 1500 may be generating the data of the map stream 210 and not uploading the map stream 210, while in other embodiments, vehicle 1500 may simply generate and upload map stream data when triggered. As such, if localization to the map results in poor planning and / or control operations, the map stream 210 may be uploaded to the cloud to update the map information through the map creation process. As a result, the map may not be updated regularly, but may only be updated when localization errors, planning errors, and / or control errors are detected.
[0033] In some examples, the system can minimize how often waypoints or frames are generated and / or included in the map stream 210. For example, instead of including every frame in the map stream 210, a distance threshold, a time threshold, or a combination thereof can be used to determine which frames should be included in the map stream 210. As such, a waypoint or frame can be included in the map stream when a particular distance (e.g., 0.5 meters, 1 meter, 2 meters, 5 meters, 10 meters, etc.) has been traveled by the vehicle 1500 and / or a particular amount of time has elapsed (e.g., 0.5 seconds, 1 second, 2 seconds, etc.). This distance or time threshold can be used based on which is met first or which is met last. For example, a first frame can be included in the map stream 210 when the distance threshold can be met, and a second frame at the distance threshold can be included in the map stream 210. When the second frame is included, the distance and time thresholds can be reset, at which time the time threshold can be met and a third frame can be included in the map stream 210, and so on. As a result, less duplicate data can be included in the map stream 210.
[0034] For example, if vehicle 1500 is present at a traffic signal for 30 seconds, the time threshold is 1 second, and the frame rate is 30 frames per second (fps), then instead of including 900 (e.g., 30 * 30) frames in map stream 210, only 30 frames may be included in map stream 210. In some embodiments, after vehicle 1500 has not moved for a threshold time - e.g., 2 seconds, 4 seconds, etc. - frame generation and / or frame inclusion in map stream 210 may be withheld (e.g., until motion is detected). As another non-limiting example, if vehicle 1500 is moving at a speed of 1 meter per second (or 3.60 kilometers (2.24 miles) per hour), the distance threshold is 2 meters, the frame rate is 30 fps, then instead of including 60 (e.g., 30 * 2) frames in map stream 210 for the distance of 2 meters, only a single frame may be included in map stream 210. As a result, the amount of data that would otherwise be sent from vehicle 1500 to the cloud for map creation 106 is reduced, but - since at least some of the data may be redundant or simply incrementally different and thus may not be necessary for accurate map creation - does not affect the accuracy of the map creation process.
[0035] In addition to, or alternatively to, sending less data in the map stream 210 (e.g., minimizing the amount of data), in embodiments, the data can be compressed - e.g., to reduce bandwidth and reduce runtime. In some embodiments, fewer points or frames (e.g., rotation / translation information of the points or frames) need to be sent in the map stream 210, and extrapolation and / or interpolation can be used to generate additional frames or points, such that extrapolation and / or interpolation can be used to determine points or frames. For example, rotation and translation information can be data intensive - e.g., requiring many bits to fully encode (x,y,z) position information and x-axis, y-axis, and z-axis rotation information - so the fewer points along a trajectory that require the rotation and / or translation information encoded therein, the less data that needs to be sent. As such, the history of the trajectory can be used to extrapolate future points within the trajectory. As another example, additional frames between frames can be generated using interpolation. As an example, when the trajectory corresponds to a vehicle 1500 driving straight at a substantially constant speed, then the first and last frames of the sequence can be included in the map stream 210, and the frames in between can be interpolated from the first and last frames of the sequence of frames. However, if the velocity changes, more frames may need to be encoded in the map stream 210 to more accurately linearly interpolate where the vehicle 1500 was at a particular point in time. In such an example, the system may not need to use all of the frames yet, but can use more frames than in the example of straight constant speed driving. In some embodiments, cubic interpolation or cubic polynomial interpolation can be used to encode the derivatives that can be used to determine the rate of change of velocity, and thus can be used to determine - or interpolate - other points along the trajectory without the need to directly encode them into the map stream 210.As such, instead of encoding the rotation, translation, position, and / or pose in each frame, the encoded interpolation and / or extrapolation information can be used instead to recreate additional frames. In some embodiments, the frequency of trajectory data generation can be adaptively updated to achieve the maximum error in the accuracy of the pose used in map making 106 interpolated between two other frames or time steps, compared to the correct pose known at the time of data reduction.
[0036] As another example, delta compression can be used to encode the difference - or delta - between the current position of vehicle 1500 (e.g., pose and position) and the previous position of vehicle 1500. For example, 64 bits can be used to encode each of the latitude, longitude, and altitude - or x, y, and z coordinates - of vehicle 1500 in a frame totaling 192 bits. However, as vehicle 1500 moves, these values may not change much from frame to frame. As a result, instead of using all 64 bits for each of the latitude, longitude, and altitude in each frame, the encoded values of the frame can instead correspond to the difference from the previous frame - encoded in fewer bits. For example, in an embodiment, the difference can be encoded using 12 bits or less, thereby resulting in a significant savings in system memory and bandwidth. Delta compression encoding can be used for relative coordinates and / or global coordinates of a local layout.
[0037] In addition to, or alternatively to, compressing position, pose, rotation, and / or translation information, sensor data 102 and / or output 204 can be compressed for inclusion in map stream 210. For example, for position, pose, dimension, and / or lane, lane marker, lane divider, sign, traffic signal, waiting condition, static and / or dynamic object, and / or other information regarding other output 204, the data can be compressed by delta encoding, extrapolation, interpolation, and / or other methods to reduce the amount of data included in map stream 210. With respect to sensor data 102 such as LiDAR data, RADAR data, ultrasonic data, and / or the like, the points represented by the data can be voxelized such that duplicate points are removed and instead volumes are represented by the data. This can enable a lower density of points to be encoded in map stream 210 while still containing sufficient information for accurate and detailed map creation 106. For example, RADAR data can be encoded using octrees. In addition, quantization of RADAR points and / or points from other sensor modalities can be performed to minimize the number of bits used to encode the RADAR information. For example, since RADAR points can have x and y positions, the geometry of lane dividers, signs, and / or other information can be used - in addition to the x and y positions - to encode RADAR points using fewer bits.
[0038] In some embodiments, a buffer or other data structure may be used to encode map stream data as serialized structured data. As such, instead of having a single field of a data structure that describes some amount of information, bit packing into a byte array may be performed - e.g., such that the data in the buffer contains only numbers instead of field names - to achieve bandwidth and / or storage savings compared to a system that includes field names in the data. As a result, a schema that associates field names with data types may be defined using integers to identify each field. For example, an interface description language may be used to describe the structure of the data, and a program may be used to generate source code from the interface description language for generating or parsing a byte stream representing the structured data. As such, a compiler may produce an application programming interface (API) - e.g., a cAPI, a pythonAPI, etc. - that receives a file and instructs the system on how to consume the information from the pin file encoded in this particular way. Additional compression may be achieved by chunking the data into smaller - e.g., 10,000 byte - chunks, converting the data to a byte array, and then sending or uploading it to the cloud for map making 106.
[0039] In some embodiments, dynamic obstacle removal may be performed on LiDAR data, RADAR data, and / or data from other sensor types to remove or filter sensor data 102 corresponding to dynamic objects (e.g., vehicles, animals, pedestrians, etc.). For example, different frames of sensor data 102 may be compared - e.g., after ego-motion compensation - to determine inconsistent points across the frames. In such an example, if a bird can fly across the field of perception of a LiDAR sensor, for instance, the points corresponding to the detected bird in one or more frames may not be present in one or more previous or subsequent frames. As such, these points corresponding to the detected bird may be filtered out or removed such that these points are not used during the generation of the final LiDAR layer of the HD map - during map making 106.
[0040] As such, sensor data 102 may be processed by the system to generate outputs corresponding to image data, LiDAR data, RADAR data, and / or trajectory data, and one or more of these outputs may undergo post-processing (e.g., the perception output may undergo fusion to generate a fused perception output, LiDAR data may undergo dynamic obstacle filtering to generate filtered LiDAR data, etc.). The resulting data may be aggregated, merged, edited (e.g., trajectory completion, interpolation, extrapolation, etc.), filtered (e.g., landmark filtering to create continuous lane boundaries and / or road boundaries), and / or otherwise processed to generate a map stream 210 representing this data generated from any number of different sensors and / or sensor modalities.
[0041] Referring now to FIG. 3, each block of method 300 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. Method 300 can also be implemented as computer-usable instructions stored on a computer storage medium. Method 300 can be provided, by way of example, as a stand-alone application, service, or hosted service (either stand-alone or in combination with another hosted service), or as a plug-in to another product. Additionally, method 300 is described, by way of example, with respect to process 104 of FIG. 2. However, this method 300 can be executed, additionally or alternatively, by any one process, or any combination of processes and systems, including but not limited to those described herein, by any one system.
[0042] FIG. 3 is a flow diagram illustrating method 300 for map stream generation, according to some embodiments of the present disclosure. Method 300 includes, at block B302, generating sensor data using sensors of a vehicle. For example, sensor data 102 can be generated.
[0043] Method 300 includes, at block B304, applying at least a first subset of the sensor data to a DNN. For example, sensor data 102 can be applied to DNN 202.
[0044] Method 300 includes, at block B306, computing an output using the DNN. For example, DNN 202 can compute output 204, and in one or more embodiments, the output can include, but is not limited to, lane divider information, road boundary information, static object information, dynamic object information, waiting condition information, intersection information, signs, poles, or traffic signal information, and / or other information corresponding to objects - static and / or dynamic - within the environment of vehicle 1500.
[0045] Method 300 includes, at block B308, converting at least an output of a first subset into a 3D coordinate system having a vehicle as an origin to generate a converted output. For example, data converter 206 can convert output 204 into a 3D coordinate system having an origin of vehicle 1500 as the origin.
[0046] Method 300 includes, at block B310, converting at least a second subset of sensor data into a 3D coordinate system to generate converted sensor data. For example, at least some of sensor data 102 can be directly used in map stream 210, but may not be generated in a 3D coordinate system related to the origin of vehicle 1500. As such, sensor data 102 can be converted into a 3D coordinate space having vehicle 1500 as the origin.
[0047] Method 300 includes, at block B312, compressing and / or minimizing sensor data, converted sensor data, converted output, and / or an output of a second subset to generate compressed data. For example, sensor data 102 and / or an output (e.g., either without conversion or after conversion if not generated in a 3D coordinate space for vehicle 1500) can be compressed and / or minimized. Compression and / or minimization can be performed using any known technique including, but not limited to, those described herein.
[0048] Method 300 includes, at block B314, encoding compressed data, sensor data, converted sensor data, converted output, and / or an output of a second subset to generate a map stream. For example, sensor data 102 (with or without conversion) and / or output 204 (with or without conversion) - e.g., after compression by data compressor 208 - can be encoded to generate map stream 210. Map stream 210 can then be transmitted to the cloud for map creation 106.
[0049] The process described with respect to FIG. 3 can be repeated for any number of driving - or segments thereof - of any number of vehicles 1500. Information from each map stream 210 can then be used for map creation 106.
[0050] Map Creation Referring to FIG. 4, FIG. 4 is a data flow diagram of the process 106 of map creation according to some embodiments of the present disclosure. In some embodiments, the process 106 can be executed in the cloud using a computing device (e.g., similar to the exemplary computing device 1600 of FIG. 16) of one or more data centers - e.g., the exemplary data center 1700 of FIG. 17. In some embodiments, the process 106 can be executed using one or more virtual machines, one or more individual computing devices (e.g., servers), or a combination thereof. For example, virtual graphics processing units (GPUs), virtual central processing units (CPUs), and / or other virtual components can be used to execute the process 106. In some embodiments, one or more of the process steps described with respect to the process 106 can be executed in parallel using one or more parallel units. For example, the registration 402 of pairs of segments - e.g., cost space sampling, aggregation, etc. - can be executed in parallel (e.g., a first pair can be registered in parallel with another pair). Additionally, within a single registration, the poses sampled to update the cost of a point in the cost space can be executed in parallel with one or more other poses. Additionally, since the map data corresponding to individual map layers can be stored in the GPU as textures, texture lookups can be executed to quickly determine the cost values in the cost space - thereby leading to a reduction in the runtime of each cost space analysis.
[0051] The map creation process 106 may include receiving a map stream 210 from one or more vehicles 1500 corresponding to any number of drives. As described herein, each map stream 210 may include various different layers of data generated using a variety of different methods - for example, tracking of self-motion (e.g., relative and global), sensor data 102 generation and processing, perception using one or more DNNs 202, etc. Each layer of the map stream 210 may correspond to a series of frames corresponding to sensor events recorded at a variable frame rate. For example, the map stream 210 layers may correspond to a camera layer, a LiDAR layer (e.g., layers of each different slice), a RADAR layer, a trajectory (or self-motion) layer, and / or other layers. The camera layer may include information obtained by performing perception - e.g., via DNN 202 - on a stream of 2D camera images, and converting the (2D and / or 3D) detections or outputs 204 of the DNN 202 into 3D landmarks and routes (e.g., by combining lane markings to define lane boundary and road boundary positions). The LiDAR and / or RADAR layer (or other sensor modality layer) may each correspond to point cloud information collected using LiDAR sensor 1564 and / or RADAR sensor 1560, respectively. As described herein, during the map stream generation process 104, preprocessing may be performed on the LiDAR data and / or RADAR data, e.g., data reduction, self-motion compensation, and / or dynamic obstacle removal. The trajectory layer may include information corresponding to the origin or rig absolute position of the vehicle 1500 and the relative self-motion over a particular time frame.
[0052] Referring to FIG. 5A, the map creation process 106 may include converting the map streams 210(1) to 210(N) - where N corresponds to the number of map streams 210 used for a particular registration process - via conversions 502(1) to 502(N) into maps 504(1) to 504(N) (e.g., converting the map stream into an operation work map format). For example, referring to FIG. 5B, for a single map stream 210(1), the conversion 502(1) may include LiDAR voxelization using a base conversion 506, a RADAR conversion 508, LiDAR altitude slicing 510, a LiDAR conversion 512, RADAR map image creation 514, LiDAR map image creation 516, and / or a LiDAR voxelizer 518. The base conversion 506 may correspond to landmarks - such as lane boundaries, road boundaries, signs, poles, trees, other vertical structures or objects, crosswalks, etc. - as determined using perception via the DNN 202. For example, the 3D landmark positions may be converted into the map format - using the base conversion 506 - to generate the base layer 520 (or, "camera layer" or "perception layer") of the map 504(1). In addition to the landmark positions, the base layer 520 may further represent the trajectory or route (e.g., global or relative) of the vehicle 1500 that generated the map stream 210(1). When generating the base layer 520, a 1:1 mapping between the aggregated input frames of the map stream 210(1) and the output base layer 520 map road segments may be maintained.
[0053] RADAR data from the map stream 210(1) - for example, when received or accessed in an unprocessed format - can be converted to the RADAR point cloud layer 522 via the RADAR conversion 508. In some embodiments, the RADAR point cloud from the RADAR point cloud layer 522 can be used to generate the RADAR map image layer 524 via the RADAR map image creation 514. For example, the RADAR point cloud can be converted to an image of the RADAR point cloud from one or more different viewpoints (e.g., top-down, sensor viewpoint, etc.). For example, a virtual camera having a top-down field of view can be used to project the RADAR point cloud into the frame of the virtual camera to generate the RADAR map image of the RADAR map image layer 524.
[0054] LiDAR data from the map stream 210(1) - for example, when received or accessed in an unprocessed format - can be converted to a LiDAR point cloud layer 526 via the LiDAR conversion 512. In some embodiments, as described herein, LiDAR data can be generated in slices (e.g., a ground slice (e.g., 5 meters to 300 meters relative to the origin of the vehicle 1500), a giraffe plane slice (e.g., 2.5 meters to 5 meters relative to the origin of the vehicle 1500), a ground plane slice (e.g., -2.5 meters to 0.5 meters with respect to the origin of the vehicle 1500), etc.). In addition to, or alternatively to, LiDAR altitude slicing in the map stream generation process 104, LiDAR altitude slicing 510 can be performed during the conversion 502(1) to determine separate slices of LiDAR data for use in generating one or more LiDAR point cloud layers 526. For example, LiDAR data corresponding to a particular slice - e.g., a giraffe slice - can be used to generate the LiDAR point cloud layer 526 of the map 504(1). In some embodiments, the LiDAR point cloud from the LiDAR point cloud layer 526 can be used to generate a LiDAR map image layer 528 via the LiDAR map image creation 516. For example, a LiDAR point cloud - e.g., corresponding to a particular slice - can be converted to an image of the LiDAR point cloud from one or more different viewpoints (e.g., top-down, sensor viewpoint, etc.). For example, a virtual camera having a top-down field of view can be used to project the LiDAR point cloud into the frame of the virtual camera to generate the LiDAR map image of the LiDAR map image layer 528. The LiDAR map image layer 528 can include a LiDAR map image encoded with elevation values determined from the point cloud (e.g., a top-down depth map), a LiDAR map image encoded with intensity values determined from the point cloud, and / or other LiDAR map image types. In some embodiments, the LiDAR point cloud layer 526 can be used - e.g., by the LiDAR voxelizer 518 - to generate the LiDAR voxel map layer 530 of the map 504(1).The LiDAR voxel map layer 530 may represent a voxelized representation of the LiDAR point cloud. The voxel map layer 530 may be used, for example, to filter out dynamic objects and for the editing of LiDAR data within an individual map 504 (as described in more detail herein) and / or within the fused HD map.
[0055] Referring back to FIG. 5A, the process described with respect to map stream 210(1) can be executed for each of map streams 210(1) - 210(N) to generate maps 504(1) - 504(N). Maps 504(1) - 504(N), or a subset thereof, can then be used for registration 402. As such, after map 504 has been generated for each map stream 210 that will be used in the current registration process, registration 402 can be executed to generate a summary map corresponding to the plurality of map streams 210. The summary map can include aggregated layers - for example, an aggregated camera or base layer, an aggregated LiDAR layer, an aggregated RADAR layer, etc. To determine which map streams 210 (and thus which maps 504) should be registered together, the position of map stream 210 - or a segment thereof - can be determined. For example, referring to FIG. 5C, a first map stream segment 540 (for example, representing the trajectory of vehicle 1500 that generated map stream 210(1)) can correspond to the first map stream 210(1), a second map stream segment 542 can correspond to the second map stream 210(2), and a third map stream segment 544 can correspond to the third map stream 210(3). Position or trajectory information - for example, GNSS data - can be used to determine that these map streams 210(1) - 210(3) correspond to similar positions or road sections. For example, GNSS coordinates from map stream 210 and / or map 504 can be used to determine whether map stream 210 - or a section thereof - is close enough in space over a distance long enough for the map streams 210 to be registered with each other. After being localized using GNSS coordinates, relative coordinates from map streams 210(1) - 210(3) can be used to determine how close the map stream segments are (for example, how close the trajectories of vehicle 1500 are). This process can result in a final list of map stream sections to be registered.For example, map stream segments 540, 542, and 544 can be determined to be on the same road - or to at least have an overlapping portion between demarcations 550A and 550B - using GNSS coordinates and / or relative coordinates from map stream 210 and / or map 504.
[0056] In some examples, an individual map stream 210 or drive - for example, map stream 210(1) corresponding to map stream segment 540 - can include a loop. For example, during map stream generation process 104, vehicle 1500 may have crossed the same road segment at two or more different times. As a result, map stream segment 540 can be split - via deduplication - into two separate segments 552A and 552B and treated as separate map stream segments for registration 402. The deduplication process can be performed as a pre - processing step to identify when vehicle 1500 - or the corresponding map stream 210 - was driven in a circular or loop pattern. After a loop is identified, the drive or map stream 210 can be split into two or more sections to ensure that each physical location is represented only once.
[0057] The layout of all map stream segments that include map stream segments determined to be within some spatial threshold relative to each other can be generated. For example, FIG. 5C may represent an exemplary visualization of the layout of map stream segments 540, 542, and 544. In the example of FIG. 5C, for registration 402, overlapping portions 552A - 552D - for example, between demarcations 550A and 550B - may be used. The layout can then be used to determine which pairs of the overlapping portions 552A - 552D to register together. In some embodiments, a minimum spanning tree algorithm is used to determine the minimum number of portions 552 that will be registered together such that each portion 552 has a connection between them (e.g., if portion 552A and portion 552B are registered together and portion 552B and 552C are registered together, then portion 552A and 552C will have a connection through portion 552B). The minimum spanning tree algorithm may include randomly selecting a pair of portions 552 and then randomly selecting another pair of portions until the minimum spanning tree is complete. In some embodiments, in addition to the minimum spanning tree algorithm, an error margin can be used to include additional pairs for robustness and accuracy.
[0058] As an example, with respect to FIG. 5D, table 560 shows a different number of connections or pairs between map streams or map stream sections using various different techniques. The number of segments, M, may correspond to the total number of map streams or map stream sections available for registration 402 of a particular segment (e.g., portions 552A - 552D of FIG. 5C). All connections may correspond to the number of pairs that would be registered if registration were performed for each possible connection between the segments of each pair. As an example, to calculate all connections, the following equation (1) may be used: All connections = M * (M - 1) / 2 (1)
[0059] As another example, to calculate minimum spanning tree connections, the following equation (2) may be used: Minimum Spanning Tree Connection = M-1 (2)
[0060] As a further example, in order to calculate a safety margin connection - for example, a safety margin added to a minimum spanning tree connection - the following equation (3) can be used: Safety Margin Connection = MIN[(Minimum Spanning Tree Connection * 2), all connections] (3)
[0061] As such, referring also to FIG. 5C, using equation (3) for portions 552A - 552D, connections 554A - 554F can be determined. For example, since there are 4 portions or segments to be registered, the safety margin will result in 6 connections being made among the 4 portions or segments of the map stream section between demarcations 550A and 550B.
[0062] In some embodiments, the quality of a section can be analyzed in order to determine which of the sections to use and / or to determine which sections to use more frequently (e.g., if some sections are registered more than once). For example, the quality can correspond to the number of sensor modalities represented in the map stream section. As such, if one section includes a camera, a track, RADAR, and LiDAR, and another section includes only a camera, a track, and RADAR, the section with LiDAR can be weighted or favored so that it is more likely to be included in more registrations 402. Additionally, in some embodiments, geometric distance can be used as a criterion when determining which sections to register together. For example, two sections that are geometrically closer can be selected over sections that are geometrically farther apart - for example, sections that are closer can correspond to the same driving lane as opposed to sections that are farther apart and can correspond to different driving lanes or opposite sides of a road.
[0063] Referring back to FIGS. 4 and 5A, pairs of sections can then be registered with each other to generate pose-links therebetween that can be used for pose optimization 404. Registration 402 can be performed to determine the geometric relationships between multiple different positions where the driving or map stream 210 overlaps - for example, portions 552A-552D of FIG. 5C - between multipole drives or between map streams 210 corresponding thereto. The output of the registration process can include relative pose-links (e.g., rotation and translation between the pose of a first map stream frame or section and the pose of a second map stream frame or section) between pairs of frames or sections within different map streams. In addition to the relative pose, the pose-links can further represent a covariance that represents the reliability at each pose-link. The relative pose-links are then used to align the maps 504 corresponding to each of the map streams 210 such that landmarks and other features - e.g., point clouds, LiDAR image maps, RADAR image maps, etc. - are aligned in the final, aggregated, HD map. As such, the registration process is performed by localizing one map 504 - or a portion thereof - of the pair to the other map 504 - or a portion thereof. Registration 402 can be performed for each sensor modality and / or for each map layer corresponding to the various sensor modalities. For example, camera-based registration, LiDAR-based registration, RADAR-based registration, and / or other registrations can be performed individually. The results can include an aggregated base layer of the HD map, an aggregated LiDAR point cloud layer of the HD map, an aggregated RADAR point cloud layer of the HD map, and an aggregated LiDAR map image layer of the HD map, etc.
[0064] The localization process used to localize one section to another for registration can be performed in the same manner as the localization process 110 used in live perception, as described in more detail herein. For example, a cost space can be sampled for a single frame or pose, an aggregated cost space can be accumulated over frames or poses, and a Kalman filter can be used in the aggregated cost space to localize one section with respect to another. As such, the pose from the trajectory of the first section can be known, and the pose of the trajectory from the second section can be sampled with respect to the first section using localization to determine the relative pose between the two - once alignment between landmarks is achieved. This process can be repeated for each of the poses of the sections such that pose links between poses are generated.
[0065] For example, with respect to the base layer or the perception layer of the map 504, during registration 402, the 3D world space position of the landmark can be projected into the virtual field of view of the virtual camera to generate an image corresponding to the 2D image space position of the landmark within the virtual image. In some embodiments, the 3D world space landmark position can be projected into the field of view or multiple virtual cameras to generate images from different viewpoints. For example, a forward virtual camera and a backward virtual camera can be used to generate two separate images for localization in the registration process. Using two or more images can enhance the robustness and accuracy of the registration process. To register to the image, the virtual image space positions of the landmarks from two sections undergoing registration 402 are compared to each other - for example, using a camera-based localization technique - to sample a cost space representing the position of the match between the sections. For example, the detection from the first section can be converted into a distance function image (e.g., as described herein with respect to localization 110), and the detection from the second section can be compared to the distance function image to generate a cost space. In addition to geometric comparison, semantic information can also be compared - for example, simultaneously or in separate projections - to calculate the cost. For example, if semantic information - such as lane boundary type, pole, sign type, etc. - does not match for a particular point, the cost for that particular point can be set to the maximum cost. The geometric cost and the semantic cost can then both be used to determine the final cost for each pose. This process can be executed over any number of frames to determine an aggregated cost space that is more refined than any individual cost space, and a Kalman filter (or other filter type) can be executed on the aggregated cost space to determine the pose - and pose links - between two segments.
[0066] With respect to FIGS. 5E and 5F, FIGS. 5E and 5F show examples of camera or base layer based registration according to some embodiments of the present disclosure. For example, visualization 570A in FIG. 5E may correspond to a forward virtual camera registration, and visualization 570B in FIG. 5F may correspond to a backward virtual camera. The 3D landmark positions from the base layer of the first map 504 may be projected into the 2D image space of the forward and backward virtual cameras. The 2D projection may then be converted into a distance function - shown as a dotted pattern in visualizations 570A and 570B - where the center line of each dotted segment 572A - 572K may correspond to a cost of zero, and as the dots move away from the center line of the section, the cost increases up to a white area that may correspond to the maximum cost. The dotted segments 572 may correspond to the distance function equivalents of the landmarks. For example, dotted segments 572A - 572C and 572J - 572K may correspond to lane boundaries and / or road boundaries, and dotted segments 572D - 572I may correspond to poles, signs, and / or other static objects or structures.
[0067] The distance function can then be compared with the 2D projection 3D landmark information from the base layer of the second map 504. For example, a solid black segment 574 (including, e.g., lines and / or dots) can represent the 2D projection of a 3D landmark of the second map 504 at the sampled pose 576. As such, this comparison of the 2D projection image with the distance function projection image of the first map 504 can correspond to a single position in the cost space corresponding to the pose 576, and any number of other poses can also be sampled to fill the cost space. As such, for example, many of the points match points along the dotted segment 572A, so each point along the solid black segment 574 can have a relatively low cost. In contrast, many of the points do not match points along the dotted segment 572J, so each point along the solid black segment 574J can have a high cost. Finally, for the pose 576, a cost corresponding to each of the points of the solid black segments 574A - 574K can be calculated - for example, using an average value - and the point in the cost space corresponding to the pose 576 can be updated to reflect the calculated cost. This process can be repeated for any number of poses - e.g., each possible pose - until the most likely relative pose of the base layer of the second map 504 with respect to the base layer of the first map 504 is determined for a particular time step or frame. This process can then be repeated for each time step or frame of each pair of sections that are registered together - e.g., for FIG. 5C, for each corresponding time step or frame from sections 552A and 552B, for each corresponding time step or frame from sections 552C and 552D, etc.
[0068] As another example, with respect to the LiDAR point cloud layer of the map 504, during registration 402, the LiDAR point cloud from the first section can be converted to a distance function point cloud and compared with the LiDAR point cloud of the second section to generate a cost space. For example, the cost can be sampled at each possible pose of the second section with respect to the distance function image corresponding to the first section. This process can be similarly performed with respect to the LiDAR intensity map image, the LiDAR elevation map image, and / or other LiDAR image types. Similarly, for RADAR, registration 402 can be performed in this manner.
[0069] The result of registration 402 between any pair of segments is a pose-link that defines the rotation, translation, and / or covariance between the poses of the segments. This registration process is performed for each determined pair from the connections - for example, the safety margin connections 554A - 554F of FIG. 5C. Using the output of the registration process, a pose-graph can be generated that represents the pose-links - or those poses - of the relative positions based on the time stamps of different operations or segments.
[0070] Referring again to FIGS. 4 and 5A, after the registration process 402 is complete and pose links are generated between the registered pair of frames or poses, pose optimization 404 can be performed to smooth the pose links and separate the group of poses from various different maneuvers into segments corresponding to the road segments used for relative localization in the final aggregated HD map. Pose optimization 404 can be performed to obtain a more desirable or optimal geometric alignment of submaps or layers based on the input absolute poses from the map stream 210, the relative pose links generated during registration 402, and the relative trajectory pose links of each individual maneuver generated using self-motion. Prior to pose optimization, a pose graph can be generated to represent the absolute and relative poses. In some embodiments, a filtering process can be performed to remove low-confidence poses - for example, using covariance as determined during the registration process 402. The (filtered) frames or poses can be clustered into road segments that may be different from the segments of the individual input maps 504 - for example, the same area may be represented by a single road segment rather than multiple overlapping road segments. Each output road segment can be associated with an absolute pose, or origin, during the pose optimization process 404. Given an initial pose graph layout, optimization can be performed to minimize the pose error of the absolute pose with respect to the observed relative poses and input absolute poses, taking into account the uncertainty information associated with each observation (e.g., maximum likelihood optimization). Optimization can be initialized using a random sample consensus (RANSAC) process to find a subset of pose links that maximally agree as measured through cycle consistency checks over multiple cycles within the pose graph. After optimization, during localization 110, each road segment can be compared to the connected road segments to determine the relative transformation between the road segments such that cost function numerical values can be transferred from road segment to road segment using the transformation so that an accurate accumulated cost space can be calculated.
[0071] For example, FIGS. 6A-6I illustrate an exemplary pose optimization process 404 corresponding to four operations or segments from the same and / or different maps 504. In some embodiments, the pose optimization process 404 may be performed separately for different sensor data formats or layers of the map 504. For example, any number of base layers of the map 504 may be registered together and then pose optimized, and any number of LiDAR point cloud layers of the map 504 may be registered together and then pose optimized separately from the pose optimization of the base layer, and so on. In other embodiments, the results of the registration processes of different layers of the map 504 may be used to generate aggregated poses and pose links, and the pose registration process may be performed for the aggregation.
[0072] With respect to FIG. 6A, registered segments from four different maneuvers 602A - 602D are shown in frame graph 600A. Frame graph 600A is an example of a portion of a larger frame graph, where the larger frame graph may correspond to any part - or all parts - of a road structure or layout. The four maneuvers 602A - 602D may all be registered to generate pose links 608 (e.g., pose links 608A - 608C) between poses 604 or positions of different maneuvers 602. Each pose 604 or position of a single maneuver 602 may be represented by a link 606. For example, link 606 may represent a translation, rotation, and / or covariance from one pose 604 to the next. As such, link 606A(6) may represent a rotation, translation, and / or covariance between the poses 604 of maneuver 602A that link 606A(5) connects. Pose link 608A may correspond to the output of registration process 402 and may encode a translation, rotation, and / or covariance between poses 604 of different maneuvers 602. For example, pose link 608A may encode a six - degree - of - freedom deformation between poses - for example, a translation (e.g., difference in (x,y,z) position), a rotation (e.g., difference in x, y, and z - axis angles), and / or a covariance between pose 604A(1) of maneuver 602A and pose 604B(2) of maneuver 602B (e.g., corresponding to the reliability in the value of the pose link).
[0073] Figure 6B shows road segment generation by which a group of poses from any number of maneuvers 602 can be combined into a single road segment. The resulting road segment can be included in the final aggregated HD map used for local or relative localization. Additionally, the generated road segment can have an origin that can have a relative position within a global coordinate system - for example, the WGS84 coordinate system. A frame graph 600B (which can correspond to frame graph 600A but can have road segment identification) shows how the frame graph can be divided into different road segments 610A - 610D. For example, a pose with a horizontal stripe pattern fill can correspond to a first road segment 610A, a pose with a vertical stripe pattern fill can correspond to a second road segment 610B, a pose with a dotted fill can correspond to a third road segment 610C, and a pose with a shaded fill can correspond to a fourth road segment 610D. To determine which poses 604 or frames will be included in each road segment, a first random pose 604 - for example, pose 604A(6) - can be selected. After the random pose 604 is selected, links 606 and / or pose links 608 can be repeatedly selected starting from the randomly selected pose 604 until a maximum distance corresponding to the maximum distance of the road segment 610 (for example, 25 meters, 50 meters, etc.) is reached or a pose 604 or frame already encoded in the road segment 610 is determined. After the road segment 610 is fully encoded, another random unencoded pose 604 can be selected and the process can be repeated. This can be repeated until each pose 604 is encoded or included in a road segment 610. Due to the probabilistic nature of this process, some very small road segments 610 can be created. As such, a post - processing algorithm can be executed to analyze the road segment size and merge road segments below a certain size threshold into adjacent road segments.
[0074] FIG. 6C shows the segment graph 620 of road segments 610A - 610D after the road segment encoding or generation process. For example, each labeled block of 610A - 610D may correspond to a collapsed representation of all the poses 604 corresponding to the respective road segment 610. After the road segment 610 is determined, the origin or seed position of each road segment may be determined. The origin or seed position may correspond to the center of the road segment 610 (e.g., where the road segment is 50 meters by 50 meters long and the origin may be (25m, 25m)). In some embodiments, the origin may correspond to the average or intermediate position of each pose 604 within the road segment 610. For example, the average value of the (x, y, z) coordinates corresponding to each pose 604 may be determined and the result may be selected as the origin. In other embodiments, the origin may be selected using different methods.
[0075] FIG. 6D shows a pose-link error correction process. For example, pose-link transformations between poses of a frame graph need not be guaranteed to be consistent with each other. As such, when three poses 604 are taken - as in FIG. 6D - a pose-link error 612 may exist. To minimize the pose-link error 612 within a road segment, a configuration of the poses within the road segment 610 having the lowest pose-link error 612 may be determined. In one or more embodiments, the layout of the poses determined during pose optimization 404 need not be globally consistent, but may be optimized such that the layout is consistent within the same road segment 610 and adjacent road segments 610. For example, pose-links 614 and 614B should be approximately the same, and pose-link 614C should be approximately equal to the combination of pose-links 614A and 614B. However, in practice, some errors - for example, localization errors during registration 402 - may result in some deltas or pose-link errors 612 that appear between the poses 604. As a result, pose optimization 404 may be used to move or shift the poses 604 or frames such that the best fit is achieved within the corresponding road segment 610. With respect to pose-links 614A and 614B, the poses 604 may be moved to distribute the error among all of the pose-links 614, rather than primarily presenting the pose-link error 612 with pose-link 614C.
[0076] FIG. 6E may represent the frame graph 600B of FIG. 6B, but shows a pose graph 630 in which the pose links connecting the road segments 610A-610D to the external road segment 610 are omitted. FIG. 6F shows the RANSAC operation performed on the pose graph 630. This process may be performed based on the reliability that the accuracy of the link 606 between poses is more accurate than the pose link 608. For example, the minimum set of pose links 608 - or the minimum spanning tree - in the pose graph 630 that maintains the connection of all nodes or poses 604 may be sampled. After the minimum set is selected, the remaining pose links 608 may be sampled to determine how many pose link errors 612 are detected. This may correspond to one iteration, and in the next iteration, another random minimum set of pose links 608 may be selected, and then all the remaining pose links 608 may be sampled for several iterations (e.g., 100, 1000, 2000, etc.) for pose link errors 612, etc. After completion of that number of iterations, the layout with the most matches may be used. For example, with respect to FIG. 6F, the error cost may be calculated for each pose link 608 other than the sampled minimum - for example, according to the following equation (4): [Number] Thus, u = log(Err), u ∈ se3, Err ∈ SE3, where Err is the difference between the pose link transformation and the transformation between the poses connected by the pose link according to the currently evaluated layout of the road segment 610, and s i is the standard deviation (e.g., transformation reliability) of the i'-th component of the pose link transformation. As such, the cost may represent the pose link error 612 with respect to the number of standard deviations. RANSAC sampling may be repeated over that number of iterations, and in each iteration, the number of pose links that fit the layout within N standard deviations may be counted. The pose links that satisfy this condition and are included in the count may be referred to as inliers.
[0077] Referring to FIG. 6G, FIG. 6G shows an updated pose graph 650 after the RANSAC process is completed - for example, the pose graph 650 shows the pose configuration with the most inliers. The pose graph 650 in FIG. 6G can undergo an optimization process, for example, a non-linear optimization process (such as a bundle adjustment process) aimed at minimizing the sum of the squared costs of the inliers - for example, using the calculated cost function described herein with respect to FIG. 6F. After bundle adjustment, the pose graph 650 can be fixed, and the road segments 610A - 610D can be fixed such that the poses 604 are at their final fixed positions of the road segments 610 - for example, as shown in FIG. 6H having the exemplary road segment 610C. As described herein, the final road segments 610 can have a relative origin (for example, for localizing with respect to a local or relative coordinate system). As such, since the relative origin has a position within the global coordinate system, localization to the local coordinate system can also localize to the global coordinate system. For example, with respect to FIG. 6H, the road segment 610C can have an origin 652C (as shown). Although not shown, the other road segments 610A, 610B, and 610D can also have their respective origins 652.
[0078] Referring to FIG. 6I, each road segment 610 can have a relative transformation calculated for each of its direct adjacent road segments 610. In the illustration of FIG. 6I, the road segment 610C has a relative transformation (for example, transformation, T 610C 610A ) determined between the road segment 610C and the road segment 610A, a relative transformation (for example, transformation, T 610C 610B ) determined between the road segment 610C and the road segment 610B, and a relative transformation (for example, transformation, T 610C 610D) may have. The aggregated cost space may include, for example, one or more cost spaces generated when localizing with respect to road segment 610A and a cost space generated when localizing with respect to road segment 610C, for example. Thus, the relative transformation may be used during localization 110 - for example, more specifically, when generating the aggregated cost function and / or when generating the local map layout used for localization. As such, when entering from road segment 610A to road segment 610C, the cost space is updated - for example, transformed, T 610C 610A using - may be updated or transformed. The relative transformation may represent six degrees of freedom, for example, rotation (e.g., differences in x, y, and z axis rotation angles) and translation (e.g., differences in (x, y, z) coordinates).
[0079] Referring back to FIG. 4, after registration 402 and pose optimization 404, the updated pose graph - for example, pose graph 650 - can be fixed, and the relative poses of the plurality of maneuvers 602 can be consequently fixed. As such, the maps 504 from each maneuver 602, whose outputs are currently aligned - for example, as a result of registration 402 and pose optimization 404 - can be used to fuse together the map layers of the individual maps 504 (e.g., via fusion 406) to form an aggregated HD map. The aggregated HD map can include the same distinct layers as the individual maps 504, but the aggregated HD map can include aggregated layers - for example, an aggregated base layer, an aggregated LiDAR point cloud layer, an aggregated RADAR point cloud layer, and an aggregated RADAR map image layer, etc. The fusion process 604 can be used to improve the data quality of the map data with respect to the data present in any single maneuver - for example, the fused map can be more accurate and reliable than a single map 504 from the map stream 210 of a single maneuver. For example, contradictions between the maps 504 can be removed, the accuracy of the map content can be improved, and a more complete representation of the world can be achieved - for example, by combining the geographic extents of two different map layers or submaps or by combining observations regarding waiting conditions across multiple maneuvers passing through the same intersection. After the final - for example, more optimal or desired - geometric layout of the poses 604 of the different maneuvers 602 is determined by the pose graph (e.g., pose graph 650), the information from the plurality of maps 504 and / or the corresponding map streams 210 can be easily transformed to the same output coordinate frame or pose associated with each road segment 610.
[0080] As an example, with respect to the 3D landmark positions represented in the base layer, the 3D landmark positions from multiple maps 504 can be fused together to generate the final representation of each lane boundary, each road boundary, each sign, each pole, etc. Similarly, for the LiDAR intensity map, the LiDAR elevation map, the LiDAR distance function image, the RADAR distance function image, and / or other layers of the map 504, separate map layers can be fused to generate an aggregated map layer. As such, if a first map layer contains data that matches - within some threshold similarity - the data of another map layer, the matching data can be used (e.g., an average value can be determined) to generate the final representation. For example, with respect to the base layer, if the data corresponds to a sign, the first representation of the sign in the first map 504 and the second representation of the sign in the second map 504 can be compared. If the first sign and the second sign are within a threshold distance of each other and / or are in the same semantic class, the final (e.g., average) representation of the sign can be included in the aggregated HD map. This same process can be performed for lane dividers, road boundaries, stop conditions, and / or other map information. In contrast, if the data from a map layer of the map 504 does not match the data from another map layer of another map, the data can be removed by a filter. For example, if the first map 504 contains data for a lane divider at a location and semantic class and one or more other maps 504 do not share this information, the lane divider from the first map 504 can be removed from consideration for the aggregated HD map or filtered out.
[0081] As described herein, the fusion process 406 can vary for different map layers of the map 504 - or different features represented therein. For example, for lane graph fusion - e.g., fusion of lane dividers, road boundaries, etc. - individually observed lanes, routes, and / or trajectories from multiple base map layers of multiple maps 504 can be fused into a single lane graph in the aggregated HD map. Lane boundaries and / or lane dividers of the various maps 504 can be fused not only for a lane graph (e.g., as a delimiter of the boundaries of the lanes in which the vehicle 1500 can proceed), but also for camera - based localization as a 3D landmark, or for 2D landmarks generated from 3D landmarks (e.g., as a set of stable semantic landmarks). In addition to lane boundaries or dividers, other road markings can also be collected from the map stream 210 and included in the map 504 for use in the fusion process 604. For example, stop lines, road text, road shoulders, and / or other markings can be fused together for use as semantic landmarks for localization and / or for updating standby condition information. With respect to standby conditions in the base layer, standby conditions from multiple maps 504 can be fused to generate the final representation of the standby conditions. Poles, signs, and / or other static objects can also be fused from multiple base layers of the map 504 to generate their aggregated representation. Poles, signs, and / or other (vertical) static objects can be used for localization.
[0082] In some embodiments, lane dividers, lane centers (e.g., rails), and / or road boundaries may not be clearly distinguishable in the perception from the base layer. In such embodiments, the base layer and / or track information from the map stream 210 may be used to infer rails and / or lane dividers. For example, after bundle adjustment, the pose graph 650 can be viewed from a top-down view to determine the pattern of tracks. Tracks determined to be in the same lane can be used to generate the lane dividers and / or rails for that particular lane. However, heuristics may be used to determine that two or more tracks are from the same driving lane. For example, the current pose or frame can be compared to determine to cluster two tracks together as belonging to the same lane. If the current frames or poses of the two tracks appear to match the same driving lane (e.g., based on some distance heuristic), the poses at some distance ahead (e.g., 25 meters) and at some distance behind (e.g., 25 meters) of the current poses of the two tracks can also be analyzed. If the current, previous, and future poses all indicate the same driving lane, the current poses can be clustered together to determine the lane divider, lane rail, and / or road boundary positions. As a result, if one track corresponds to a vehicle 1500 that changes lanes while another vehicle 1500 remains stable within the lane, the combination of the two tracks will result in an inaccurate representation of the actual lane divider, lane rail, and / or road boundary. In some embodiments, in addition to analyzing the distance between the poses of the tracks, or alternatively, the angle formed between the tracks can be analyzed. For example, if the difference in angles is greater than some threshold (e.g., 15 degrees, 40 degrees, etc.), the two tracks can be considered as not being from the same lane and may not be clustered together.
[0083] Regarding the LiDAR and / or RADAR map layers, the fusion of LiDAR points and RADAR points from multiple maps 504 can ensure that the fused or aggregated map contains a more complete point cloud than any individual map 504 (e.g., due to occlusion, dynamic objects, etc.). The fusion process 604 can also reduce the redundancy of the point cloud coverage by removing redundancies in order to reduce the amount of data required to store and / or transmit the point cloud information of the aggregated HD map. For example, if the point cloud points of one or more drives do not match exactly with the point cloud points from another drive, the points from the non-matching drive can be removed - for example, via dynamic object removal. The resulting aggregated map layer of LiDAR can include an aggregated slice (e.g., a ground plane slice, a giraffe plane slice, etc.) point cloud layer, a LiDAR elevation map image layer that stores the average altitude per pixel, and / or a surface reflectivity or intensity map image layer that stores the average intensity value per pixel. For RADAR, the resulting aggregated map layer can include a RADAR cross section (RCS: RADAR cross section) map image layer and / or a RADAR point cloud layer (which may or may not be sliced). In an embodiment, the in-memory representation of all LiDAR and RADAR map image layers can use a floating-point representation. For the elevation model, surface reflectivity model, and / or RCS, pixels without valid data can be encoded by a not-a-number (NaN) data type. Regarding the surface reflectivity and elevation map image layers, the map image generator can search for the peak density in the altitude distribution of each point. Only these points can be considered as ground detections, and noisy measurement results can be removed by a filter (e.g., measurement results corresponding to obstacles). In some embodiments, a median filter can be applied to fill in missing measurement results that would otherwise have gaps.
[0084] Regarding LiDAR, a voxel-based fusion strategy can be implemented. For example, if a voxel is observed only by a minority of a large number of drives, the voxel can be determined to be a noisy detection. In an embodiment, this LiDAR fusion can be performed for each road segment 610. In a non-limiting embodiment, the LiDAR fusion can be performed according to the following process: (1) finding a tight 3D bounding box of the LiDAR points, (2) given a preset voxel resolution and 3D points, (x, y, z), calculating the index of the 3D points, (3) using a low-density representation of the 3D volume instead of a high-density representation for memory usage reduction, (4) saving the ID of the drive or map stream 210 that observed the voxel and the average color in the voxel data, and (5) after updating the voxel volume at all points in the drive, setting a threshold on the number of drives for each voxel such that only voxels with a number of drives greater than the threshold are retained (for example, the threshold can be set to half of all drives 602 in road segment 610). In addition to removing noisy 3D points, the LiDAR fusion can generate and save point clouds in the giraffe plane and the ground plane slice. The 3D points within the corresponding plane or slice can be used later to generate different types of LiDAR map images for LiDAR localization. For example, the giraffe plane may include points within the range of z = [a1, b1], and the ground plane points may be within the range of z = [a2, b2], where a1, b1, a2, and b2 may be preset parameters. Thus, the points are selected based on the z-coordinate (altitude) of the points within their frame coordinate system. It should be noted that the filtering process is performed based on frame coordinates instead of road segment coordinates, and thus the points may not be filtered based on their z-coordinates after they are converted to the road segment coordinate system. To solve this problem, it can be noted that the points belong to the giraffe plane or the ground plane before they are transformed, and the information can be saved inside the data structure related to the voxels. If a voxel passes through the noise filtering process and belongs to one of the two planes, the voxel can be saved to the corresponding point cloud file.Similar processes may be performed on RADAR data.
[0085] In some embodiments, for example, if the fusion process 406 is performed during the health check 112, older LiDAR data, RADAR data, and / or map data may be weighted more negatively compared to more recent data. As such, if a discrepancy is determined between a newer map 504 and an older map 504, data from the older map 504 may be filtered out and data from the newer map 504 may be retained. This may be the result of changes in road conditions, construction, and / or the like, and more recent or current data may be more useful for navigation of the road segment 610.
[0086] After the fusion process 406, map data 108 - for example, representing an aggregated HD map including an aggregation layer - may have been generated. The map data 108 may then be used for localization 110 as described in more detail herein.
[0087] Referring now to FIG. 7, each block of the method 700 described herein includes a computational process that may be performed using any combination of hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in memory. The method 700 may also be implemented as computer-usable instructions stored on a computer storage medium. The method 700 may be provided, by way of example, as a stand-alone application, service, or hosted service (either stand-alone or in combination with another hosted service), or as a plug-in to another product. Additionally, the method 700, by way of example, is described with respect to the process 106 of FIG. 4. However, this method 700 may be performed, additionally or alternatively, within any one system, by any one process, or by any combination of processes and systems, including but not limited to those described herein.
[0088] Figure 7 is a flowchart showing a method 700 for map creation according to some embodiments of the present disclosure. Method 700 includes, at block B702, receiving data representing a plurality of map streams corresponding to a plurality of drives from a plurality of vehicles. For example, map stream 210 may be received for the map creation 106 process, where map stream 210 may correspond to any number of drives from any number of vehicles 1500.
[0089] Method 700 includes, at block B704, converting each map stream into a respective map including a plurality of layers to generate a plurality of individual maps. For example, each received map stream 210 may undergo a conversion 502 to generate map 504 such that a plurality of maps 504 (e.g., maps 504(1) to 504(N)) are generated. Each map 504 may include a plurality of layers, such as but not limited to the layers described with respect to FIG. 5B.
[0090] Method 700 includes, at block B706, geometrically registering a pair of segments from two or more individual maps to generate a frame graph representing pose links between the pair of segments. For example, segments of various map streams 210 and / or maps 504 (e.g., portions 552A to 552D) may be determined - for example, using trajectory information and GNSS data indicating proximity - and the segments may be registered with each other to generate a pose graph (e.g., frame graph 600A of FIG. 6A) (e.g., in an embodiment, using a safety margin in addition to a minimum spanning tree). The pose graph may include pose links between poses or frames of different drives 602. In some embodiments, different registration processes may be performed for different map layers and / or sensor modalities - for example, since the frame rates of different sensors may be different, the poses of vehicle 1500 in each frame may be different for different sensor modalities and thus also for different map layers.
[0091] Method 700 includes, at block B708, assigning a group of poses from a frame graph to a road segment to generate a pose graph that includes a pose corresponding to the road segment. For example, the process described herein with respect to FIG. 6B can be executed to determine poses 604 from different maneuvers 602 corresponding to a single road segment 610. After road segment 610C is determined, a pose graph that includes poses 604 from road segment 610C and each adjacent road segment (e.g., 610A, 610B, and 610D) can be generated to generate a pose graph 630 - for example, similar to a larger frame graph 600A, but without omitting external pose connections (e.g., pose connections to poses 604 outside of road segments 610A - 610D).
[0092] Method 700 includes, at block B710, executing one or more pose optimization algorithms on the pose graph to generate an updated pose graph. For example, a RANSAC operation (e.g., FIG. 6F), pose graph optimization (e.g., FIG. 6G), and / or other optimization algorithms can be executed to generate an updated pose graph 650.
[0093] Method 700 includes, at block B712, fusing map layers from a plurality of individual maps based on updated poses from an updated pose graph to generate a fused map. For example, after the relative pose 604 of each frame of a map layer is known, data from the map layer can be fused to generate an aggregated or fused HD map layer. For example, 3D landmark positions from two or more base layers of each map 504 can be fused according to an average position aligned to their poses. As such, if a pole in a first base layer of a first map 504 has a position (x1, y1, z1) and a pole in a second base layer of another map 504 has a position (x2, y2, z2), the pole can have a single final position in the fused base layer of the HD map as ((x1 + x2) / 2, (y1 + y2) / 2, (z1 + z2) / 2). Additionally, since the origin 652 of the corresponding road segment 610 of the pole can be known, the final pole position can be determined relative to the origin 652 of the road segment. This process can be repeated for all 3D landmarks in the base layer and can be repeated for each other layer type (e.g., RADAR layer, LiDAR layer, etc.) of the map 504.
[0094] Method 700 includes, at block B714, transmitting data representing the fused map to one or more vehicles for use in the execution of one or more operations. For example, map data 108 representing the final fused HD map can be transmitted to one or more vehicles 1500 for localization, path planning, control decisions, and / or other operations. For example, after being localized in the fused HD map, information from the planning and control layers of the autonomous driving stack can be obtained - for example, lane graphs, waiting conditions, static obstacles, etc. - and this information can be provided to the planning and control subsystems of the autonomous vehicle 1500. In some embodiments, the selection layer of the fused HD map can be transmitted to each vehicle 1500 based on the configuration of the vehicle 1500. For example, if vehicle 1500 does not have a LiDAR sensor, the LiDAR layer of the fused map may not be transmitted to vehicle 1500 for storage and / or use. As a result of this adapted approach, the bandwidth requirements for transmitting map data 108 and / or the storage requirements for vehicle 1500 for storing map data 108 can be reduced.
[0095] Localization Referring to FIG. 8A, FIG. 8A is a data flow diagram of a localization process 110 according to some embodiments of the present disclosure. In some embodiments, process 110 may be executed using vehicle 1500. In some embodiments, one or more of the processes described with respect to process 110 may be executed in parallel using one or more parallel processing devices. For example, localizations using different map layers may be executed in parallel with one or more other map layers. Within a single map layer, cost space sampling 802, cost space aggregation 804, and / or filtering 806 may be executed in parallel. For example, different poses may be sampled in parallel during cost space sampling 802 to more efficiently generate the cost space for the current frame or time step. Additionally, since the map data corresponding to the map layer may be stored in the GPU as a texture, texture lookups may be executed to quickly determine the cost values of the cost space - thereby leading to a runtime reduction for each cost space analysis. Additionally, as described herein, the localization process 110 - for example, cost space sampling 802, cost space aggregation 804, and / or filtering 806 - may be executed during registration 402 in the map creation process 106. For example, the localization process 110 may be used to geometrically register poses from a pair of segments, and cost space sampling 802, cost space aggregation 804, and / or filtering 806 may be used to align a map layer from one segment with a map layer from another segment.
[0096] The goal of the localization process 110 can be used to localize the origin 820 of the vehicle 1500 with respect to the local origin 652 of the road segment 610 of the fused HD map represented by the map data 108. For example, an ellipsoid corresponding to the fused localization 1308 - described in more detail herein with respect to FIGS. 9A-9C - can be determined for the vehicle 1500 at a particular time step or frame using the localization process 110. The origin 820 of the vehicle 1500 can correspond to a reference point or origin of the vehicle 1500, such as the center of the rear axle of the vehicle 1500. The vehicle 1500 can be localized with respect to the origin 652 of the road segment 610, and the road segment 610 can have a corresponding position 824 within the global coordinate system. As such, after the vehicle 1500 is localized to the road segment 610 of the HD map, the vehicle 1500 can likewise be globally localized. In some embodiments, an ellipsoid 910 can be determined for each individual sensor modality - such as LiDAR localization, camera localization, RADAR localization, etc. - and the outputs of each localization technique can be fused via the localization fusion 810. A final origin position can be generated - such as the origin 820 of the vehicle 1500 - and used as the localization result of the vehicle 1500 at the current frame or time step.
[0097] At the start of operation, the current road segment 610 of vehicle 1500 can be determined. In some embodiments, the current road segment 610 can be known from the last operation - for example, when vehicle 1500 was stopped - and the last known road segment where vehicle 1500 was localized can be stored. In other embodiments, the current road segment 610 can be determined. To determine the current road segment 610, GNSS data can be used to globally localize vehicle 1500 and then to determine the road segment 610 corresponding to the global localization result. If the result returns two or more road segments 610, the road segment 610 having the origin 652 closest to the origin of vehicle 1500 can be determined to be the current road segment 610. After the current road segment 610 is determined, the road segment 610 can be determined to be the seed road segment for the lateral search. The lateral search can be executed to generate a local layout such as road segments 610 adjacent to the current road segment 610 at the first level and then a second level of road segments 610 adjacent to the road segments 610 from the first level. Understanding the road segments 610 adjacent to the current road segment 610 can be useful for the localization process 110 because when vehicle 1500 moves from one road segment 610 to another road segment 610, the relative transformation between the road segments 610 can be used to update the sampled cost space generated for the previous road segment 610 before being used in the aggregated cost space for localization. After vehicle 1500 moves from the seed road segment to an adjacent road segment, another lateral search can be executed for the new road segment to generate an updated local layout, and this process can be repeated as vehicle 1500 traverses the map from road segment to road segment. Additionally, as described herein, when vehicle 1500 moves from one road segment 610 to another road segment 610, a previously calculated cost space (e.g., some number of the previous cost space within a buffer, e.g., 50, 100, etc.) can be updated to reflect the same cost space with respect to the origin 652 of the new road segment 610.As a result, the calculated cost space can be carried over through the road segments to generate an aggregated cost space via cost space aggregation 804.
[0098] The localization process 110 may localize the vehicle 1500 at each time step or frame using sensor data 102 - e.g., real - time sensor data 102 generated by the vehicle 1500 - map data 108, and / or output 204. For example, sensor data 102, output 204, and map data 108 may be used to perform cost space sampling 802. The cost space sampling 802 may vary for different sensor modalities corresponding to different map layers. For example, the cost space sampling 802 may be performed separately for a LiDAR map layer (e.g., LiDAR point cloud layer, LiDAR map image layer, and / or LiDAR voxel map layer), a RADAR map layer (e.g., RADAR point cloud layer and / or RADAR map image layer), and / or a base layer (e.g., for a landmark or camera - based map layer). Within each sensor modality, cost sampling may be performed using one or more different techniques, and the costs across different techniques may be weighted to generate the final cost of the sampled pose. This process may be repeated for each pose in the cost space to generate the final cost space of the frame during localization. For example, LiDAR intensity cost, LiDAR elevation cost, and LiDAR (slice) point cloud cost (e.g., using a distance function) may be calculated, then averaged or otherwise weighted and used for the final cost of a pose or point in the cost space. Similarly, for camera or landmark - based cost space sampling, semantic cost may be calculated, geometric cost (e.g., using a distance function) may be calculated, then averaged or otherwise weighted and used for the final cost of a pose or point in the cost space. As a further example, RADAR point cloud cost (e.g., using a distance function) may be calculated and used for the final cost of a pose or point in the cost space. As such, the cost space sampling 802 may be performed to sample the cost of each different pose within the cost space.The cost space may correspond to some regions in a map that can include only a portion of the current road segment 610, the entire current road segment 610, the current road segment 610 and one or more adjacent road segments 610, and / or some other regions of the entire fused HD map. As such, the size of the cost space may be a programmable parameter of the system.
[0099] The result of the cost space sampling 802 for any individual sensor modality may be a cost space that represents a geometric match or likelihood that the vehicle 1500 can be positioned with respect to each particular pose. For example, a point in the cost space may have a corresponding relative position with respect to the current road segment 610, and the cost space may indicate the likelihood or probability that the vehicle 1500 is currently at each particular pose (e.g., the (x,y,z) position with respect to the origin of the road segment 610 and the axis angles with respect to each of the x, y, and z axes).
[0100] Referring to FIG. 9A, cost space 902 may represent a sensor-style cost space - for example, a camera-based cost space generated according to FIGS. 10A-10C. For example, cost space 902 may represent the likelihood that vehicle 1500 is currently at each of a plurality of poses - for example, represented by points in cost space 902 - in the current frame. Although cost space 902 is represented in 2D in FIG. 9A, it may correspond to a 3D cost space (for example, having (x, y, z) positions and / or axis angles with respect to each of the x, y, and z axes). As such, sensor data 102 - for example, before or after preprocessing - and / or output 204 (for example, detection of landmark positions in a 2D image space and / or 3D image space) may be compared with map data 108 for each of the plurality of poses. If a pose does not match well with map data 108, the cost may be high, and the point in the cost space corresponding to the pose may be represented as such - for example, represented in red, or in the case of FIG. 9A, represented by the non-dotted or white portion. If a pose matches well with map data 108, the cost may be low, and the point in the cost space corresponding to the pose may be represented as such - for example, represented in green, or in the case of FIG. 9A, represented by the dotted points. For example, referring to FIG. 9A, if cost space 902 corresponds to visualization 1002 of FIG. 10A, the dotted portion 908 may correspond to the low cost of poses along the diagonal where landmark 1010 may match well with the prediction or output 204 of DNN 202. For example, at the pose at the lower left bottom of the dotted portion of the cost space, the prediction of the landmark can be aligned well with landmark 1010 from map data 108, and similarly at the upper right portion of the dotted portion, the prediction of the landmark from the corresponding pose can also be aligned well with landmark 1010. As such, these points may be represented with low cost. However, due to noise and the large number of low-cost poses, a single cost space 902 may not be accurate for localization - for example, vehicle 1500 may not be able to be located at each of the poses represented by dotted portion 908.As such, the aggregated cost space 904 can be generated via the cost space aggregation 804 as described herein.
[0101] Cost space sampling 802 can be performed separately for different sensor modalities, as described herein. For example, with respect to FIGS. 10A - 10D, a camera - or landmark - based cost space (e.g., corresponding to the base layer of a fused HD map) can be generated using geometric cost and / or semantic cost analysis at each pose of the cost space. For example, at a given time step or frame, outputs 204 - e.g., landmark positions such as lane dividers, road boundaries, signs, poles, etc. - can be calculated with respect to an image, e.g., the image represented in visualization 1002. 3D landmark information from map data 108 can be projected into the 2D image space so as to correspond to the position of the 3D landmark in the 2D image space relative to the current prediction of vehicle 1500 at the current pose being sampled in the cost space. For example, sign 1010 can correspond to a 2D projection from map data, and sign 1012 can correspond to the current prediction or output 204 from one or more DNNs 202. Similarly, lane divider 1014 can correspond to a 2D projection from map data 108, and lane divider 1016 can correspond to the current prediction or output 204 from one or more DNNs 202. To calculate the cost of the current pose - e.g., represented by pose indicator 1018 - the current output 204 from DNN 202 can be transformed into a distance function corresponding to the geometry of the prediction as represented in visualization 1004 (e.g., where the prediction is split into points, each point having zero cost at its center and the cost increasing outward from the center until the maximum cost is achieved as represented by the white area of visualization 1004), and the current output 204 can be individually transformed into the semantic labels of the prediction as represented in visualization 1006. Additionally, the 2D projection of the 3D landmark can be projected into the image space, and each point from the 2D projection can be compared to the portion of the distance - function representation where the projected point lands to determine the associated cost. As such, the dot portion can correspond to the distance - function representation of the current prediction of DNN 202, and the black solid line or dots can represent the 2D projection from map data 108.Costs can be calculated for each point of the 2D projection, and an average cost can be determined using the relative costs from each point. For example, the cost at point 1020A can be high or maximum, and the cost at point 1020B can be low - for example, the cost at point 1020B aligns with the center of the lane divider's distance function representation. This cost can correspond to the geometric cost of the current pose of the current frame. Similarly, the semantic label corresponding to the 2D projection point from the map data 108 can be compared with the semantic information of the projection, as shown in FIG. 10C. As such, if the points do not semantically match, the cost can be set to the maximum value, and if the points match, the cost can be set to the minimum value. These values can be averaged - or weighted in other ways - to determine the final semantic cost. The final semantic cost and the final geometric cost can be weighted to determine the final overall cost for updating the cost space (e.g., cost space 902). For example, for each point, the semantic cost may need to be low for the corresponding geometric cost to vote. As such, if the semantic information does not match, the cost for that particular point can be set to the maximum value. If the semantic information matches, the cost can be set to the minimum value or zero for the semantic cost, and the final cost for that point can represent the geometric cost. Finally, the points within the cost space corresponding to the current pose can be updated to reflect the final costs of all the 2D projection points.
[0102] As another example, with respect to FIGS. 11A - 11B, a RADAR - based cost space (e.g., corresponding to the RADAR layer of a fused HD map) can be generated using a distance function corresponding to map data 108 (e.g., the top - down projection of a RADAR point cloud having each point converted to a distance - function representation). For example, at a given time step or frame, map data 108 corresponding to that represented in the RADAR point cloud - visualization 1102 can be converted to a distance function as represented in visualization 1104 - for example, where each RADAR point can have zero cost at its center and the cost increases to a maximum cost as the distance from the center increases. For example, with respect to visualization 1104, the white portions of visualization 1102 can correspond to the maximum cost. To calculate the cost of the current pose - e.g., represented by pose indicator 1106 - the RADAR data from sensor data 102 can be converted to a RADAR point cloud and compared to the distance - function representation of the RADAR point cloud (e.g., as represented in visualization 1104). The hollow circles in visualizations 1102 and 1104 can correspond to the current RADAR point cloud prediction of vehicle 1500. As such, for each current RADAR point, the cost can be determined by comparing the distance - function RADAR value to which the current RADAR point corresponds or lands with each current RADAR point. As such, current RADAR point 1108A can have a maximum cost while current RADAR point 1108B can have a low cost - e.g., because point 1108B lands near the center of the point from the RADAR point cloud in map data 108. Finally, an average value or other weighting of the costs from each current RADAR point can be calculated and the final cost value can be used to update the cost space of the currently sampled pose.
[0103] As another example, with respect to FIGS. 12A-12D, a LiDAR-based cost space (e.g., corresponding to the LiDAR layer of a fused HD map) can be generated using a distance function on a (sliced) LiDAR point cloud, a LiDAR intensity map, and / or a LiDAR elevation map. For example, for that as indicated by a given pose-pose indicator 1210, current or real-time LiDAR data (e.g., corresponding to sensor data 102) can be generated, converted to a value for comparison with an intensity map generated from map data 108 (e.g., as shown in visualization 1202), converted to a value for comparison with an elevation map generated from map data 108 (e.g., as shown in visualization 1204), and the LiDAR point cloud of map data 108 (e.g., as shown in visualization 1206) can be converted to a distance function representation of the same (e.g., as shown in visualization 1208) for comparison with the current LiDAR point cloud corresponding to sensor data 102. In an example, the LiDAR point cloud can correspond to a slice of the LiDAR point cloud, and one or more separate slices can be converted to a distance function representation and used to calculate cost. The costs from the elevation comparison, intensity comparison, and distance function comparison can be averaged or otherwise weighted to determine the final cost corresponding to the current pose on the LiDAR-based cost map.
[0104] For example, with respect to FIG. 12A, the LiDAR layer of the fused HD map represented by map data 108 may include a LiDAR intensity (or reflectivity) image (e.g., a top-down projection of intensity values from fused LiDAR data). For example, painted surfaces such as lane markers may have higher reflectivity, and this reflected intensity can be captured and used to compare map data 108 with current LiDAR sensor data. Current LiDAR sensor data from vehicle 1500 can be converted to a LiDAR intensity representation 1212A and compared with the LiDAR intensity image from map data 108 at the current pose. For points in the current LiDAR intensity representation 1212A that have intensity values similar to or matching points from map data 108, the cost can be low, and if the intensity values do not match, the cost can be high. For example, a zero difference in the intensity of points may correspond to zero cost, a difference above a threshold may correspond to maximum cost, and between zero difference and the threshold difference, the cost can increase from zero cost to maximum cost. The cost of each point in the current LiDAR intensity representation 1212A can be averaged with each other's points or weighted in other ways to determine the cost of the LiDAR intensity comparison.
[0105] As another example, with respect to FIG. 12B, the LiDAR layer of the fused HD map represented by map data 108 may include a LiDAR elevation image (e.g., a top-down projection of elevation values that results in a top-down depth map). Current LiDAR sensor data from vehicle 1500 can be converted to a LiDAR elevation representation 1212B and compared to the LiDAR elevation image generated from map data 108 at the current pose. For points in the current LiDAR elevation representation 1212B that have elevation values similar to or matching points from the map data 108, the cost can be low, and if the elevation values do not match, the cost can be high. For example, a zero difference in the elevation of points can correspond to zero cost, a difference above a threshold can correspond to maximum cost, and between the zero difference and the threshold difference, the cost can increase from zero cost to maximum cost. The cost for each point in the current LiDAR elevation representation 1212B can be averaged or otherwise weighted among the points to determine the cost of the LiDAR elevation comparison.
[0106] In some embodiments, the elevation value from the map data 108 can be determined with respect to the origin 652 of the current road segment 610, and the elevation value from the current LiDAR elevation representation 1212B can correspond to the origin 820 or reference point of the vehicle 1500, so the conversion can be performed to compare the LiDAR elevation image from the map data 108 with the values from the LiDAR elevation representation 1212B. For example, if a point from the LiDAR elevation representation 1212B has an elevation value of 1.0 meter (e.g., 1.0 meter above the origin of the vehicle 1500), the point in the map data 108 corresponding to the point from the representation 1212B has a value of 1.5 meters and an elevation difference between the origin 820 of the vehicle 1500 and the origin 652 of the road segment 610 of 0.5 meter (e.g., the origin 652 of the road segment is 0.5 meter higher than the origin 820 of the vehicle 1500), and the actual difference between the point from the map data 108 and the representation 1212B can be 0.0 meter (e.g., 1.5 meters - 0.5 meters = 1 meter as the final value of the point from the map data 108 with respect to the origin 820 of the vehicle 1500). Depending on the embodiment, the conversion between the values of the map data 108 or the representation 1212B can correspond to a conversion from the road segment origin 652 to the vehicle origin 820, from the vehicle origin 820 to the road segment origin 652, or a combination thereof.
[0107] As a further example, with respect to FIGS. 12C-12D, the LiDAR layer of the fused HD map represented by map data 108 may include sliced LiDAR point clouds (e.g., corresponding to a ground plane slice, a giraffe plane slice, another defined slice, e.g., a 1-meter thick slice extending from 2 meters to 3 meters from the ground plane, etc.). In some embodiments, the point cloud may not be sliced and instead may represent the entire point cloud. The sliced LiDAR point cloud (e.g., as shown in visualization 1206) may be converted to a distance function representation of the same (e.g., as shown in visualization 1208). For example, each point from the LiDAR point cloud may be converted such that the center of the point has zero cost and the cost increases further from the center of the point to some maximum cost (e.g., as represented by the white region of visualization 1208). To calculate the cost of the current pose - e.g., as represented by pose indicator 1210 - the LiDAR data from sensor data 102 is converted to a LiDAR point cloud (or its corresponding slice) and can be compared to the distance function representation of the LiDAR point cloud (e.g., as represented in visualization 1208). The hollow circles within visualizations 1206 and 1208 may correspond to the current LiDAR point cloud prediction of vehicle 1500. As such, for each current LiDAR point, the cost can be determined by comparing the respective current LiDAR point to the distance function LiDAR value that the current LiDAR point corresponds to or lands on. As such, current LiDAR point 1214A may have a maximum cost, while current LiDAR point 1214B may have a low cost - e.g., because point 1214B lands near the center of the points from the LiDAR point cloud within map data 108. Finally, an average value or other weighting of the respective costs from the current LiDAR points can be calculated and the final cost value can be used - in addition to the cost values from the elevation and intensity comparisons - to update the cost space of the currently sampled pose.
[0108] In some embodiments, at least one of the LiDAR-based cost space sampling 802, cost space aggregation 804, and / or filtering 806 can be executed on a GPU - e.g., an individual GPU, virtual GPU, etc. - and / or using one or more parallel processing devices. For example, for a camera-based cost space (e.g., as described with respect to FIGS. 10A-10C), the detection information and map data 108 projection can be stored as a texture in memory on or accessible to the GPU, and the comparison can correspond to a texture search executed using the GPU. Similarly, with respect to LiDAR and / or RADAR, the comparison can correspond to a texture search. Additionally, in some embodiments, parallel processing can be used to execute two or more cost spaces in parallel - e.g., a first cost space corresponding to LiDAR and a second cost space corresponding to RADAR can be generated in parallel using different GPU and / or parallel processing device resources. For example, individual localizations 808 can be computed in parallel such that the system execution time for fused localization is reduced. As a result, these processes can be executed more efficiently than when executed by the CPU alone.
[0109] Referring back to FIG. 8A, after cost space sampling 802 has been performed for a single frame or time step and for any number of sensor modalities, cost space aggregation 804 can be performed. Cost space aggregation 804 can be performed separately for each sensor modality - for example, LiDAR-based cost space aggregation, RADAR-based cost space aggregation, camera-based cost space aggregation, etc. For example, cost space aggregation 804 can aggregate cost spaces calculated for any number of frames (e.g., 25 frames, 80 frames, 100 frames, 300 frames, etc.). To aggregate the cost spaces, each previously calculated cost space can be ego-motion compensated to correspond to the current frame. For example, the rotation and / or translation of vehicle 1500 from each previous frame included in the aggregation relative to the current pose of vehicle 1500 can be determined and used to warp the cost space values from the previous frame. In addition to transforming the previous cost space based on ego-motion, the cost space can also be transformed using a transformation from one road segment to the next - such as that described with respect to FIG. 6I - so that each cost space corresponds to the origin 652 of the current road segment 610 of the fused HD map. For example, while vehicle 1500 is localized with respect to a first road segment 610, some cost spaces that will be aggregated may be generated, and while vehicle 1500 is localized with respect to a second road segment 610, some other cost spaces that will be aggregated may be generated. As such, the cost space from the first or previous road segment 610 can be transformed such that the cost space values are relative to the second or current road segment 610. The cost spaces can be aggregated once within the same reference frame corresponding to the current frame and the current road segment 610. As a result, and referring to FIG. 9B, the ego-motion of vehicle 1500 over time can help eliminate the ambiguity of the individual cost spaces such that an aggregated cost space 904 can be generated.
[0110] The aggregated cost space 904 can then undergo a filtering process 806 - for example, using a Kalman filter or another filter type - to determine an ellipsoid 910 corresponding to the calculated position of the vehicle 1500 with respect to the current road segment 610. Similar to what was described above regarding the transformation of the aggregated cost space 904, the filtered cost space 906 can also undergo a transformation to correct for self-motion and road segment switching. The ellipsoid 910 can indicate the current position of the vehicle 1500 with respect to the current road segment 610, and this process can be repeated for each new frame. The result may be an individual localization based on the calculated sensor modality of the ellipsoid 910, or multiple ellipsoids 910 can be calculated for each frame - for example, one for each sensor modality.
[0111] The localization fusion 810 can then be performed in the individual localization 808 to generate a final localization result. For example, referring to FIG. 13, the individual localization 808 may correspond to a LiDAR-based localization 1302 (e.g., represented by an ellipse and an origin in the visualization 1300), a RADAR-based localization 1304, a camera-based localization 1306, other sensor modality localizations (not shown), and / or a fused localization 1308. Only a single localization for each sensor modality is described herein, but this is not intended to be limiting. In some embodiments, there may be two or more localization results for different sensor modalities. For example, the vehicle 1500 may be localized with respect to a first camera (e.g., a front-facing camera) and may be separately localized with respect to a second camera (e.g., a rear-facing camera). In such an example, the individual localization 808 may include the first camera-based localization and the second camera-based localization. In some embodiments, the fused localization 1308 may correspond to the fusion of the individual localizations 808 in the current frame and / or may correspond to the previous fused localization result from one or more previous frames carried over to the current frame - e.g., based on self-motion. As such, in an embodiment, the fused localization 1308 of the current frame can consider the individual localizations 808 and the previous fused localization result to advance the current localization state through the frames.
[0112] To calculate the fused localization 1308 of the current frame, a match / mismatch analysis can be performed on the individual localizations 808. For example, in some embodiments, a distance threshold can be used to determine clusters of the individual localizations 808, and the cluster with the least intra-cluster covariance can be selected for fusion. The individual localizations 808 within the selected cluster can then be averaged or otherwise weighted to determine the fused localization 1308 of the current frame. In some embodiments, a filter - for example, a Kalman filter - can be used to generate the fused localization 1308 of the clustered individual localizations 808 of the current frame. For example, the Kalman filter, when used, may not handle outliers well, and thus a large number of outliers can have an undesirable effect on the final result. As such, a clustering technique can help filter out or remove outliers so that the Kalman filter - based fusion is more accurate. In some embodiments, for example, if the previous fusion result is carried over to the current frame as an individual localization 808, the fusion result from the previous frame can drift. For example, if the current individual localization 808 is significantly different from the fusion result (e.g., if the fusion result can be filtered out of the cluster), the fusion result can be re-initialized for the current frame, and the re-initialized fusion result can then be carried over to subsequent frames until a certain amount of drift is detected again.
[0113] In some embodiments, the fused localization 1308 can be determined by factoring at each individual localization 808. For example, instead of grouping the results into clusters, each individual localization 808 can be weighted based on a distance assessment. For example, covariance can be calculated for the individual localizations 808, and the individual localization 808 having the highest covariance (e.g., corresponding to the largest outlier) can be weighted less in determining the fused localization 1308. This can be performed using a robustified mean value so that outliers do not have an undesirable effect on the fused localization 1308. For example, the distance of each individual localization 808 to the robustified mean value can be calculated, and the greater the distance, the less weight the individual localization 808 can have in determining the fused localization 1308.
[0114] The fused localization 1308 of the current frame can then be used to localize the vehicle 1500 with respect to the road segment 610 of the fused HD map (represented by the map data 108) and / or with respect to the global coordinate system. For example, since the road segment 610 can have a known global position, the localization of the vehicle 1500 to the road segment 610 can have a corresponding global localization result. Additionally, since the local or relative localization to the road segment 610 is more accurate than just the global or GNSS localization result, the planning and control of the vehicle 1500 can be more reliable and safer than in a GNSS-based localization system alone.
[0115] Referring now to FIG. 14, each block of method 1400 described herein may include a computing process that is executed using any combination of hardware, firmware, and / or software. For example, various functions may be implemented by a processor that executes instructions stored in memory. Method 1400 may also be implemented as computer-usable instructions stored on a computer storage medium. Method 1400 may be provided, by way of example, as a stand-alone application, service, or hosted service (either stand-alone or in combination with another hosted service), or as a plug-in to another product. Further, method 1400 is described, by way of example, with respect to process 110 of FIG. 8A. However, this method 1400 may be executed within any one process, or any combination of processes and systems, including but not limited to those described herein, by any one system.
[0116] FIG. 14 is a flow diagram illustrating method 1400 for localization, according to some embodiments of the present disclosure. Method 1400 includes, at block B1402, generating, in each of a plurality of frames, one or more of sensor data or DNN outputs based on sensor data. For example, sensor data 102 and / or output 204 may be computed in each of some buffered number of frames or time steps (e.g., 25, 50, 70, 100, etc.). For example, if LiDAR-based localization or RADAR-based localization is used, sensor data 102 may be generated and / or processed in each frame or time step (e.g., to generate an elevation representation, an intensity representation, etc.). If camera-based localization is used, sensor data 102 may be applied to DNN 202 to generate output 204 corresponding to landmarks.
[0117] Method 1400 includes, at block B1404, comparing, in each frame, one or more of the sensor data or DNN outputs with the map data to generate a cost space representing the likelihood that the vehicle is located at each of a plurality of poses. For example, cost space sampling 802 may be performed in each frame to generate a cost space corresponding to a particular sensor data format. The comparison of the map data 108 with the sensor data 102 and / or the output 204 may include comparing LiDAR elevation information, LiDAR intensity information, LiDAR point cloud (slice) information, RADAR point cloud information, camera landmark information (e.g., comparing the 2D projection of the 3D landmark positions from the map data 108 with the current real-time prediction of the DNN 202), and the like.
[0118] Method 1400 includes, at block B1406, aggregating each cost space from each frame to generate an aggregated cost space. For example, cost space aggregation 804 may be performed on some buffered or previous cost spaces to generate the aggregated cost space. The aggregation may include self-motion transformation of the previous cost space and / or road segment transformation of the previous cost space generated for road segments other than the current road segment 610 of the vehicle 1500.
[0119] Method 1400 includes, at block B1408, applying a Kalman filter to the aggregated cost space to calculate the final position of the vehicle in the current frame of a plurality of frames. For example, filtering process 806 may be applied to the aggregated cost space to generate an ellipse or other representation of the estimated position of the vehicle 1500 for a particular sensor format in the current frame or time step.
[0120] Blocks B1402 - B1408 of process 1408 may be repeated for any number of different sensor formats - in an embodiment, in parallel - such that two or more ellipses or position predictions of the vehicle 1500 are generated.
[0121] Method 1400 includes, at block B1410, using the final position of the vehicle, in addition to one or more other final positions of the vehicle, to determine the fused position of the vehicle. For example, ellipsoids of other representations output after filtering 806 of different sensor modalities may undergo localization fusion 810 to produce a final fused localization result. In some embodiments, as described herein, previous fused localization results from previous frames or time steps may also be used when determining the current fused localization result.
[0122] Exemplary autonomous vehicle FIG. 15A is a diagram of an exemplary autonomous vehicle 1500 according to some embodiments of the present disclosure. The autonomous vehicle 1500 (or referred to herein as "vehicle 1500") can include, but is not limited to, a passenger vehicle, such as a car, truck, bus, first responder vehicle, shuttle, electric or motorized bicycle, motorcycle, fire truck, police vehicle, ambulance, boat, construction vehicle, submarine, drone, and / or another type of vehicle (e.g., unmanned and / or carrying one or more passengers). The autonomous vehicle is generally described in terms of the level of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicle" (Standard No. J3016 - 201806 published on June 15, 2018, Standard No. J3016 - 201609 published on September 30, 2016, and previous and future versions of this standard). The moving vehicle 1500 can have the ability to function according to one or more of automation levels 3 to 5 of the autonomous driving level. For example, the moving vehicle 1500 can have the ability of conditional automation (level 3), high automation (level 4), and / or full automation (level 5) depending on the embodiment.
[0123] The moving vehicle 1500 can include components such as a chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the moving vehicle. The moving vehicle 1500 can include a propulsion system 1550, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. The propulsion system 1550 can be connected to the drive train of the moving vehicle 1500 and can include a transmission to enable the propulsion force of the moving vehicle 1500. The propulsion system 1550 can be controlled in response to receiving a signal from a throttle / acceleration device 1552.
[0124] The steering system 1554, which may include a steering wheel, can be used to steer the moving vehicle 1500 (e.g., along a desired path or route) when the propulsion system 1550 is operating (e.g., when the moving vehicle is in motion). The steering system 1554 can receive signals from a steering actuator 1556. The steering wheel may be an option for a fully automated (level 5) function.
[0125] The brake sensor system 1546 can be used to operate the vehicle brakes in response to receiving signals from a brake actuator 1548 and / or a brake sensor.
[0126] The controller 1536, which may include one or more system-on-chips (SoCs) 1504 (FIG. 15C) and / or GPUs, can provide signals (e.g., representations of commands) to one or more components and / or systems of the mobile vehicle 1500. For example, the controller can send signals to operate the mobile vehicle brakes via one or more brake actuators 1548, to operate the steering system 1554 via one or more steering actuators 1556, and to operate the propulsion system 1550 via one or more throttle / acceleration devices 1552. The controller 1536 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist the driver in operating the mobile vehicle 1500. The controller 1536 may include a first controller 1536 for autonomous driving functions, a second controller 1536 for functional safety functions, a third controller 1536 for artificial intelligence functions (e.g., computer vision), a fourth controller 1536 for infotainment functions, a fifth controller 1536 for redundancy in emergencies, and / or other controllers. In some examples, a single controller 1536 can handle two or more of the aforementioned functions, and two or more controllers 1536 can handle a single function and / or any combination thereof.
[0127] Controller 1536 can provide signals for controlling one or more components and / or systems of the moving vehicle 1500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data can be received, for example and without limitation, from a global navigation satellite system sensor 1558 (e.g., a global positioning system sensor), a RADAR sensor 1560, an ultrasonic sensor 1562, a LIDAR sensor 1564, an inertial measurement unit (IMU) sensor 1566 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 1596, a stereo camera 1568, a wide view camera 1570 (e.g., a fish-eye camera), an infrared camera 1572, a surround camera 1574 (e.g., a 360-degree camera), a long-range and / or mid-range camera 1598, a speed sensor 1544 (e.g., for measuring the speed of the moving vehicle 1500), a vibration sensor 1542, a steering sensor 1540, a brake sensor (e.g., as part of a brake sensor system 1546), and / or other sensor types.
[0128] One or more of the controllers 1536 of the mobile vehicle 1500 can receive an input (e.g., represented by input data) from the instrument cluster 1532 of the mobile vehicle 1500 and provide an output (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 1534, an audible annunciator, a loudspeaker, and / or other components of the mobile vehicle 1500. The output can include information such as mobile vehicle velocity, speed, time, map data (e.g., the HD map 1522 of FIG. 15C), position data (e.g., the position of the mobile vehicle 1500 on a map, etc.), direction, the positions of other mobile vehicles (e.g., occupancy grids), information regarding objects and the situation of objects as perceived by the controller 1536, and the like. For example, the HMI display 1534 can display information regarding the presence of one or more objects (e.g., road signs, warning signs, changes in traffic signals, etc.) and / or driving operations that the mobile vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at Exit 34B within 3.22 km (2 miles), etc.).
[0129] The mobile vehicle 1500 further includes a network interface 1524 that can communicate via one or more networks using one or more wireless antennas 1526 and / or a modem. For example, the network interface 1524 can have the ability to communicate via LTE, WCDMA (registered trademark), UMTS, GSM, CDMA2000, and the like. The wireless antenna 1526 can also use local area networks such as Bluetooth (registered trademark), Bluetooth (registered trademark) LE, Z-Wave, ZigBee, and / or low power wide-area networks (LPWANs) such as LoRaWAN, SigFox to enable communication between objects (e.g., mobile vehicles, mobile devices, etc.) in the environment.
[0130] FIG. 15B is an example of the camera positions and fields of view of the exemplary autonomous vehicle 1500 of FIG. 15A according to some embodiments of the present disclosure. The cameras and their respective fields of view are one exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be placed at different positions on the moving vehicle 1500.
[0131] The camera type of the camera may include, but is not limited to, a digital camera adapted to be used with components and / or systems of the moving vehicle 1500. The camera can operate at automotive safety integrity level (ASIL) B and / or at another ASIL. The camera type may have the ability to capture images at any rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may have the ability to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include an RCCC (red clear clear clear) color filter array, an RCCB (red clear clear blue) color filter array, an RBGC (red blue green clear) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to increase light sensitivity.
[0132] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional mono-camera can be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more of the cameras (e.g., all of the cameras) can record and provide image data (e.g., video) simultaneously.
[0133] One or more of the cameras can be mounted in mounting components such as custom-designed (3D printed) parts to remove stray light and reflections from inside the vehicle that can interfere with the camera's image data capture ability (e.g., reflections from the dashboard reflected in the front windshield mirror). Referring to the side mirror mounting component, the side mirror component can be custom 3D printed such that the camera mounting plate conforms to the shape of the side mirror. In some examples, the camera can be integrated within the side mirror. For side view cameras, the camera can also be integrated within four struts located at each corner of the cabin.
[0134] A camera having a field of view that includes a portion of the environment in front of the moving vehicle 1500 (e.g., a forward-facing camera) can be used for surround view to assist in identifying the forward path and obstacles and, with the assistance of one or more controllers 1536 and / or a control SoC, assist in providing information essential for the generation of an occupancy grid and / or the determination of a preferred moving vehicle path. The forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera can also be used for ADAS functions and systems including other functions such as lane departure warning ("LDW (Lane Departure Warning)"), autonomous cruise control ("ACC (Autonomous Cruise Control)"), and / or traffic sign recognition.
[0135] Various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imaging device. Another example may be a wide-view camera 1570 that can be used to capture objects entering the view from the surroundings (e.g., pedestrians, intersecting traffic, or bicycles). Although only one wide-view camera is shown in FIG. 15B, any number of wide-view cameras 1570 may be present on the moving vehicle 1500. Additionally, a long-range camera 1598 (e.g., a long-view stereo camera pair) can be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. The long-range camera 1598 can also be used for object detection and classification, as well as for basic object tracking.
[0136] One or more stereo cameras 1568 can also be included in the forward-facing configuration. The stereo camera 1568 can include an integrated control unit with an expandable processing unit that can provide a programmable logic (FPGA) and a multi-core microprocessor with a CAN or Ethernet® interface integrated on a single chip. Such a unit can be used to generate a 3D map of the environment of the moving vehicle that includes distance estimates for all points in the image. An alternative stereo camera 1568 can include a compact stereo vision sensor that includes two camera lenses (one each on the left and right) and an image processing chip that can measure the distance from the moving vehicle to the target object and activate autonomous emergency braking and lane departure warning functions using the generated information (e.g., metadata). Other types of stereo cameras 1568 may be used in addition to, or in place of, those described herein.
[0137] A camera (e.g., a side view camera) having a field of view that includes a portion of the environment relative to the side of the moving vehicle 1500 can be used for surround view to provide information for creating and updating an occupancy grid and for generating a side impact collision warning. For example, surround cameras 1574 (e.g., four surround cameras 1574 as shown in FIG. 15B) can be positioned on the moving vehicle 1500. The surround cameras 1574 can include wide view cameras 1570, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be disposed in front of, behind, and on the sides of the moving vehicle. In an alternative arrangement, the moving vehicle may use three surround cameras 1574 (e.g., left, right, and rear), and one or more other cameras (e.g., a forward-facing camera) can be utilized as a fourth surround view camera.
[0138] A camera (e.g., a rear view camera) having a field of view that includes a portion of the environment relative to the rear of the moving vehicle 1500 can be used for parking assistance, surround view, rear collision warning, and for creating and updating an occupancy grid. As described herein, a variety of cameras can be used, including but not limited to cameras suitable as forward-facing cameras (e.g., long range and / or mid-range cameras 1598, stereo cameras 1568), infrared cameras 1572, etc.
[0139] FIG. 15C is a block diagram of an exemplary system architecture of the exemplary autonomous vehicle 1500 of FIG. 15A, according to some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be excluded altogether. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by an entity may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory.
[0140] Each of the components, features, and systems of the moving vehicle 1500 of FIG. 15C is illustrated as being connected via a bus 1502. The bus 1502 may include a Controller Area Network (CAN) data interface (or referred to as a "CAN bus"). CAN may be a network within the moving vehicle 1500 used to assist in the control of various features and functions of the moving vehicle 1500, such as the operation of brakes, acceleration, brakes, steering, windshield wipers, etc. The CAN bus may be configured to have dozens or hundreds of nodes, each having its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button position, and / or other moving vehicle status indicators. The CAN bus may be ASIL B compliant.
[0141] Bus 1502 is described herein as being a CAN bus, but this is not intended to be limiting. For example, in addition to, or as an alternative to, a CAN bus, FlexRay and / or Ethernet® may be used. Additionally, a single line is used to represent bus 1502, but this is not intended to be limiting. For example, any number of buses 1502 may exist that include one or more CAN buses, one or more FlexRay buses, one or more Ethernet® buses, and / or one or more other types of buses that use different protocols. In some examples, two or more buses 1502 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 1502 may be used for a collision avoidance function, and a second bus 1502 may be used for actuation control. In any example, each bus 1502 may communicate with any of the components of the moving vehicle 1500, and two or more buses 1502 may communicate with the same component. In some examples, each SoC 1504, each controller 1536, and / or each computer within the moving vehicle may have access to the same input data (e.g., an input from a sensor of the moving vehicle 1500) and may be connected to a common bus such as a CAN bus.
[0142] The moving vehicle 1500 may include one or more controllers 1536, such as those described herein with respect to FIG. 15A. The controller 1536 may be used for various functions. The controller 1536 may be coupled to any of the various other components and systems of the moving vehicle 1500 and may be used for the control of the moving vehicle 1500, the artificial intelligence of the moving vehicle 1500, the infotainment for the moving vehicle 1500, and / or the like.
[0143] The mobile vehicle 1500 may include a system-on-chip (SoC) 1504. The SoC 1504 may include a CPU 1506, a GPU 1508, a processor 1510, a cache 1512, an accelerator 1514, a data store 1516, and / or other components and features not shown. The SoC 1504 may be used to control the mobile vehicle 1500 within various platforms and systems. For example, the SoC 1504 may be coupled in a system (such as the system of the mobile vehicle 1500) having an HD map 1522 that can obtain map refreshes and / or updates via a network interface 1524 from one or more servers (such as the server 1578 of FIG. 15D).
[0144] The CPU 1506 may include a CPU cluster or CPU complex (or also referred to as a "CCPLEX"). The CPU 1506 may include a plurality of cores and / or an L2 cache. For example, in some embodiments, the CPU 1506 may include 8 cores within a coherent multiprocessor configuration. In some embodiments, the CPU 1506 may include 4 dual-core clusters, each cluster having a dedicated L2 cache (such as a 2MB L2 cache). The CPU 1506 (such as a CCPLEX) may be configured to support simultaneous cluster operation that allows any combination of clusters of the CPU 1506 to become active at any given time.
[0145] CPU 1506 can implement power management capabilities that include one or more of the following features: individual hardware blocks can be automatically clock-gated when in an idle state to conserve dynamic power; each core clock can be gated when the core is not actively executing instructions by the execution of WFI / WFE instructions; each core can be independently power-gated; each core cluster can be independently clock-gated when all cores are clock-gated or power-gated, and / or each core cluster can be independently power-gated when all cores are power-gated. CPU 1506 can further implement an enhanced algorithm for managing power states, where the allowed power states and expected wake-up times are specified and the hardware / microcode determines the best power state to input to the cores, clusters, and CCPLEX. The processing cores can support a simplified power state input sequence in software where the work is offloaded to the microcode.
[0146] GPU 1508 may include an integrated GPU (or referred to herein as "iGPU"). GPU 1508 can be programmable and can be efficient for parallel workloads. In some examples, GPU 1508 can use an enhanced tensor instruction set. GPU 1508 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache having at least 96 KB of storage capacity), and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache having 512 KB of storage capacity). In some embodiments, GPU 1508 may include at least eight streaming microprocessors. GPU 1508 can use a compute application programming interface (API). Additionally, GPU 1508 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0147] The GPU 1508 can be power optimized for the best performance in automotive and embedded use cases. For example, the GPU 1508 can be manufactured on FinFET (Fin field-effect transistor). However, this is not intended to be limiting, and the GPU 1508 can be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate several mixed-precision processing cores partitioned into multiple blocks. By way of example and not limitation, for instance, 64 PF32 cores and 32 PF64 cores may be partitioned into 4 processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, 2 mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths for efficient execution of workloads having a mix of compute and addressing operations. The streaming microprocessor can include independent thread scheduling capabilities to enable higher fine-grained synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to simplify programming while improving performance.
[0148] In some examples, GPU 1508 may include high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide for a peak memory bandwidth of 900 GB / second. In some examples, in addition to, or instead of, HBM memory, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used.
[0149] GPU 1508 can include unified memory technology that includes access counters to enable more accurate movement of those memory pages to the processor that most frequently accesses the memory pages, thereby improving the efficiency of the memory ranges shared among the processors. In some examples, address translation service (ATS) support may be used to enable the GPU 1508 to directly access the CPU 1506 page table. In such examples, when the GPU 1508 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU 1506. In response, the CPU 1506 can examine its page table for the virtual-to-physical mapping of the address and send the translation back to the GPU 1508. As such, the unified memory technology can enable a single unified virtual address space for the memory of both the CPU 1506 and the GPU 1508, thereby simplifying GPU 1508 programming and porting of applications to the GPU 1508.
[0150] In addition, GPU 1508 may include an access counter that can record the frequency of access of GPU 1508 to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses that page most frequently.
[0151] SoC 1504 may include any number of caches 1512, including those described herein. For example, cache 1512 may include an L3 cache that is available to both CPU 1506 and GPU 1508 (e.g., connected to both CPU 1506 and GPU 1508). Cache 1512 may include a write-back cache that can record the state of lines, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include more than 4 MB, although smaller cache sizes may be used, depending on the embodiment.
[0152] SoC 1504 may include an arithmetic logic unit (ALU) that can be utilized when executing processing for any of the various tasks or operations of vehicle 1500 (e.g., processing DNN). In addition, SoC 1504 may include a floating point unit (FPU) (or other math co-processor or numeric co-processor type) for performing mathematical operations within the system. For example, SoC 104 may include one or more FPUs integrated as execution units within CPU 1506 and / or GPU 1508.
[0153] SoC 1504 may include one or more acceleration devices 1514 (e.g., hardware acceleration devices, software acceleration devices, or combinations thereof). For example, SoC 1504 may include a hardware acceleration cluster that may include optimized hardware acceleration devices and / or large on-chip memories. A large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other operations. The hardware acceleration cluster may be used to complement the GPU 1508 and to offload some of the tasks of the GPU 1508 (e.g., to free up more cycles of the GPU 1508 for other tasks). As an example, the acceleration device 1514 may be used for target workloads that are stable enough to be suitable for acceleration (e.g., perception, convolutional neural networks (CNNs)). In this specification, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., as used for object detection).
[0154] Accelerator 1514 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) configured to provide an additional 10 trillion operations per second for deep learning applications and inferences. The TPU may be an accelerator configured and optimized to execute image processing functions (e.g., CNN, RCNN, etc.). The DLA may further be optimized for a specific set of neural network types and floating point operations, as well as inferences. The design of the DLA can provide more performance per millimeter than a general-purpose GPU and greatly exceed the performance of the CPU. The TPU can execute several functions, including, for example, a single instance convolution function and a post-processor function that support INT8, INT16, and FP16 data types for both features and weights.
[0155] The DLA can quickly and efficiently execute neural networks, particularly CNNs, with processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object identification and detection using data from a camera sensor, CNNs for distance estimation using data from a camera sensor, CNNs for emergency vehicle detection and identification and detection using data from a microphone, CNNs for face recognition and moving vehicle owner identification using data from a camera sensor, and / or CNNs for security and / or safety related events.
[0156] The DLA can execute any function of the GPU 1508, and by using the inference accelerator, for example, a designer can target either the DLA or the GPU 1508 for any function. For example, a designer can focus on processing CNNs and floating point operations on the DLA and leave other functions to the GPU 1508 and / or other accelerators 1514.
[0157] The acceleration device 1514 (e.g., a hardware acceleration cluster) may include, or be referred to herein as, a programmable vision accelerator (PVA) and a computer vision acceleration device. The PVA can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0158] The RISC core can interact with an image sensor (e.g., the image sensor of any of the cameras described herein), an image signal processor, and / or the like. Each RISC core may include any amount of memory. The RISC core can use any of several protocols depending on the embodiment. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0159] DMA can enable the components of the PVA to access the system memory independent of the CPU1506. DMA can support any number of features used to provide optimization for the PVA, including but not limited to supporting multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing up to six dimensions or more, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0160] The vector processor may be a programmable processor designed to efficiently and flexibly execute the programming of computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. The vector processing subsystem can operate as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.
[0161] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each vector processor may be configured to execute independently from other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may be able to simultaneously execute different computer vision algorithms on the same image, or even execute different algorithms sequentially on an image or portions of an image. In particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system safety.
[0162] The accelerator 1514 (e.g., a hardware acceleration cluster) may include a computer vision network-on-chip and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 1514. In some examples, the on-chip memory may include at least 4MB of SRAM consisting of, for example and without limitation, 8 field-configurable memory blocks that may be accessible by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed access to the PVA and the DLA to the memory. The backbone may include a computer vision network-on-chip that interconnects the PVA and the DLA to the memory (e.g., using an APB).
[0163] The computer vision network-on-chip may include an interface that determines that both the PVA and the DLA are operable and provide valid signals before any control signal / address / data transmission. Such an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. This type of interface can comply with the ISO26262 or IEC61508 standard, although other standards and protocols may be used.
[0164] In some examples, the SoC 1504 may include a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and scale of objects (e.g., within a world model) for generating real-time visualization simulations for RADAR signal interpretation, for acoustic propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulation, for comparison against LIDAR data for localization and / or other functions, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing related operations.
[0165] The acceleration device 1514 (e.g., a hardware acceleration device cluster) has various applications for autonomous driving. The PVA may be a programmable vision acceleration device that can be used in extremely important processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are suitable for areas of algorithms that require predictable processing at low power and low latency. In other words, the PVA functions well with low latency and low power and predictable execution time, even on small data sets, with semi-dense or dense normal calculations. Therefore, since the PVA is efficient in object detection and integer calculation operations, in the context of a platform for autonomous vehicles, the PVA is designed to execute classic computer vision algorithms.
[0166] For example, according to one embodiment of the present technology, the PVA is used to execute computer stereo vision. A semi-global matching-based algorithm may be used in some examples, but this is not intended to be limiting. Many applications for level 3-5 autonomous driving require motion estimation / stereo matching on the fly (e.g., SFM (structure from motion), pedestrian recognition, lane detection, etc.). The PVA can execute computer stereo vision functions with inputs from two monocular cameras.
[0167] In some examples, the PVA may be used to execute high-density optical flow. By processing raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used in the time of flight depth processing, for example, by processing the raw time of flight data to provide the processed time of flight data.
[0168] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a measure of the reliability of each object detection. Such reliability values can be interpreted as probabilities or as providing the relative "weight" of each detection compared to other detections. This reliability value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a reliability threshold and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, a false positive detection would cause the moving vehicle to automatically execute emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network that regresses reliability values. The neural network can receive as its inputs at least some subsets of parameters such as bounding box dimensions, ground plane estimations obtained (for example, from another subsystem), the azimuth, distance, and 3D position estimations of objects obtained from the neural network and / or other sensors (such as LIDAR sensor 1564 or RADAR sensor 1560), and the output of an inertial measurement unit (IMU) sensor 1566 that correlates with the above, among others.
[0169] SoC1504 may include a data store 1516 (e.g., memory). The data store 1516 may be the on-chip memory of SoC1504 and can store neural networks to be executed by the GPU and / or DLA. In some examples, the data store 1516 may have a capacity large enough to store multiple instances of the neural network for redundancy and security. The data store 1512 may include an L2 or L3 cache 1512. References to the data store 1516 may include references to memory related to the PVA, DLA, and / or other accelerators 1514 as described herein.
[0170] SoC1504 may include one or more processors 1510 (e.g., embedded processors). The processor 1510 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management capabilities and related security enforcement. The boot and power management processor may be part of the SoC1504 boot sequence and can provide runtime power management services. The boot power and management processor can provide clock and voltage programming, assistance with system low-power state transitions, management of the SoC1504 heat and temperature sensors, and / or management of the SoC1504 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and SoC1504 can use the ring oscillator to detect the temperature of the CPU1506, GPU1508, and / or accelerator 1514. If the temperature is determined to exceed a threshold, the boot and power management processor can enter a temperature fault routine and place SoC1504 in a lower power state and / or put the vehicle 1500 in a chauffeur safe stop mode (e.g., safely stop the vehicle 1500).
[0171] Processor 1510 may further include a set of embedded processors that can perform the functions of an audio processing engine. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio via multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core that has a digital signal processor with dedicated RAM.
[0172] Processor 1510 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripheral devices (such as timers and interrupt controllers), various I / O controller peripheral devices, and routing logic.
[0173] Processor 1510 may further include a safety cluster engine that includes a dedicated processor subsystem for processing the safety management of automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripheral devices (such as timers, interrupt controllers, etc.), and / or routing logic. In the safety mode, two or more cores can operate in a lockstep mode and function as a single core with comparison logic for detecting any differences between their operations.
[0174] Processor 1510 may further include a real-time camera engine that may include a dedicated processor subsystem for processing real-time camera management.
[0175] Processor 1510 may further include a high dynamic range signal processor that includes an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0176] The processor 1510 may include a video image synthesizer, which may be a processing block (e.g., implemented in a microprocessor) that implements the post-video processing functions required by the video playback application to produce the final image for the player window. The video image synthesizer can perform lens distortion correction with the wide-view camera 1570, the surround camera 1574, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of a high-level SoC configured to identify in-cabin events and respond appropriately. The in-cabin system can perform lip reading to activate cellular service and make calls, compose emails, change the destination of the moving vehicle, activate or change the infotainment system and settings of the moving vehicle, or provide voice-activated web surfing. Certain functions are available to the driver only when operating in autonomous mode and are otherwise disabled.
[0177] The video image synthesizer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, when motion occurs within the video, the noise reduction reduces the weight of the information provided by adjacent frames and appropriately weights the spatial information. When an image or a portion of an image does not contain motion, the temporal noise reduction performed by the video image synthesizer can reduce the noise in the current image using information from the previous image.
[0178] The video image synthesizer may also be configured to perform stereo rectification on the input stereo lens frame. The video image synthesizer can further be used for user interface synthesis when the operating system desktop is in use, and the GPU 1508 is not required to continuously render new surfaces. Even when the GPU 1508 is powered on and actively performing 3D rendering, the video image synthesizer can be used to offload the GPU 1508 to improve performance and responsiveness.
[0179] SoC 1504 may further include a mobile industry processor interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for the camera and related pixel input functions to receive video and inputs from the camera. SoC 1504 may further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.
[0180] SoC1504 may further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC1504 can be used to process data from cameras (connected via, for example, Gigabit Multimedia Serial Link and Ethernet®), sensors (such as LIDAR sensor 1564, RADAR sensor 1560, etc., which can be connected via Ethernet®), data from bus 1502 (such as the speed and steering wheel position of moving vehicle 1500), and data from GNSS sensor 1558 (connected via, for example, Ethernet® or CAN bus). SoC1504 may include its own DMA engine and may further include a dedicated high-performance large-capacity storage controller that can be used to free CPU1506 from routine data management tasks.
[0181] SoC1504 may be an end-to-end platform with a flexible architecture that extends to automation levels 3 - 5, thereby leveraging and efficiently using computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible, reliable driving software stack together with deep learning tools, providing an integrated functional safety architecture. SoC1504 can be faster, more reliable, more energy-efficient, and more space-efficient than conventional systems. For example, when accelerator 1514 is coupled with CPU1506, GPU1508, and data store 1516 can provide a fast and efficient platform for level 3 - 5 autonomous vehicles.
[0182] Therefore, the present technology provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be implemented to run on a CPU using a high-level programming language such as the C programming language to execute a variety of processing algorithms across a variety of visual data. However, a CPU often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. Specifically, many CPUs cannot execute real-time composite object detection algorithms, which are requirements for in-vehicle ADAS applications and actual level 3-5 autonomous vehicles.
[0183] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the technology described herein enables multiple neural networks to be executed simultaneously and / or sequentially and enables the results to be combined to enable level 3-5 autonomous driving functionality. For example, a CNN running on a DLA or a dGPU (e.g., GPU 1520) can include text and word recognition that enables a supercomputer to read and understand traffic signs, including signs that the neural network has not been specifically trained for. The DLA can further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass the semantic understanding to a route planning module running on the CPU complex.
[0184] As another example, multiple neural networks can be executed simultaneously as required for level 3, 4, or 5 operation. For example, along with the electro-optical, a warning sign consisting of "Attention: Flashing light indicates a frozen state" can be interpreted independently or collectively by several neural networks. The sign itself can be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing light indicates a frozen state" can be interpreted by a second deployed neural network that informs the route planning software of the moving vehicle (preferably running on a CPU complex) that a frozen state exists when a flashing light is detected. The flashing light can be identified by informing the route planning software of the moving vehicle of the presence (or absence) of the flashing light and operating a third deployed neural network through multiple frames. All three neural networks can be executed simultaneously, such as within the DLA and / or on the GPU 1508.
[0185] In some examples, a CNN for face recognition and moving vehicle owner identification can use data from a camera sensor to identify the presence of a regular driver and / or owner of the moving vehicle 1500. A always-on sensor processing engine can be used to unlock and illuminate the moving vehicle when the owner approaches the driver's side door, and also to stop the operation of the moving vehicle when the owner leaves the moving vehicle in security mode. In this way, the SoC 1504 provides security against theft and / or carjacking.
[0186] In another example, a CNN for emergency vehicle detection and identification can detect and identify an emergency vehicle siren using data from microphone 1596. In contrast to conventional systems that use a general classifier to detect sirens and manually extract features, SoC 1504 uses a CNN for environmental and urban sound classification, as well as for visual data classification. In a preferred embodiment, the CNN executed on the DLA is trained to identify the relative end velocity of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area where the moving vehicle is operating, as identified by GNSS sensor 1558. Thus, for example, when operating in Europe, the CNN will attempt to detect European sirens, and when in the United States, the CNN will attempt to identify only North American sirens. After an emergency vehicle is detected, a control program can be used to execute an emergency vehicle safety routine to decelerate the moving vehicle, stop it at the side of the road, park the moving vehicle, and / or idle the moving vehicle, with the assistance of ultrasonic sensor 1562 until the emergency vehicle has passed.
[0187] The moving vehicle may include a CPU 1518 (e.g., a separate CPU or dCPU) that can be coupled to SoC 1504 via a high-speed interconnect (e.g., PCIe). CPU 1518 may include, for example, an X86 processor. CPU 1518 can be used to perform any of a variety of functions, including, for example, mediating potential inconsistencies between ADAS sensors and SoC 1504, and / or monitoring the status and health of controller 1536 and / or infotainment SoC 1530.
[0188] The mobile vehicle 1500 may include a GPU 1520 (e.g., an individual GPU or a dGPU) that can be coupled to the SoC 1504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 1520 can provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of the mobile vehicle 1500.
[0189] The mobile vehicle 1500 may further include a network interface 1524 that may include one or more wireless antennas 1526 (e.g., one or more wireless antennas for different communication protocols such as cellular antennas, Bluetooth® antennas, etc.). The network interface 1524 can be used to enable wireless connections with the cloud via the Internet (e.g., with servers 1578 and / or other network devices), with other mobile vehicles, and / or with computing devices (e.g., a passenger's client device). To communicate with other mobile vehicles, a direct link can be established between two mobile vehicles and / or an indirect link can be established (e.g., through a network and via the Internet). A direct link can use and provide a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 1500 information regarding mobile vehicles in proximity to the mobile vehicle 1500 (e.g., mobile vehicles in front of, beside, and / or behind the mobile vehicle 1500). This functionality may be part of the cooperative adaptive cruise control function of the mobile vehicle 1500.
[0190] The network interface 1524 may include a SoC that provides modulation and demodulation functions and enables the controller 1536 to communicate via a wireless network. The network interface 1524 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless functions for communicating via LTE, WCDMA®, UMTS, GSM, CDMA2000, Bluetooth®, Bluetooth® LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0191] The mobile vehicle 1500 may further include a data store 1528 that may include storage outside the chip (e.g., outside the SoC 1504). The data store 1528 may include one or more storage elements including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least 1 bit of data.
[0192] The vehicle 1500 may further include a GNSS sensor 1558. The GNSS sensor 1558 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) supports mapping, perception, occupancy grid generation, and / or route planning functions. For example, any number of GNSS sensors 1558 may be used, including but not limited to GPS with an Ethernet® to serial (RS-232) bridge via a USB connector.
[0193] The moving vehicle 1500 may further include a RADAR sensor 1560. The RADAR sensor 1560 can be used by the moving vehicle 1500 for long-range moving vehicle detection even in darkness and / or severe weather conditions. The RADAR functional safety level may be ASIL B. In some examples, the RADAR sensor 1560 can use CAN and / or bus 1502 for control and to access object tracking data (e.g., to transmit data generated by the RADAR sensor 1560) using access to Ethernet (registered trademark) to access raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 1560 may be suitable for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.
[0194] The RADAR sensor 1560 may include different configurations such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for an adaptive cruise control function. The long-range RADAR system can provide a wide field of view realized by two or more independent scans, such as within a range of 250 m. The RADAR sensor 1560 can help distinguish between static and moving objects and can be used by an ADAS system for emergency brake assist and forward collision warning. The long-range RADAR sensor may include a monostatic multi-modal RADAR having a plurality (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In one example having six antennas, the four central antennas can create a focused beam pattern designed to record the surroundings of the moving vehicle 1500 at high speed while minimizing interference from traffic in adjacent lanes. The other two antennas can widen the field of view and enable rapid detection of moving vehicles entering or leaving the lane of the moving vehicle 1500.
[0195] As an example, a mid-range RADAR system can include a range up to 1560 m (front) or 80 m (rear), and a field of view up to 42 degrees (front) or 1550 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be installed at both ends of the rear bumper. When installed at both ends of the rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and adjacent to the moving vehicle.
[0196] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assist.
[0197] The moving vehicle 1500 can further include ultrasonic sensors 1562. The ultrasonic sensors 1562, which can be positioned at the front, rear, and / or sides of the moving vehicle 1500, can be used for parking assist and / or for creating and updating occupancy grids. A wide variety of ultrasonic sensors 1562 can be used, and different ultrasonic sensors 1562 can be used for different ranges of detection (e.g., 2.5 m, 4 m). The ultrasonic sensors 1562 can operate at a functional safety level of ASIL B.
[0198] The moving vehicle 1500 can include a LIDAR sensor 1564. The LIDAR sensor 1564 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 1564 can also be at a functional safety level of ASIL B. In some examples, the moving vehicle 1500 can include multiple (e.g., two, four, six, etc.) LIDAR sensors 1564 that can use Ethernet (registered trademark) (e.g., to provide data to a gigabit Ethernet (registered trademark) switch).
[0199] In some examples, the LIDAR sensor 1564 may have the ability to provide a list of objects and their distances in a 360-degree field of view. Commercially available LIDAR sensors 1564 may have, for example, an accuracy of 2 cm to 3 cm, support for 1500 Mbps Ethernet® connections, and may have a advertised range of about 1500 m. In some examples, one or more non-protruding LIDAR sensors 1564 may be used. In such examples, the LIDAR sensor 1564 may be implemented as a small device that can be incorporated into the front, rear, sides, and / or corners of the moving vehicle 1500. In such examples, the LIDAR sensor 1564 may have a range of 200 m even for low-reflectivity objects and can provide up to a 120-degree horizontal and 35-degree vertical field of view. The LIDAR sensor 1564 attached to the front may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0200] In some examples, LIDAR technologies such as 3D flash LIDAR can also be used. 3D flash LIDAR uses a laser flash as a transmitter to illuminate the area around a moving vehicle up to about 200 m. The flash LIDAR unit includes a receptor that records the laser pulse travel time and the reflected light on each pixel, corresponding in sequence to the range from the moving vehicle to the object. Flash LIDAR can enable high-precision and distortion-free images of the surroundings to be generated with every laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the moving vehicle 1500. Available 3D flash LIDAR systems include solid-state 3D steering array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than a blower. The flash LIDAR device can use class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser light in the form of 3D range point clouds and co-recorded intensity data. By using flash LIDAR, also, since the flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 1564 can be less susceptible to the effects of motion blur, vibration, and / or shock.
[0201] The moving vehicle can further include an IMU sensor 1566. In some examples, the IMU sensor 1566 can be positioned at the center of the rear axle of the moving vehicle 1500. The IMU sensor 1566 can include, for example, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types, but is not limited thereto. In some examples, in a 6-axis application, etc., the IMU sensor 1566 can include an accelerometer and a gyroscope, while in a 9-axis application, the IMU sensor 1566 can include an accelerometer, a gyroscope, and a magnetometer.
[0202] In some embodiments, the IMU sensor 1566 may be implemented as a miniature, high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. As such, in some examples, the IMU sensor 1566 may enable the moving vehicle 1500 to estimate its heading without the need for input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 1566. In some examples, the IMU sensor 1566 and the GNSS sensor 1558 may be combined in a single integrated unit.
[0203] The moving vehicle may include a microphone 1596 disposed within and / or around the moving vehicle 1500. The microphone 1596 may be used, among other things, for emergency vehicle detection and identification.
[0204] The mobile vehicle may further include any number of camera types, including stereo camera 1568, wide view camera 1570, infrared camera 1572, surround camera 1574, long and / or medium distance camera 1598, and / or other camera types. The cameras can be used to capture image data around the entire outer surface of the mobile vehicle 1500. The type of camera used is determined according to the embodiments and requirements of the mobile vehicle 1500, and any combination of camera types can be used to achieve the necessary coverage around the mobile vehicle 1500. Additionally, the number of cameras can vary according to the embodiments. For example, the mobile vehicle may include 6 cameras, 7 cameras, 10 cameras, 12 cameras, and / or another number of cameras. As an example, the cameras can support, but are not limited to, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet (registered trademark). Each camera is described in more detail herein in relation to FIGS. 15A and 15B.
[0205] The mobile vehicle 1500 may further include a vibration sensor 1542. The vibration sensor 1542 can measure the vibration of components of the mobile vehicle, such as the axle. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 1542 are used, the difference in vibration can be used to determine the friction or slipperiness of the road surface (for example, when the difference in vibration is between a power-driven axle and a freely rotating axle).
[0206] The moving vehicle 1500 may include an ADAS system 1538. In some examples, the ADAS system 1538 may include a SoC. The ADAS system 1538 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.
[0207] The ACC system may use a RADAR sensor 1560, a LIDAR sensor 1564, and / or a camera. The ACC system may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the moving vehicle immediately in front of the moving vehicle 1500 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance holding and advises the moving vehicle 1500 to change lanes when necessary. Lateral ACC is related to other ADAS applications such as LCA and CWS.
[0208] CACC can use information from other moving vehicles that can be received via a wireless link from other moving vehicles via network interface 1524 and / or wireless antenna 1526, or indirectly via a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the immediately preceding moving vehicle (e.g., the moving vehicle immediately in front of moving vehicle 1500 in the same lane as moving vehicle 1500), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of an I2V information source and a V2V information source. Given information about the moving vehicle ahead of moving vehicle 1500, CACC can be made more reliable and has the potential to make traffic flow smoother and reduce road congestion.
[0209] The FCW system is designed to warn the driver of a hazard so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or RADAR sensor 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide an alert in the form of an acoustic, visual alert, vibration, and / or quick brake pulse.
[0210] The AEB system can detect an imminent frontal collision with another moving vehicle or other object and automatically apply the brakes if the driver does not take corrective action within the specified time or distance parameters. The AEB system can use a forward-facing camera and / or RADAR sensor 1560 connected to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a danger, it usually first warns the driver to take corrective action to avoid the collision. If the driver does not take corrective action, the AEB system can automatically apply the brakes as part of an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system may include techniques such as dynamic brake support and / or imminent collision braking.
[0211] The LDW system provides visual, audible, and / or tactile warnings, such as vibrations of the steering wheel or seat, to warn the driver when the moving vehicle 1500 crosses the lane dividing line. The LDW system does not activate when the driver indicates an intentional lane departure by activating the turn indicator. The LDW system can use a forward-facing camera connected to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to driver feedback such as a display, speaker, and / or vibrating component.
[0212] The LKA system is a modified form of the LDW system. The LKA system provides steering input or brakes to correct the moving vehicle 1500 if the moving vehicle 1500 begins to veer out of the lane.
[0213] The BSW system detects and warns the driver of a moving vehicle in the blind spot of a motor vehicle. The BSW system can provide visual, audible, and / or tactile warnings to indicate that a merge or lane change is not safe. The system can provide additional warnings when the driver uses the turn indicator. The BSW system can use a rear-facing camera and / or RADAR sensor 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0214] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 1500 is backing up. Some RCTW systems include AEB to ensure that vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0215] Conventional ADAS systems enable warning the driver and allowing the driver to determine whether a safety condition actually exists and act accordingly. Thus, conventional ADAS systems, while not usually catastrophic, tend to produce false judgment results that can trouble and distract the driver. However, in the autonomous vehicle 1500, when the results conflict, the moving vehicle 1500 itself must decide whether to accept the results from the primary computer or the secondary computer (e.g., the first controller 1536 or the second controller 1536). For example, in some embodiments, the ADAS system 1538 may be a backup and / or secondary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor can execute diverse software that is redundant in hardware components to detect malfunctions in perception and dynamic driving tasks. The output from the ADAS system 1538 can be provided to the supervisory MCU. When the outputs from the primary computer and the secondary computer conflict, the supervisory MCU needs to determine how to reconcile the conflict to ensure safe operation.
[0216] In some examples, the primary computer can be configured to provide a reliability score to the supervisory MCU that indicates the reliability of the primary computer in the selected result. If the reliability score exceeds a threshold, the supervisory MCU can follow the instructions of the primary computer regardless of whether the secondary computer gives conflicting or inconsistent results. If the reliability score does not meet the threshold and the primary and secondary computers indicate different (e.g., conflicting) results, the supervisory MCU can mediate between the computers to determine an appropriate result.
[0217] The supervisory MCU may be configured to execute a neural network trained and configured to determine, based on the outputs from the primary computer and the secondary computer, a state in which the secondary computer provides a false alarm. Thus, the neural network within the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot be trusted. For example, when the secondary computer is a RADAR-based FCW system, the neural network within the supervisory MCU can learn when the FCW is identifying metallic objects such as manhole covers or gratings in a drain that are not actually dangerous but trigger an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network within the supervisory MCU can learn to ignore the LDW when a person on a bicycle or a pedestrian is present and lane departure is actually the safest operation. In embodiments including a neural network executing on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for executing the neural network with associated memory. In a preferred embodiment, the supervisory MCU may comprise components of the SoC1504 and / or be included as components of the SoC1504.
[0218] In other examples, the ADAS system 1538 may include a secondary computer that executes ADAS functions using conventional rules of computer vision. As such, the secondary computer can use classical computer vision rules (if-then), and the presence of a neural network within the supervisory MCU can improve reliability, safety, and performance. For example, the various implementation forms and intentional non-identities make the overall system more fault-tolerant, especially against faults caused by software (or software-hardware interface) functions. For example, if there is a software bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a critical error.
[0219] In some examples, the output of the ADAS system 1538 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 1538 indicates a forward collision warning due to an object immediately ahead, the perception block can use this information when identifying the object. In other examples, the secondary computer may have its own neural network that is trained as described herein and thus reduces the risk of misjudgment.
[0220] The mobile vehicle 1500 may further include an infotainment SoC 1530 (e.g., an in-vehicle infotainment system (IVI) in the mobile vehicle). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more individual components. The infotainment SoC 1530 may include a combination of hardware and software used to provide the mobile vehicle 1500 with audio (e.g., music, mobile devices, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connection (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, wireless data systems, fuel level, total mileage, brake fuel level, oil level, opening / closing doors, air filter information, etc., mobile vehicle-related information). For example, the infotainment SoC 1530 may include radio, disk player, navigation system, video player, USB and Bluetooth (registered trademark) connections, car computer, in-vehicle entertainment, Wi-Fi, steering wheel audio control device, hands-free voice control, heads-up display (HUD), HMI display 1534, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 1530 may be further used to provide information (e.g., visual and / or audible) to the user of the mobile vehicle, such as information from the ADAS system 1538, autonomous driving information such as planned mobile vehicle operations, trajectories, surrounding environment information (e.g., intersection information, mobile vehicle information, road information, etc.), and / or other information.
[0221] The infotainment SoC 1530 may include GPU functionality. The infotainment SoC 1530 can communicate with other devices, systems, and / or components of the moving vehicle 1500 via a bus 1502 (e.g., CAN bus, Ethernet®, etc.). In some examples, the infotainment system's GPU can execute some self-driving functions in the event that the primary controller 1536 (e.g., the primary and / or backup computer of the moving vehicle 1500) fails, and thus the infotainment SoC 1530 can be coupled to the supervisory MCU. In such examples, the infotainment SoC 1530 can put the moving vehicle 1500 into the chauffeur's safe stop mode as described herein.
[0222] The moving vehicle 1500 may further include an instrument cluster 1532 (e.g., digital dash, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 1532 may include a controller and / or a supercomputer (e.g., an individual controller or supercomputer). The instrument cluster 1532 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, gear shift position indicator, seat belt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting control device, safety system control device, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1530 and the instrument cluster 1532. In other words, the instrument cluster 1532 may be included as part of the infotainment SoC 1530, and vice versa.
[0223] FIG. 15D is a system diagram of communication between a cloud-based server of FIG. 15A and an exemplary autonomous vehicle 1500, according to some embodiments of the present disclosure. System 1576 may include a server 1578, a network 1590, and a moving vehicle including the moving vehicle 1500. Server 1578 may include a plurality of GPUs 1584(A)-1584(H) (collectively referred to herein as GPU 1584), PCIe switches 1582(A)-1582(H) (collectively referred to herein as PCIe switch 1582), and / or CPUs 1580(A)-1580(B) (collectively referred to herein as CPU 1580). The GPUs 1584, CPUs 1580, and PCIe switches may be interconnected by high-speed interconnects, such as, but not limited to, an NVLink interface 1588 and / or a PCIe connection 1586 developed by NVIDIA, for example. In some examples, the GPUs 1584 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 1584 and the PCIe switches 1582 are connected via a PCIe interconnect. Eight GPUs 1584, two CPUs 1580, and two PCIe switches are shown, but this is not intended to be limiting. Depending on the embodiment, each server 1578 may include any number of GPUs 1584, CPUs 1580, and / or PCIe switches. For example, the server 1578 may include eight, sixteen, thirty-two, and / or more GPUs 1584, respectively.
[0224] Server 1578 can receive, via network 1590, from a moving vehicle, image data representing an image indicating an unexpected or changed road condition, such as recently started road construction. Server 1578 can transmit, via network 1590 to the moving vehicle, neural network 1592, an updated neural network 1592, and / or map information 1594 including information regarding traffic and road conditions. The update of map information 1594 may include an update of HD map 1522, such as information regarding a construction site, a depression, a detour, a flood, and / or other obstacles. In some examples, neural network 1592, the updated neural network 1592, and / or map information 1594 may result from new training and / or experience represented in data received from any number of moving vehicles in the environment and / or based on training performed in a data center (e.g., using server 1578 and / or other servers).
[0225] Server 1578 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by a moving vehicle and / or (e.g., using a game engine) generated in a simulation. In some examples, the training data is tagged (e.g., when the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., when the neural network does not require supervised learning). The training can be performed according to any one or more classes of machine learning techniques, including but not limited to, for example, the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, associative learning, transfer learning, feature learning (including principal component and cluster analysis), multi-linear subspace learning, manifold learning, representation learning (including pre-dictionary learning), rule-based machine learning, anomaly detection, and their variants or combinations. After the machine learning model is trained, the machine learning model can be used by the moving vehicle (e.g., transmitted to the moving vehicle via network 1590), and / or the machine learning model can be used by server 1578 to remotely monitor the moving vehicle.
[0226] In some examples, server 1578 can receive data from a moving vehicle and apply the data to an up-to-date real-time neural network for real-time intelligent inference. Server 1578 can include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 1584, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1578 can include a deep learning infrastructure that uses only CPU-powered data centers.
[0227] The deep learning infrastructure of server 1578 can have the ability of high-speed real-time inference and use that ability to evaluate and verify the condition of the processors, software, and / or related hardware within moving vehicle 1500. For example, the deep learning infrastructure can receive periodic updates from moving vehicle 1500, such as sequences of images and / or objects in which moving vehicle 1500 is located within the sequence of images (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can execute its own neural network to identify objects and compare them with those identified by moving vehicle 1500. If the results do not match and the infrastructure concludes that the AI within moving vehicle 1500 is not functioning properly, server 1578 can send a signal to moving vehicle 1500 that infers control, notifies the passengers, and commands the fail-safe computer of moving vehicle 1500 to complete a safe parking operation.
[0228] For inference, server 1578 can include GPU 1584 and one or more programmable inference acceleration devices (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time responsiveness. In other examples, such as when less performance is required, servers powered by CPUs, FPGAs, and other processors can be used for inference.
[0229] Exemplary computing device FIG. 16 is a block diagram of an example of a computing device 1600 suitable for use in the implementation of some embodiments of the present disclosure. The computing device 1600 may include an interconnect system 1602 that indirectly or directly connects the following devices: a memory 1604, one or more central processing units (CPUs) 1606, one or more graphics processing units (GPUs) 1608, a communication interface 1610, an input / output (I / O) port 1612, an input / output component 1614, a power supply 1616, one or more presentation components 1618 (e.g., a display), and one or more logic units 1620. In at least one embodiment, the computing device 1600 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). By way of non-limiting example, one or more of the GPUs 1608 may include one or more vGPUs, one or more of the CPUs 1606 may include one or more vCPUs, and / or one or more of the logic units 1620 may include one or more virtual logic units. As such, the computing device 1600 may include individual components (e.g., all GPUs dedicated to the computing device 1600), virtual components (e.g., a portion of a GPU dedicated to the computing device 1600), or a combination thereof.
[0230] The various blocks of FIG. 16 are shown as being connected via an interconnect system 1602 by lines, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 1618 such as a display device may be considered an I / O component 1614 (e.g., if the display is a touch screen). As another example, the CPU 1606 and / or GPU 1608 may include memory (e.g., memory 1604 may represent a storage device in addition to the memory of the GPU 1608, CPU 1606, and / or other components). In other words, the computing device of FIG. 16 is merely illustrative. Categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types are all intended to be within the scope of the computing device of FIG. 16 and thus are not distinguished.
[0231] The interconnect system 1602 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 1602 may include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 1606 may be directly connected to the memory 1604. Additionally, the CPU 1606 may be directly connected to the GPU 1608. If there are direct or point-to-point connections between components, the interconnect system 1602 may include a PCIe link for making the connections. In these examples, a PCI bus need not be included in the computing device 1600.
[0232] The memory 1604 may include any of a variety of computer-readable media. A computer-readable media may be any available media that can be accessed by the computing device 1600. Computer-readable media may include both volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, computer-readable media may comprise computer storage media and communication media.
[0233] A computer storage medium can include both volatile and nonvolatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules, and / or other data types. For example, memory 1604 can store computer readable instructions (such as representing a program and / or program elements) such as an operating system. A computer storage medium can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 1600. In this specification, a computer storage medium does not include a signal per se.
[0234] A computer storage medium can implement computer readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristic sets or changes in such a manner as to encode information in the signal. By way of example, and not limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Any combination of the foregoing should also be included within the scope of computer readable media.
[0235] The CPU 1606 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1600 to execute one or more of the methods and / or processes described herein. The CPU 1606 may each include one or more (e.g., 1, 2, 4, 8, 28, 72, etc.) cores having the ability to process multiple software threads simultaneously. The CPU 1606 may include any type of processor and, depending on the type of computing device 1600 implemented, may include different types of processors (e.g., a processor having fewer cores for a mobile device and a processor having more cores for a server). For example, depending on the type of computing device 1600, the processor may be an Advanced RISC Machines (ARM) processor implemented using reduced instruction set computing (RISC), or an x86 processor implemented using complex instruction set computing (CISC). The computing device 1600 may include one or more CPUs 1606 within one or more microprocessors or auxiliary coprocessors, such as a computing coprocessor.
[0236] In addition to or instead of CPU 1606, GPU 1608 may be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 1600 to execute one or more of the methods and / or processes described herein. One or more of GPU 1608 may be an integrated GPU (e.g., it may be one or more of CPU 1606, and / or one or more of GPU 1608 may be a discrete GPU. In an embodiment, one or more of GPU 1608 may be one or more coprocessors of CPU 1606. GPU 1608 may be used by computing device 1600 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU 1608 may be used for general-purpose computing on GPU (GPGPU). GPU 1608 may include hundreds or thousands of cores having the ability to process hundreds or thousands of software threads simultaneously. GPU 1608 can generate pixel data for an output image in response to rendering commands (e.g., rendering commands from CPU 1606 received via a host interface). GPU 1608 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory may be included as part of memory 1604. GPU 1608 may include two or more GPUs operating in parallel (e.g., via a link). The link can be directly connected to the GPUs (e.g., using NVLINK), or the GPUs can be connected via a switch (e.g., using NVSwitch). When coupled together, each GPU 1608 can generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.
[0237] In addition to and / or instead of CPU 1606 and / or GPU 1608, logic unit 1620 may be configured to execute at least some of the computer-readable instructions to control one or more of computing devices 1600 to execute one or more of the methods and / or processes described herein. In an example, CPU 1606, GPU 1608, and / or logic unit 1620 may execute any combination of methods, processes, and / or portions thereof discretely or in parallel. One or more of logic units 1620 may be part of and / or integrated with one or more of CPU 1606 and / or GPU 1608, and / or one or more of logic units 1620 may be discrete components with respect to and / or otherwise external to CPU 1606 and / or GPU 1608. In an example, one or more of logic units 1620 may be one or more coprocessors of one or more of CPU 1606 and / or one or more of GPU 1608.
[0238] Examples of the logic unit 1620 include one or more processing cores and / or their components, such as tensor cores (TC), tensor processing units (TPU), pixel visual cores (PVC), vision processing units (VPU), graphics processing clusters (GPC), texture processing clusters (TPC), streaming multiprocessors (SM), tree traversal units (TTU), artificial intelligence accelerators (AIA), deep learning accelerators (DLA), arithmetic logic units (ALU), application specific integrated circuits (ASIC), floating point units (FPU), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like.
[0239] The communication interface 1610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1600 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 1610 may include components and functions for enabling communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth®, Bluetooth® LE, ZigBee, etc.), wired networks (e.g., communicating via Ethernet® or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0240] The I / O port 1612 can enable the computing device 1600 to be logically connected to other devices, including some of which can be built in (e.g., integrated) into the computing device 1600, such as I / O components 1614, presentation components 1618, and / or other components. Exemplary I / O components 1614 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, etc. The I / O components 1614 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be sent to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, face recognition, biometric recognition, gesture recognition on and adjacent to the screen, air gestures, head and gaze tracking, and touch recognition related to the display of the computing device 1600 (as will be described in more detail later). The computing device 1600 can include a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. Additionally, the computing device 1600 can include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertia measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 1600 to render immersive augmented reality or virtual reality.
[0241] The power supply device 1616 can include a hard-wired power supply device, a battery power supply device, or a combination thereof. The power supply device 1616 can provide power to the computing device 1600 to enable the components of the computing device 1600 to operate.
[0242] The presentation component 1618 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 1618 can receive data from other components (e.g., the GPU 1608, the CPU 1606, etc.) and output the data (e.g., as an image, a video, an audio, etc.).
[0243] Exemplary data center FIG. 17 shows an exemplary data center 1700 that may be used in at least one embodiment of the present disclosure. The data center 1700 may include a data center infrastructure layer 1710, a framework layer 1720, a software layer 1730, and / or an application layer 1740.
[0244] As shown in FIG. 17, the data center infrastructure layer 1710 may include a resource orchestrator 1712, grouped computing resources 1714, and node computing resources (“node C.R.”) 1716(1) to 1716(N), where “N” represents a natural number of any integer. In at least one embodiment, the node C.R. 1716(1) to 1716(N) may include any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors or graphics processing units (GPU), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and / or cooling modules, etc., but are not limited thereto. In some embodiments, one or more of the node C.R. 1716(1) to 1716(N) may correspond to a server having one or more of the aforementioned computing resources. Additionally, in some embodiments, the node C.R. 1716(1) to 17161(N) may include one or more virtual components, such as vGPU, vCPU, and / or the like, and / or one or more of the node C.R. 1716(1) to 1716(N) may correspond to a virtual machine (VM).
[0245] In at least one embodiment, the grouped computing resources 1714 can include a separate group of node C.R.s 1716 housed within one or more racks (not shown), or multiple racks housed in data centers at various geographical locations (also not shown). The separate group of node C.R.s 1716 within the grouped computing resources 1714 can include grouped computing, network, memory, or storage resources that can be configured or assigned to support one or more workloads. In at least one embodiment, some node C.R.s 1716 that include CPUs, GPUs, and / or other processors can be grouped within one or more racks to provide computing resources for supporting one or more workloads. One or more racks can also include any number of power modules, cooling modules, and / or network switches in any combination.
[0246] The resource orchestrator 1722 can configure or otherwise control one or more node C.R.s 1716(1)-1716(N) and / or the grouped computing resources 1714. In at least one embodiment, the resource orchestrator 1722 can include a software design infrastructure ("SDI") management entity of the data center 1700. The resource orchestrator 1722 can include hardware, software, or some combination thereof.
[0247] In at least one embodiment, as shown in FIG. 17, the framework layer 1720 may include a job scheduler 1732, a configuration manager 1734, a resource manager 1736, and / or a distributed file system 1738. The framework layer 1720 may include a framework to support software 1732 of the software layer 1730 and / or one or more applications 1742 of the application layer 1740. The software 1732 or application 1742 may each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1720 may be of a type of free and open source software web application framework, such as Apache Spark (trademark) (hereinafter “Spark”), that may use the distributed file system 1738 for large-scale data processing (e.g., “big data”), but is not limited thereto. In at least one embodiment, the job scheduler 1732 may include a Spark driver to facilitate scheduling of workloads supported by the various layers of the data center 1700. The configuration manager 1734 may have the ability to configure different layers, such as the software layer 1730 and the framework layer 1720 including Spark and the distributed file system 1738 to support large-scale data processing. The resource manager 1736 may have the ability to manage the mapped or assigned clustered or grouped computing resources for the support of the distributed file system 1738 and the job scheduler 1732. In at least one embodiment, the clustered or grouped computing resources may include the grouped computing resources 1714 in the data center infrastructure layer 1710. The resource manager 1036 may coordinate with the resource orchestrator 1712 to manage these mapped or assigned computing resources.
[0248] In at least one embodiment, the software 1732 included in the software layer 1730 may include at least a portion of nodes C.R. 1716(1) - 1716(N), the grouped computing resources 1714, and / or software used by the distributed file system 1738 of the framework layer 1720. One or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.
[0249] In at least one embodiment, the application 1742 included in the application layer 1740 may include at least a portion of nodes C.R. 1716(1) - 1716(N), the grouped computing resources 1714, and / or one or more types of applications used by the distributed file system 1738 of the framework layer 1720. One or more types of applications may include, but are not limited to, machine learning applications including any number of genomics applications, cognitive computing, and training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0250] In at least one embodiment, any one of the configuration manager 1734, the resource manager 1736, and the resource orchestrator 1712 can implement any number and type of self - rewriting actions based on any amount and type of data obtained in any technically possible manner. Self - rewriting actions can free the data center operator of the data center 1700 by making potentially bad configuration decisions and possibly avoiding under - utilized and / or poorly performing parts of the data center.
[0251] Data center 1700 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters by a neural network architecture that uses the software and / or computing resources described above with respect to data center 1700. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1700 by using weight parameters calculated via one or more training techniques, which are not limited to, for example, those described herein.
[0252] In at least one embodiment, data center 1700 can use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) for training and / or performing inferences using the aforementioned resources. Further, the aforementioned one or more software and / or hardware resources may be configured as a service, such as an image recognition, speech recognition, or other artificial intelligence service, to enable a user to train or perform inferences on information.
[0253] Exemplary network environment A network environment suitable for use in implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented as one or more instances of the computing device 1600 of FIG. 16, and for example, each device may include similar components, features, and / or functionality of the computing device 1600. Additionally, when a backend device (e.g., a server, NAS, etc.) is implemented, the backend device may be included as part of the data center 1700, an example of which is described in further detail herein with respect to FIG. 17.
[0254] The components of the network environment may communicate with each other via a network, which may be wired, wireless, or both. The network may include multiple networks, or a network of networks. By way of example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the Internet and / or the public switched telephone network (PSTN), and / or one or more private networks. When the network includes a wireless telecommunications network, components, such as base stations, communication towers, or access points (and other components), may provide a wireless connection.
[0255] A compatible network environment may include one or more peer-to-peer network environments (in which case, a server may or may not be included in the network environment) and one or more client-server network environments (in which case, one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein with respect to a server may be implemented on any number of client devices.
[0256] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more of the servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework to support software of the software layer and / or one or more applications of the application layer. The software or application may each include web-based service software or an application. In an embodiment, one or more of the client devices may use web-based service software or an application (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be of a type of free and open-source software web application framework that may use a distributed file system for large-scale data processing (e.g., "big data"), but is not limited thereto.
[0257] A cloud-based network environment may provide cloud computing and / or cloud storage that implements any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions may be distributed to multiple locations from a central or core server (such as one or more data centers that may be distributed across a state, region, country, the world, etc.). When the connection to a user (such as a client device) is relatively close to an edge server, the core server may delegate at least a portion of the functionality to the edge server. The cloud-based network environment may be private (such as restricted to a single organization), public (such as available to multiple organizations), and / or a combination thereof (such as a hybrid cloud environment).
[0258] The client device may include at least some of the components, features, and functionality of the exemplary computing device 1600 described herein with respect to FIG. 16. By way of example, and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smart watch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.
[0259] The present disclosure may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer or other machine, such as a mobile information terminal or other handheld device. Generally, program modules include routines, programs, objects, components, data structures, etc., and refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be implemented in a variety of configurations including handheld devices, household appliances, general purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.
[0260] As used herein, a description of "and / or" with respect to two or more elements should be construed to mean either only one element, or a combination of elements. For example, "element A, element B, and / or element C" can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Additionally, "at least one of element A or element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0261] The subject matter of this disclosure has been described with specificity in order to meet statutory requirements. However, the description itself is not intended to limit the scope of the disclosure. Rather, the inventors intend that the claimed subject matter be implemented in other forms, including in combination with other current or future technologies, in different steps or combinations of steps similar to those described in this document. Further, the terms "step" and / or "block" may be used herein to imply different elements of a method being used, but these terms should not be construed as implying any particular order among the various steps disclosed herein unless the order of individual steps is explicitly recited and when so recited.
[0262] Exemplary clause In one or more embodiments, the method includes generating sensor data using one or more vehicle sensors for at least two of a plurality of frames corresponding to operation of a vehicle, calculating an output indicative of a position of a landmark using one or more neural networks (NN) and at least partially based on the sensor data, converting the position to a three-dimensional (3D) world space position relative to an origin of the vehicle, encoding the 3D world space position to generate encoded data corresponding to the frame, and transmitting the encoded data to a server to cause the server to generate a map including the landmark.
[0263] In one or more embodiments, the sensor data includes image data representing an image, one or more of the outputs are calculated in a two-dimensional (2D) image space, and converting the position includes converting a 2D image space position to a 3D world space position at least partially based on at least one of intrinsic camera parameters or extrinsic camera parameters corresponding to a camera that generated the image data.
[0264] In one or more embodiments, a landmark includes one or more of a lane divider; a road boundary, a sign, a pole, a waiting condition, a vertical structure, another road user, a static object, or a dynamic object.
[0265] In one or more embodiments, the output further represents at least one of a pose of the landmark or a geometry of the landmark.
[0266] In one or more embodiments, the output further represents semantic information corresponding to the landmark.
[0267] In one or more embodiments, the landmark includes a lane divider, and the method further includes combining detections of at least two lane dividers to generate a continuous lane divider representation for the at least two lane dividers.
[0268] In one or more embodiments, the method further includes determining, for each of at least two frames, a translation and a rotation of the vehicle with respect to a previous frame of the frames, based at least in part on sensor data, wherein encoding further includes encoding the rotation and the translation.
[0269] In one or more embodiments, the method further includes compressing data representing the rotation and the translation using a delta compression algorithm to generate delta-compressed data, wherein encoding the rotation and the translation further includes encoding the delta-compressed data.
[0270] In one or more embodiments, the sensor data includes global navigation satellite system (GNSS) data, and the method further includes determining, for each of at least two frames, a global position of the vehicle based at least in part on the GNSS data, wherein encoding includes further encoding the global position.
[0271] In one or more embodiments, the sensor data includes LiDAR data, and the method further includes removing dynamic objects from the LiDAR data with a filter to generate filtered LiDAR data, where the encoded data further represents the filtered LiDAR data.
[0272] In one or more embodiments, the encoded data further represents sensor data, and the sensor data includes one or more of LiDAR data, RADAR data, ultrasonic data, GNSS data, image data, or inertial measurement unit (IMU) data.
[0273] In one or more embodiments, the method further includes receiving, from a server, data representing a request for a landmark's position, where transmitting the encoded data is at least partially based on the request.
[0274] In one or more embodiments, the method further includes determining to generate each of at least two frames based at least in part on one or more of a time threshold being met or a distance threshold being met.
[0275] In one or more embodiments, the encoded data is encoded as structured data serialized using a protocol buffer.
[0276] In one or more embodiments, the server includes at least one of a data center server, a cloud server, or an edge server.
[0277] In one or more embodiments, the method includes, for at least two frames out of a plurality of frames corresponding to the driving of a vehicle, using one or more sensors of the vehicle to generate one or more of LiDAR data or RADAR data; removing points corresponding to one or more of the LiDAR data or RADAR data with a filter to generate filtered data; determining the pose of the vehicle with respect to a previous frame; encoding the filtered data and the pose to generate encoded data; and transmitting the encoded data to a cloud server to cause the cloud server to generate at least one of a LiDAR layer of a map or a RADAR layer of the map using the pose.
[0278] In one or more embodiments, removing with a filter includes at least one of dynamic object filtering or duplicate elimination.
[0279] In one or more embodiments, the method further includes using one or more sensors of the vehicle to generate sensor data, and calculating an output indicating the position of a landmark using one or more neural networks (NN) and at least partially based on the sensor data, wherein transmitting further includes transmitting data representing the position of the landmark.
[0280] In one or more embodiments, encoding the filtered data includes encoding the filtered data using an octree.
[0281] In one or more embodiments, the method further includes compressing the filtered data using quantization.
[0282] In one or more embodiments, the system includes one or more sensors, one or more processors, and one or more memory devices storing instructions that, when executed by the one or more processors, cause the one or more processors to use the one or more sensors to generate sensor data in a current frame, calculate an output indicating a landmark's position using one or more neural networks (NN) and based at least in part on the sensor data, convert the position to a three-dimensional (3D) world space position relative to an origin of the vehicle, determine a pose of the current frame relative to a previous pose of a previous frame, encode the 3D world space position and the pose to generate encoded data corresponding to the current frame, and transmit the encoded data.
[0283] In one or more embodiments, the system is included in at least one of a control system of an autonomous or semi-autonomous machine, a perception system of an autonomous or semi-autonomous machine, a system for performing simulation operations, a system for performing deep learning operations, a system implemented using edge devices, a system implemented using robots, a system incorporating one or more virtual machines (VMs), a system implemented at least in part in a data center, or a system implemented at least in part using cloud computing resources.
[0284] In one or more embodiments, the method includes calculating, using one or more neural networks (NN) and at least in part based on sensor data generated by one or more sensors of a vehicle, an output indicating a position in a two-dimensional (2D) image space corresponding to a detected landmark; generating a distance function representation of the detected landmark based at least in part on the position; for at least two of a plurality of poses of the vehicle represented in a cost space, projecting map landmarks corresponding to a map into the 2D image space to generate projected map landmarks, comparing the projected map landmarks with the distance function representation, calculating a cost based at least in part on the comparison, and generating a cost space by updating points in the cost space corresponding to each of the at least two poses based at least in part on the cost; and localizing the vehicle with respect to the map based at least in part on the cost space.
[0285] In one or more embodiments, localizing comprises localizing to an origin of one of a plurality of road segments of a map, and the method further includes localizing the vehicle in a global coordinate system based at least in part on localizing the vehicle to the origin of the road segment.
[0286] In one or more embodiments, the cost space corresponds to a current frame, and the method further includes generating a plurality of additional cost spaces corresponding to a plurality of previous frames, generating an aggregated cost space corresponding to the cost space and the plurality of additional cost spaces, generating the aggregated cost space including using ego-motion compensation for the current frame, and applying a filter to the aggregated cost space to determine a position representation of the vehicle for the current frame, wherein localizing further comprises being based at least in part on the position representation.
[0287] In one or more embodiments, the filter includes a Kalman filter.
[0288] In one or more embodiments, the position representation includes an ellipsoid.
[0289] In one or more embodiments, each of the plurality of cost spaces is generated based at least in part on respective outputs calculated based at least in part on respective sensor data for each of the plurality of frames corresponding to the plurality of frames.
[0290] In one or more embodiments, one or more of the plurality of additional cost spaces are generated to correspond to a first road segment, and generating an aggregated cost space includes transforming one or more of the additional cost spaces to correspond to a second road segment corresponding to the cost space of the current frame.
[0291] In one or more embodiments, transforming includes a translation transformation and a rotation transformation.
[0292] In one or more embodiments, localizing corresponds to a localization mode, and the method further includes localizing the vehicle using one or more additional modes of localization and fusing the localization mode with one or more additional modes of localization to determine a final localization result.
[0293] In one or more embodiments, one or more additional modes of localization include a LiDAR mode, a RADAR mode, or a fusion mode.
[0294] In one or more embodiments, the localization mode corresponds to a camera mode.
[0295] In one or more embodiments, fusing includes applying a Kalman filter to the localization results from the localization mode and the additional modes of localization.
[0296] In one or more embodiments, one or more of the map landmarks or detected landmarks correspond to at least one of lane dividers, road boundaries, signs, poles, vertical structures, waiting conditions, static objects, or dynamic actors.
[0297] In one or more embodiments, generating the cost space further includes, for at least two of a plurality of poses of a vehicle represented in the cost space, projecting a map landmark corresponding to the map into another 2D image space to generate an additional projected map landmark, comparing map semantic information of the additional projected map landmark with detected semantic information of the detected landmark, and calculating another cost based at least in part on comparing the map semantic information with the detected semantic information.
[0298] In one or more embodiments, the method includes generating LiDAR data using one or more LiDAR sensors of a vehicle, projecting, for at least two of a plurality of poses of the vehicle represented in the cost space, points corresponding to the LiDAR data into a distance function representation of a LiDAR point cloud corresponding to a LiDAR layer of the map, comparing the points with the distance function representation, calculating a cost based at least in part on the comparison, and updating points in the cost space corresponding to each of the at least two poses based at least in part on the cost to generate the cost space, and localizing the vehicle on the map based at least in part on the cost space.
[0299] In one or more embodiments, the method further includes determining a point from a larger set of points represented by the LiDAR data based at least in part on an elevation value corresponding to the point being within an elevation range, where the LiDAR point cloud corresponds to the elevation range.
[0300] In one or more embodiments, generating a cost space further includes, for each of at least two poses of a vehicle represented in the cost space, projecting an elevation representation corresponding to LiDAR data to a map elevation representation corresponding to the LiDAR layer of the map, comparing the elevation representation to the map elevation representation, and calculating another cost based at least in part on comparing the elevation representation to the map elevation representation.
[0301] In one or more embodiments, comparing the elevation representation to the map elevation representation includes adjusting at least one of the elevation representation or the map elevation representation based at least in part on a difference between the map elevation representation and an origin of a vehicle origin or a road segment of the corresponding vehicle.
[0302] In one or more embodiments, generating a cost space further includes, for each of at least two poses of a vehicle represented by the cost space, projecting an intensity representation corresponding to LiDAR data to a map intensity representation corresponding to the LiDAR layer of the map, comparing the intensity representation to the map intensity representation, and calculating another cost based at least in part on comparing the intensity representation to the map intensity representation.
[0303] In one or more embodiments, localizing is to an origin of one of a plurality of road segments of a map, and the method further includes localizing the vehicle in a global coordinate system based at least in part on localizing the vehicle to the origin of the road segment.
[0304] In one or more embodiments, the cost space corresponds to the current frame, and the method further includes generating a plurality of additional cost spaces corresponding to a plurality of previous frames, generating an aggregated cost space corresponding to the cost space and the plurality of additional cost spaces, generating an aggregated cost space including using self-motion compensation for the current frame, and applying a filter to the aggregated cost space to determine a position representation of the vehicle in the current frame, where localizing is further at least partially based on the position representation.
[0305] In one or more embodiments, the filter includes a Kalman filter.
[0306] In one or more embodiments, the position representation includes an ellipsoid.
[0307] In one or more embodiments, one or more of the plurality of additional cost spaces are generated to correspond to a first road segment, and generating the aggregated cost space includes transforming one or more of the additional cost spaces to correspond to a second road segment corresponding to the cost space of the current frame.
[0308] In one or more embodiments, localizing corresponds to a localization mode and the method further includes localizing the vehicle using one or more additional modes of localization, and fusing the localization mode with one or more additional modes of localization to determine a final localization result.
[0309] In one or more embodiments, the localization mode includes a LiDAR mode, and one or more additional modes of localization include a RADAR mode, an image mode, or a fusion mode.
[0310] In one or more embodiments, the system includes one or more sensors, one or more processors, and one or more memory devices storing instructions that, when executed by the one or more processors, cause the one or more processors to use the one or more sensors to generate sensor data, project points corresponding to the sensor data into a distance function representation of a point cloud corresponding to a map for at least two poses out of a plurality of poses represented in a cost space, compare the points to the distance function representation, calculate a cost based at least in part on the comparison, and generate a cost space by updating points in the cost space corresponding to each of the at least two poses based at least in part on the cost, and localize a vehicle on the map based at least in part on the cost space.
[0311] In one or more embodiments, the system is included in at least one of a control system of an autonomous or semi-autonomous machine, a perception system of an autonomous or semi-autonomous machine; a system for performing simulation operations, a system for performing deep learning operations, a system implemented using an edge device, a system incorporating one or more virtual machines (VMs), a system implemented using a robot, a system implemented at least in part in a data center, or a system implemented at least in part using cloud computing resources.
[0312] In one or more embodiments, localizing is to an origin of one of a plurality of road segments of a map, and the method further includes localizing the vehicle in a global coordinate system based at least in part on localizing the vehicle to the origin of the road segment.
[0313] In one or more embodiments, the cost space corresponds to the current frame, and the method further includes generating a plurality of additional cost spaces corresponding to a plurality of previous frames, generating an aggregated cost space corresponding to the cost space and the plurality of additional cost spaces, generating an aggregated cost space that includes using ego-motion compensation for the current frame, applying a filter to the aggregated cost space to determine a position representation of the vehicle in the current frame, where localizing is further based at least in part on the position representation.
[0314] In one or more embodiments, the sensor data corresponds to one or more of LiDAR data or RADAR data, and the one or more sensors include one or more of a LiDAR sensor or a RADAR sensor.
Claims
1. receiving first sensor data generated using one or more first sensors of a machine and second sensor data generated using one or more second sensors of the machine, the first sensor data being associated with a first sensor modality and the second sensor data being associated with a second sensor modality; localizing the machine based at least on a first comparison of the first sensor data to first map data associated with the first sensor modality and a second comparison of the second sensor data to second map data associated with the second sensor modality; causing the machine to perform one or more navigation operations based at least on the localization; A method comprising:
2. analyzing one or more first points represented by the first sensor data with respect to one or more second points associated with the first map data; analyzing one or more third points represented by the second sensor data with respect to one or more fourth points associated with the second map data; Further comprising:
2. The method of claim 1, wherein the localizing the machine is further based at least on an analysis of the one or more first points with respect to the one or more second points and an analysis of the one or more third points with respect to the one or more fourth points.
3. determining one or more first distances between one or more first points represented by the first sensor data and one or more second points associated with the first map data; determining one or more second distances between one or more third points represented by the second sensor data and one or more fourth points associated with the second map data; Further comprising: The method of claim 1 , wherein the localizing of the machine is further performed based at least on the one or more first distances and the one or more second distances.
4. determining one or more first positions of one or more first landmarks using the first sensor data; determining one or more second locations of the one or more first landmarks within an environment based at least on the first map data; determining one or more first positions of one or more second landmarks using the second sensor data; determining one or more second locations of the one or more second landmarks within the environment based at least on the second map data; Further comprising:
2. The method of claim 1, wherein the localizing of the machine is based at least on the one or more first positions of the one or more first landmarks, the one or more second positions of the one or more first landmarks, the one or more first positions of the one or more second landmarks, and the one or more second positions of the one or more second landmarks.
5. determining a first cost based at least on the first sensor data and the first map data; determining a second cost based at least on the second sensor data and the second map data; Further comprising: The method of claim 1 , wherein the localizing the machine is further based at least on the first cost and the second cost.
6. the first cost and the second cost are associated with a first pose of the machine; The method further comprises: determining a third cost associated with a second pose of the machine based at least on the first sensor data and the first map data; determining a fourth cost associated with the second pose of the machine based at least on the second sensor data and the second map data; and The method of claim 5 , wherein the localizing of the machine is further performed based at least on the third cost and the fourth cost.
7. The localizing the machine includes: determining a first final cost associated with the first pose of the machine based at least on the first cost and the second cost; determining a second final cost associated with the second pose of the machine based at least on the third cost and the fourth cost; determining a location within the environment associated with the first pose based at least on the first final cost and the second final cost; The method of claim 6, comprising:
8. the first map data indicating one or more first locations of one or more first landmarks within an environment determined using third sensor data associated with the first sensor modality; 2. The method of claim 1 , wherein the second map data indicates one or more second locations of one or more second landmarks within the environment determined using fourth sensor data associated with the second sensor modality.
9. 1. A system comprising one or more processors, The one or more processors: receiving first sensor data generated using one or more first sensors of a machine and second sensor data generated using one or more second sensors of the machine, the first sensor data associated with a first sensor modality and the second sensor data associated with a second sensor modality; localizing the machine based at least on a first comparison of the first sensor data to first map data associated with the first sensor modality and a second comparison of the second sensor data to second map data associated with the second sensor modality; The system performs one or more navigation operations on the machine based at least on the localization.
10. the first sensor data is analyzed at least with respect to a first portion of a map, analyzing one or more first points represented by the first sensor data with respect to the one or more second points associated with the first portion of the map; 10. The system of claim 9, wherein the second sensor data is analyzed at least with respect to a second portion of the map, analyzing one or more third points represented by the second sensor data with respect to one or more fourth points associated with the second portion of the map.
11. the first sensor data is analyzed with respect to at least a first portion of the map to determine one or more first distances between one or more first points represented by the first sensor data and one or more second points associated with the first portion of the map; 10. The system of claim 9, wherein the second sensor data is analyzed with respect to at least a second portion of the map to determine one or more second distances between one or more third points represented by the second sensor data and one or more fourth points associated with the second portion of the map.
12. the first sensor data is analyzed with respect to at least a first portion of the map to determine one or more first distances between one or more first positions of one or more first landmarks determined using the first sensor data and one or more second positions of the one or more first landmarks determined using the first portion of the map; 10. The system of claim 9, wherein the second sensor data is analyzed with respect to at least a second portion of the map to determine one or more second distances between one or more first positions of one or more second landmarks determined using the second sensor data and one or more second positions of the one or more second landmarks determined using the second portion of the map.
13. The one or more processors further include: determining a first cost based at least on the analyzed first sensor data for the first portion of the map; determining a second cost based at least on the analyzed second sensor data with respect to a second portion of the map; Do the following:
10. The system of claim 9, wherein the localization of the machine is based at least on the first cost and the second cost.
14. the first cost and the second cost are associated with a first pose of the machine; The one or more processors: determining a third cost based at least on the first sensor data analyzed at least with respect to a first portion of the map, the third cost being associated with a second pose of the machine; determining a fourth cost based at least on the second sensor data analyzed at least with respect to a second portion of the map, the fourth cost being associated with the second pose of the machine; 14. The system of claim 13, wherein the localizing of the machine is further performed based at least on the third cost and the fourth cost.
15. The localizing of the machine includes: determining a first final cost associated with the first pose of the machine based at least on the first cost and the second cost; determining a second final cost associated with the second pose of the machine based at least on the third cost and the fourth cost; determining a position within the environment relative to the first pose based at least on the first final cost and the second final cost. The system of claim 14 , comprising:
16. a first portion of the map illustrating one or more first locations of one or more first landmarks within the environment determined using third sensor data associated with the first sensor modality; 10. The system of claim 9, wherein a second portion of the map illustrates one or more second locations of one or more second landmarks within the environment determined using fourth sensor data associated with the second sensor modality.
17. Control systems for autonomous or semi-autonomous machines; Autonomous or semi-autonomous machine perception systems; A system for performing a simulation operation; A system for performing deep learning operations, A system implemented using edge devices; A system implemented using a robot, A system incorporating one or more virtual machines (VMs); A system implemented at least in part in a data center; or The system of claim 9 , including at least one of the systems that is implemented at least in part using cloud computing resources.
18. 1. A processor comprising one or more processing units configured to cause a machine to perform one or more control operations based at least on a localization of the machine, the localization of the machine being performed based at least on first sensor data associated with a first sensor modality, second sensor data associated with a second sensor modality, a first map associated with the first sensor modality, and a second map associated with the second sensor modality.
19. The one or more processing units further include: determining a first cost based at least on the first sensor data and the first map; determining a second cost based at least on the second sensor data and the second map; 20. The processor of claim 18, wherein the localization of the machine is performed based at least on the first cost and the second cost.
20. The processor, Control systems for autonomous or semi-autonomous machines; Autonomous or semi-autonomous machine perception systems; A system for performing a simulation operation; A system for performing deep learning operations, A system implemented using edge devices; A system implemented using a robot, A system incorporating one or more virtual machines (VMs); A system implemented at least in part in a data center; or 20. The processor of claim 18, comprising at least one of the systems that is implemented at least in part using cloud computing resources.
Citation Information
Patent Citations
Landmark recognizing system
JP2008051612A
Mobile object position measurement device
JP2012127896A
Position estimation device of moving body and position estimation method
JP2019070983A
Information processing device, information processing method, program, and movable body
WO2019082669A1