Semantic segmentation of aggregated sensor data

By using lidar data and machine learning models, the road network maps of autonomous driving vehicles are automatically generated and updated, and the resource consumption and accuracy of map creation and maintenance in the prior art is solved, and efficient and accurate map updates are achieved.

CN120051814APending Publication Date: 2025-05-27ZOOX INC
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202380072396.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-10-12
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In existing autonomous vehicle navigation systems, the initial creation and maintenance of road maps require a large amount of resources and manual input, and the existence of unrelated elements in the sensor data leads to inaccurate segmentation and classification of map elements.

Method used

By using lidar data to generate and update semantic markered road network maps, machine learning models are used to filter, aggregate and classify lidar data, and automatically identify and segment map elements.

Benefits of technology

An automated road map creation and update process is realized, reducing maintenance costs and time, and improving the accuracy of segmentation and classification of map elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051814A_ABST
    Figure CN120051814A_ABST
Patent Text Reader

Abstract

Techniques for segmenting and classifying representations of aggregated sensor data from a scene are discussed herein. Sensor data may be collected during multiple traversal of the same scene, and the sensor data may be filtered to remove portions of the sensor data that are independent of a road network map. In some examples, the filtered data may be aggregated and represented in voxels of a three-dimensional voxel space from which an image representing a top-down view of the scene may be generated, despite other views may be envisaged. Operations may include segmenting and / or classifying the image, e.g., through a trained machine learning model to associate category tags indicative of map elements (e.g., driving lanes, stop lines, turning lanes, etc.) with segments identified in the image. In addition, the techniques may create or update a road network map based on the segmented and semantically labeled image (s) of the various portions of the environment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This PCT international patent application claims the priority of the US patent application with serial number No. 17 / 965,686, filed on October 13, 2022, the entire content of which is incorporated herein by reference. Background Art

[0003] Autonomous vehicles typically use maps of roads to navigate within an environment. The maps can indicate semantic map elements such as roads, lanes, sidewalks, parking spaces, crosswalks, etc. However, the initial creation of such maps and the maintenance of the maps to keep them up - to - date by incorporating changes in the environment may require significant resources and manual input. Additionally, the presence of irrelevant elements in sensor data can lead to inaccurate segmentation and classification of map elements related to the road network map of the environment. Brief Description of the Drawings

[0004] The detailed description is described with reference to the accompanying drawings. In the drawings, the left - most (one or more) digits of the reference numerals identify the drawing in which the reference numeral first appears. The same reference numerals are used in different drawings to indicate similar or identical components or features.

[0005] Figure 1 is a schematic diagram illustrating an example process for generating a road network map with semantic tags according to an example of the present disclosure.

[0006] Figure 2 depicts an example vehicle that captures sensor data of the environment.

[0007] Figure 3 includes text and visual flowcharts to illustrate an example method for using LiDAR data to determine the semantic segmentation of a top - down view of a scene, as described herein.

[0008] Figure 4 is a block diagram illustrating an example computing system for implementing the techniques described herein.

[0009] Figure 5 illustrates an example process for creating or updating a road network map for controlling an autonomous vehicle, as described herein. Detailed Description

[0010] The techniques described herein relate to generating and updating a semantic-tagged road network map based on semantic segmentation of sensor data. In an example, the road network map can include map elements associated with a drivable surface and / or ground plane of an environment, indicating both a semantic tag associated with the map element and a spatial extent of the map element. For example, map elements can include lane elements indicating the extent of a lane or driving corridor, stop line elements indicating a stop line at an intersection, crosswalk elements indicating a crosswalk area, and the like. Additional non-limiting examples of map elements can include bike lane elements, parking space elements, intersection elements, lane divider elements, yield line elements, and the like. As can be appreciated, such a road network map can be used for route planning or generating a trajectory for an autonomous vehicle to traverse the environment.

[0011] In the examples described herein, sensor data associated with the same portion of an environment or scene can be captured by one or more vehicles during multiple traversals of the environment (e.g., along a route within a geofenced area or within an operating area of an autonomous vehicle). For example, sensors mounted on the vehicle(s) can capture sensor data of the same portion of the environment or scene as the vehicle(s) traverse the environment at different times. In an example, the sensor data can include data captured by different vehicles or the same vehicle at different times (e.g., during different trips). The sensor data can be associated with the portion of the environment or scene based on the known location and / or trajectory of the vehicle(s) at the time the sensor data is captured. The present disclosure generally relates to systems configured to receive sensor data representative of an environment and generate a segmented output of semantic tags suitable for creating and updating a road network map. However, other applications of the output are also envisioned (e.g., path planning, obstacle avoidance, and platooning operations in the context of autonomous vehicle operation).

[0012] Although mapping and navigation systems using various sensor modalities (e.g., radar or 4D radar, range cameras, stereo cameras, time-of-flight sensors, etc.) can benefit from the techniques described herein, example systems implementing the techniques described herein can use lidar data as sensor data to interpret the environment. For example, a (one or more) vehicle equipped with one or more lidar sensors can capture lidar data, each lidar sensor configured to generate lidar echoes about a 360-degree field of view around the (one or more) vehicle. Any combination of sensor information can be used. Generally, lidar echoes can include both position information and intensity information. For example, the position information can include distance (e.g., depth or range from the sensor) and azimuth (e.g., angle relative to a reference line or position). In some examples, a scanning lidar sensor can rotate (e.g., about a vertically oriented axis) to scan the entire 360-degree field of view around the sensor, although other forms of lidar are also contemplated (including flash lidar, solid-state lidar, etc.). Lidar data can include a high-resolution point cloud of the environment, including three-dimensional (3D) data points and associated information such as intensity information, signal-to-noise ratio, velocity information, etc.

[0013] In some examples, the techniques described herein can filter lidar data to remove echoes that are not relevant to a road network map. For example, lidar echoes from surfaces that are above a vertical distance threshold from the ground plane can be filtered out, leaving only lidar data that is related to features near or on the ground plane. In some examples, one or more techniques described in U.S. Patent No. 10,444,759 (issued on October 15, 2019, titled "Voxel-Based Ground Plane Estimation and Object Segmentation") and U.S. Patent Application Serial No. 16 / 698,055 (filed on November 27, 2019, titled "Height Estimation Using Sensor Data") can be used to determine the ground plane, the entire contents of both of which are incorporated herein by reference for all purposes. In some cases, the vertical distance can be specified relative to the origin or a virtual origin of the lidar system that collects the lidar data. Additionally, transient elements (e.g., vehicles, pedestrians, construction equipment, etc.) can be identified and removed from the lidar data so that the remaining lidar data mainly corresponds to time-invariant elements in the environment. In some examples, object tracking can be used to identify transient and static objects, using techniques such as those described in U.S. Patent Application Serial No. 16 / 866,865 (filed on May 5, 2020, titled "Object Velocity and / or Yaw Rate Detection and Tracking"), and U.S. Patent Application Serial No. 17 / 364,603 (filed on June 30, 2021, titled "Tracking Objects Using Radar Data"), the entire contents of both of which are incorporated herein by reference for all purposes. In some examples, sensor data associated with a low confidence score can also be removed. For example, lidar data can include confidence scores associated with an individual or a group of data points, the confidence scores indicating the accuracy of the data, the confidence scores can be based on the signal-to-noise ratio (e.g., data points with a low SNR may result in a correspondingly lower confidence score) and / or sensor performance (e.g., environmental conditions such as adverse weather conditions may affect sensor performance).

[0014] In the examples described herein, lidar data captured by one or more vehicles at different times and / or during different trips and associated with the same portion of the environment or scene can be aggregated into a common scene representation. Various techniques can be used to aggregate points in individual instances of lidar data into a common scene representation. For example, aggregated lidar data associated with the same scene can be mapped to cells in a three-dimensional (3D) volumetric element ("voxel") space that represents the physical volume of the scene. In the voxel space, each cell or voxel can be associated with cell data that represents the statistical accumulation of lidar data corresponding to echoes from the physical volume in the scene represented by the cell or voxel. In some examples, the cell data can include an average intensity, number of echoes, average x value, average y value, average z value, and / or covariance matrix based on the lidar data.

[0015] In an example, the techniques described herein can generate a two-dimensional (2D) image representing a top-down perspective or a planar view of the 3D voxel space. Specifically, the 2D image can represent an average intensity value and / or a weighted average intensity value weighted by a plurality of data points. As can be appreciated, lidar echoes corresponding to time-invariant features of the scene are expected to occur in a large proportion of instances in the lidar data, resulting in higher values of the weighted average intensity or average intensity. Accordingly, lidar echoes representing transient elements in the scene are expected to occur in a small proportion of instances in the lidar data, resulting in low intensity values in the 2D image. Thus, the 2D image effectively represents the voxel space as seen from above and emphasizes the time-invariant features of the scene. Additionally, as described above, the aggregated lidar data has been filtered to remove data from surfaces at a vertical distance threshold above the ground plane, and the 2D image generated based on the aggregated lidar data captures the time-invariant features at or near the ground plane that are most relevant to a road network map. Subsequently, such a 2D image can be input into a trained machine learning (ML) model to receive an output containing pixel classification and / or segmentation data associated with the 2D image.

[0016] Embodiments of the present disclosure may maintain a previously trained ML model for classifying pixels of an image depicting a top - down view of a road network. The ML model may be trained using training data that includes images of the top - down view of the road network, where the pixels have been labeled with class labels from a specified set of class labels. The labeling of the training data may be manual, automatically generated and then manually verified, or provided by a third - party map provider. The specified set of class labels may include class labels related to the road network corresponding to the map elements discussed above (e.g., "road lane", "bike lane", "left - turn lane", "sidewalk", "crosswalk", etc.). The specified set of class labels may also include other common elements such as "buildings", "vegetation", "water bodies", etc., as well as an all - encompassing class label for labeling pixels that do not belong to one of the other classes in the set of class labels (e.g., "other" or "background").

[0017] The techniques discussed herein may include providing a 2D image representing a top - down view as an input image to a trained ML model and receiving from the ML model an output indicating the classification associated with an individual "pixel" of the input image, where a "pixel" may also represent a group of adjacent data points. The classes may include a class label associated with the pixel and a confidence score indicating the level of certainty of the class label. As a non - limiting example, the trained ML model may include a fully convolutional network (FCN) that returns an output of the same size as the input image, where the output at coordinates (x, y) includes the (one or more) class labels associated with the pixel at coordinates (x, y) of the input image and a pixel - level confidence score (or class - label probability). Conventional machine - learning classifiers may output a global class label given an input image, while the output produced by the FCN provides the additional advantage of localization of the class label since the class label is associated with an individual pixel at a known (x, y) coordinate within the image.

[0018] In some examples of the systems described herein, a segmented output image can be generated from the output (e.g., as generated by an FCN) by clustering spatially contiguous pixels with the same class label (e.g., by using connected component analysis techniques known in the art). In such examples, a segment-level confidence score can be calculated for each segment based on pixel-level confidence scores obtained as output from an ML model. For example, the segment-level confidence score can be the mean or median of the pixel-level confidence scores of the pixels in the segment. The output image of the segment can indicate the spatial extent of the segment(s) (e.g., using a bounding box for 2D elements or a spline for linear elements), the class label(s) associated with the segment(s), and / or the corresponding segment-level confidence score(s).

[0019] In other examples, a 2D image representing a top-down view can first be segmented using image segmentation techniques known in the art. Subsequently, machine learning techniques that output global class labels (instead of pixel-level class labels provided by an FCN), such as deep convolutional neural networks (CNNs), conditional random fields (CRFs), rule-based systems, probabilistic frameworks, etc., can be applied to the individual segments to receive class labels and confidence scores (or class probabilities) as output. In this example, the segmented image, where the individual segments are associated with class labels and confidence scores, is similar to the output image of segments generated by applying a trained ML model, such as an FCN, and then clustering similarly labeled pixels to generate segments, as described above.

[0020] The techniques discussed herein can improve the functionality of a computing device in a number of ways. For example, in prior art, the initial task of associating class labels corresponding to map elements with three-dimensional sensor data (e.g., lidar data) captured by a vehicle's sensors may need to be performed manually by a cartographer or skilled technician. Prior art may also involve a cartographer or skilled technician manually updating a road network map in response to updated sensor data indicating a change, which may take several hours to several weeks of work, depending on the size of the environment represented by the road network map. The techniques discussed herein can automate such map creation and update operations, thereby significantly reducing the processing and time required to maintain the road network map for operating an autonomous vehicle in an environment.

[0021] As an example, in some applications of the techniques described herein, the segmented output image can be used to create and / or update a road network map in the context of a vehicle (such as an autonomous vehicle). For example, during the initial creation of a road network map (e.g., for a new environment that an autonomous vehicle needs to traverse), sensor data sets collected by one or more vehicles (which may or may not be operating in autonomous mode) can be automatically processed using the techniques described herein to generate a segmented output image having class labels corresponding to map elements, as described above. Also as described above, the segmentation can be associated with a segmentation-level confidence score. In an example, if the segmentation-level confidence score is equal to or greater than a minimum threshold, the class label associated with the segmentation and the spatial extent of the segmentation can be transferred to the newly created road network map. However, if the segmentation-level confidence score is less than the minimum threshold, the segmentation can be flagged for further inspection (e.g., by a human operator) to determine its class label and / or extent. Alternatively or additionally, the segmentation can be compared with data from a third-party map provider to determine a correspondence between the segmentation and map elements in the third-party map data. When a match is determined, the segmentation can be added to the new road network map, and if a mismatch is determined, the segmentation can be flagged for manual verification.

[0022] In another example, the systems described herein can generate the segmented output image based on sensor data sets collected at regular intervals or when a change is expected (e.g., due to known construction work, traffic flow changes, etc.). The system can compare the class labels and spatial extents of the segments of the segmented output image with an existing road network map. In cases where the system determines that there are differences between the class labels and / or spatial extents of one or more segments of the segmented output image and the existing road network map, (one or more) different segments having a segmentation-level confidence score equal to or greater than a minimum threshold can be automatically updated in the road network map. Different segments having a segmentation-level confidence score less than the minimum threshold can be submitted for further inspection (e.g., verified by a human operator). Thus, the road network map can be maintained or kept up-to-date using the techniques described herein.

[0023] The techniques described herein can be implemented in a number of ways. The following is referenced with respect to the following figures ( Figures 1-5) provides example implementation manners. Although discussed in the context of autonomous vehicles, the methods, apparatuses, and systems described herein can be applied to various systems (e.g., sensor systems or robotic platforms) and are not limited to autonomous vehicles. In another example, the techniques can be used in an aviation or nautical context, or in any system that uses maps. Additionally, although discussed in the context of lidar data, the sensor data can include any two-dimensional, three-dimensional, or multi-dimensional data, such as image data (e.g., stereo cameras, time-of-flight data, etc.), radar data, sonar data, etc. Further, the techniques described herein can be used with real data (e.g., data captured using one or more sensors), simulated data (e.g., data generated by a simulator), or any combination of the two.

[0024] Figure 1 is a schematic diagram illustrating an example process 100 for generating a road network map with semantic markings of an environment. As shown in Figure 1 As shown, vehicles 102(1), 102(2), …, 102(N) (collectively referred to as vehicles 102) can traverse the same portion of an environment or scene 104, capturing a set of sensor data 106 (e.g., lidar data) representing the scene 104 using sensors (not shown) mounted on the vehicles. In an example, the scene 104 is associated with a known location (e.g., a point and heading on a map, latitude and longitude with a heading direction, etc.), and the sensor data representing the scene 104 is recorded at a specific moment. For example, the first vehicle 102(1) can traverse the scene 104 at a first time and / or during a first trip, as shown in 104(1), the second vehicle 102(2) can traverse the same scene 104 at a second time and / or during a second trip, as shown in 104(2), and the Nth vehicle 102(N) can traverse the same scene 104 at the Nth time and / or during the Nth trip, as shown in 104(N). In an example, the different times can include different times of day (including at night) and / or during different weather conditions. The vehicles 102 can be the same vehicle (e.g., traversing the scene 104 at different times) or different vehicles (e.g., equipped with similar sensor configurations). Additionally, during the process of traversing the environment to collect the sensor data 106, the vehicles 102 can each operate in an automatic mode or be driven by a human driver.

[0025] As shown in Figure 1As shown in, in addition to time-invariant background features (such as building 112), scene 104 may present many permanent or time-invariant road map elements, such as crosswalk 108 and stop line 110. Other non-limiting examples of time-invariant road map elements may include (one or more) driving lanes, (one or more) bicycle lanes, (one or more) parking spaces, (one or more) turning lanes, (one or more) intersections, (one or more) lane dividers, (one or more) center lines, (one or more) yield lines, etc. Road map elements can generally be characterized by highly reflective markings applied to drivable surfaces. Similarly, other examples of time-invariant background features may also include trees, telephone / utility poles, signs, etc. Scene 104 may also present different transient elements at different times, such as pedestrians 114A and vehicle 116A at 104(1), and pedestrians 114B and vehicle 116B at 104(2). Sensor data 106 collected by vehicle 102 during multiple traversals of scene 104 (e.g., during different trips) includes the time-invariant features (including both road map elements and the background) and transient features of the scene. Of course, although described herein as permanent or time-invariant, such features may change (e.g., by moving their location, construction, growth, etc.), and the techniques described herein can ensure that such changes are reflected in the data set (e.g., map) upon which safe navigation in the environment depends.

[0026] Vehicle 102 may include multiple sensors of different modalities (e.g., radar or 4D radar, range cameras, stereo cameras, time-of-flight sensors, etc.). In an example implementation, vehicle 102 may capture a set of sensor data (e.g., lidar data) 106 using the vehicle's lidar sensor, as further described in reference to Figure 2 in further detail. The set of lidar data 106 may include a high-resolution point cloud of the environment, including three-dimensional (3D) data points and associated information, such as intensity information, signal-to-noise ratio, velocity information, etc. The lidar data 106 may be represented as points in a three-dimensional voxel space, represented as a three-dimensional grid (or other representation, model, etc.). In an implementation, the set of lidar data 106 may represent a 360-degree field of view, where scene 104 may represent only a portion of the entire field of view.

[0027] As in Figure 1As shown in, example process 100 may include a sensor data filter 118 configured to: receive a lidar data set 106, determine portions of the lidar data set 106 relevant to the generation of a road network map, and output a corresponding filtered sensor data set 120. For example, lidar echoes from surfaces above a vertical distance threshold from the ground plane may be filtered out such that the filtered lidar data set 120 includes lidar data related to features near or on the ground plane, as further described in reference Figure 2 As can be appreciated, the filtered lidar data 120 retains the features most relevant to the road network map on or near the ground plane. In some examples, the sensor data filter 118 may also identify transient elements (e.g., vehicles, pedestrians, construction equipment, etc.) in the lidar data set 106 and remove data associated with the identified transient elements such that the filtered lidar data 120 primarily corresponds to time-invariant elements in the environment. Further details of the functionality implemented by the sensor data filter 118 will be discussed with reference to Figure 2 In some examples, the sensor data filter 118 may be implemented on a computing system of the vehicle 102, and the computing system of the vehicle 102 may locally store the filtered sensor data 120 and / or transmit the filtered sensor data 120 to an external computing system for further processing.

[0028] As in Figure 1As shown, example process 100 may also include an aggregated view generator 122. In an example, individual instances of lidar data from the filtered lidar data set 120 may be represented as a set of 3D data points in an associated 3D voxel space. Thus, the filtered lidar data set 120 may be associated with a corresponding set of 3D voxel spaces. The aggregated view generator 122 may implement a function to align with the set of 3D voxel spaces such that the set of 3D data points from individual instances of the filtered lidar data set 120 may be mapped to cells of a common three-dimensional voxel space representing the physical volume of the scene 104. The set of 3D voxel spaces may be aligned by using the position and / or trajectory information of the vehicle(s) 102 at the time the corresponding lidar data was captured. As an example, the aggregated view generator 122 may use the techniques described below to align the set of 3D voxel spaces: U.S. Patent No. 10,983,199, titled "Vehicle Sensor Calibration and Positioning," filed on April 20, 2021, which describes algorithms for identifying previously visited locations and continuously determining the position and / or orientation of a vehicle within a map; and U.S. Patent No. 10,782,136, titled "Modifying Map Elements Related to Map Data," filed on September 22, 2020, which describes algorithms for aligning sensor data collected when traversing the same location(s) while following different trajectories. Both patents are hereby incorporated by reference in their entirety and for all purposes. In some cases, the edges of the individual voxel spaces may be filled (e.g., with null data or baseline data) to align with the common 3D voxel space (e.g., when there is an incomplete overlap between a single voxel space and the common 3D voxel space).

[0029] As can be appreciated, lidar echoes from time-invariant features of the scene 104 are expected to occur in most instances from the filtered lidar data set 120. Additionally, data points associated with the time-invariant features will be mapped to the same voxel(s) of the common 3D voxel space after alignment. In the common 3D voxel space, each cell or voxel may be associated with cell data representing the statistical accumulation of the filtered lidar data set 120, which corresponds to echoes from the physical volume in the scene 104 represented by the cell or voxel. In some examples, the cell data may include the average intensity, number of echoes, average x value, average y value, average z value, and / or covariance matrix based on the filtered lidar data set 120. In various examples, such data may also include the distribution of semantic classes and / or their relative confidence. Refer to Figure 3 The common 3D voxel space is described in further detail.

[0030] The aggregated view generator 122 may also include functionality to generate a two-dimensional (2D) image 124 that represents a top-down view of a common 3D voxel space filled with data points from the filtered lidar data set 120. For example, the pixels of the 2D image 124 may represent cumulative intensity values, maximum intensity values, or average intensity values in the projection of each column of voxels onto a two-dimensional plane (e.g., the ground plane), as illustrated in the example image 126, which may correspond to a portion of the environment including the scene 104. In some examples, the 2D image 124 may represent a weighted average intensity value (e.g., weighted by the total number of data points in the voxels of the corresponding column). As described, voxels that contain time-invariant features of the scene 104 may accumulate data points from multiple individual instances of the filtered lidar data set 120 and thus produce a higher weighted average intensity value or average intensity value, as illustrated in the example image 126. However, intensity values from transient elements of the scene 104 are less likely to recur at consistent positions or voxels in the common 3D voxel space, resulting in lower average or weighted average intensity values. Thus, the 2D image 124 of the common 3D voxel space as seen from above (e.g., in a top-down view) effectively emphasizes the time-invariant features of the scene 104. Additionally, road map elements are typically characterized by high reflectivity and thus will correspond to regions of higher intensity in the 2D image 124. Although a top-down view is used in the context of this application, the aggregated view generator 122 is capable of generating 2D images from other viewpoints (such as side views, bottom-up, or any arbitrary viewpoint), and subsequent processing can be adapted to the viewpoint(s) used.

[0031] As shown in Figure 1 the example process 100 may include a semantic classifier 128 configured to accept the 2D image 124 as input and output a semantic labeled image 130. In an example, the labeled image 130 may include (one or more) segments corresponding to map elements of a road network map, indicating the (one or more) semantic labels associated with the (one or more) segments and the spatial extent of the (one or more) segments. For example, the map elements may include lane elements to indicate the extent of a lane or travel corridor, stop line elements to indicate a stop line at an intersection, crosswalk elements to indicate a pedestrian crossing area, etc. Additional non-limiting examples of map elements may include bicycle lane elements, parking space elements, intersection elements, lane divider elements, yield line elements, etc. In some examples, the (one or more) semantic labels may indicate the map element by name (e.g., "crosswalk" indicates a crosswalk element).

[0032] In an example, the implementation of the semantic classifier 128 can include a previously trained machine learning (ML) model 132 that is used to classify pixels in an image (such as the 2D image 124), where the semantic labels correspond to map elements associated with a road network map. For example, the trained ML model 132 can be previously trained using training data that includes images of top-down views of road networks, where the pixels have been labeled with class labels from a specified set of class labels (e.g., manually labeled, class labels provided by a third-party map provider, and / or class labels generated automatically and verified manually). The specified set of class labels can include class labels related to the road network that correspond to the map elements discussed above (e.g., "lane", "bike lane", "left turn lane", "sidewalk", "crosswalk", etc.). The specified set of class labels can also include other common elements, such as "building", "vegetation", "water body", etc., as well as an all-inclusive class label for labeling pixels of one of the other classes that do not belong to the set of class labels (e.g., "other" or "background").

[0033] The semantic classifier 128 can provide the 2D image 124, which represents a top-down view associated with the scene 104, as an input image to the trained ML model 132 and receive from the ML model 132 an output indicating the class associated with an individual pixel of the 2D image 124. The class can include a class label associated with the individual pixel and a confidence score indicating the level of certainty of the class label. For example, the confidence score can be a value between 0 and 1, where 0 indicates that the ML model 132 is completely unsure about the applied class label, while a confidence score close to 1 indicates strong confidence in the applied class label. In some examples, the class can alternatively include a set of probabilities indicating the likelihood that an individual pixel belongs to each class label in the specified set of class labels. As a non-limiting example, the trained ML model can include a fully convolutional network (FCN). As is well known in the art, an FCN is a specialized convolutional neural network (CNN) architecture that incorporates multiple upsampling convolutional layers into a standard CNN, where feature maps from intermediate layers are merged back during upsampling (e.g., using "skip connections"). Thus, the FCN generates an output of the same size and resolution as the input image, where the output includes class labels and class label probabilities (which correspond to confidence scores) associated with individual pixels. Although conventional machine learning classifiers can output a global class label given an input image, the output produced by the FCN provides the additional advantage of localization of the class labels because the class labels are associated with individual pixels located at known (x, y) coordinates within the input image.

[0034] In addition, the semantic classifier 128 can determine segments (e.g., groups of contiguous pixels with the same class label) based on the output received from the trained ML model 132 (e.g., the trained FCN). For example, the semantic classifier 128 can cluster spatially contiguous pixels with the same class label (e.g., by using connected component analysis) to form segments. The semantic classifier 128 can also determine a segment-level confidence score for individual segments of the determined segments based on the pixel-level confidence scores output by the ML model 132. For example, the segment-level confidence score can be an average or a median of the confidence scores associated with the pixels in the segment. The determined segments can be overlaid on the 2D image 124 to generate a labeled image 130, as illustrated by the example labeled image 134, including map data indicating the extent of the determined segments and their associated semantic class labels. In the example labeled image 134, segment 136 is associated with the "crosswalk" class label, segment 138 is associated with the "driving lane" class label, and segment 140 is associated with the "parking lane" class label. Additionally, each of segments 136, 138, 140 can also be associated with a confidence score indicating the certainty of the identified class label, as described above. In some examples, additional class labels can also be included in the labeled image 130, such as the region 142 corresponding to the "building" class label in the example labeled image 134, to indicate permanent structures in the scene 104 that may not be directly related to the road network.

[0035] In an alternative implementation, the labeled image 130 can be generated by the following technique, where: the ML model 132 is trained to output (one or more) masks or segments representing objects in a given top-down view (such as the 2D image 124) as input. In such an example, the semantic classifier 128 can implement one or more techniques described in U.S. Patent No. 10,649,459, titled "Data Segmentation Using Masks," issued on May 12, 2020, the entire text of which is incorporated herein by reference for all purposes. In yet another alternative implementation, the trained ML model 132 can be trained using three-dimensional data (e.g., 3D data points represented in a 3D voxel space), and the 3D voxel space filled with data points from the filtered lidar data set 120 can be directly input into such an ML model 132 to determine (one or more) pixel-level or segment-level classifications.

[0036] As in Figure 1As shown in, process 100 may include a road map generator 144. The road map generator 144 may include the function of generating a road network map of the environment based on map data and / or marker images obtained for different parts of the environment, such as the parts shown in the example marker image 134. The road network map may be, for example, a two-dimensional (2D) or three-dimensional (3D) representation that indicates one or more of the following: driving lane elements, bicycle lane elements, parking lane elements, crosswalk elements, intersection elements, lane separation line elements, stop line elements, yield sign elements, yield line elements, lane elements, speed bump elements, crosswalk elements, etc. In some examples, the road network map can be encoded with information indicating the attributes of specific parts of the road. For example, the driving lanes in the road network map can be encoded with information indicating speed limits, one-way streets, parking restrictions, etc. In at least some cases, the road network map may also include an underlying road grid that includes 3D tiles onto which 2D road network data can be projected. Other details related to such road network data 108 and / or road grid 110 are described in U.S. Patent No. 10,699,477, titled "Generating Shadowless Maps," filed on June 30, 2020, and U.S. Patent No. 11,188,091, titled "Grid Extraction Based on Semantic Information," filed on November 30, 2021, the entire contents of both of which are incorporated herein by reference.

[0037] In some examples, the road network map may include class labels (as illustrated in the example marker image 134) indicated in segments of the (one or more) marker images 130 and the spatial extent of the segments, which are located in a common two-dimensional reference system. For example, during the initial creation of the road network map (e.g., for a new environment that an autonomous vehicle needs to traverse), the (one or more) marker images 130 (such as the example marker image 134) can be generated based on sensor data 106 corresponding to different parts of the environment using the above techniques. In an example, if the segment-level confidence score of a segment is equal to or greater than a minimum confidence threshold, the class label associated with the segment and the spatial extent of the segment can be transferred to the corresponding area of the newly created road network map. However, if the segment-level confidence score is less than the minimum confidence threshold, the segment can be marked for further inspection (e.g., by a human operator) to determine its class label and / or extent. Alternatively or additionally, the segment can be compared with data from a third-party map provider to determine the correspondence between the segment and map elements in the third-party map data. When a match is determined, the segment can be added to the new road network map without additional manual verification.

[0038] The road map generator 144 can also compare the class labels and spatial extents of the segments of the labeled image 130 with an existing road network map to determine the (one or more) differences between the class labels and / or spatial extents of one or more segments of the labeled image 130 and the existing road network map. In some examples, the (one or more) existing road network maps can be previously generated (e.g., by applying process 100 to previously obtained sensor data), and / or obtained from a third-party map provider. As a non-limiting example, the difference in spatial extent between two segments with the same class label can be calculated as the intersection between the two segments (e.g., the number of pixels in the overlapping area) divided by the union of the two segments (e.g., the number of pixels in the area covered by the segments together). For different (one or more) segments (e.g., having different class labels and / or spatial extents that differ by more than a minimum difference threshold), if the (one or more) segment-level confidence scores are equal to or greater than a minimum confidence threshold, the (one or more) segments in the road network map can be automatically updated by replacing the segments in the existing road network map with the segments from the labeled image 130. And if the (one or more) segment-level confidence scores are less than the minimum confidence threshold, the (one or more) different segments can be submitted for further verification (e.g., submitted to a human operator).

[0039] The techniques described herein can improve the functionality of a computing device by providing a framework for efficiently creating and updating road network maps for autonomous vehicle navigation. For example, when a road network map of a new area is needed, process 100 can be executed, and in areas with an existing road network map, the process can be executed at fixed intervals to keep the map up-to-date. When road changes are expected (e.g., due to known construction work, traffic flow changes, etc.), process 100 can also be executed. In some cases, the road network map generated by process 100 can be provided to a planning system to generate a trajectory for an autonomous vehicle to traverse the environment covered by the (one or more) road network maps.

[0040] Figure 2 An example environment 200 is illustrated, in which an example vehicle 202 is traversing the environment while capturing sensor data, as described above with reference to Figure 1 the description. The example environment 200 shows the vehicle 202 traveling on a road surface 204. The vehicle 202 can be an autonomous vehicle and / or one of the vehicles 102 discussed above. The vehicle 202 is driven by wheels 206 (the vehicle has four wheels, two of which are Figure 2(shown in). Although the exemplary autonomous vehicle 202 has four wheels, the systems and methods described herein can be incorporated into vehicles having fewer or greater numbers of wheels and / or tracks. The exemplary vehicle 202 can have four-wheel steering and can operate with the same performance characteristics in all directions. For example, when traveling in a first direction 210, a first end 208 of the vehicle 202 can be the front end of the vehicle 202, and when traveling in an opposite second direction 212, the first end 208 can become the rear end of the vehicle 202, as shown in Figure 2 shown. Similarly, when traveling in the second direction 212, a second end 214 of the vehicle 202 can be the front end of the vehicle 202, and when traveling in the first direction 210, the second end 214 can become the rear end of the vehicle 202. These exemplary features can facilitate greater maneuverability, such as in tight spaces or congested environments, such as parking lots and urban areas. Also as illustrated, a 3D coordinate system is associated with the vehicle 202, where the Y-axis is along the direction of motion of the vehicle 202 (e.g., the first direction 210 or the second direction 212), the X-axis is perpendicular to the direction of motion and in a horizontal plane, and the Z-axis is along the vertical height.

[0041] The exemplary vehicle 202 can be a driverless vehicle, such as an autonomous vehicle, configured to operate according to a level 5 classification issued by the National Highway Traffic Safety Administration, which describes a vehicle capable of performing all safety-critical functions throughout a trip without the driver (or occupant) needing to control the vehicle at any time. In such an example, since the vehicle 202 can be configured to control all functions from the start to the end of a trip, including all parking functions, it may not include a driver and / or control devices for driving the vehicle 202, such as a steering wheel, an accelerator pedal, and / or a brake pedal. This is merely an example, and the systems and methods described herein can be incorporated into any ground, air, or water vehicle, including vehicles that require constant manual control by a driver to vehicles with partial or full autonomous control. In some cases, the techniques can be implemented in any system that uses machine vision to navigate an environment and are not limited to vehicles.

[0042] Vehicle 202 may include a plurality of sensors, including one or more lidar sensors 216a, 216b. In an example, lidar sensors 216a, 216b may be substantially the same, e.g., except for their positions on vehicle 202. In some examples, vehicle 202 may include one or more additional lidar sensors (not shown), which may be substantially the same as sensors 216a, 216b and are arranged to cover a substantially 360-degree field of view around vehicle 202. Alternatively or additionally, vehicle 202 may also include one or more radar or 4D radar sensors 218a, 218b. As described above, lidar sensors 216a, 216b generate a 3D point cloud of the environment based on time-of-flight measurements of reflected light returning to the (one or more) receivers of lidar sensors 216a, 216b. In some examples, the 3D point cloud may be generated by aggregating reflected echoes over a short time period (e.g., 1 / 30 second). Although the techniques described herein apply to point clouds generated by lidar sensors, radar sensors 218a, 218b (such as 4D radar sensors) may also generate a 3D point cloud of the environment, to which the techniques described herein may apply. Although in Figure 2 two modalities of sensors 216, 218 are illustrated, vehicle 202 may include any number of additional sensors, having any number of different modalities. Without limitation, additional sensors (not shown) may be one or more of the following: inertial sensors (e.g., inertial measurement units, accelerometers, magnetometers, gyroscopes, etc.), imaging sensors (e.g., cameras, including stereo cameras and range cameras), time-of-flight sensors, sonar sensors, thermal imaging sensors, or any other sensor modality.

[0043] As illustrated, scene 220 represents a portion of environment 200 in which vehicle 202 operates. Scene 220 illustrates various time-invariant or permanent elements typically represented in a road network map of environment 200, such as (one or more) crosswalks 222, stop lines 224, sidewalks 226, etc., as well as other permanent elements as part of the background, such as (one or more) trees 228, buildings 230, etc. Scene 220 also illustrates various transient elements, such as (one or more) vehicles 232, (one or more) pedestrians 234, etc.

[0044] As in Figure 2As illustrated in the block diagrams, the sensor system(s) 236 of vehicle 202 may generate sensor data, such as lidar data 238 captured by lidar sensors 216a, 216b, radar data 240 captured by radar sensors 218a, 218b, and / or image data 242 captured by an imaging sensor (not shown). In some examples, there may be less or more sensor modes of sensor data than shown. Lidar data 238 may include measurements for computing a 3D cloud of the scene 220 (e.g., 3D positions of points in a local or global coordinate system), and information associated with the points in the 3D point cloud, such as intensity information, signal-to-noise ratio (SNR), velocity information, confidence scores indicating the accuracy of the data associated with the data, etc. Alternatively or additionally, radar data 240 and / or image data 242 may also be used to compute a 3D point cloud of the scene 220. Although using a 3D point cloud as an example, lidar and other data may be represented in other ways and still benefit from the teachings of this disclosure.

[0045] The sensor system(s) 236 may provide sensor data, such as lidar data 238, to the vehicle computing system(s) 244. When the vehicle 202 is operating in the environment 200, the sensor system 236 may provide the sensor data 238, 240, 242 continuously or at discrete time intervals. In some examples, the discrete time intervals between providing the sensor data 238, 240, 242 may be based on the speed of the vehicle 202. For example, when the speed of the vehicle 202 is high, the time interval may be small.

[0046] As shown in Figure 2 the vehicle computing system(s) 244 may implement a sensor data filtering component 246. In some examples, the sensor data filtering component 246 may alternatively be implemented on an external computing system accessible via a network by the vehicle computing system(s) 244. In an example, the sensor data filtering component 246 may filter out portions of the lidar data 238 that are not relevant to the road network map and generate filtered data 248 that retains the remaining portions of the lidar data 238. For example, a road network map typically includes map elements (e.g., road surface 204) located on or near the ground plane on which the vehicle is traveling. In an example, the lidar sensors 216a, 216b may be associated with a virtual origin 250 that may be at a height 252 above the ground plane or road surface 204, as shown in Figure 2As shown in. The virtual origin 250 can form the center of the coordinate system in which the data points of the lidar data 238 are located. The sensor data filtering component 246 can filter the lidar data 238 using a first threshold distance 254a below the height 252 and a second threshold distance 254b above the height 252. For example, lidar echoes from surfaces at a height below the first threshold distance 254a from the height 252, and lidar echoes from surfaces at a height above the second threshold distance 254b from the height 252, can both be filtered out from the lidar data 238. In a non-limiting example, the first threshold distance 254a can be set to 50 cm, and the second threshold distance 254b can be set to 30 cm. In some examples, the first threshold distance 254a can correspond to the height of the vehicle 202's axle from the ground. As shown in the 2D representation of the lidar data points 256 captured from the scene 220, the lidar data 230L corresponding to the surface of the building 230 is filtered out by the sensor data filtering component 246 because its height is above the second threshold distance 254b from the virtual origin 250. While the lidar data 258 covering lidar echoes from the road surface 204, crosswalk 222, stop line 224, sidewalk 226, and other areas located on or near the ground plane is within the threshold distances 254a, 254b from the virtual origin 250 and is thus retained in the filtered data 248. Additionally, the lidar data 222L and 224L corresponding to the crosswalk 222 and stop line 224 respectively can be characterized by high-intensity features in the lidar data 258 due to their high-intensity markings. In some examples, lidar data associated with a specific classification of sensor data from other sensor modalities (e.g., visual cameras) can be enhanced to display easily recognizable intensity features. For example, regions corresponding to road markings of yellow, red, blue, or green can be detected from the camera image, and the lidar data in these regions can be enhanced based on the color of the region.

[0047] In some examples, the sensor data filtering component 246 can identify data points corresponding to transient elements (e.g., vehicles, pedestrians, construction equipment, etc.) in the lidar data 238 and filter out these data points such that the filtered data 248 mainly corresponds to the time-invariant elements of the scene 220. In some examples, object tracking and machine learning-based object recognition can be used to identify transient (or dynamic) and static objects in the scene, e.g., using techniques such as those described in U.S. Patent Application No. 16 / 866,865 filed on May 5, 2020, and U.S. Patent Application No. 17 / 364,603 filed on June 30, 2021, as described above, the entire contents of both of which are hereby incorporated by reference herein for all purposes. InFigure 2 In the example 2D representation of lidar data 256 shown in Figure 2 , the clusters of data points 232L corresponding to vehicle 232 and the clusters of data points 234L corresponding to pedestrian 234 can be identified as transient elements and filtered out from the lidar data 238. The clusters of points 232L, 234L can be spatial clusters or groupings identified in the lidar data 238. Although the representation of lidar data 256 is shown as two-dimensional for illustrative purposes, it should be understood that the lidar data 238 is three-dimensional and can be represented by 3D data points in a 3D voxel space, as described in further detail below with reference to Figure 3 . Thus, the clusters of points 222L, 224L, 228L, 230L, 232L, 234L are also three-dimensional. Figure 3 which is described in further detail below. Thus, the clusters of points 222L, 224L, 228L, 230L, 232L, 234L are also three-dimensional.

[0048] Additionally, the sensor data filtering component 246 can also remove data associated with low confidence scores. As described above, in some cases, lidar data can include confidence scores associated with an individual or a group of data points, which indicate the accuracy of the data. For example, the confidence score can be based on the functionality of the sensor (e.g., if a fault is detected, the confidence score may be low), the signal-to-noise ratio (e.g., data points with a low SNR may result in corresponding low confidence scores), and / or environmental conditions that may affect the performance of the sensor (e.g., under adverse weather conditions, the overall confidence score of the lidar data 238 may be low). As an example, in the spatial map of lidar data 256 shown in Figure 2 , the cluster of points 228L corresponding to tree 228 may be filtered out due to a low confidence score, which may be caused by a low SNR value because lidar echoes may be irregularly reflected from the leaves of tree 228. Figure 2 in the spatial map of lidar data 256 shown in Figure 2 , the cluster of points 228L corresponding to tree 228 may be filtered out due to a low confidence score, which may be caused by a low SNR value because lidar echoes may be irregularly reflected from the leaves of tree 228.

[0049] In some examples, the vehicle computing system(s) 244 can store the filtered data 248 and / or the sensor data 238, 240, 242 on the local memory 260 on the vehicle 202. The vehicle computing system(s) 244 can also transmit the filtered data 248 and / or the sensor data 238, 240, 242 or a compressed representation thereof to an external computing system for further processing or storage. The filtered data 248 and / or the sensor data 238, 240, 242 can be transmitted when they are generated and / or acquired, or periodically at fixed intervals. In some examples, the transmission of the filtered data 248 and / or the sensor data 238, 240, 242 can be triggered by events such as the amount of unused space remaining in the memory 260 being below a threshold, an abnormal situation being detected in the environment 200 that the vehicle 202 is traversing, in response to a debugging request, etc.

[0050] Although the above techniques are described in the context of lidar data 238, in some examples, the vehicle computing system(s) 244 may process sensors of other modalities or combinations of modalities carried by vehicle 202. For example, the vehicle computing system(s) 244 may process each of the sensor data 238, 240, 242 individually (e.g., the vehicle computing system(s) 244 may implement a separate sensor data filtering component 246 for each sensor modality), and combine the results in the filtered data 248 (e.g., using a “late sensor fusion” technique). In other examples, the sensor data 238, 240, 242 may be combined before applying the sensor data filtering component 246 to the combined or merged sensor data (e.g., using an “early sensor fusion” technique).

[0051] Figure 3 Includes text and visual flowcharts for illustrating example process 300, which is for determining a semantic segmentation of a top-down view of a scene generated from three-dimensional (3D) sensor data (e.g., lidar data). In the examples described herein, the sensor data may be obtained by a lidar sensor disposed on an autonomous vehicle (such as vehicle 202). In this example, process 300 generates a two-dimensional (2D) image (e.g., a top-down view) from an aggregated representation of the 3D sensor data, and determines a semantic segmentation of the image indicative of map elements related to a road network map.

[0052] In operation 302, process 300 includes receiving filtered sensor data 304 of a portion of the environment. The filtered sensor data 304 can be obtained from raw sensor data (such as captured by sensors of the vehicle(s)), as the output of a sensor data filtering component 246, for example, and may primarily include echoes associated with the drivable surface and / or features near the ground plane, such as driving lanes, turning lanes, stop lines, crosswalks, etc. As shown, the filtered sensor data 304 can include sensor data instances 304(1), 304(2), 304(N) (collectively referred to as instances 304 or individually as an instance or each instance 304) associated with a portion of the environment, where each instance 304 can be captured by a vehicle traversing the portion of the environment during different trips and / or at different times. In an example, each instance 304 can be associated with a corresponding 3D voxel space 306 (where a first voxel space 306(1), a second voxel space 306(2) through an Nth voxel space 306(N) are illustrated, and the origin(s) of each voxel space 306 can correspond to a virtual origin associated with the sensor that captured the sensor data). Individual data points 308 of the filtered sensor data 304 can be represented in the voxel space 306 corresponding to the respective instance 304 during which the individual data point 308 was captured. In some instances, the voxels of the voxel space 306 can represent any amount of data, including but not limited to: covariance matrices, position information, classification information, segmentation information, number of observations, whether the voxel is occupied, etc.

[0053] In operation 310, process 300 includes aggregating the filtered sensor data 304 in a common 3D voxel space 312. Operation 310 may include determining an alignment between the individual voxel space 306 and the voxel space 312 based on the position and / or trajectory information of the vehicle(s) at the time the corresponding sensor data was captured. For example, a combination of inertial data from inertial sensors on the vehicle, global positioning data from the vehicle GPS system, and / or pose estimation based on matching landmarks in the scene (e.g., based on camera image features) may be used to generate an optimized 3D pose and position of the vehicle at each instance of capturing the filtered sensor data 304. Examples of techniques for determining the alignment between a first trajectory and a second trajectory (e.g., using algorithms such as CLAMS (Concurrent Calibration, Localization, and Mapping) or SLAM (Simultaneous Localization and Mapping)) are discussed in U.S. Patent No. 10,983,199 issued on April 20, 2021 and U.S. Patent No. 10,782,136 issued on September 22, 2020, the entire contents of which are incorporated herein by reference as described above. For example, the alignment between an individual in the voxel space 306 and the voxel space 312 may indicate a 3D affine transformation (e.g., rotation, translation, scaling, and shearing) to be applied to the voxel space to align it spatially with the voxel space 312. As can be appreciated, each voxel space 306 may require a different 3D affine transformation.

[0054] As illustrated, the voxel space 312 can extend over a three-dimensional space (e.g., X, Y, Z) and may contain any number of voxels along each dimension. In some cases, the voxel space 312 may correspond to the physical volume of a portion of the environment. For example, the voxel space 312 may represent a physical volume that is 100 centimeters wide, 100 centimeters long, and 100 centimeters high. Additionally, each voxel (e.g., voxel 314) in the voxel space 312 may represent a physical volume, such as 5 centimeters in each dimension. In some cases, the voxels may be of uniform size throughout the voxel space 312, while in some cases, the volume of the voxels may vary based on the position of the voxel relative to the origin of the voxel space 312. In some cases, a ground plane voxel corresponding to the ground plane may be determined in the voxel space 312, and the volume or size of the voxels in the voxel space 312 may be smaller when closer to the ground plane voxel to capture more detail and may increase proportionally with the distance from the voxel to the ground plane voxel.

[0055] In an example, data points 308 of filtered sensor data 304 can be mapped to voxels in a common voxel space 312 to obtain aggregated data 316. The individual data points of data points 308 can be mapped to the voxel space 312 by applying a 3D affine transformation to align between their corresponding voxel spaces 306 and voxel space 312. Although depicted as multiple data points for illustrative purposes, each voxel can store a single data point obtained by integrating all data (e.g., lidar echoes) from the physical volume corresponding to the voxel. For example, data points can be statistically accumulated in individual voxels, and an individual voxel can include data representing the number of echoes, average intensity, average x value of the data, average y value of the data, average z value of the data, and / or a covariance matrix of the sensor data associated with the individual voxel. Since road map elements are typically characterized by high reflectivity, voxels containing such road map elements can be associated with higher average intensity values.

[0056] In an example, some or all of the voxels in the voxel space 312 can be pre-initialized with data representing previously captured sensor data. In some examples, the voxel space 312 can be initialized as a blank space, and sensor data can be added to the voxel space 312 when the sensor data is captured. In other examples, voxels within the voxel space 312 can be instantiated when sensor data is to be associated with such voxels, thereby reducing or minimizing the amount of memory associated with the voxel space. As a non-limiting example, this can be performed using techniques such as voxel hashing. In still other examples, voxels that do not contain data or contain a number of points below a threshold number can be discarded or omitted to create a sparse voxel space.

[0057] At operation 318, process 300 includes generating a top-down view of the aggregated data 316 in the voxel space 312. The top-down view can be represented as a 2D image 320 in the (X, Y) plane of the (X, Y, Z) axes of the 3D voxel space 312. As referenced Figure 2 above, the Z axis is associated with vertical height, and thus, the image data 322 (e.g., pixel 324) in the pixels of the image 320 represents a top-down perspective or planar view of the voxel space 312, where the height is folded or flattened. In some examples, the pixels of the image 320 can represent cumulative intensity values obtained by summing the intensity values contributed by instances 304 of the filtered sensor data at each pixel position.

[0058] In other examples, the 2D image 320 can represent the average intensity value of the intensity values of each column along the Z direction of the voxel space 312 and / or a weighted average intensity value weighted by a plurality of data points. As can be understood, lidar echoes corresponding to time-invariant features of portions of the environment are expected to appear in most instances of the filtered sensor data 304, resulting in higher weighted average intensity values or average intensity values. Correspondingly, even if lidar echoes representing transient elements are not completely filtered out, such echoes are expected to appear in a small portion of the instances of the filtered sensor data 304, and thus, the corresponding average intensity values in the image 320 will be low. Thus, the 2D image 320 effectively represents the voxel space 312 as seen from above and emphasizes time-invariant features at or near the ground plane of portions of the environment.

[0059] At operation 326, process 300 includes determining a pixel class associated with a pixel of the 2D image 320. Operation 326 can determine the pixel class by inputting the 2D image 320 into a machine learning (ML) model that is trained to determine a class label from a specified set of class labels and receives an output that includes the class label corresponding to an individual pixel of the image 320. As described above, the specified set of class labels can include class labels related to a road network, corresponding to map elements such as "road lane", "bike lane", "left turn lane", "sidewalk", "crosswalk", etc. The specified set of class labels can also include permanent elements of the environment such as "building", "vegetation", "water body", etc., and an all-inclusive class label (e.g., "other" or "background") for labeling pixels that do not belong to one of the other classes in the set of class labels. The class can include a confidence score indicating the correctness of the class label. For example, the trained ML model can determine the confidence score based on the prior probability of the class label and the class label probability calculated for an individual pixel during classification. As a non-limiting example, the trained ML model can include a fully convolutional network (FCN) that returns an output image of the same size as the input image (e.g., image 320), where the output image includes the class label and a pixel-level confidence score associated with an individual pixel. Conventional machine learning classifiers can output a global class label given an input image, while the output image produced by the FCN provides localization of the class label because an individual pixel located at a known (x, y) coordinate within the output image corresponds to the same (x, y) coordinate within the input image. In Figure 3 the example illustrated in Figure 1For the 2D image 126, the pixel 328 can be classified as the "crosswalk" class label with a confidence score of 0.8, and the adjacent pixel 330 can also be classified as the "crosswalk" class label with a confidence score of 0.9.

[0060] At operation 332, the process 300 includes determining the semantic segmentation of the 2D image 320, which includes map data that identifies class labels associated with (one or more) portions or (one or more) segments of the 2D image 320, the (one or more) spatial extents (e.g., (one or more) bounding boxes encapsulating the (one or more) two-dimensional segments or splines fitted to linear segments) of the (one or more) portions or (one or more) segments within the 2D image 320, and, in some examples, (one or more) segment-level confidence scores associated with the identified (one or more) segments. In some examples, determining the semantic segmentation can include clustering spatially contiguous pixels with the same class label in the output generated by an ML model (e.g., FCN) by using connected component analysis techniques known in the art. In such examples, a neighborhood can be defined around each pixel, and the clustering can be iteratively expanded from "seed" pixels by adding pixels with the same class label within the defined neighborhood. Additionally, the segment-level confidence score for each segment can be calculated based on the pixel-level confidence scores obtained as the output from the ML model. For example, the segment-level confidence score can be the average or median of the pixel-level confidence scores of the pixels in the segment.

[0061] In Figure 3 In the example illustrated in 332, the adjacent pixels 328 and 330 share the same class label "crosswalk". Thus, during operation 332, the pixels 328, 330 can be clustered into a segment 334, as shown in the example 2D image 126. Additionally, the segment 334 can be assigned the segment-level label "crosswalk" and a segment-level confidence score of 0.85 (e.g., the average of the pixel-level confidence scores). The map data corresponding to the segment 334 can include a bounding box indicating the extent of the segment 334, the class label (e.g., "crosswalk"), and the associated segment-level confidence score. In some examples, the bounding box can be non-rectangular and can alternatively include a polygon representation, an elliptical representation, a set of contour points, etc. Although pixels 328 and 330 are used as an example, it can be understood that a typical segment will include more pixels. Operation 332 can also compare the segment size to a minimum size threshold and eliminate segments whose size (e.g., the number of pixels in the segment) is less than the minimum size threshold.

[0062] In an alternative implementation, operation 332 can be directly performed after operation 310 by using an ML model that is trained to segment and classify three-dimensional data, such as the aggregated data in voxel space 312. In some examples, the techniques described in U.S. Patent Application No. 17 / 127,196, titled "Multi-Resolution Top-Down Segmentation," filed on December 18, 2020, which is incorporated herein by reference in its entirety for all purposes, can be used to identify an object by determining the portion of lidar data corresponding to the object. Alternatively or additionally, machine learning models and techniques for determining semantic segmentation of sensor data, as discussed in U.S. Patent Application No. 10,535,138, titled "Sensor Data Segmentation," published on January 14, 2020, which is incorporated herein by reference in its entirety for all purposes, can be used to associate class labels with segments without generating a top-down view.

[0063] Figure 4 is a block diagram of an example system for implementing the techniques described herein. In at least one example, system 400 can include a vehicle 402 that can be similar to vehicle 102(1, 2,..., N). In the illustrated example system 400, vehicle 402 can be an autonomous vehicle, such as vehicle 202; however, vehicle 402 can be any other type of vehicle.

[0064] Vehicle 402 can include one or more computing devices 404, one or more sensor systems 406, one or more transmitters 408, one or more communication connections 410 (also referred to as communication devices and / or modems), at least one direct connection 412 (e.g., for physically coupling to vehicle 402 to exchange data and / or provide power), and one or more drive systems 414. One or more sensor systems 406 can be configured to capture sensor data related to the environment.

[0065] One or more sensor systems 406 may include time-of-flight sensors, position sensors (e.g., GPS, compass, etc.), inertial sensors (e.g., inertial measurement unit (IMU), accelerometer, magnetometer, gyroscope, etc.), lidar sensors, radar sensors, sonar sensors, infrared sensors, cameras (e.g., RGB, IR, intensity, depth, etc.), microphone sensors, environmental sensors (e.g., temperature sensor, humidity sensor, light sensor, pressure sensor, etc.), ultrasonic transducers, wheel encoders, etc. One or more sensor systems 406 may include multiple instances of each of these or other types of sensors. For example, the time-of-flight sensors may include individual time-of-flight sensors located at the corners, front, rear, sides, and / or top of the vehicle 402. As another example, the camera sensors may include multiple cameras arranged at various locations outside and / or inside the vehicle 402. One or more sensor systems 406 may provide input to the computing device 404.

[0066] The vehicle 402 may also include one or more transmitters 408 for emitting light and / or sound. The one or more transmitters 408 in this example include internal auditory and visual transmitters for communicating with the passengers of the vehicle 402. By way of example and not limitation, the internal transmitters may include speakers, lights, signs, displays, touchscreens, haptic transmitters (e.g., vibration and / or force feedback), mechanical actuators (e.g., seatbelt tensioners, seat positioners, headrest positioners, etc.), etc. The one or more transmitters 608 in this example also include external transmitters. By way of example and not limitation, the external transmitters in this example include lights (e.g., indicator lights, signs, light arrays, etc.) for emitting signals of the driving direction or other indicators of vehicle actions, and one or more auditory transmitters (e.g., speakers, speaker arrays, horns, etc.) for auditory communication with pedestrians or other nearby vehicles, one or more of which may include beam control technology.

[0067] The vehicle 402 may also include one or more communication connections 410 for enabling communication between the vehicle 402 and one or more other local or remote computing devices (e.g., remotely controlled computing devices) or remote services. For example, the one or more communication connections 410 may facilitate communication with other local computing devices on the vehicle 402 and / or one or more drive systems 414. Similarly, the one or more communication connections 410 may allow the vehicle 402 to communicate with other nearby computing devices (e.g., other nearby vehicles, traffic lights, etc.).

[0068] One or more communication connections 410 may include physical and / or logical interfaces for connecting computing device 404 to another computing device or one or more external networks 440 (e.g., the Internet). For example, one or more communication connections 410 may implement Wi-Fi-based communication, such as via frequencies defined by the IEEE 802.11 standard, short-range wireless frequencies (e.g., Bluetooth), cellular communication (e.g., 2G, 3G, 4G, 4G LTE, 5G, etc.), satellite communication, dedicated short-range communication (DSRC), or any suitable wired or wireless communication protocol that enables the corresponding computing device to interact with other computing devices.

[0069] In at least one example, vehicle 402 may include one or more drive systems 414. In some examples, vehicle 402 may have a single drive system 414. In at least one example, if vehicle 402 has multiple drive systems 414, the individual drive systems 414 may be located at opposite ends of vehicle 402 (e.g., front and rear, etc.). In at least one example, (one or more) drive systems 414 may include one or more sensor systems 406 to (one or more) detect conditions of the drive systems 414 and / or the surrounding environment of vehicle 402. By way of example and not limitation, (one or more) sensor systems 406 may include one or more wheel encoders (e.g., rotary encoders) to sense the rotation of the wheels of the drive system, inertial sensors (e.g., inertial measurement units, accelerometers, gyroscopes, magnetometers, etc.) to measure the orientation and acceleration of the drive system, cameras or other image sensors, ultrasonic sensors to acoustically detect objects around the drive system, lidar sensors, radar sensors, etc. Some sensors (such as wheel encoders) may be unique to (one or more) drive systems 414. In some cases, (one or more) sensor systems 406 on (one or more) drive systems 414 may overlap or supplement corresponding systems of vehicle 402 (e.g., sensor system 406).

[0070] (One or more) drive systems 414 may include a number of vehicle systems, including a high-voltage battery, an electric motor for propelling the vehicle, an inverter for converting direct current from the battery into alternating current for use by other vehicle systems, a steering system including a steering motor and a steering rack (which may be electric), a braking system including a hydraulic or electric actuator, a suspension system including hydraulic and / or pneumatic components, a stability control system for distributing braking force to mitigate traction loss and maintain control, an HVAC system, lighting (e.g., lighting for illuminating the vehicle's external environment, such as headlights / taillights), and one or more other systems (e.g., a cooling system, a safety system, an on-board charging system, other electrical components, such as a DC / DC converter, high-voltage connectors, high-voltage cables, a charging system, a charging port, etc.). Additionally, (one or more) drive systems 414 may include a drive system controller that may receive and preprocess data from the sensor system 406 and control the operation of the various vehicle systems. In some examples, the drive system controller may include one or more processors and a memory communicatively coupled to the one or more processors. The memory may store one or more components to perform the various functions of (one or more) drive systems 414. Further, (one or more) drive systems 414 also include one or more communication connections such that the respective drive system can communicate with one or more other local or remote computing devices.

[0071] The computing device 404 may include one or more processors 416 and a memory 418 communicatively coupled to the one or more processors 416. In the illustrated example, the memory 418 of the computing device 404 stores a localization component 420, a perception component 422 (including a sensor data filtering component 424 and a scene representation component 426), a prediction component 428, a planning component 430, a map component 432, and one or more system controllers 434. Although depicted as residing in the memory 418 for illustrative purposes, it is contemplated that the localization component 420, the perception component 422, the prediction component 428, the planning component 430, the map component 432, and the one or more system controllers 434 may alternatively or additionally be accessible by the computing device 404 (e.g., stored in different components of the vehicle 402) and / or accessible to the vehicle 402 (e.g., remotely stored).

[0072] In the memory 418 of the computing device 404, the localization component 420 can include the functionality of receiving data from the sensor system(s) 406 to determine the location of the vehicle 402. For example, the localization component 420 can include and / or request / receive a three-dimensional map of the environment and can continuously determine the location of the autonomous vehicle within the map. In some examples, the localization component 420 can use SLAM (Simultaneous Localization and Mapping) or CLAMS (Concurrent Calibration, Localization, and Mapping) to receive time-of-flight data, image data, lidar data, radar data, sonar data, IMU data, GPS data, wheel encoder data, or any combination thereof, etc., to accurately determine the location of the autonomous vehicle (e.g., a point and heading on a map, latitude and longitude with a heading direction, etc.). In an example, the localization component 420 can provide data to various components of the vehicle 402 to determine the precise location of the autonomous vehicle 402 within the environment that the vehicle is traversing, as described herein. In some examples, the location of the autonomous vehicle 402 can be associated with the sensor data captured by the sensor system(s) 406 at that location.

[0073] The perception component 422 can include the functionality of performing object detection, segmentation, and / or classification. In some examples, the perception component 422 can include a sensor data filtering component 424, which can implement a functionality similar to that of the sensor data filtering component 246. In an example, the sensor data filtering component 424 can process the sensor data from the sensor system to filter out portions of the sensor data that may be irrelevant to the road network map. For example, the sensor data filtering component 424 can filter out portions of the sensor data associated with surfaces at an elevation above a vertical distance threshold from the ground plane in the environment, portions of the sensor data where transient elements (e.g., vehicles, pedestrians, etc.) are present in the environment, and / or portions of the sensor data associated with a low confidence score, as referenced Figure 2 as described.

[0074] The perception component 422 may also include a scene representation component 426, which may generate a scene representation of a portion of the environment that the vehicle 402 is traversing. The scene representation component 426 may represent the filtered sensor data output by the sensor data filtering component 424 in a 3D voxel space. As described above, the 3D voxel space may represent the physical volume in the portion of the environment around an origin corresponding to the (one or more) sensor systems 406. The physical volume represented by the 3D voxel space may be uniquely identified using the position information of the vehicle 402 (e.g., as provided by the positioning component 420), and the scene representation may be associated with the unique global position. In additional and / or alternative examples, the perception component 422 may provide processed sensor data that indicates one or more features associated with the environment, which may include but are not limited to: the presence of another entity in the environment, the state of another entity in the environment, the time of day, the day of the week, the season, the weather conditions, dark / light indication, etc.

[0075] The planning component 428 may determine a path to follow when the vehicle 402 traverses the environment. For example, the planning component 428 may determine various routes and paths at various levels of detail. In some examples, the planning component 428 may determine a travel route from a first location (e.g., the current location) to a second location (e.g., the target location). For the sake of discussion, a route may be a series of waypoints for traveling between two locations. As a non-limiting example, waypoints include streets, intersections, Global Positioning System (GPS) coordinates, etc. Additionally, the planning component 428 may generate instructions for guiding the autonomous vehicle to travel along at least a portion of the route from the first location to the second location. In at least one example, the planning component 428 may determine how to guide the autonomous vehicle to travel from a first waypoint in the sequence of waypoints to a second waypoint in the sequence of waypoints. In some examples, the instructions may be a path or a portion of a path. The planning component 428 may utilize a map (such as the road network map discussed herein) to determine the waypoints or routes for traveling between two locations.

[0076] The memory 418 may also include one or more maps 430, such as the road network maps discussed herein, which the vehicle 402 may use to navigate in the environment. However, the maps may be any number of data structures modeled in two dimensions, three dimensions (e.g., a three-dimensional grid of the environment), or N dimensions, which are capable of providing information about the environment, such as but not limited to: topology (such as intersections), streets, mountains, roads, terrain, and the general environment. In some examples, the (one or more) maps 430 may be stored in a tiled format such that individual tiles of the map represent discrete portions of the environment and may be loaded into working memory as needed, as described herein. In some examples, the vehicle 402 may be controlled at least in part based on the (one or more) maps 430. That is, the (one or more) maps 430 may be used in conjunction with the localization component 420, the perception component 422 (and sub-components), and / or the planning component 428 to determine the current position of the vehicle 402 and / or to generate a route and / or trajectory for navigating in the environment.

[0077] In at least one example, the computing device 404 may include one or more system controllers 432, which may be configured to control the steering, propulsion, braking, safety, transmitter, communication, and other systems of the vehicle 402. These system controllers 432 may communicate with and / or control the corresponding systems of the (one or more) drive systems 414 and / or other components of the vehicle 402, which may be configured to operate in accordance with a path provided by the planning component 430.

[0078] The vehicle 402 may be connected to the (one or more) computing devices 434 via the network 436 and may include one or more processors 438 and a memory 440 communicatively coupled to the one or more processors 438. In at least one instance, the one or more processors 438 may be similar to the (one or more) processors 416, and the memory 440 may be similar to the memory 418. In the illustrated example, the memory 440 of the (one or more) computing devices 434 stores a data aggregation component 442, a segmentation component 444, and / or a road network map component 446. In at least one instance, the segmentation component 444 may store a pre-trained ML model 448 for classifying and / or segmenting map elements from a representation of sensor data, as described herein. Although depicted as residing in the memory 440 for illustrative purposes, it is contemplated that the data aggregation component 442 and / or the segmentation component 444 may alternatively or additionally be accessible by the computing device 434 (e.g., stored in different components of the computing device 434 and / or may be accessible by the computing device 434 (e.g., remotely stored on cloud storage).

[0079] The data aggregation component 442 may include functionality for receiving sensor data or representations thereof (e.g., the output of the scene representation component 426) from one or more vehicles (e.g., similar to vehicle 402 and vehicle 102(1, 2,..N)) during multiple traversals of the environment, which may include an indication of the precise location where the sensor data was captured. In an example, sensor data may be collected at regular time intervals or when an expected change occurs (e.g., due to known construction work, traffic flow changes, etc.). In some examples, the data aggregation component 442 may accumulate data associated with the same portion of the environment or scene in voxels of a common 3D voxel space, as described herein. In some examples, in addition to average intensity or weighted average intensity information (e.g., from lidar data, radar data, etc.), the common 3D voxel space may also include texture or color information (e.g., in the RGB, Lab, HSV / HSL color spaces, etc.), individual "facets" associated with individual voxels of the 3D voxel space (e.g., polygons associated with a single color and / or intensity), reflectivity information (e.g., specularity information, retroreflectivity information, BRDF information, BSSRDF information, etc.), and the like. The data aggregation component 442 may also compute projections of the aggregated 3D data to generate a two-dimensional view of the scene (e.g., a top-down view), as described herein.

[0080] In some examples, the data aggregation component 442 may maintain a 3D voxel space filled with previously captured data, which may be indexed by the "capture location of the sensor data". When additional data is received (e.g., from vehicle 402), the 3D voxel space may be updated. In some examples, for instance, voxels within the common 3D voxel space may only be instantiated when sensor data is to be associated with such voxels, thereby using techniques such as voxel hashing to maintain a sparse voxel space representation.

[0081] The segmentation component 444 can include functionality to segment and classify the aggregated data representation (e.g., a common 3D voxel space or its 2D projected image) generated by the data aggregation component 442 to generate map data indicating the extent of the (one or more) segments and corresponding class labels. In some examples, the segmentation component 444 can implement the functionality of the semantic classifier 128 and / or operations 326 and 332 to generate a semantically labeled segmented image and / or map data, as described above. In other examples, the segmentation component 444 can apply 3D segmentation techniques to directly identify segments in the 3D voxel space. The segmentation component 444 can include a machine learning (ML) model 448 trained for segmentation and / or classification tasks in 2D or 3D. As an example, the ML model 448 can include a fully convolutional network (FCN) that is trained to identify class labels (e.g., map elements such as "driving lane", "crosswalk", "stop line", etc.) associated with individual pixels of a 2D input image (e.g., a top-down view of a road network). In such an example, the segmentation component 444 can also determine segments (e.g., groups of contiguous pixels with the same class label) based on the output received from the trained ML model 448 (e.g., the FCN). In other examples, the ML model 448 can be trained to directly determine segments and corresponding class labels in a 2D input image (e.g., a top-down view) or 3D data (e.g., a common 3D voxel space). Although described for illustrative purposes as residing in the memory 440, it is contemplated that the ML model 448 can alternatively or additionally be accessed by the computing device 434 via a software as a service (SaaS) interface of a cloud computing platform. The segmentation component 444 can also determine a segment-level confidence score for the identified segments based on the class probabilities output by the ML model 448.

[0082] The memory 440 can also include a road network map component 446, which can include functionality to create or update a road network map based on the output of the segmentation component 444 (e.g., the identified segments with class labels and corresponding confidence scores). In an example, if the confidence score associated with the identified segment is equal to or greater than a minimum threshold, the class label associated with the segment and / or the spatial extent of the segment can be transferred to the road network map of the environment. In an example, the road network map can include, but is not limited to, a top-down view of the environment, where semantically labeled map elements show the spatial extent of the map elements. In other examples, the road network map 446 can include a three-dimensional grid of the environment, where semantic labels can be applied to the polygonal surfaces of the grid based on the class labels applied to the identified segments (e.g., as an output of the segmentation component 444).

[0083] The road network map component 446 can also compare the identified segment with an existing road network map of the same part of the environment to determine differences. In some examples, the existing road network map can be obtained from a third-party map provider. In an example where the currently identified segment and the corresponding segment in the existing road network map have the same category label, the difference can be calculated within the spatial extent of the two segments. As a non-limiting example, the difference in spatial extent can be calculated as the intersection between the two segments (e.g., the number of pixels in the overlapping area) divided by the union of the two segments (e.g., the number of pixels in the area covered by these segments together). In some examples, the difference in category label or spatial extent can be marked for manual verification (e.g., if the confidence score of the segment is less than a minimum threshold). In other examples, the category label and / or spatial extent of the segment can be automatically updated in the road network map (e.g., if the confidence score of the segment is greater than or equal to the minimum threshold).

[0084] In some examples, the (one or more) existing road network maps can be stored in a repository remote from the (one or more) computing devices 434, and a portion of the existing road network map at or near the capture location of the received sensor data can be transmitted to the memory 440.

[0085] The (one or more) processors 416 of the computing device 404 and the (one or more) processors 438 of the (one or more) computing devices 434 can be any suitable processors capable of executing instructions to process data and perform the operations described herein. By way of example and not limitation, processors 416 and 438 can include one or more central processing units (CPUs), graphics processing units (GPUs), or any other device or portion of a device that processes electronic data to transform that electronic data into other electronic data that can be stored in registers and / or memory. In some examples, integrated circuits (e.g., ASICs, etc.), gate arrays (e.g., FPGAs, etc.), and other hardware devices can also be considered processors so long as they are configured to implement encoded instructions.

[0086] The memories 418 of computing device 404 and 440 of computing device 434 are examples of non-transitory computer-readable media. Memories 418 and 440 may store an operating system and one or more software applications, instructions, programs, and / or data to implement the methods and functions attributed to the various systems described herein. In various implementations, memories 418 and 440 may be implemented using any suitable storage technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash-type memory, or any other type of memory capable of storing information. The architectures, systems, and individual elements described herein may include many other logical, program, and physical components, and the components shown in the figures are only examples relevant to the discussion herein.

[0087] In some examples, aspects of some or all of the components discussed herein may include any model, algorithm, and / or machine learning algorithm. For example, in some examples, the components in memories 418 and 440 may be implemented as neural networks. As described herein, an exemplary neural network is an algorithm that passes input data through a series of connected layers to produce an output. Each layer in a neural network may also include another neural network, or may include any number of layers (whether convolutional or not). As can be understood in the context of the present invention, neural networks may utilize machine learning, which may refer to a broad class of algorithms that generate outputs based on learned parameters.

[0088] Although discussed in the context of neural networks, any type of machine learning consistent with the present invention can be used. For example, machine learning or machine learning algorithms can include, but are not limited to, regression algorithms (e.g., ordinary least squares regression (OLSR), linear regression, logistic regression, stepwise regression, multivariate adaptive regression splines (MARS), locally estimated scatterplot smoothing (LOESS)), instance-based algorithms (e.g., ridge regression, least absolute shrinkage and selection operator (LASSO), elastic net, least angle regression (LARS)), decision tree algorithms (e.g., classification and regression tree (CART), iterative dichotomiser 3 (ID3), chi-squared automatic interaction detection (CHAID), decision stump, conditional decision tree), Bayesian algorithms (e.g., naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, averaged one-dependence estimators (AODE), Bayesian belief network (BNN), Bayesian network), clustering algorithms (e.g., k-means, k-medians, expectation maximization (EM), hierarchical clustering), association rule learning algorithms (e.g., perceptron, backpropagation, hopfield network, radial basis function network (RBFN)), deep learning algorithms (e.g., deep Boltzmann machine (DBM), deep belief network (DBN), convolutional neural network (CNN), stacked autoencoder), dimensionality reduction algorithms (e.g., principal component analysis (PCA), principal component regression (PCR), partial least squares regression (PLSR), Sammon mapping, multidimensional scaling (MDS), projection pursuit, linear discriminant analysis (LDA), mixture discriminant analysis (MDA), quadratic discriminant analysis (QDA), flexible discriminant analysis (FDA)), ensemble algorithms (e.g., boosting, bagging, AdaBoost, stacked generalization (blending), gradient boosting machine (GBM), gradient boosting regression tree (GBRT), random forest), SVM (support vector machine), supervised learning, unsupervised learning, semi-supervised learning, etc. Other examples of architectures include neural networks such as ResNet-50, ResNet-101, VGG, DenseNet, PointNet, etc.

[0089] Figure 1 , Figure 3 and Figure 5illustrates an example process according to an example of the present invention. The processes are shown as logical flowcharts, each operation of which represents a sequence of operations that can be implemented in hardware, software, or a combination thereof. In the context of software, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific abstract data types. The order in which the operations are described should not be construed as limiting, and any number of the described operations may be omitted or implemented in any order and / or in parallel combination to implement the processes.

[0090] Figure 5 is an example process 500 for updating a map (e.g., of a road network) based on lidar data from a scene in an environment and subsequently controlling an autonomous vehicle based on the environmental map. For example, as described herein, some or all of process 500 may be performed by Figure 4 one or more components among. For example, some or all of process 500 may be performed by the sensor data filtering component 424, the scene representation component 426, the data aggregation component 442, the segmentation component 444, and / or the road network map component 446.

[0091] At operation 502, the process may include receiving a set of lidar data of the same scene of the environment captured at different times (e.g., during different trips). In some examples, operation 502 may include receiving lidar data, time-of-flight data, image data, radar data, etc. of a scene in the environment. In some examples, the set of sensor data of the scene may be captured by a vehicle (e.g., an autonomous vehicle) during multiple different traversals of the environment and may be associated with the exact location of the scene based on the position and / or heading of the vehicle when the corresponding sensor data was captured.

[0092] At operation 504, the process may include determining a filtered set of sensor data of the scene. During operation 504, portions of the set of sensor data that are not relevant to the road network map may be filtered out. For example, the filtered portion of the sensor data may include data associated with surfaces at elevations above a vertical distance threshold from the ground plane in the environment, data associated with transient elements in the environment (e.g., vehicles, pedestrians, etc.), and / or data associated with low confidence scores (e.g., due to sensor failures, adverse environmental conditions, low signal-to-noise ratio, and / or surface characteristics such as reflectivity). Thus, the filtered set of sensor data mainly corresponds to time-invariant elements in the scene that are near or located on the ground plane or a drivable surface.

[0093] At operation 506, the process can include associating aggregated sensor data with a 3D voxel space. As described with reference to Figure 3 , the individual data of the filtered sensor data set can be represented in associated 3D voxel spaces that are not spatially aligned in the global 3D coordinate system. Operation 506 can include aligning the individual 3D voxel spaces with the 3D voxel space and aggregating the individual data of the filtered sensor data set into a common 3D voxel space after alignment. Each voxel of the 3D voxel space can be associated with a specific physical volume of the scene. Additionally, each voxel can store a single data point obtained by integrating all data (e.g., lidar echoes) from the physical volume corresponding to the voxel. For example, the data points can be statistically accumulated in a single voxel, and a single voxel can include data representing the number of echoes, average intensity, average x value of the data, average y value of the data, average z value of the data, and / or covariance matrix of the sensor data associated with the single voxel.

[0094] At operation 508, the process can include generating an image representing a top-down view of the 3D voxel space. The image can be generated by projecting the data in the 3D voxel space (with X, Y, and Z axes) onto a two-dimensional plane (with X and Y axes). The pixels of the image can represent the maximum value, cumulative value, average intensity value, and / or weighted average intensity value (e.g., weighted by multiple data points) of the data in the corresponding column of the 3D voxel space. Of course, other views can be used, and the techniques herein are not limited to the top-down view.

[0095] At operation 510, the process can include determining labeled segments and corresponding confidence scores in the image (e.g., map data). Operation 510 can include inputting the image data into a machine learning model to determine the segments and / or classifications associated with the image. In one example, the machine learning model may include a fully convolutional network (FCN) that is trained to classify the image pixels of the top-down view of the road network into one of a set of specified semantic class labels that represent map elements associated with the map of the road network. The individual pixels can be clustered into segments based on a connected component analysis of the labeled image. Operation 510 can also include determining the confidence score associated with the identified segments. For example, it can be calculated based on the obtained pixel-level confidence scores. For example, the segment-level confidence score can be the average or median of the pixel-level confidence scores of the pixels in the segment (e.g., obtained as the output of the ML model).

[0096] At operation 512, the process may include determining whether the segments identified in the map data generated at operation 510 match segments corresponding to map elements in an existing map of the same scene. For example, if the existing map includes candidate segments having the same class label and matching spatial extent as the segments identified at operation 510, a match may be determined. As a non-limiting example, the spatial extents between two segments may be compared by calculating a match rate, e.g., the intersection between the two segments (e.g., the number of pixels in the overlapping region) divided by the union of the two segments (e.g., the number of pixels in the combined region covered by the two segments), and if the match rate is higher than a threshold (indicating significant overlap), the spatial extents may be considered to match. For each segment identified at operation 510, operation 512 may be performed on a per-segment basis.

[0097] If it is determined that the segments match (yes at 512), e.g., the segments have not changed and no update to the existing map is required, process 500 may control the autonomous vehicle at operation 514 at least in part based on the existing map. In some examples, operation 514 may include generating a route (e.g., a series of waypoints) between the current location and the destination on the existing map. Operation 514 may also include determining the vehicle's trajectory or otherwise controlling the vehicle to safely traverse the environment at least in part based on map elements (e.g., driving lanes, stop lines, turning lanes, etc.) indicated on the existing map. In an example, operation 514 may be performed by planning component 428.

[0098] Alternatively, if at operation 512 it is determined that one or more segments do not match the segments corresponding to map elements in the existing map, or the previous map does not exist (e.g., in the case where this is the first instance of data collection for the scene) (no at 512), then at operation 516, process 500 may determine whether a confidence score associated with the non-matching segments is greater than a minimum confidence threshold. If at operation 516 it is determined that the confidence score associated with the non-matching segments (e.g., segments that have changed) is greater than the minimum confidence threshold (yes at 516), then at operation 518, the existing map may be automatically updated by incorporating the changed segments into the existing map. Operation 518 may include changing the spatial extent of a segment, changing the semantic label of a segment, adding new segments and corresponding labels, and / or removing segments from the existing map. In some examples, there may be no existing map (e.g., for a new area of the environment), and all identified segments associated with a confidence score greater than the minimum confidence threshold may be added to a newly created map, which may become the existing map for future executions of process 500.

[0099] Alternatively, if at operation 516 it is determined that the confidence score associated with the mismatched segment is less than or equal to a minimum confidence threshold (No at 516), then at operation 520, the existing map can be updated based on manual inspection or other verification. In some cases, if it is determined that the identified segment is incorrect, the existing map may not be updated (e.g., the update is a no-op). In other cases, the existing map may be updated after manual inspection or verification, which may include modifying the identified segment (e.g., class label or spatial extent) prior to the update. Additionally, after automatically updating the existing map at operation 518, or after updating based on manual inspection at operation 520, process 500 can use the updated map to control the autonomous vehicle, thereby performing operation 514.

[0100] Example clause

[0101] A: An example system includes: one or more processors; and one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations including: receiving lidar data representing a portion of an environment; identifying a subset of the lidar data having an elevation above a ground plane threshold distance; generating filtered lidar data that does not include the subset of the lidar data; associating intensity information of the filtered lidar data with a voxel space; determining image data representing a top-down view of the voxel space based on the intensity information; inputting the image data into a machine learning (ML) model; receiving class information from the ML model, the class information including classes associated with individual pixels of the image data, wherein the classes indicate road map elements in a set of road map elements; determining a segment of the image data corresponding to a first road map element in the set of road map elements, the first road map element corresponding to an element on a driving surface, at least in part based on the class information; and updating a road network map of the environment at least in part based on the segment of the image data.

[0102] B: The system of example A, wherein: the lidar data is captured during multiple traversals of the environment by a lidar sensor mounted on a vehicle traversing the environment.

[0103] C: The system of example A or example B, wherein the intensity information includes an aggregation of the lidar data captured during the multiple traversals based on one of: cumulative intensity, maximum intensity, average intensity, or weighted average intensity.

[0104] D: A system according to any one of Examples A to C, wherein the ML model includes a fully convolutional network (FCN) previously trained using training data, the training data including images of a top-down view of a road network, wherein pixels of the images are labeled as road map elements belonging to a set of the road map elements.

[0105] E: A system according to any one of Examples A to D, wherein the set of road map elements includes one or more of the following: driving lane elements; turning lane elements; bicycle lane elements; crosswalk elements; sidewalk elements; intersection elements; lane separator elements; stop line elements; or yield line elements.

[0106] F: A system according to any one of Examples A to E, wherein updating the road network map includes: receiving an existing road network map for the portion of the environment, the existing road network map including an area indicating the first road map element; determining a confidence score associated with the segment of the image data corresponding to the first road map element; determining that the segment does not match the area based on comparing a spatial extent of the segment with the area; and updating the area indicating the first road map element at least in part based on the confidence score being greater than a threshold.

[0107] G: An example method includes: receiving sensor data related to a portion of an environment, wherein the sensor data is captured by sensors mounted on a plurality of vehicles traversing the portion of the environment; associating a portion of the sensor data with a voxel space representing the portion of the environment; determining a two-dimensional (2D) image representing a top-down view of the voxel space; determining a segmented image at least in part based on the 2D image and a plurality of semantic labels; and generating a road network map for the portion of the environment at least in part based on the segmented image.

[0108] H: The method according to Example G, wherein the portion of the sensor data corresponds to data associated with one or more of the following: a surface of the portion of the environment, the surface of the portion of the environment being at an elevation less than or equal to a vertical threshold from a ground plane, static elements of the portion of the environment, or high confidence values.

[0109] I: The method according to Example G or Example H, wherein: individual voxels of the voxel space correspond to respective physical volumes of the portion of the environment, and an individual voxel is associated with one of the following: an average intensity value, a cumulative intensity value, a maximum intensity value, or a weighted average intensity value of the sensor data associated with the respective physical volume.

[0110] J: A method according to any one of Examples G to I, wherein the plurality of semantic labels includes one or more road map elements on a drivable surface characterized by reflection markers.

[0111] K: A method according to any one of Examples G to J, wherein determining the segmented image includes: inputting the 2D image into a machine learning (ML) model; determining a confidence score associated with the segmentation of the segmented image based at least in part on an output from the ML model; and updating a road network map associated with the portion of the environment based on the confidence score being greater than a threshold.

[0112] L: A method according to any one of Examples G to K, wherein the ML model is a fully convolutional network (FCN) previously trained using training data, the training data including images of a top-down view of a road network labeled with the plurality of semantic labels.

[0113] M: A method according to any one of Examples G to L, further comprising: labeling the segmentation for verification based on the confidence score being less than the threshold; presenting the segmentation to an operator; receiving an input from the operator indicating the validity of the segmentation; and updating the road network map based on receiving the input.

[0114] N: A method according to any one of Examples G to M, wherein the sensor data includes one or more of the following: lidar data or radar data.

[0115] O: A method according to any one of Examples G to N, further comprising: transmitting the road network map to an autonomous vehicle configured to use the road network map to navigate in the environment.

[0116] P: An exemplary one or more non-transitory computer-readable media storing instructions executable by a processor, wherein the instructions when executed cause the processor to perform operations including: receiving a set of sensor data representing a portion of an environment, individual sensor data in the set of sensor data being associated with a corresponding voxel space at a corresponding 3D position; generating an aggregation of the set of sensor data of voxels mapped to a common voxel space; generating an image representing a two-dimensional (2D) projection of the common voxel space; inputting the image into a machine learning (ML) model, the machine learning (ML) model being trained to output one or more semantic labels associated with pixels of the input image; determining a segmented image associated with a segmentation of one or more semantic markers based on an output of the ML model; and controlling an autonomous vehicle traversing the portion of the environment based at least in part on the segmented image.

[0117] Q: According to one or more non-transitory computer-readable media of Example P, the operations further include: updating a road network map of the portion of the environment based on the segmented image, wherein the updating occurs at predefined intervals and is at least partially based on a difference between the segmented image and an existing road network map of the portion of the environment.

[0118] R: According to one or more non-transitory computer-readable media of Example P or Example Q, wherein the segmentation of the one or more semantic tags is associated with a corresponding confidence score, the confidence score being at least partially based on the output of the ML model, and based on the confidence score being less than a threshold, the updating is further based on input from a human operator.

[0119] S: According to one or more non-transitory computer-readable media of any one of Examples P to R, wherein: the sensor data includes one or more of the following: lidar data, radar data, or range data, the set of sensor data is captured by sensors mounted on a plurality of vehicles traversing the portion of the environment, and the 3D position of the respective voxel space of the individual sensor data in the set of sensor data is based on positioning information associated with the sensor at the time of capture of the sensor data.

[0120] T: According to one or more non-transitory computer-readable media of any one of Examples P to S, wherein the aggregation includes one of the following: an average intensity value, a cumulative intensity value, a maximum intensity value, or a weighted average intensity value of the set of sensor data.

[0121] Conclusion

[0122] Although one or more examples of the techniques described herein have been described, various changes, additions, permutations, and equivalents are included within the scope of the techniques described herein.

[0123] In the example description, reference is made to the accompanying drawings that form a part of this document, and the drawings show, by way of example, specific examples of the claimed subject matter. It should be understood that other examples may be used and changes or alterations may be made, such as structural changes. These examples, changes or alterations do not necessarily depart from the scope of the claimed subject matter. Although the steps in this document may be presented in a particular order, in some cases the order may be changed so as to provide certain inputs at different times or in a different order, without changing the function of the system and method. The disclosed procedures may also be executed in a different order. In addition, the various calculations described in this document do not have to be executed in the disclosed order, and other examples using an alternative order of calculations can be easily implemented. In addition to reordering, in some cases, a calculation can also be broken down into sub-calculations with the same result.

Claims

1. A method, comprising: receiving a set of sensor data representing a portion of an environment, wherein individual sensor data in the set of sensor data is associated with a corresponding voxel space at a corresponding 3D position; generating an aggregation of the set of sensor data of voxels mapped to a common voxel space; generating an image representing a two-dimensional (2D) projection of the common voxel space; inputting the image into a machine learning (ML) model, the machine learning (ML) model being trained to output one or more semantic labels associated with pixels of the input image; determining a segmented image associated with segmentation of one or more semantic labels based on the output of the ML model; and updating a road network map of the portion of the environment based on the segmented image.

2. The method according to claim 1, wherein the segmentation of the one or more semantic labels is associated with a corresponding confidence score, the confidence score being at least partially based on the output of the ML model, and based on the confidence score being less than a threshold, the updating is further based on input from a human operator.

3. The method according to claim 1 or claim 2, wherein: the sensor data includes one or more of the following: lidar data, radar data, or range data, and the 3D position of the corresponding voxel space of the individual sensor data in the set of sensor data is based on positioning information associated with the sensor at the time of capture of the sensor data.

4. The method according to any one of claims 1-3, wherein the aggregation includes one of the following: an average intensity value, a cumulative intensity value, a maximum intensity value, or a weighted average intensity value of the set of sensor data.

5. The method according to any one of claims 1-4, wherein the ML model is a fully convolutional network (FCN), the fully convolutional network (FCN) having been previously trained using training data including images of a top-down view of a road network labeled with the one or more semantic labels.

6. The method according to any one of claims 1-5, wherein the set of sensor data is captured by sensors mounted on a plurality of vehicles traversing the portion of the environment.

7. The method according to any one of claims 1-6, wherein the one or more semantic labels include road map elements on a drivable surface, and the road map elements include one or more of the following: driving lane elements; turning lane elements; bike lane elements; crosswalk elements; sidewalk elements; intersection elements; lane separation line elements; stop line elements; or yield line elements.

8. The method according to claim 7, wherein one or more of the road map elements are characterized by reflective markers.

9. The method according to any one of claims 1-8, wherein updating the road network map includes: Receive an existing road network map for a portion of the environment, the existing road network map including an area indicating a first road map element; Determine a confidence score associated with a segmentation of the segmented image corresponding to the first road map element; Determine that the segmentation does not match the area based on comparing a spatial extent of the segmentation with the area; and Update the area indicating the first road map element at least in part based on the confidence score being greater than a threshold.

10. The method according to any one of claims 1-9, wherein, Generating the aggregation of the set of sensor data includes: Associating a subset of the set of sensor data with a voxel space representing a portion of the environment, wherein the subset corresponds to sensor data associated with one or more of: A surface of a portion of the environment, the surface of the portion of the environment being at an elevation less than or equal to a vertical threshold from a ground plane, Static elements of the environment, or A confidence value equal to or higher than a threshold confidence value.

11. The method according to any one of claims 1-10, wherein, The sensor data includes lidar data, and the aggregation includes aggregating intensity information of the lidar data based on one of: cumulative intensity, maximum intensity, average intensity, or weighted average intensity.

12. The method according to any one of claims 1-11, wherein, The updating of the road network map occurs at a predefined interval and is at least in part based on a difference between the segmented image and an existing road network map of a portion of the environment.

13. The method according to any one of claims 1-12, further comprising: Controlling an autonomous vehicle traversing a portion of the environment at least in part based on the updated road network map.

14. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1-13.

15. A system, comprising: One or more processors; and One or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, configure the system to perform the method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Voxel based ground plane estimation and object segmentation

    US10444759B2

  • Sensor data segmentation

    US10535138B2

  • Data segmentation using masks

    US10649459B2

  • Generating maps without shadows

    US10699477B2

  • Modifying map elements associated with map data

    US10782136B2