Three-dimensional intersection structure prediction for autonomous driving application
By employing a DNN that utilizes 2D ground truth data and 3D geometric consistency constraints, the system effectively predicts 3D intersection structures from 2D image data, addressing the limitations of conventional methods and enhancing the accuracy and efficiency of autonomous driving systems.
Patent Information
- Application Number
- JP2025014556
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-12-09
- Filing Date
- 2025-01-31
- Publication Date
- 2025-06-10
AI Technical Summary
Conventional systems for autonomous driving struggle with accurately predicting 3D intersection structures from 2D image data, often relying on pre-stored HD maps or deep neural networks that require extensive manual labeling and annotation, which is time-consuming and labor-intensive.
A system and method using a deep neural network (DNN) to predict 3D intersection structures directly from 2D image data, employing a combination of 2D ground truth data and 3D geometric consistency constraints to improve accuracy and reduce the need for manual labeling.
The proposed solution enables accurate and efficient prediction of 3D intersection structures, improving scalability and reducing the complexity of map generation and annotation, while also addressing inaccuracies related to flat ground assumptions.
Smart Images

Figure 2025087679000001_ABST
Abstract
Description
Background Art
[0001] Autonomous driving systems and advanced driver assistance systems (ADAS) can utilize sensors such as cameras to perform various tasks - for example, lane keeping, lane changing, lane designation, camera calibration, steering, path planning, and localization. For example, in order for autonomous and ADAS systems to operate independently and efficiently, an understanding of the vehicle's surrounding environment can be made in real-time or near real-time. This understanding can include information regarding the position of objects, obstacles, lanes, and / or intersections within the environment relative to various demarcations, such as lanes, road boundaries, intersection lines, and / or the like. The information on the surrounding environment can be used by the vehicle when making decisions such as which path or trajectory to follow.
[0002] As an example, information regarding the position and layout of intersections within the environment of an autonomous or semi-autonomous vehicle can prove beneficial when making decisions regarding path planning, obstacle avoidance, and / or control - for example, where to stop, which path to use to safely cross an intersection, where other vehicles or pedestrians may be present, and / or the like. When the vehicle is operating in an urban and / or suburban driving environment, in such an environment, understanding the intersection scene and path planning become essential due to an increase in the number of variables compared to a highway driving environment, and information regarding the position and layout of intersections is particularly important. For example, when a vehicle is making a left turn at an intersection in a multi-lane, two-way driving environment, determining the position and directionality of other lanes, as well as determining the position of crosswalks or bike lanes, is essential for safe and efficient autonomous and / or semi-autonomous driving.
[0003] In conventional systems, intersections can be interpolated from a pre-stored high-definition (HD) three-dimensional (3D) map of the driving surface of a vehicle. For example, from the HD map, the structure and orientation of intersections and their surrounding areas can be collected. However, to use an HD map for intersection identification and navigation, it is necessary to pre-identify each intersection that a vehicle may encounter and record it in the HD map, which is a time-consuming and laborious task. For example, if manual labeling is required for a larger geographical area (e.g., a city, state, country) to enable a vehicle to drive efficiently independently in various regions, the map update process can become more logistically complex.
[0004] In other conventional systems, a deep neural network (DNN) can be trained to predict intersection information in a two-dimensional (2D) image space, and these predictions can be converted to 3D world space coordinates. However, the coordinate conversion from 2D to 3D is inherently inaccurate because it requires a flat ground assumption. Many roads - especially in more urban environments - are not flat, and if the slope or gradient of the driving surface is not considered, 2D predictions may not be accurately mapped to the 3D world space. To account for this, some DNNs are trained to predict intersection information in the 3D world space using 3D ground truth data. However, generating a sufficient amount of accurate and reliable (e.g., from LIDAR data) 3D ground truth data to effectively train a DNN is costly, and annotating LIDAR data to generate ground truth information is a difficult task. For example, identifying intersection lines, boundaries, and / or other information from LIDAR point clouds is unreliable and requires extensive human labeling and annotation for each instance of training data. As a result, the end-to-end training process for these conventional DNNs is quite time-consuming, and as a result, it may result in an inaccurate DNN that takes time to optimize for deployment in vehicles.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Means for Solving the Problems
[0006] Embodiments of the present disclosure relate to three-dimensional (3D) intersection structure prediction for autonomous driving applications. A system and method are disclosed for predicting 3D intersection structures in world space - such as entry lines, exit lines, pedestrian crossings, bicycle lanes, etc. - from two-dimensional (2D) image data in image space using a deep neural network (DNN). As such, in contrast to conventional approaches, the present system and method use the DNN to detect the 3D position of the intersection structure from 2D input data using the live perception capabilities of the sensor platform and utilize this information to provide techniques for generating a route for navigating an intersection, for example, for a vehicle or machine. This approach improves the scalability for handling various types of intersections without the burden of manually labeling each intersection individually for HD map generation, the inaccuracies resulting from calculating 3D intersection structures from 2D perception results using flat ground assumptions, and the challenges of generating and accurately labeling LIDAR point clouds for 3D ground truth generation. In some embodiments, to provide redundancy and further verify the results of the perception-based approach - for example, when high-quality map data is available - the live perception technique may be executed in combination with a map-based technique.
[0007] For example, to train a DNN to accurately predict a 3D intersection structure from a 2D image, the DNN can be trained using a first loss function corresponding to 2D ground truth data and a second loss function corresponding to 3D geometric consistency. The first loss function may compare the 3D output of the DNN, after being transformed into the 2D image space using the intrinsic and / or extrinsic parameters of the sensor, with the 2D ground truth data. The second loss function may analyze the 3D prediction of the DNN in consideration of one or more geometric constraints. For example, using the geometric knowledge of intersections, a penalty can be imposed on the predictions of the DNN that do not match the known geometric shapes. As such, a penalty can be imposed on the output if the 3D output of the DNN is not smooth, for example, because quantum leaps are impossible in the physical 3D real world. As another example, a straight-line constraint may be imposed such that a penalty is imposed on the prediction that the approach or exit line of an intersection is not straight or is within some threshold of being straight. Further, a lane width constraint may be imposed such that the range of possible known lane widths is forced onto the output of the DNN to impose a penalty on the prediction that it is outside some threshold range of the lane width. As such, when trained and deployed in a vehicle, the DNN can accurately predict the 3D intersection structure from 2D sensor data.
[0008] The present system and method for 3D intersection structure prediction for an autonomous driving application will be described in detail below with reference to the accompanying drawings.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8A
Figure 8B
Figure 8C
Figure 8D
Figure 9
Figure 10
DETAILED DESCRIPTION OF THE INVENTION
[0010] A system and method for three-dimensional (3D) intersection structure prediction for an autonomous driving application are disclosed. Although the present disclosure may be described with respect to an exemplary autonomous vehicle 800 (or referred to as "vehicle 800" or "ego vehicle 800", an example of which is described with respect to FIGS. 8A-8D), this is not intended to be limiting. For example, the systems and methods described herein can be used, without limitation, in non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robot platforms, warehouse vehicles, off-road vehicles, flying vessels, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, submarines, drones, and / or other vehicle types. Additionally, although the present disclosure may be described with respect to intersection structures for vehicle applications, this is not intended to be limiting, and the systems and methods described herein can be used in any other technical space where augmented reality, virtual reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or detection of intersection or other environmental structures and / or poses can be used.
[0011] During deployment, sensor data (e.g., 2D images, videos, etc.) can be received and / or generated using sensors (e.g., cameras, RADAR sensors, LIDAR sensors, etc.) located on autonomous or semi-autonomous vehicles or otherwise arranged. The sensor data is applied to a neural network (e.g., a deep neural network (DNN) such as a convolutional neural network (CNN)) trained to identify areas of interest related to road signs, road boundaries, intersections, and / or the like (e.g., raised pavement markings, rumble strips, colored lane dividers, sidewalks, crosswalks, side roads, etc.), as well as semantic information and / or directional information related thereto. More specifically, the DNN is trained to calculate a 3D intersection pose from two-dimensional (2D) input data (e.g., an image, a distance image, or other 2D sensor data representation) including one or more key points corresponding to line segments defining the 3D intersection pose, and also to generate an output that identifies semantic information (e.g., crosswalk, crosswalk entry, crosswalk exit, intersection entry, intersection exit, etc., or a combination thereof) and / or other information corresponding to the key points and / or line segments generated therefrom. In some examples, the calculated key points may be 3D world space positions where the center points and / or endpoints of intersection features or line segments such as pedestrian crossing (e.g., crosswalk) entry lines, intersection entry lines, intersection exit lines, stop lines, bicycle lanes, and / or pedestrian crossing exit lines are located. As such, the 3D structure and pose of an intersection can be determined using key points, line segments, and / or related (e.g., semantic) information, and the 3D structure can include representations of lanes, lane types, key points, line segments, crosswalks, directions, directions of travel, and / or other information corresponding to the intersection.
[0012] During training, the DNN can be trained using 2D images or other sensor data representations labeled or annotated with line segments representing lanes, crosswalks, entry lines, exit lines, bicycle lanes, intersection areas, etc., and may further include corresponding semantic information. In some examples, the key points can be labeled to include the corner points or end points of the lane (e.g., the line segments representing the lane segments of the intersection structure) that can be inferred from the center point and / or width information. As a result, since information can be determined using line segment annotations and semantic information, the intersection structure can be encoded using these key points and / or line segments that require only limited labeling. In some examples, the ground truth of the 2D intersection structure can be defined by a set of polygons corresponding to potential areas of interest within the intersection and the corresponding semantic information (e.g., inside the intersection, outside the intersection, etc.). Then, the labeled line segments and semantic information can be used to compare with the calculated output of the DNN (e.g., after conversion from 3D world space to 2D image space) - for example, using a loss function.
[0013] To train the DNN to predict the 3D intersection structure from a 2D input image, the 3D intersection structure prediction of the DNN can be projected into the 2D image space - for example, using the intrinsic and / or extrinsic parameters of the sensor that generated the 2D image. Then, a loss function can be used to measure the distance between the pixels of the 2D intersection ground truth data and the 2D projection of the 3D prediction. This distance can be minimized using a loss function to improve the accuracy or precision of the DNN. As such, the DNN can be trained to learn the mapping between the 2D image and the 3D intersection structure in the world space.
[0014] Furthermore, one or more 3D geometric consistency constraints can be applied to the 3D predictions of the DNN - for example, via a loss function - to train the DNN to more accurately predict the 3D intersection structure considering the geometric constraints. Knowledge about real-world intersections and their design can be used to determine the 3D geometric consistency constraints. For example, smoothness constraints, straight-line constraints, and / or statistical variability of lane widths may be used as geometric constraints for evaluating the output of the DNN. A loss function corresponding to the 3D geometric consistency constraints can be used to penalize outputs that do not conform to these constraints. In that case, the total loss can include the loss of 3D geometric consistency and the loss of 2D ground truth. The DNN can be trained to minimize the total loss to accurately and efficiently predict the 3D intersection structure based on the 2D image input.
[0015] Once the DNN is trained, the DNN can predict an output in the 3D world space corresponding to line segments, key points, intersection areas, 3D intersection structures, and / or other outputs corresponding to the 2D input data. In some embodiments, the 3D position of a key point can be determined as one or more endpoints of each line segment, the center point of a predicted line segment, or a combination thereof. The 3D positions of these key points can then be used - for example, by a post-processor - to construct a 3D intersection structure for use when a vehicle navigates or crosses an intersection.
[0016] In some embodiments, once the 3D positions of the key points are determined, any number of additional post - processing operations can be performed to finally "connect the dots (or key points)" and generate a path for navigating the intersection. For example, a polyline can be generated that connects the key points and represents a potential path for crossing the intersection. A path type that can be determined in relation to the vehicle's location, the positions of the key points, and / or the like can be assigned to the final path. Potential non - limiting path types include left turns, right turns, lane changes, and / or lane continuations. Curve fitting can also be implemented to determine the final shape that most accurately reflects the natural driving curve of the potential path. Curve fitting can be performed using polyline fitting, polynomial fitting, clothoid fitting, and / or other types of curve - fitting algorithms. The shape of the potential path can be determined based on the position, the direction - of - travel vector (e.g., angle), semantic information, and / or other information related to the connected key points. The curve - fitting process can be repeated for all key points that can potentially be connected to each other to generate all possible paths that the vehicle can follow to navigate the intersection. In some instances, infeasible paths can be excluded from consideration based on traffic rules and physical limitations associated with such paths.
[0017] In some embodiments, a matching algorithm can be used to connect the key points and generate a potential path for the vehicle to navigate the intersection. In such instances, a matching score can be determined for each pair of key points based on the position of the key points, the direction - of - travel vector, semantic information, and / or the shape of the approximate curve between the pair of key points. In some instances, a linear matching algorithm such as the Hungarian matching algorithm can be used. In other instances, a non - linear matching algorithm such as a spectral matching algorithm can be used to connect pairs of key points.
[0018] In any example, once the route passing through the intersection is determined, this information can be used to perform one or more operations by the vehicle. For example, the world model manager can update the world model to assist in navigating the intersection, the route planning layer of the autonomous driving software stack can use the intersection information to determine a route through the intersection (e.g., along one of the determined potential routes), and / or the control component can determine the control of the vehicle to navigate through the intersection according to the determined route.
[0019] Training of the DNN for calculating the 3D intersection structure Referring to FIG. 1, FIG. 1 is an exemplary data flow diagram showing an exemplary process 100 for training a deep neural network (DNN) 104 to detect intersections according to some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be excluded altogether. Further, many of the elements described herein are functional entities that can be implemented as individual or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be implemented by hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in a memory. In some embodiments, the training of the neural network by process 100 can be implemented using at least in part components, features, and / or functions similar to those described herein with respect to the exemplary computing device 900 of FIG. 9 and / or the exemplary data center 1000 of FIG. 10.
[0020] Process 100 may include generating and / or receiving sensor data 102 from one or more sensors. Sensor data 102 may be received from one or more sensors of a vehicle (e.g., vehicle 800 of FIGS. 8A - 8C described herein) as a non-limiting example. Sensor data 102 may be used within process 100 by a vehicle to train one or more DNNs 104 to detect 3D intersections from at least partially 2D input data (e.g., camera images). During training, sensor data 102 may be generated using one or more data collection vehicles that generate sensor data for training a DNN such as DNN 104, and / or may be pre-generated and included in a training data set. Sensor data 102 used during training may alternatively or additionally be generated using simulated sensor data (e.g., sensor data generated using one or more virtual sensors of a virtual vehicle within a virtual environment). When trained and deployed within vehicle 800, sensor data 102 may be generated by one or more sensors of vehicle 800 and processed by DNN 104 to compute 3D intersection structures, semantic information, and / or direction information corresponding to intersections.
[0021] As such, the sensor data 102 can include, but is not limited to, sensor data 102 from any of the vehicle's sensors such as, referring to FIGS. 8A - 8C, RADAR sensor 860, ultrasonic sensor 862, LIDAR sensor 864, stereo camera 868, wide view camera 870 (e.g., fisheye camera), infrared camera(s) 872, surround camera 874 (e.g., 360 - degree camera), long - range and / or mid - range camera 878, and / or other sensor types. As another example, the sensor data 102 can include virtual (e.g., simulated or augmented) sensor data generated from any number of sensors of a virtual vehicle or other virtual objects in a virtual (e.g., test) environment. In such an example, the virtual sensors can correspond to virtual vehicles or other virtual objects within a simulated environment (e.g., used to test, train, and / or validate neural network performance), and the virtual sensor data can represent sensor data captured by the virtual sensors within the simulated or virtual environment. As such, by using the virtual sensor data, the DNN 104 described herein can be tested, trained, and / or validated using simulated data or augmented data within a simulated environment, thereby enabling testing of more extreme scenarios outside of the real - world environment where such tests may be less safe.
[0022] In some embodiments, sensor data 102 may include sensor data representing image data representing an image, image data representing a video (e.g., a snapshot of a video), and / or a representation of the sensor's field of perception (e.g., a depth map of a LIDAR sensor, a value graph of an ultrasonic sensor, etc.). When the sensor data 102 includes image data, any type of image data format may be used, such as, without limitation, JPEG (Joint Photographic Experts Group) or luminance / chrominance (YUV) format, compressed images such as frame streaming resulting from compressed video formats such as H.264 / AVC (Advanced Video Coding) or H.265 / HEVC (High Efficiency Video Coding), raw images derived from RCCB (Red Clear Blue), RCCC (Red Clear), or other types of image sensors, and / or other formats. Additionally, in some instances, the sensor data 102 may be used within process 100 without preprocessing (e.g., in the raw or captured format), and in other instances, the sensor data 102 may receive preprocessing (e.g., noise balancing, demosaicing, scaling, cropping, expanding, white balancing, tone curve adjustment, etc., e.g., using a sensor data preprocessor (not shown)). As used herein, the sensor data 102 may refer to unprocessed sensor data, preprocessed sensor data, or a combination thereof.
[0023] The sensor data 102 used for training may include the first image (e.g., as captured by one or more image sensors), the downsampled image, the upsampled image, the cropped or region of interest (ROI) image, the image extended in other ways, and / or a combination thereof. The DNN 104 can be trained using the image (and / or other sensor data 102) and the corresponding ground truth data - for example, 2D ground truth data 116. The ground truth data may include annotations, labels, masks, and / or the like. For example, in some embodiments, the 2D ground truth data may correspond to the annotation of the intersection area or line segment, the classification of the intersection area or line segment, and / or the related direction information - as described, for example, with respect to FIGS. 2A-2B and FIGS. 3A-3B. In some embodiments, the 2D ground truth data 116 may be similar to the ground truth data described in Patent Document 1 filed on April 14, 2020, Patent Document 2 filed on June 24, 2020, and / or Patent Document 3 filed on March 10, 2020, which are each incorporated herein by reference in their entirety.
[0024] When annotations are used to generate 2D ground truth data 116, in some examples, the annotations can be generated within a drawing program (e.g., an annotation program), a computer aided design (CAD) program, a labeling program, or another type of program suitable for generating annotations, and / or can be hand-drawn. In any example, the 2D ground truth data 116 can be synthetically generated (e.g., generated from a computer model or rendering), can be actually generated (e.g., designed and generated from real-world data), can be machine automated (e.g., using feature analysis and learning to extract features from data and then generate labels), can be manually annotated by humans (e.g., a labeler or annotation expert defines the position of the label), and / or can be a combination thereof (e.g., a human identifies the center point or origin and the dimensions of the area, and a machine generates the polygon and / or label of the intersection area).
[0025] The intersection area can include an annotation corresponding to a boundary shape (e.g., a polygon) depicting the area of interest of the intersection, or other label types. In some examples, the intersection area can be depicted in the sensor data 102 - e.g., within the sensor data representation of the sensor data 102 - by one or more polygons corresponding to, for example, a pedestrian intersection area, an intersection entry area, an intersection exit area, an ambiguous area, a lane-less area, an intersection interior area, a partially visible area, an overall visible area, etc. The polygon can be generated as a bounding box. The semantic information (e.g., classification) corresponding to the 2D ground truth data 116 can be generated for each image represented by the sensor data 102 used to train the DNN 104, and / or for each one or more polygons within the image. The number of classifications can correspond to the number and / or type of features that the DNN 104 is trained to predict, or the number and / or type of features of the intersection area within each respective image.
[0026] Depending on the embodiment, the semantic information may correspond to a classification or tag corresponding to a feature type or class of intersection areas, such as, but not limited to, a pedestrian intersection area, an intersection entry area, an intersection exit area, an ambiguous area, an area without lanes, an area inside the intersection, a partially visible area, an entirely visible area, and / or the like. In some examples, the classification may first correspond to an area inside the intersection and / or an area outside the intersection. The area inside the intersection classification may refer to an intersection area that includes an area within the intersection where the paths of vehicles crossing the intersection in various directions may intersect. The area outside the intersection classification may refer to an intersection area that includes an area outside the area inside the intersection.
[0027] An intersection area classified as an intersection external area may be further labeled with a classification corresponding to an attribute corresponding to a characteristic type of an intersection exit area, including but not limited to a pedestrian crosswalk area, an intersection entry area, an intersection exit area, an ambiguous area, a no-lane area, and / or the like. Specifically, the attribute of intersection entry may correspond to an intersection area where one or more vehicles are attempting to enter the corresponding intersection from various different directions. The intersection exit area may correspond to an intersection area where one or more vehicles that have most recently exited the intersection in various directions may be located. It should be understood that information regarding the intersection exit area may be particularly important because vehicle 800 must safely cross the intersection exit area in order to safely cross the intersection. Similarly, the pedestrian crosswalk area may refer to an intersection area corresponding to a pedestrian crosswalk located outside the intersection interior area. An area classified as a "no-lane area" may correspond to an intersection area where vehicles are not permitted to cross, such as a bicycle lane, a pedestrian path, and / or the like. The "ambiguous area" attribute may correspond to an intersection area where the driving direction of a vehicle is ambiguous. In addition, the classification of the intersection interior area and the intersection external area class may also include one of the attributes of an area that is entirely visible and / or an area that is partially visible. In an example, when the classification includes an attribute or class label of an area that is entirely visible, the corresponding intersection area may include, for example, an entirely visible surface without obstacles. In contrast, when the classification includes an attribute or class label of an area that is partially visible, the corresponding intersection area may include obstacles such as a shield where the driving surface within the area is only partially visible in the corresponding sensor data 102. The labeling ontology described herein is for illustrative purposes only, and additional and / or alternative class labels may be used without departing from the scope of the present disclosure.
[0028] Referring to FIGS. 2A-2B as non-limiting examples, FIGS. 2A-2B show exemplary annotations corresponding to sensor data 102 for use in ground-truth generation to train a DNN to detect intersections, according to some embodiments of the present disclosure. For example, FIG. 2A shows an exemplary labeling of an image 200A (e.g., corresponding to semantic information) that can be used to generate ground-truth data according to the training process 100 of FIG. 1. The intersection area or crossing region within the image can be annotated by intersection areas (e.g., areas 204A, 204B, 206, 208, 210A, and 210B) and corresponding classifications (e.g., "inside intersection", "partially visible", "vehicle exit", "vehicle entry", "partially visible", "pedestrian crosswalk", etc.). For example, intersection area 204A can be labeled using a polygon and classified as having one or more attributes such as "intersection entry" and "partially visible". Similarly, intersection areas 204B, 206, 208, 210A, and 210B can also be labeled using polygons, intersection 204B can be classified as having one or more attributes such as "vehicle entry" and "partially visible", intersection area 206 can be classified as having one or more attributes such as "pedestrian crosswalk" and "partially visible", intersection area 208 can be classified as having one or more attributes such as "inside intersection" and "partially visible", intersection area 210A can be classified as having one or more attributes such as "vehicle exit" and "partially visible", and intersection area 210B can be classified as having one or more attributes such as "vehicle exit" and "fully visible". In some examples, each intersection area belonging to a common class or classification can also be annotated with a polygon of a matching color (or some other visual representation of the semantic information). For example, the polygons of intersection areas 204A and 204B can be of the same color and / or style since they are both classified as vehicle entry classifications.Similarly, the polygons of intersection areas 210A and 210B can be annotated using the same color and / or style since they are both classified as vehicle exit classifications. These labelings or annotation styles can be recognized by system 100 as corresponding to specific classes, and this information can be used to generate encoded ground truth data for training DNN 104.
[0029] Referring now to FIG. 2B, FIG. 2B shows another example of annotations applied to sensor data to train DNN 104 to detect intersection areas, according to some embodiments of the present invention. As shown herein, intersection areas 222A-222C, 224A-224C, 226A-226B, 228A-228B, and 230 can be annotated with polygons and corresponding classifications (e.g., "inside intersection", "partially visible", "vehicle exit", "vehicle entry", "partially visible", "pedestrian crosswalk", etc.). For example, intersection areas 222A, 222B, and 222C can be labeled using polygons of the same color and / or style and classified as one or more of "vehicle entry" and "partially visible". Similarly, intersection areas 224A, 224B, and 224C can be labeled using polygons of the same color and / or style and classified as one or more of "pedestrian crosswalk", "fully visible", and "partially visible". Intersection areas 226A and 226B can be labeled using polygons of the same color and / or style and classified as one or more of "no lane", "fully visible", and "partially visible". Intersection areas 228A and 228B can be labeled using polygons of the same color and / or style and classified as one or more of "vehicle exit", "fully visible", and "partially visible". Intersection area 230 can be labeled using a polygon and classified as one or more of "inside intersection" and "partially visible".
[0030] Annotations can be in a similar visual representation for the same classification. As shown in the illustration, the intersection areas 222A, 222B, and 222C can be classified as vehicle exit areas. In this way, features similarly classified in the image can be annotated in a similar manner. Note that the classification can be a compound noun. In Figure 2B, different classification labels can be represented by solid lines, dashed lines, etc. to represent different classifications. Further, the different classification labels can be nouns and / or compound nouns. This is not intended to be limiting, and any naming convention for the classification can be used to explain the differences in the classification labels of features (e.g., intersection areas) within the image.
[0031] As an additional or alternative option for generating the 2D ground truth data 116, lane (or line) labels can be generated in relation to the sensor data 102. For example, annotations or other label types corresponding to features corresponding to intersections or areas of interest can be generated. In some examples, the intersection structure can be defined as a set of line segments corresponding to lanes, crosswalks, entry lines, exit lines, bicycle lanes, etc. within the sensor data 102. The line segments can be generated as polylines, and the center of each polyline is defined as the center of the corresponding line segment. Semantic information (e.g., classification) can be generated for each image (and / or other sensor data representation) represented by the sensor data 102 used to train the DNN 104, and / or for one or more of the line segments and centers within the image. Similar to the above, the number of classifications can correspond to the number and / or type of features that the DNN 104 is trained to predict, or the number and / or type of lanes and / or features within each image. Depending on the embodiment, the classification can correspond to a classification or tag corresponding to a feature type, such as, but not limited to, crosswalk, crosswalk entry, crosswalk exit, intersection entry, intersection exit, and / or bicycle lane.
[0032] In some examples, the intersection structure can be determined based on the annotations. In such examples, a set of key points can be determined from the lane labels, where each key point corresponds to the center (or left end, or right end, etc.) of a corresponding line segment extending across the lane. The key points are mainly described with respect to the center points of the lane segments, but this is not intended to be limiting, and in some examples, the corners or endpoints of each lane can also be determined as key points for each instance of the sensor data 102. For example, the corners or endpoints of each lane can be inferred from the center key points and the directionality of the lane, or from the lane label or the annotation itself. Additionally, the number of lanes or line segments, and the direction of travel, directionality, width, and / or other geometry corresponding to each line segment from each lane can be determined from the annotation - for example, from the lane label and classification. As a result, even if the annotation does not directly indicate a specific intersection structure or pose information - such as direction of travel, lane width, and / or lane directionality - the annotation can be analyzed or processed to determine this information.
[0033] For example, if the first lane label extends along the width of the lane and includes the classification "pedestrian crossing entry", and the second lane label extends along the same width of the lane and includes the classification "pedestrian crossing exit_intersection entry", this information can be used to determine the direction of travel of the lane (e.g., from the first lane label to the second lane label) (e.g., the vehicle travels in the direction across the first lane label towards the second lane label). Additionally, the lane label may indicate the directionality (e.g., angle) of the lane, and from this information, the direction of travel (e.g., angle) can be determined - such as by calculating the normal of the lane label.
[0034] Referring now to FIGS. 3A-3B, FIGS. 3A-3B show exemplary annotations applied to sensor data 102 for use in generating 2D ground truth data 116 for training DNN 104 to detect intersection structures and poses according to some embodiments of the present disclosure. For example, FIG. 3A shows an exemplary labeling (e.g., corresponding to an annotation) of an image 300A that can be used to generate 2D ground truth data 116 according to the training process 100 of FIG. 1. Lanes in the image can be annotated with lane labels (e.g., lanes 304, 306, 310, and 312) and corresponding classifications (e.g., pedestrian entry, intersection exit, intersection entry, pedestrian exit, no lane). For example, lane 304 can be labeled using line segments and classified as one or more of intersection entry and pedestrian entry. Similarly, lanes 306, 308, and 310 can also be labeled using line segments, lane 306 can be classified as one or more of intersection entry and pedestrian exit, lane 308 can be classified as one or more of intersection entry and pedestrian exit, and lane 312 can be classified as no lane.
[0035] Furthermore, the labels of lanes 304, 306, and 308 can be further annotated with the corresponding direction of travel, as indicated by arrows 302A-302V. The direction of travel can represent the direction of traffic associated with a particular lane. In some examples, the direction of travel can be associated with the center (or key) point of its corresponding lane label. For example, direction of travel 302S can be associated with the center point of lane 304. In FIG. 3A, different classification labels can be represented by different line types - e.g., solid lines, dashed lines, etc. - to represent different classifications. However, this is not intended to be limiting, and any visualization of lane labels and their classifications can include different shapes, patterns, fillings, colors, symbols, and / or other identifiers to indicate differences in the classification labels of features (e.g., lanes) within the image.
[0036] Referring now to FIG. 3B, FIG. 3B shows another example of annotations applied to sensor data for training a machine learning model to detect intersection structures and poses, according to some embodiments of the present invention. Here, lanes 322A and 322B within image 300B can be annotated with line segments. The line segment corresponding to lane 322B can be annotated to extend onto the vehicle. This can help train DNN 104 to predict the position of key points even when the actual position may be occluded. As a result, the presence of vehicles or other objects within the lane may not compromise the ability of the system to generate a proposed path through the intersection. Annotations can be similar visual representations for the same classification. As shown, lanes 322A and 322B can be classified as intersection approach line stop lines. In this way, features similarly classified in the image can be annotated in a similar manner. Further, note that the classification can be represented using compound nouns. In FIG. 3B, different classification labels can be represented by solid lines, dashed lines, etc. to represent different classifications. Further, different classification labels can be nouns and / or compound nouns. This is not intended to be limiting, and any naming convention for classification may be used to explain the differences in classification labels for features (e.g., lanes) within the image.
[0037] The encoder may be configured to encode 2D ground truth data 116 corresponding to an intersection structure and pose, using annotations. For example, as described herein, even if the annotations may be limited to lane labels and classifications, information such as key points, the number of lanes, the direction of travel, directionality, and / or other structural and pose information may be determined from the annotations. Once this information is determined, the information may be encoded by the encoder to generate 2D ground truth data 116. For example, the direction-of-travel angle corresponding to a lane may be determined using the normal of the direction of the line segment corresponding to the lane, and the direction of travel may be determined using semantic or classification information (e.g., if a line segment corresponds to a crosswalk entry and the next line segment after that line segment corresponds to a crosswalk exit and intersection entry, the direction of travel may be determined to be towards the intersection, from the line segment to the next line segment). Once the directionality (e.g., the angle of a lane or other geometry) is determined from the annotations, a direction vector may be determined and attributed to key points representing line segment-line segment (e.g., a center key point). Similarly, once the direction of travel is determined (e.g., from the directionality - as its normal - or otherwise), a direction-of-travel vector may be determined and attributed to key points representing line segment-line segment (e.g., a center key point). For example, once the direction of travel (e.g., the angle corresponding to the direction of travel of a vehicle along a lane) is determined, a direction-of-travel vector may be determined and attributed to key points representing line segment-line segment (e.g., a center key point).
[0038] For each instance of sensor data 102 (e.g., for each image if the sensor data 102 includes image data), when 2D ground truth data 116 is generated, the DNN 104 can be trained to directly compute a 3D intersection structure (including the position of the intersecting line segments and the associated semantic information) using the 2D ground truth data 116. For example, the DNN 104 can generate an output 106, and the output 106 can be compared with the 2D ground truth data 116 corresponding to each instance of the sensor data 102 and / or the 3D geometric consistency constraint 118 using a loss function 114. For example, with respect to the 2D ground truth data 116, the 3D output 108 of the DNN 104 can be converted to 2D space (e.g., using a 3D-to-2D converter 112) for comparison with the 2D ground truth data 116 using one or more loss functions 114.
[0039] DNN 104 can initially output arbitrarily initialized values of (x, y, z) coordinates in the 3D world space (e.g., 3D output 108), which will be improved over time using the loss function 114. For example, the 3D output 108 may correspond to the 3D world space positions of key points corresponding to line segments depicting intersections depicted in the sensor data 102 (e.g., 2D camera images, or other 2D image data representations). As such, to generate line segments corresponding to intersections, center key points and / or end key points - each including the corresponding (x, y, z) positions in the 3D world space based on an origin on the vehicle 800 (e.g., a point on the vehicle such as the center of the axle, the foremost point of the vehicle, the uppermost point of the vehicle, etc.) - may be used (e.g., connected). For example, the 3D output 108 may be calculated as the reliability of each point in the 3D world space (e.g., each (x, y, z) coordinate in the target space or design space) as to whether the point corresponds to a key point (and / or line segment). When incorporating semantic information 110, each point in the 3D world space may have a reliability associated with each type of semantic class, and a threshold reliability may be used to exclude points that do not have a semantic class exceeding the threshold reliability. As such, if a point has a reliability exceeding the threshold for any class, it can be predicted that there is a key point (or line segment) at that position and the key point (or line segment) can be of the class exceeding the threshold reliability.
[0040] However, since it may be difficult to generate a sufficiently large amount of accurate and reliable 3D ground truth data, process 100 can train DNN 104 using 2D ground truth data 116. As such, the predicted positions in the 3D world space of the cross key points - or the corresponding line segments - can be transformed into 2D space using a 3D-to-2D converter. For example, the 3D-to-2D converter uses the intrinsic parameters (e.g., optical center, focal length, asymmetry coefficient, etc.) and / or extrinsic parameters (e.g., sensor position (e.g., relative to the origin of vehicle 800), rotation, translation, transformation from 3D world space to 3D camera coordinate system, etc.) of the sensor that generated the instance of sensor data 102 to transform values from 3D world space to 2D image space.
[0041] Next, the position of the intersection key point (or the line segment constructed therefrom) converted from 3D to 2D can be compared with the known 2D position of the key point (or line segment) in the 2D ground truth data 116. For example, the 2D ground truth data 116 may include a set of polyline segments with corresponding semantic information and / or orientation information (as described with respect to FIGS. 3A-3B, for example), and / or a set of polygon areas indicating potential conflict areas at intersections with corresponding semantic information (as described with respect to FIGS. 2A-2B, for example). In some embodiments, in addition to or alternatively to the 2D ground truth data 116 format described herein, other 2D ground truth data 116 formats or styles can be used without departing from the scope of the present disclosure. In any embodiment, the 3D coordinate prediction of the DNN 104 (e.g., the 3D output 108) can be projected into the same 2D (e.g., image) space as the 2D ground truth data 116, and a loss function 114 can be used to compare the distance between the 2D ground truth data 116 and the converted 2D position of the intersection feature. By minimizing the distance, the accuracy of the prediction of the DNN 104 is improved, and over time, the DNN 104 learns the mapping between the 2D intersection structure information extracted from the (e.g., 2D) sensor data 102 and the 3D intersection structure in the 3D world space. Additionally, for each key point and / or line segment, the semantic information 110 output by the DNN 104 can be compared with the semantic information from the 2D ground truth data 116. Further, the 3D output 108 can include a direction vector, and / or the semantic information 110 can indicate the direction corresponding to the key point and / or line segment. As such, this information from the output 106 can be used to determine the direction associated with the key point and / or line segment that can be compared with the direction vector and / or the forward direction information from the 2D ground truth data 116 (as described herein with respect to FIGS. 3A-3B, etc.).
[0042] Referring to FIG. 4 as an example, FIG. 4 is an exemplary illustration of converting a 3D output 108 into a 2D space for training a neural network according to some embodiments of the present disclosure. For example, the 3D output 108A may correspond to a visualization of the 3D output 108 of the DNN 104 corresponding to an instance of the sensor data 102 - for example, an instance corresponding to the 2D ground truth data 116 of the images 300A in FIGS. 3A and 4. The 3D output 108A may include an origin (0, 0, 0) corresponding to a certain origin on the vehicle, and may extend horizontally from left to right up to the maximum range of the design space or target space, longitudinally back and forth up to the range of the design space, and vertically up and down up to the range of the design space (not depicted as it is a top-down view). In a non-limiting example, the design space may be 40 meters wide (e.g., the x-axis), 100 meters long (e.g., the y-axis), and 3 meters high (e.g., the z-axis). However, any design space may be used. The design space may be selected based on an analysis of a plurality (e.g., hundreds, thousands, etc.) of intersections to determine a representative value, average value, or other value corresponding to the dimensions of the intersection, such that the design space is likely to be at least as large as any intersection the vehicle may cross. In the exemplary visualization of FIG. 4, the shown horizontal (x-axis) range extends from -16 meters to +16 meters, the shown vertical (y-axis) range extends from 0 meters to 28 meters, and the vertical range (not depicted as it is a top-down visualization) extends from 0 to 3 meters. As such, the 3D output 108 of the DNN 104 can be predicted within the design space. In the example of FIG. 4, as shown in the visualization 400, the 3D output can be converted into a 2D image space using a 3D-to-2D converter 112. For example, the visualization 400 may include the 2D-converted prediction of the DNN 104 overlaid on the image 300A, and the image 300A may include the 2D ground truth data 116 corresponding to the image. As such, the loss function 114 corresponding to the 2D ground truth data 116 can be used to compare the distance and / or semantic information of the converted output of the DNN 104 with the 2D ground truth data 116.As shown in FIG. 4, the 2D-transformed output indicates that there is inaccuracy within the dashed region 402. For example, line segments 404A-404D in the 3D-to-2D transformed output of visualization 400 do not match the respective line segments 404A-404D in image 300A in terms of direction. Similarly, the semantic information associated with line segment 404B does not match the semantic information of line segment 404B from the 2D ground truth data 116. As such, the loss function 114 can be used to calculate the distance and / or difference between the position, orientation, semantic information, and / or other information of the 3D-to-2D transformed output and the 2D ground truth data 116.
[0043] In some embodiments, in addition to or alternatively to the loss function of the 2D ground truth data 116, the loss function of the 3D geometric consistency constraint 118 can be used. For example, based on the empirical knowledge of real-world intersections and intersection designs, various 3D geometric constraints can be determined and used for comparison with the computed 3D output 108 of the DNN 104. For example, at a macro scale, since there is no quantum leap to the physical 3D real world, the 3D geometric consistency constraint 118 can include a smoothness term. As such, 3D outputs 108 that include non-smooth predictions - for example, completely separated line segments, line segments offset from each other in any of the x, y, and / or z directions, and / or other non-smooth line segments - can be penalized using the loss function 114 corresponding to 3D geometric consistency. Another 3D geometric consistency constraint 118 can include a straight-line constraint based on the knowledge that real-world intersections are designed to have straight entry / exit lines and that the line segments of intersections are generally linear. As such, 3D outputs 108 that include non-straight segments (for example, after generating line segments using key points) can be penalized using the loss function 114 corresponding to 3D geometric consistency. Similarly, the statistical variability of lane widths can be used as a 3D geometric consistency constraint 118, such that in an example, penalties can be imposed on line segments longer than a threshold length due to the statistical variability of lane widths, penalties can be imposed on line segments shorter than the threshold length due to the statistical variability of lane widths, and the greater the difference, the more penalties can be imposed. For example, in the United States, the average lane width may be 3.5 meters, and thus, as a non-limiting example, the tolerance range can be from 3.25 meters to 3.75 meters due to statistical variability. As such, if the 3D output 108 shows lanes of 2.5 meters or 5 meters, these outputs can be penalized using the loss function 114 of the 3D geometric consistency constraint 118. The 3D geometric consistency constraints 118 described herein are not intended to be limiting, and additional or alternative constraints can be used without departing from the scope of the present disclosure.
[0044] As such, until the DNN 104 converges to an acceptable or desirable accuracy, the feedback from the loss function 114 can be used to update the parameters (e.g., weights and biases) of the DNN 104 considering the 2D ground truth data 116 and / or the 3D geometric consistency constraints 118. By using the process 100, the DNN 104 can be trained to accurately predict the output 106 - e.g., the 3D output 108 and / or the semantic information 110 - from the sensor data 102 using the loss function 114, the 2D ground truth data 116, and the 3D geometric consistency constraints. As described herein, in some examples, the DNN 104 can be trained to predict different outputs 106 using different loss functions 114. For example, the first loss function 114 can be used to compare the output 106 converted from 3D to 2D with the 2D ground truth data, and the second loss function 114 can be used to compare the 3D output with the 3D geometric consistency constraints 118. In such examples, two or more loss functions 114 can be weighted (e.g., similarly, in different ways, etc.) to generate the final loss value that can be used to update the parameters of the DNN 104.
[0045] Referring now to FIG. 5, each block of method 500 described herein may include a computing process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in memory. Method 500 may also be implemented as computer-usable instructions stored on a computer storage medium. Method 500 may be provided, for example, as a stand-alone application, service, or hosted service (either stand-alone or in combination with another hosted service), or as a plug-in to another product. Further, method 500 is described, by way of example, with respect to process 100 of FIG. 1. However, these methods may be performed additionally or alternatively by any one process and / or any one system, or any combination of processes and / or systems, including but not limited to those described herein.
[0046] FIG. 5 is a flow diagram illustrating an exemplary method 500 for training a neural network to detect intersections, according to some embodiments of the present disclosure. Method 500 includes, at block B502, applying sensor data representing a sensor data representation in a 2D space to a neural network. For example, sensor data 102 representing an image (or other sensor data representation) may be applied to DNN 104.
[0047] Method 500 includes, at block B504, using the neural network to compute data representing a 3D world space position, at least in part based on the sensor data. For example, DNN 104 may compute an output 106 - specifically, a 3D output 108 in the example - after processing sensor data 102.
[0048] Method 500 includes, at block B506, converting a 3D world space position into a 2D space to generate a 2D space position. For example, the 3D-to-2D converter 112 can convert the 3D output 108 into a 2D (image) space using the intrinsic and / or extrinsic parameters of the sensor that generated the sensor data 102.
[0049] Method 500 includes, at block B508, using a first loss function to compare the 2D space position with a 2D space ground truth position associated with the sensor data. For example, the 3D-to-2D converted output of the DNN 104 can be compared with the 2D ground truth data 116 using the first loss function 114.
[0050] Method 500 includes, at block B510, using a second loss function to compare the 3D world space position with one or more geometric consistency constraints. For example, the 3D output 108 of the DNN 104 can be compared with the 3D geometric consistency constraints 118 using the second loss function.
[0051] Method 500 includes, at block B512, updating one or more parameters of the neural network based at least in part on the comparison of the 2D space position with the 2D space ground truth position and the comparison of the 3D world space position with one or more geometric consistency constraints. For example, the DNN 104 can be updated based on the outputs of the first loss function and the second loss function using a training engine or optimizer. This process can be repeated for each instance of the sensor data 102 used to train the DNN 104 until the DNN converges to an acceptable level of accuracy.
[0052] Expansion of the DNN for calculating the 3D intersection structure Referring now to FIG. 6A, FIG. 6 is a data flow diagram showing an exemplary process 600 for detecting an intersection structure in a 3D world space using a neural network, according to some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be excluded altogether. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory. In some embodiments, the training of the neural network 100 by the process may be implemented using at least partially the same components, features, and / or functions as described herein with respect to the exemplary computing device 900 of FIG. 9 and / or the exemplary data center 1000 of FIG. 10.
[0053] Process 600 may include receiving and / or generating sensor data 102. For example, sensor data 102 may be similar to the sensor data 102 described with respect to FIG. 1. For example, sensor data 102 may be generated during operation of vehicle 800 using one or more sensor types of vehicle 800. Sensor data 102 - for example, 2D sensor data - may be applied to DNN 104, which may compute output 106. Output 106 may include, as described herein with respect to FIG. 1, 3D output 108, semantic information corresponding to the 3D output, and / or directional information corresponding to the 3D output (e.g., direction and / or direction of travel associated with each keypoint and / or line segment).
[0054] DNN 104 can calculate an output 106 using sensor data 102, and the output 106 can ultimately be applied to a decoder or one or more other post - processing components to generate key points, classifications, number of lanes, lane travel directions, lane orientations, and / or other information. In this specification, examples are described with respect to the use of deep neural networks (DNNs), specifically convolutional neural networks (CNNs), but this is not intended to be limiting. For example, without limitation, DNN 104 can be a linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k - nearest neighbor (Knn), K - means clustering, random forest, dimensionality reduction algorithm, gradient boosting algorithm, neural network (e.g., auto - encoder, convolutional, recurrent, perceptron, long / short term memory / LSTM, Hopfield, Boltzmann, deep belief, de - convolutional, adversarial generation, liquid state machine, etc.), lane detection algorithm, computer vision algorithm - using machine learning model, and / or other types of machine learning models, including any type of machine learning model.
[0055] As an example, for instance, if DNN 104 includes a CNN, DNN 104 can include any number of layers. One or more layers can include an input layer. The input layer can hold values related to sensor data 102 (e.g., before or after post - processing). For example, if sensor data 102 is an image, the input layer can hold values representing the raw pixel values of the image as a volume (e.g., width, height, and color channels (e.g., RGB) such as 32×32×3).
[0056] One or more layers may include a convolutional layer. The convolutional layer can compute the output of neurons connected to local regions within the input layer, and each neuron computes the inner product between its weights and the small region within the input volume to which it is connected. The result of the convolutional layer can be another volume where one of the dimensions is based on the number of filters applied (e.g., if the number of filters is 12, the width, height, and number of filters such as 32×32×12).
[0057] One or more layers may include a deconvolutional layer (or transposed convolutional layer). For example, the result of the deconvolutional layer can be another volume having a higher dimension than the input dimension of the data received by the deconvolutional layer.
[0058] One or more layers may include a rectified linear unit (ReLU) layer. The ReLU layer can apply an element-wise activation function such as max(0,x) and threshold at zero. The resulting volume of the ReLU layer can be the same as the volume of the input to the ReLU layer.
[0059] One or more layers may include a pooling layer. The pooling layer can perform a down-sampling operation along the spatial dimensions (e.g., height and width), resulting in a volume smaller than the input to the pooling layer (e.g., an input volume of 32×32×12 to 16×16×12).
[0060] One or more layers may include one or more fully-connected layers. Each neuron within a fully-connected layer may be connected to each neuron in the previous volume. The fully-connected layer can compute class scores, and the resulting volume can be 1×1×number of classes. In some examples, the CNN may include a fully-connected layer such that the output of one or more layers of the CNN can be provided as input to the fully-connected layer of the CNN. In some examples, one or more convolutional streams may be implemented by DNN104, and some or all of the convolutional streams may each include a respective fully-connected layer.
[0061] In some non-limiting examples, DNN104 may include a series of convolutional layers and max-pooling layers to facilitate image feature extraction, followed by multi-scale extended convolutional layers and upsampling layers to facilitate global context feature extraction.
[0062] Although an input layer, convolutional layer, pooling layer, ReLU layer, and fully-connected layer are described herein with respect to DNN104, this is not intended to be limiting. For example, additional or alternative layers such as normalization layers, softmax layers, and / or other layer types may be used in DNN104.
[0063] In embodiments where DNN104 includes a CNN, different orders and numbers of layers of the CNN may be used depending on the embodiment. In other words, the order and number of layers of DNN104 are not limited to any one architecture.
[0064] In addition, some layers such as the convolutional layer and the fully connected layer may contain parameters (e.g., weights and / or biases), while other layers such as the ReLU layer and the pooling layer may not contain parameters. In some examples, the parameters can be learned by the DNN 104 during training. Further, some layers such as the convolutional layer, the fully connected layer, and the pooling layer may contain additional hyperparameters (e.g., learning rate, stride, epoch, etc.), while other layers such as the ReLU layer may not contain additional hyperparameters. The parameters and hyperparameters are not limited and may vary according to the embodiments.
[0065] As an example of a DNN, FIG. 6B shows an exemplary DNN 104A for use in calculating an intersection structure according to some embodiments of the present disclosure. For example, DNN 104A may include an encoder-decoder type DNN 104 - for example, a first 2D encoder network 608 and a second 3D decoder network 612. For example, during training, sensor data 102 - for example, input image 602 - may pass through a set of convolutional layers of a 2D encoder network 608 that learns latent variables (e.g., descriptors of the entire image or other sensor data representations) represented in a latent space representation 610 (e.g., a vector in the latent space). In an embodiment, the latent space representation 610 may correspond to a high-dimensional space vector having a number of members (e.g., 512, 1024, etc.). Similar to the image reconstruction task, the latent space representation 610 may represent an overall instance of the sensor data 102 (e.g., not just features identified from pixels). As such, the convolutional layers of the encoder network 608 may extract latent variables. This intermediate result - for example, latent space representation 610 - may then be deconvolved into the target space or design space of the 3D output 108 using a 3D decoder network 612. As a result, the 3D output may have a 1:1 mapping in the 3D world space and thus can directly represent the intersection structure in the target space or design space. Thus, in contrast to conventional systems where 2D outputs are converted to 3D space - a difficult and inaccurate process that relies on a flat ground assumption - DNN 104A can directly calculate the 3D output 108. For example, similar to the 3D output 108A of FIG. 4, the 3D intersection structure 614 may correspond to an accurate representation of the output 106 of DNN 104A after processing the 2D image space input image 102A. As described herein, the different line labels in the visualization of the 3D intersection structure 614 may correspond to different semantic information 110 associated with each of the line segments generated using the calculated 3D output 108 (e.g., using key point positions).
[0066] Only a single 2D encoder network 608 that processes a single instance of sensor data 102A is shown, but this is not intended to be limiting. For example, any number of 2D encoder networks 608 may process any number of instances of sensor data 102 from any number of sensors. For example, at any time instance or frame, a first sensor (e.g., a first camera, LIDAR sensor, etc.) can generate first sensor data 102 that is processed by a first instance of a 2D encoder network 608 (e.g., trained for a particular type of sensor data input) to compute a first latent space representation 610, a second sensor (e.g., a second camera, LIDAR sensor, RADAR sensor, etc.) can generate second sensor data 102 that is processed by a second instance of a 2D encoder network 608 (e.g., trained for a particular type of sensor data input) to compute a second latent space representation 610, and so on. Then, a plurality of latent space representations 610 for a given time instance or frame can be combined (e.g., concatenated), and the single combined latent space representation 610 can be processed by a 3D decoder network 612 to compute a 3D intersection structure 614. In some embodiments, two or more of the 2D encoder networks 608 can be processed in parallel using parallel processing units of the vehicle 800. As a result, the processing time of any number of 2D encoder networks 608 can be similar or the same as the processing of only a single instance of the 2D encoder network 608, thereby enabling the processing of additional information (e.g., sensor data 102 from multiple sources) in generating the 3D intersection structure 614.
[0067] Referring back to FIG. 6A, the output 106 can be processed using the post-processor 602. For example, temporal post-processing can be used to further enhance the robustness and accuracy of the prediction. In such an example, the output 106 from one or more previous instances of the DNN 104 can be compared (e.g., weighted) with the current output 106 from the current instance of the DNN 104 to generate an updated temporally smoothed result.
[0068] The output 106, before or after post-processing, can be applied to a path generator 604 that generates a path and / or trajectory for the vehicle 800 to follow to navigate the intersection. For example, semantic information and / or direction information (e.g., direction vectors, travel directions, etc.) associated with line segments from the 3D intersection structure can be used to determine potential paths through the intersection. In an embodiment, the path generator 604 can connect (center) key points corresponding to the line segments according to their 3D world space coordinates to generate a polyline representing a potential path for the vehicle 800 to cross the intersection in real-time or near real-time. The final path can be assigned a path type determined in relation to the position of the vehicle, the position of the key points, and / or the travel direction (e.g., angle) of the lane. Potential path types can include, but are not limited to, left turns, right turns, lane changes, and / or lane continuations.
[0069] In some examples, the route generator 604 may implement curve fitting to determine a final shape that most accurately reflects the natural curve of the potential route. Any known curve fitting algorithm may be used, such as, but not limited to, polyline fitting, polynomial fitting, and / or cycloid fitting. The shape of the potential route may be determined based on the positions of the key points and the corresponding lane travel directions associated with the connected key points. In some examples, the shape of the potential route may be aligned with the tangent of the travel direction vector at the positions of the connected key points. The curve fitting process may be repeated for all key points that may potentially be connected to each other to generate all possible routes that the vehicle 800 can follow to navigate through the intersection. In some examples, infeasible routes may be excluded from consideration based on traffic rules and physical limitations associated with such routes. The remaining potential routes may be determined to be feasible 3D routes or trajectories that the vehicle 800 can follow to cross the intersection.
[0070] In some embodiments, the route generator 604 may use a matching algorithm to connect key points and generate potential routes for the vehicle 800 to navigate through the intersection. In such examples, a matching score may be determined for each pair of key points based on the positions of the key points, the lane travel directions corresponding to the key points (e.g., two key points corresponding to different travel directions are not connected), and the shape of the approximate curve between the pair of key points. Each key point corresponding to intersection entry may be connected to a plurality of key points corresponding to intersection exit, thereby generating a plurality of potential routes for the vehicle 800. In some examples, a linear matching algorithm, such as the Hungarian matching algorithm, may be used. In other examples, a non-linear matching algorithm, such as a spectral matching algorithm, may be used to connect pairs of key points.
[0071] The route passing through the intersection can be used by the autonomous driving software stack 606 of vehicle 800 (or, referred to herein as "driving stack 606") to perform one or more operations. In some examples, the lane graph can be extended by information related to the route. The lane graph can be input to one or more control components of vehicle 800 to perform various planning tasks and control tasks. For example, the world model manager can update the world model to assist in navigating the intersection, the route planning layer of the driving stack 606 can use the intersection information to determine a route through the intersection (e.g., along one of the determined potential routes), and / or the control component can determine the control of the vehicle to navigate the intersection according to the determined route.
[0072] Referring now to FIG. 7, each block of method 700 described herein can include a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor that executes instructions stored in memory. Method 700 can also be implemented as computer-usable instructions stored on a computer storage medium. Method 700 can be provided, by way of example, as a stand-alone application, service, or hosted service (stand-alone or in combination with another hosted service), or as a plug-in to another product. Further, method 700 is described, by way of example, with respect to process 600 of FIG. 6. However, these methods can be performed additionally or alternatively by any one process and / or any one system, or any combination of processes and / or systems, including but not limited to those described herein.
[0073] Figure 7 is a flowchart showing an exemplary method 700 for detecting an intersection structure in 3D world space using a neural network, according to some embodiments of the present disclosure. Method 700 includes, at block B702, applying sensor data representing a 2D sensor data representation of an intersection within the field of view of at least one sensor of an autonomous machine to a neural network. For example, sensor data 102 (e.g., an image depicting an intersection) can be applied to DNN 104.
[0074] Method 700 includes, at block B704, using a neural network to calculate data representing a 3D world space position corresponding to the intersection and a reliability value corresponding to semantic information, based at least in part on the sensor data. For example, DNN 104 can use sensor data 102 to calculate an output 106 - e.g., a 3D output 108 and semantic information 110.
[0075] Method 700 includes, at block B706, decoding data representing the 3D world space position to determine a line segment position of a line segment associated with the intersection. For example, 3D output 108 can correspond to the position of key points, which can indicate the 3D position of the line segments of the intersection.
[0076] Method 700 includes, at block B708, decoding data representing the reliability value to determine the associated semantic information for each line segment. For example, semantic information 110 can be calculated as a reliability (or probability) corresponding to a plurality of different class types (such as, but not limited to, those described herein), and the reliability value can be used to determine semantic information 110 - e.g., the class having the highest reliability value (or a reliability value exceeding a threshold) can be attributed to a particular key point and / or line segment.
[0077] Method 700 includes, at block B710, executing one or more operations by an autonomous machine, based at least in part on a line segment position and associated semantic information. For example, output 106 may be used by path generator 604 to generate a path for vehicle 800 through an intersection, and / or the output 106 or path may be used by driving stack 606 to perform world model management, path planning, control, and / or other operations.
[0078] Exemplary Autonomous Vehicle FIG. 8A is an illustration of an exemplary autonomous vehicle 800, according to some embodiments of the present disclosure. Autonomous vehicle 800 (or referred to herein as "vehicle 800") may include, but is not limited to, a passenger vehicle, such as a car, truck, bus, first responder vehicle, shuttle, electric or motorized bicycle, motorcycle, fire truck, police vehicle, ambulance, boat, construction vehicle, submarine, drone, and / or another type of vehicle (e.g., unmanned and / or carrying one or more passengers). Autonomous vehicles are generally described with respect to an automation level as defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicle" (Standard No. J3016 - 201806, published on Jun. 15, 2018, Standard No. J3016 - 201609, published on Sep. 30, 2016, and previous and future versions of this standard). Vehicle 800 may have the ability to function at one or more of automation levels 3 - 5 of the autonomous driving level. For example, vehicle 800 may have the ability for conditional automation (level 3), highly automated (level 4), and / or fully automated (level 5), depending on the embodiment.
[0079] Vehicle 800 may include components such as the vehicle chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components. Vehicle 800 may include a propulsion system 850, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. Propulsion system 850 may be connected to the drive train of vehicle 800 and may include a transmission to enable the propulsion of vehicle 800. Propulsion system 850 may be controlled in response to receiving a signal from throttle / acceleration device 852.
[0080] Steering system 854, which may include a steering wheel, may be used to steer vehicle 800 (e.g., along a desired path or route) when propulsion system 850 is operating (e.g., when the vehicle is in motion). Steering system 854 may be able to receive a signal from steering actuator 856. The steering wheel may be an option for a fully automated (Level 5) function.
[0081] Brake sensor system 846 may be used to operate vehicle brakes in response to receiving a signal from brake actuator 848 and / or a brake sensor.
[0082] The controller 836, which may include one or more system-on-chips (SoCs) 804 (FIG. 8C) and / or GPUs, can provide signals (e.g., representations of commands) to one or more components and / or systems of the vehicle 800. For example, the controller can send signals to operate the vehicle brakes via one or more brake actuators 848, to operate the steering system 854 via one or more steering actuators 856, and to operate the propulsion system 850 via one or more throttle / acceleration devices 852. The controller 836 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or to assist the driver in operating the vehicle 800. The controller 836 may include a first controller 836 for autonomous driving functions, a second controller 836 for functional safety functions, a third controller 836 for artificial intelligence functions (e.g., computer vision), a fourth controller 836 for infotainment functions, a fifth controller 836 for redundancy in an emergency, and / or other controllers. In some examples, a single controller 836 can process two or more of the aforementioned functions, and two or more controllers 836 can process a single function and / or any combination thereof.
[0083] Controller 836 can provide signals for controlling one or more components and / or systems of vehicle 800 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data can be received, for example and without limitation, from a global navigation satellite system sensor 858 (e.g., a global positioning system sensor), a RADAR sensor 860, an ultrasonic sensor 862, a LIDAR sensor 864, an inertial measurement unit (IMU) sensor 866 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 896, a stereo camera 868, a wide view camera 870 (e.g., a fish-eye camera), an infrared camera 872, a surround camera 874 (e.g., a 360-degree camera), a long range and / or mid-range camera 898, a speed sensor 844 (e.g., for measuring the speed of vehicle 800), a vibration sensor 842, a steering sensor 840, a brake sensor (e.g., as part of a brake sensor system 846), and / or other sensor types.
[0084] One or more of the controllers 836 of the vehicle 800 may receive inputs (e.g., represented by input data) from the instrument cluster 832 of the vehicle 800 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 834, an audible annunciator, a loudspeaker, and / or other components of the vehicle 800. The output may include information such as vehicle velocity, speed, time, map data (e.g., the HD map 822 of FIG. 8C), position data (e.g., the position of the vehicle 800 on a map, etc.), direction, the position of other vehicles (e.g., occupancy grid), information regarding objects and the situation of objects as perceived by the controller 836, etc. For example, the HMI display 834 may display information regarding the presence of one or more objects (e.g., road signs, warning signs, changes in traffic signals, etc.) and / or driving operations that the vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at Exit 34B within 3.22 km (2 miles), etc.).
[0085] The vehicle 800 further includes a network interface 824 that can communicate via one or more networks using one or more wireless antennas 826 and / or a modem. For example, the network interface 824 may have the ability to communicate via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. The wireless antenna 826 may also use local area networks such as Bluetooth, Bluetooth LE, Z-Wave, ZigBee, and / or low power wide-area networks (LPWAN) such as LoRaWAN, SigFox to enable communication between objects (e.g., vehicles, mobile devices, etc.) in the environment.
[0086] FIG. 8B is an example of the camera position and field of view of the exemplary autonomous vehicle 800 of FIG. 8A according to some embodiments of the present disclosure. The cameras and their respective fields of view are one exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be placed at different positions on the vehicle 800.
[0087] The camera type of the camera may include, but is not limited to, a digital camera adapted to be used with the components and / or systems of the vehicle 800. The camera can operate at automotive safety integrity level (ASIL) B and / or at another ASIL. The camera type may have the ability to capture images at any rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may have the ability to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include an RCCC (red clear clear clear) color filter array, an RCCB (red clear clear blue) color filter array, an RBGC (red blue green clear) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a clear pixel camera, such as a camera with an RCCC, RCCB, and / or RBGC color filter array, may be used in efforts to increase light sensitivity.
[0088] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional mono-camera can be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more of the cameras (e.g., all of the cameras) can record and provide image data (e.g., video) simultaneously.
[0089] One or more of the cameras can be mounted in mounting parts such as custom-designed (3D printed) parts to remove stray light and reflections from inside the vehicle that may interfere with the camera's image data capture ability (e.g., reflections from the dashboard reflected in the front windshield mirror). Referring to the side mirror mounting part, the side mirror part can be custom 3D printed so that the camera mounting plate conforms to the shape of the side mirror. In some examples, the camera can be integrated within the side mirror. For side view cameras, the camera can also be integrated within four struts located at each corner of the cabin.
[0090] A camera having a field of view that includes a portion of the environment in front of the vehicle 800 (e.g., a forward-facing camera) can be used for surround view to assist in identifying the forward path and obstacles and, with the help of one or more controllers 836 and / or a control SoC, assist in providing information essential for generating an occupancy grid and / or determining a preferred vehicle path. The forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera can also be used for ADAS functions and systems including other functions such as lane departure warning ("LDW"), autonomous cruise control ("ACC"), and / or traffic sign recognition.
[0091] Various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imaging device. Another example may be a wide-view camera 870 that can be used to capture objects entering the view from the surroundings (e.g., pedestrians, intersecting traffic, or bicycles). Although only one wide-view camera is shown in FIG. 8B, any number of wide-view cameras 870 may be present on the vehicle 800. Additionally, a long-range camera 898 (e.g., a long-view stereo camera pair) can be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. The long-range camera 898 can also be used for object detection and classification, as well as for basic object tracking.
[0092] One or more stereo cameras 868 can also be included in the forward-facing configuration. The stereo camera 868 can include an integrated control unit with an expandable processing unit that can provide a programmable logic (FPGA) and a multi-core microprocessor with a CAN or Ethernet® interface integrated on a single chip. Such a unit can be used to generate a 3D map of the vehicle's environment that includes distance estimates for all points in the image. An alternative stereo camera 868 can include a compact stereo vision sensor that includes two camera lenses (one each for left and right) and an image processing chip that can measure the distance from the vehicle to the target object and activate autonomous emergency braking and lane departure warning functions using the generated information (e.g., metadata). Other types of stereo cameras 868 may be used in addition to, or instead of, those described herein.
[0093] A camera (e.g., a side view camera) having a field of view that includes a portion of the environment relative to the side of the vehicle 800 can be used for surround view to provide information for creating and updating an occupancy grid and for generating a side impact collision warning. For example, surround cameras 874 (e.g., four surround cameras 874 as shown in FIG. 8B) can be positioned on the vehicle 800. The surround cameras 874 can include wide view cameras 870, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be disposed at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 874 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround view camera.
[0094] A camera (e.g., a rearview camera) having a field of view that includes a portion of the environment relative to the rear of the vehicle 800 can be used for parking assistance, surround view, rear collision warning, and for creating and updating an occupancy grid. A wide variety of cameras can be used, including but not limited to cameras suitable as forward-facing cameras (e.g., long range and / or midrange cameras 898, stereo cameras 868), infrared cameras 872, etc., as described herein.
[0095] Figure 8C is a block diagram of an exemplary system architecture of the exemplary autonomous vehicle 800 of some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) may be used in addition to or instead of those shown, and some elements may be excluded altogether. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components and in any suitable combination and location. The various functions described herein as being performed by an entity may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor that executes instructions stored in a memory.
[0096] Each of the components, features, and systems of the vehicle 800 in FIG. 8C is illustrated as being connected via a bus 802. The bus 802 may include a Controller Area Network (CAN) data interface (or referred to as a "CAN bus"). CAN may be a network within the vehicle 800 used to assist in controlling various features and functions of the vehicle 800, such as the operation of brakes, acceleration, brakes, steering, front windshield wipers, etc. The CAN bus may be configured to have dozens or hundreds of nodes, each having its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button position, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0097] Bus 802 is described herein as being a CAN bus, but this is not intended to be limiting. For example, in addition to or as an alternative to the CAN bus, FlexRay and / or Ethernet® may be used. Additionally, a single line is used to represent bus 802, but this is not intended to be limiting. For example, any number of buses 802 may exist that include one or more CAN buses, one or more FlexRay buses, one or more Ethernet® buses, and / or one or more other types of buses that use different protocols. In some examples, two or more buses 802 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 802 may be used for a collision avoidance function and a second bus 802 may be used for actuation control. In any example, each bus 802 may communicate with any of the components of vehicle 800, and two or more buses 802 may communicate with the same component. In some examples, each SoC 804, each controller 836, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 800) and may be connected to a common bus such as a CAN bus.
[0098] Vehicle 800 may include one or more controllers 836, such as those described herein with respect to FIG. 8A. Controller 836 may be used for various functions. Controller 836 may be coupled to any of the various other components and systems of vehicle 800 and may be used for the control of vehicle 800, the artificial intelligence of vehicle 800, the infotainment for vehicle 800, and / or the like.
[0099] Vehicle 800 may include a system-on-chip (SoC) 804. SoC 804 may include a CPU 806, a GPU 808, a processor 810, a cache 812, an accelerator 814, a data store 816, and / or other components and features not shown. SoC 804 may be used to control vehicle 800 within various platforms and systems. For example, SoC 804 may be coupled in a system (e.g., a system of vehicle 800) having an HD map 822 that can obtain map refreshes and / or updates via a network interface 824 from one or more servers (e.g., server 878 of FIG. 8D).
[0100] CPU 806 may include a CPU cluster or CPU complex (or also referred to as a "CCPLEX"). CPU 806 may include a plurality of cores and / or an L2 cache. For example, in some embodiments, CPU 806 may include 8 cores within a coherent multi-processor configuration. In some embodiments, CPU 806 may include 4 dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2MB L2 cache). CPU 806 (e.g., CCPLEX) may be configured to support simultaneous cluster operation that allows any combination of clusters of CPU 806 to become active at any given time.
[0101] The CPU 806 can implement a power management capability that includes one or more of the following features: individual hardware blocks can be automatically clock-gated when in an idle state to conserve dynamic power, each core clock can be gated when the core is not actively executing instructions by the execution of WFI / WFE instructions, each core can be independently power-gated, each core cluster can be independently clock-gated when all cores are clock-gated or power-gated, and / or each core cluster can be independently power-gated when all cores are power-gated. The CPU 806 can further implement an enhanced algorithm for managing power states, where the allowed power states and the expected wake-up times are specified and the hardware / microcode determines the best power state to input to the cores, clusters, and CCPLEX. The processing cores can support a simplified power state input sequence in software where the work is offloaded to microcode.
[0102] The GPU 808 may include an integrated GPU (or referred to herein as "iGPU"). The GPU 808 can be programmable and can be efficient for parallel workloads. In some examples, the GPU 808 can use an enhanced tensor instruction set. The GPU 808 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache having a storage capacity of at least 96 KB), and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache having a storage capacity of 512 KB). In some embodiments, the GPU 808 may include at least eight streaming microprocessors. The GPU 808 can use a compute application programming interface (API). Additionally, the GPU 808 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0103] The GPU 808 can be power-optimized for the best performance in automotive and embedded use cases. For example, the GPU 808 can be manufactured on FinFET (Fin field-effect transistor). However, this is not intended to be limiting, and the GPU 808 can be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate several mixed-precision processing cores divided into multiple blocks. By way of non-limiting example, for instance, 64 PF32 cores and 32 PF64 cores may be divided into 4 processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, 2 mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads having a mix of compute and addressing operations. The streaming microprocessor can include independent thread scheduling capabilities to enable higher fine-grained synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to simplify programming while improving performance.
[0104] In some examples, the GPU 808 may include high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide for a peak memory bandwidth of 900 GB / second. In some examples, in addition to, or instead of, HBM memory, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used.
[0105] The GPU 808 can include unified memory technology that includes access counters to enable more accurate movement of those memory pages to the processor that most frequently accesses the memory pages, thereby improving the efficiency of the storage ranges shared between processors. In some examples, address translation service (ATS) support may be used to enable the GPU 808 to directly access the CPU 806 page table. In such examples, when the GPU 808 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU 806. In response, the CPU 806 can examine its page table for the virtual-to-physical mapping of the address and send the translation back to the GPU 808. As such, the unified memory technology can enable a single unified virtual address space for the memory of both the CPU 806 and the GPU 808, thereby simplifying GPU 808 programming and porting of applications to the GPU 808.
[0106] In addition, GPU 808 may include an access counter that can record the frequency of access of GPU 808 to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses that page most frequently.
[0107] SoC 804 may include any number of caches 812, including those described herein. For example, cache 812 may include an L3 cache that is available to both CPU 806 and GPU 808 (e.g., connected to both CPU 806 and GPU 808). Cache 812 may include a write-back cache that can record the state of lines, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include more than 4 MB, although smaller cache sizes may be used, depending on the embodiment.
[0108] SoC 804 may include an arithmetic logic unit (ALU) that can be utilized when executing processing for any of the various tasks or operations of vehicle 800 (e.g., processing DNN). In addition, SoC 804 may include a floating point unit (FPU) (or other mass coprocessor or numeric coprocessor type) for performing mathematical operations within the system. For example, SoC 104 may include one or more FPUs integrated as execution units within CPU 806 and / or GPU 808.
[0109] The SoC 804 may include one or more accelerators 814 (e.g., a hardware accelerator, a software accelerator, or a combination thereof). For example, the SoC 804 may include a hardware acceleration cluster that may include an optimized hardware accelerator and / or a large on-chip memory. The large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other operations. The hardware acceleration cluster may be used to complement the GPU 808 and to offload a portion of the tasks of the GPU 808 (e.g., to free up more cycles of the GPU 808 for other tasks). As an example, the accelerator 814 may be used for target workloads that are stable enough to be suitable for acceleration (e.g., perception, convolutional neural network (CNN)). As used herein, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., as used for object detection).
[0110] The acceleration device 814 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inferences. The TPU may be an accelerator configured and optimized to execute image processing functions (e.g., those of CNN, RCNN, etc.). The DLA may further be optimized for a specific set of neural network types and floating-point operations, as well as for inferences. The design of the DLA can provide more performance per millimeter than a general-purpose GPU and greatly exceed the performance of the CPU. The TPU can execute several functions, including, for example, a single-instance convolution function that supports INT8, INT16, and FP16 data types for both features and weights, as well as a post-processor function.
[0111] The DLA can quickly and efficiently execute neural networks, particularly CNNs, with processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object identification and detection using data from a camera sensor, CNNs for distance estimation using data from a camera sensor, CNNs for emergency vehicle detection and identification and detection using data from a microphone, CNNs for face recognition and vehicle owner identification using data from a camera sensor, and / or CNNs for security and / or safety-related events.
[0112] The DLA can execute any function of the GPU 808, and by using the inference accelerator, for example, a designer can target either the DLA or the GPU 808 for any function. For example, the designer can focus on processing CNNs and floating-point operations on the DLA and leave other functions to the GPU 808 and / or other acceleration devices 814.
[0113] The acceleration device 814 (e.g., a hardware acceleration cluster) may include, or be referred to herein as, a programmable vision accelerator (PVA) and a computer vision acceleration device. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0114] The RISC core can interact with an image sensor (e.g., the image sensor of any of the cameras described herein), an image signal processor, and / or the like. Each RISC core may include any amount of memory. The RISC core can use any of several protocols depending on the embodiment. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0115] DMA can enable the components of the PVA to access the system memory independent of the CPU806. DMA can support any number of features used to optimize for the PVA, including but not limited to supporting multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing up to six dimensions or more, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0116] The vector processor may be a programmable processor designed to efficiently and flexibly execute the programming of computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. The vector processing subsystem can operate as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.
[0117] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each vector processor may be configured to execute independently from other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may be able to execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms sequentially on an image or portions of an image. In particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system safety.
[0118] The accelerator 814 (e.g., a hardware acceleration cluster) may include a computer vision network on chip and SRAM to provide high bandwidth, low latency SRAM for the accelerator 814. In some examples, the on-chip memory may include at least 4MB of SRAM consisting of, for example and without limitation, 8 field-configurable memory blocks that may be accessible by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed access to the PVA and the DLA to the memory. The backbone may include a computer vision network on chip that interconnects the PVA and the DLA to the memory (e.g., using the APB).
[0119] A computer vision network-on-chip may include an interface that determines that both the PVA and the DLA are operable and provide valid signals before any control signal / address / data transmission. Such an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. This type of interface can comply with the ISO26262 or IEC61508 standard, although other standards and protocols may be used.
[0120] In some examples, the SoC804 may include a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and scale of objects (e.g., within a world model) for generating real-time visualization simulations for RADAR signal interpretation, for acoustic propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulation, for comparison against LIDAR data for localization and / or other functions, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing related operations.
[0121] The acceleration device 814 (e.g., a hardware acceleration device cluster) has various applications for autonomous driving. The PVA may be a programmable vision acceleration device that can be used in extremely important processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are suitable for areas of algorithms that require predictable processing at low power and low latency. In other words, the PVA functions well with low latency and low power and a predictable execution time, even on small data sets, for semi-dense or dense normal calculations. Therefore, since the PVA is efficient in object detection and integer calculations, in the context of a platform for autonomous vehicles, the PVA is designed to execute classic computer vision algorithms.
[0122] For example, according to one embodiment of the present technology, the PVA is used to execute computer stereo vision. A semi-global matching-based algorithm may be used in some examples, but this is not intended to be limiting. A number of applications for level 3-5 autonomous driving require motion estimation / stereo matching on the fly (e.g., SFM (structure from motion), pedestrian recognition, lane detection, etc.). The PVA can execute computer stereo vision functions with inputs from two monocular cameras.
[0123] In some examples, the PVA can be used to execute high-density optical flow. By processing raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used in flight depth processing time, for example, by processing the raw time of flight data to provide the processed time of flight data.
[0124] DLA can be used to execute any type of network to enhance control and driving safety, including, for example, a neural network that outputs a measurement of the reliability of each object detection. Such reliability values can be interpreted as probabilities or as providing the relative "weight" of each detection compared to other detections. This reliability value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a reliability threshold and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, a false positive detection would cause the vehicle to automatically execute an emergency brake, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can execute a neural network that regresses reliability values. The neural network can receive as its inputs at least some subset of parameters such as bounding box dimensions, ground plane estimation obtained (e.g., from another subsystem), vehicle 800 orientation, distance, 3D position estimation of an object obtained from the neural network and / or other sensors (e.g., LIDAR sensor 864 or RADAR sensor 860), and the output of an inertial measurement unit (IMU) sensor 866 that correlates with the above, among others.
[0125] SoC 804 can include a data store 816 (e.g., a memory). The data store 816 can be the on-chip memory of the SoC 804 and can store neural networks that are to be executed by the GPU and / or DLA. In some examples, the data store 816 can have a capacity large enough to store multiple instances of the neural network for redundancy and safety. The data store 812 can include an L2 or L3 cache 812. References to the data store 816 can include references to memory related to PVA, DLA, and / or other accelerators 814 as described herein.
[0126] SoC804 may include one or more processors 810 (e.g., embedded processors). The processor 810 may include a boot and power management processor that may be a dedicated processor and subsystem for handling boot power and management capabilities and related security enforcement. The boot and power management processor may be part of the SoC804 boot sequence and can provide runtime power management services. The boot power and management processor can provide clock and voltage programming, assistance with system low-power state transitions, management of the SoC804 heat and temperature sensors, and / or management of the SoC804 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC804 can use the ring oscillator to detect the temperature of the CPU806, GPU808, and / or accelerator 814. If the temperature is determined to exceed a threshold, the boot and power management processor can enter a temperature fault routine and place the SoC804 in a lower power state and / or put the vehicle 800 in a chauffeur-safe stop mode (e.g., safely stop the vehicle 800).
[0127] The processor 810 may further include a set of embedded processors that can perform the functions of an audio processing engine. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio via multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor having dedicated RAM.
[0128] The processor 810 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0129] The processor 810 may further include a safety cluster engine that includes a dedicated processor subsystem for handling the safety management of automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (such as timers, interrupt controllers, etc.), and / or routing logic. In the safety mode, two or more cores can operate in a lockstep mode and function as a single core with comparison logic to detect any differences during their operations.
[0130] The processor 810 may further include a real-time camera engine that may include a dedicated processor subsystem for handling real-time camera management.
[0131] The processor 810 may further include a high dynamic range signal processor that may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0132] The processor 810 may include a video image synthesizer, which may be a processing block (e.g., implemented in a microprocessor) that implements the post-video processing functions required by the video playback application to produce the final image for the player window. The video image synthesizer can perform lens distortion correction with the wide-view camera 870, with the surround camera 874, and / or with the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC configured to identify in-cabin events and respond appropriately. The in-cabin system can perform lip reading to enable cellular service and make calls, compose emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are available to the driver only when operating in autonomous mode and are otherwise disabled.
[0133] The video image synthesizer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if motion occurs within the video, the noise reduction reduces the weight of the information provided by adjacent frames and appropriately weights the spatial information. If an image or a portion of the image does not contain motion, the temporal noise reduction performed by the video image synthesizer can reduce the noise in the current image using information from the previous image.
[0134] The video image synthesizer may also be configured to perform stereo rectification on the input stereo lens frame. The video image synthesizer can further be used for user interface synthesis when the operating system desktop is in use, and the GPU 808 is not required to continuously render new surfaces. Even when the power of the GPU 808 is turned on and 3D rendering is actively performed, the video image synthesizer can be used to offload the GPU 808 to improve performance and responsiveness.
[0135] SoC 804 may further include a mobile industry processor interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for receiving video and inputs from a camera, for the camera and related pixel input functions. SoC 804 may further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.
[0136] SoC 804 may further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC 804 can be used to process data from a camera (e.g., connected via gigabit multimedia serial link and Ethernet (R)), sensors (e.g., LIDAR sensor 864, RADAR sensor 860 that can be connected via Ethernet (R)), data from bus 802 (e.g., speed, steering wheel position of vehicle 800), data from GNSS sensor 858 (e.g., connected via Ethernet (R) or CAN bus). SoC 804 may include its own DMA engine and may further include a dedicated high-performance large-capacity storage controller that can be used to free the CPU 806 from routine data management tasks.
[0137] The SoC804 may be an in-vehicle platform with a flexible architecture that extends to automation levels 3 to 5, thereby leveraging and efficiently using computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible and reliable driving software stack together with deep learning tools, providing a comprehensive functional safety architecture. The SoC804 can be faster, more reliable, more energy-efficient, and more space-efficient than conventional systems. For example, when the accelerator 814 is coupled to the CPU 806, the GPU 808, and the data store 816 can provide a fast and efficient platform for level 3 to 5 autonomous vehicles.
[0138] Accordingly, the present technology provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be implemented using a high-level programming language such as the C programming language to execute various processing algorithms over a wide variety of visual data and can be executed on a CPU. However, the CPU often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. Specifically, many CPUs cannot execute real-time multi-object detection algorithms, which are requirements for in-vehicle ADAS applications and actual level 3 to 5 autonomous vehicles.
[0139] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the techniques described herein enable multiple neural networks to be executed simultaneously and / or sequentially and enable results to be combined to enable level 3-5 autonomous driving functions. For example, a CNN executed on a DLA or a dGPU (e.g., GPU820) can include text and word recognition that enables a supercomputer to read and understand traffic signs, including signs that the neural network has not specifically been trained on. The DLA can further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass the semantic understanding to a route planning module executed on the CPU complex.
[0140] As another example, multiple neural networks can be executed simultaneously as required for level 3, 4, or 5 driving. For example, along with the electro-optical, a warning sign consisting of "Caution: Flashing light indicates frozen condition" can be interpreted independently or collectively by several neural networks. The sign itself can be identified as a traffic sign by a first deployed neural network (e.g., a neural network that has been trained), and the text "Flashing light indicates frozen condition" can be interpreted by a second deployed neural network that informs the vehicle's route planning software (preferably executed on the CPU complex) that a frozen condition exists when a flashing light is detected. The flashing light can be identified by informing the vehicle's route planning software of the presence (or absence) of the flashing light and operating a third deployed neural network over multiple frames. All three neural networks can be executed simultaneously, e.g., within the DLA and / or on the GPU808.
[0141] In some examples, the CNN for face recognition and vehicle owner identification can use data from the camera sensor to identify the presence of the authorized driver and / or owner of vehicle 800. The always-on sensor processing engine can be used to unlock and turn on the vehicle when the owner approaches the driver's side door and, in security mode, to stop the vehicle's operation when the owner leaves the vehicle. In this way, SoC 804 provides security against theft and / or carjacking.
[0142] In another example, the CNN for emergency vehicle detection and identification can use data from microphone 896 to detect and identify emergency vehicle sirens. In contrast to conventional systems that use a general classifier to detect sirens and manually extract features, SoC 804 uses a CNN for environmental and urban sound classification as well as for visual data classification. In a preferred embodiment, the CNN executed on the DLA is trained to identify the relative end speed of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by GNSS sensor 858. Thus, for example, when operating in Europe, the CNN will attempt to detect European sirens, and when in the United States, the CNN will attempt to identify only North American sirens. After an emergency vehicle is detected, a control program can be used to execute an emergency vehicle safety routine to slow down the vehicle, stop it at the side of the road, park the vehicle, and / or idle the vehicle, with the assistance of ultrasonic sensor 862 until the emergency vehicle has passed.
[0143] The vehicle may include a CPU 818 (e.g., an individual CPU, or dCPU) that can be connected to the SoC 804 via a high-speed interconnect (e.g., PCIe). The CPU 818 may include, for example, an X86 processor. The CPU 818 may be used to perform any of a variety of functions, including, for example, mediating potential inconsistencies between ADAS sensors and the SoC 804, and / or monitoring the status and health of the controller 836 and / or the infotainment SoC 830.
[0144] The vehicle 800 may include a GPU 820 (e.g., an individual GPU, or dGPU) that can be connected to the SoC 804 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 820 can provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of the vehicle 800.
[0145] The vehicle 800 may further include a network interface 824 that can include one or more wireless antennas 826 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 824 can be used to enable wireless connections with the cloud via the Internet (e.g., with the server 878 and / or other network devices), with other vehicles, and / or with computing devices (e.g., the passenger's client device). To communicate with other vehicles, a direct link can be established between two vehicles, and / or an indirect link can be established (e.g., through a network and via the Internet). The direct link can use a vehicle-to-vehicle communication link and can be provided. The vehicle-to-vehicle communication link can provide vehicle 800 information regarding vehicles in proximity to the vehicle 800 (e.g., vehicles in front of, beside, and / or behind the vehicle 800). This functionality may be part of the cooperative adaptive cruise control function of the vehicle 800.
[0146] Network interface 824 may include a SoC that provides modulation and demodulation functions and enables controller 836 to communicate via a wireless network. Network interface 824 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. Frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0147] Vehicle 800 may further include a data store 828 that may include storage outside the chip (e.g., outside SoC 804). Data store 828 may include one or more storage elements including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least 1 bit of data.
[0148] Vehicle 800 may further include a GNSS sensor 858. GNSS sensor 858 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) supports mapping, perception, occupancy grid generation, and / or route planning functions. For example, any number of GNSS sensors 858 may be used, including but not limited to GPS having an Ethernet® to serial (RS-232) bridge USB connector.
[0149] Vehicle 800 may further include a RADAR sensor 860. The RADAR sensor 860 can be used by the vehicle 800 for long-range vehicle detection even in darkness and / or severe weather conditions. The RADAR functional safety level may be ASIL B. In some examples, the RADAR sensor 860 can use CAN and / or bus 802 for control and to access object tracking data (e.g., to transmit data generated by the RADAR sensor 860) using access to Ethernet (registered trademark) to access raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 860 may be suitable for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.
[0150] The RADAR sensor 860 may include different configurations such as long range with a narrow field of view, short range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for an adaptive cruise control function. The long-range RADAR system can provide a wide field of view realized by two or more independent scans, such as within a range of 250 m. The RADAR sensor 860 can help distinguish between static and moving objects and can be used by an ADAS system for emergency brake assist and forward collision warning. The long-range RADAR sensor may include a monostatic multi-modal RADAR with a plurality (e.g., six or more) of fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In one example with six antennas, the four central antennas can create a focused beam pattern designed to record the surroundings of the vehicle 800 at high speed while minimizing interference from traffic in adjacent lanes. The other two antennas can widen the field of view and enable rapid detection of vehicles entering or leaving the lane of the vehicle 800.
[0151] As an example, a mid-range RADAR system may include a range up to 860 m (front) or 80 m (rear), and a field of view up to 42 degrees (front) or 850 degrees (rear). A short-range RADAR system may include, but is not limited to, RADAR sensors designed to be installed at both ends of the rear bumper. When installed at both ends of the rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and adjacent to the vehicle.
[0152] The short-range RADAR system may be used in an ADAS system for blind spot detection and / or lane change assist.
[0153] Vehicle 800 may further include ultrasonic sensors 862. The ultrasonic sensors 862, which may be positioned at the front, rear, and / or sides of vehicle 800, may be used for parking assist and / or for creating and updating occupancy grids. A variety of ultrasonic sensors 862 may be used, and different ultrasonic sensors 862 may be used for different ranges of detection (e.g., 2.5 m, 4 m). The ultrasonic sensors 862 can operate at a functional safety level of ASIL B.
[0154] Vehicle 800 may include a LIDAR sensor 864. The LIDAR sensor 864 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 864 may also be at a functional safety level of ASIL B. In some examples, vehicle 800 may include multiple (e.g., two, four, six, etc.) LIDAR sensors 864 that can use Ethernet (registered trademark) (e.g., to provide data to a gigabit Ethernet (registered trademark) switch).
[0155] In some examples, the LIDAR sensor 864 may have the ability to provide a list of objects and their distances in a 360-degree field of view. Commercially available LIDAR sensors 864 may, for example, have an accuracy of 2 cm to 3 cm, support an 800 Mbps Ethernet® connection, and have an advertised range of approximately 800 m. In some examples, one or more non-protruding LIDAR sensors 864 may be used. In such examples, the LIDAR sensor 864 may be implemented as a small device that can be incorporated into the front, rear, sides, and / or corners of the vehicle 800. In such examples, the LIDAR sensor 864 may have a range of 200 m even for low-reflectivity objects and can provide up to a 120-degree horizontal and 35-degree vertical field of view. The LIDAR sensor 864 mounted on the front may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0156] In some examples, LIDAR technologies such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a flash of laser as a transmitter to illuminate the area around the vehicle up to about 200 m. The flash LIDAR unit includes a receptor that records the laser pulse travel time and the reflected light on each pixel corresponding in sequence to the range from the vehicle to the object. Flash LIDAR may enable a high-precision and distortion-free image of the surroundings to be generated with every laser flash. In some examples, four flash LIDAR sensors may be deployed, one on each side of the vehicle 800. Available 3D flash LIDAR systems include solid-state 3D steering array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than a blower. The flash LIDAR device can use a 5-nanosecond class I (eye-safe) laser pulse per frame and can capture the reflected laser light in the form of a 3D range point cloud and co-described intensity data. By using flash LIDAR, and since flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 864 may be less susceptible to the effects of motion blur, vibration, and / or shock.
[0157] The vehicle may further include an IMU sensor 866. In some examples, the IMU sensor 866 may be positioned at the center of the rear axle of the vehicle 800. The IMU sensor 866 may include, for example, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types, but is not limited thereto. In some examples, in a six-axis application, etc., the IMU sensor 866 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 866 may include an accelerometer, a gyroscope, and a magnetometer.
[0158] In some embodiments, the IMU sensor 866 may be implemented as a miniature, high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. As such, in some examples, the IMU sensor 866 may enable the vehicle 800 to estimate its heading without the need for input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 866. In some examples, the IMU sensor 866 and the GNSS sensor 858 may be combined in a single integrated unit.
[0159] The vehicle may include a microphone 896 placed inside and / or around the vehicle 800. The microphone 896 may be used, among other things, for emergency vehicle detection and identification.
[0160] The vehicle may further include any number of camera types, including a stereo camera 868, a wide view camera 870, an infrared camera 872, a surround camera 874, a long range and / or mid range camera 898, and / or other camera types. The cameras can be used to capture image data around the entire outer surface of the vehicle 800. The type of cameras used depends on the embodiment and requirements of the vehicle 800, and any combination of camera types can be used to achieve the necessary coverage around the vehicle 800. Additionally, the number of cameras can vary depending on the embodiment. For example, the vehicle may include 6 cameras, 7 cameras, 10 cameras, 12 cameras, and / or another number of cameras. As an example, the cameras can support, but are not limited to, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet (registered trademark). Each camera is described in more detail herein in relation to FIGS. 8A and 8B.
[0161] Vehicle 800 may further include a vibration sensor 842. The vibration sensor 842 can measure the vibration of components of the vehicle, such as an axle. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 842 are used, the difference in vibration can be used to determine the friction or slip of the road surface (for example, when the difference in vibration is between a power driven axle and a free rotating axle).
[0162] Vehicle 800 may include an ADAS system 838. In some examples, ADAS system 838 may include a SoC. ADAS system 838 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.
[0163] The ACC system may use a RADAR sensor 860, a LIDAR sensor 864, and / or a camera. The ACC system may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle directly in front of vehicle 800 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and advises vehicle 800 to change lanes when necessary. Lateral ACC is related to other ADAS applications such as LCA and CWS.
[0164] CACC can use information from other vehicles that can be received via a wireless link from the network interface 824 and / or the wireless antenna 826 of other vehicles, or indirectly via a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. In general, the V2V communication concept provides information about the vehicle immediately ahead (e.g., the vehicle immediately in front of vehicle 800 that is in the same lane as vehicle 1100), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of an I2V information source and a V2V information source. Given information about the vehicles ahead of vehicle 800, CACC can be made more reliable and has the potential to make the traffic flow smoother and reduce road congestion.
[0165] The FCW system is designed to warn the driver of a danger so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or RADAR sensor 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide warnings in the form of acoustic, visual alerts, vibrations, and / or quick brake pulses.
[0166] The AEB system can detect an imminent forward collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within the specified time or distance parameters. The AEB system can use a forward-facing camera and / or RADAR sensor 860 connected to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a danger, the AEB system usually first warns the driver to take corrective action to avoid the collision. If the driver does not take corrective action, the AEB system can automatically apply the brakes as part of an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system may include techniques such as dynamic brake support and / or imminent collision braking.
[0167] The LDW system provides visual, audible, and / or tactile warnings, such as vibrations of the steering wheel or seat, to warn the driver when vehicle 800 crosses a lane demarcation line. The LDW system does not activate when the driver indicates an intentional lane departure by activating the turn indicator. The LDW system can use a forward-facing camera connected to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically connected to driver feedback, such as a display, speaker, and / or vibrating component.
[0168] The LKA system is a modified form of the LDW system. The LKA system provides steering input or brakes to correct vehicle 800 if vehicle 800 begins to veer out of a lane.
[0169] The BSW system detects and warns the vehicle driver in the blind spots of the vehicle. The BSW system can provide visual, audible, and / or tactile warnings to indicate that a merge or lane change is not safe. The system can provide additional warnings when the driver uses the turn indicator. The BSW system can use a rear-facing camera and / or RADAR sensor 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0170] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 800 is backing up. Some RCTW systems include AEB to ensure that vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0171] Conventional ADAS systems allow the driver to be warned and to determine whether a safety condition actually exists and act accordingly. Thus, conventional ADAS systems, while not usually catastrophic, tend to produce false positive results that can annoy and distract the driver. However, in the case of an autonomous vehicle 800, if the results conflict, the vehicle 800 itself must determine whether to accept the results from the primary computer or the secondary computer (e.g., the first controller 836 or the second controller 836). For example, in some embodiments, the ADAS system 838 may be a backup and / or secondary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor can execute diverse software that is redundant in hardware components to detect malfunctions in perception and dynamic driving tasks. The output from the ADAS system 838 can be provided to the supervisory MCU. If the outputs from the primary and secondary computers conflict, the supervisory MCU needs to determine how to reconcile that conflict to ensure safe operation.
[0172] In some examples, the primary computer can be configured to provide a reliability score to the supervisory MCU that indicates the reliability of the primary computer in the selected result. If the reliability score exceeds a threshold, the supervisory MCU can follow the instructions of the primary computer regardless of whether the secondary computer gives conflicting or inconsistent results. If the reliability score does not meet the threshold and the primary and secondary computers indicate different (e.g., conflicting) results, the supervisory MCU can arbitrate between the computers to determine an appropriate result.
[0173] The supervisory MCU may be configured to execute a neural network trained and configured to determine, based on outputs from the primary computer and the secondary computer, a state in which the secondary computer provides a false alarm. Thus, the neural network within the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot be trusted. For example, when the secondary computer is a RADAR-based FCW system, the neural network within the supervisory MCU can learn when the FCW identifies metallic objects such as manhole covers or gratings in a drain that are not actually dangerous but trigger an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network within the supervisory MCU can learn to ignore the LDW when a person on a bicycle or a pedestrian is present and lane departure is actually the safest maneuver. In an embodiment that includes a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for executing the neural network with associated memory. In a preferred embodiment, the supervisory MCU may comprise components of the SoC1104 and / or may be included as components of the SoC804.
[0174] In other examples, the ADAS system 838 may include a secondary computer that executes ADAS functions using conventional rules of computer vision. As such, the secondary computer can use classical computer vision rules (if-then), and the presence of the neural network within the supervisory MCU can improve reliability, safety, and performance. For example, the various implementation forms and intentional non-identities make the overall system more fault-tolerant against faults caused particularly by software (or software-hardware interface) functions. For example, if there is a software bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have a greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer has not caused a critical error.
[0175] In some examples, the output of the ADAS system 838 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 838 indicates a forward collision warning due to an object immediately ahead, the perception block can use this information when identifying the object. In other examples, the secondary computer may have its own neural network that is trained as described herein and thus reduces the risk of misjudgment.
[0176] Vehicle 800 may further include an infotainment SoC 830 (e.g., an in-vehicle infotainment system (IVI) within the vehicle). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more individual components. The infotainment SoC 830 may include a combination of hardware and software used to provide the vehicle 800 with audio (e.g., music, mobile devices, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calls), network connections (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, wireless data systems, fuel level, total mileage, brake fuel level, oil level, open / close doors, air filter information, and other vehicle-related information). For example, the infotainment SoC 830 may include radio, disk player, navigation system, video player, USB and Bluetooth connections, car computer, in-vehicle entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, heads-up display (HUD), HMI display 834, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 830 may be further used to provide information (e.g., visual and / or audible) to the vehicle's user, such as information from the ADAS system 838, autonomous driving information such as planned vehicle operations, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0177] The Infotainment SoC 830 may include GPU functionality. The Infotainment SoC 830 can communicate with other devices, systems, and / or components of the vehicle 800 via a bus 802 (e.g., CAN bus, Ethernet®, etc.). In some examples, the Infotainment SoC 830 can be coupled to a supervisory MCU such that the GPU of the Infotainment system can perform some self-driving functions in the event that the primary controller 836 (e.g., the primary and / or backup computer of the vehicle 800) fails. In such examples, the Infotainment SoC 830 can put the vehicle 800 into the chauffeur's safe stop mode as described herein.
[0178] The vehicle 800 may further include an instrument cluster 832 (e.g., digital dash, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 832 may include a controller and / or a supercomputer (e.g., an individual controller or supercomputer). The instrument cluster 832 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, gear shift position indicator, seat belt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting control device, safety system control device, navigation information, etc. In some examples, information may be displayed and / or shared between the Infotainment SoC 830 and the instrument cluster 832. In other words, the instrument cluster 832 may be included as part of the Infotainment SoC 830, and vice versa.
[0179] FIG. 8D is a system diagram of communication between a cloud-based server of FIG. 8A and an exemplary autonomous vehicle 800 according to some embodiments of the present disclosure. System 876 may include a server 878, a network 890, and vehicles including vehicle 800. Server 878 may include a plurality of GPUs 884(A)-884(H) (collectively referred to herein as GPU 884), PCIe switches 882(A)-882(H) (collectively referred to herein as PCIe switch 882), and / or CPUs 880(A)-880(B) (collectively referred to herein as CPU 880). The GPUs 884, CPUs 880, and PCIe switches may be interconnected by high-speed interconnects, such as, but not limited to, NVLink interface 888 and / or PCIe connection 886 developed by NVIDIA, for example. In some examples, the GPUs 884 are connected via NVLink and / or NVSwitch SoC, and the GPUs 884 and PCIe switches 882 are connected via PCIe interconnects. Eight GPUs 884, two CPUs 880, and two PCIe switches are shown, but this is not intended to be limiting. Depending on the embodiment, each server 878 may include any number of GPUs 884, CPUs 880, and / or PCIe switches. For example, server 878 may include eight, sixteen, thirty-two, and / or more GPUs 884, respectively.
[0180] Server 878 can receive, via network 890, image data representing an image indicating an unexpected or changed road condition, such as a recently started road construction, from a vehicle. Server 878 can transmit, via network 890, neural network 892, an updated neural network 892, and / or map information 894 including information regarding traffic and road conditions to the vehicle. The update of the map information 894 may include an update of the HD map 822, such as information regarding a construction site, a depression, a detour, a flood, and / or other obstacles. In some examples, the neural network 892, the updated neural network 892, and / or the map information 894 may result from new training and / or experience represented in data received from any number of vehicles in the environment and / or based on training performed in a data center (e.g., using server 878 and / or other servers).
[0181] Server 878 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by a vehicle and / or (e.g., using a game engine) generated in a simulation. In some examples, the training data is tagged (e.g., when the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., when the neural network does not require supervised learning). The training can be performed according to any one or more classes of machine learning techniques, including but not limited to, for example, the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, associative learning, transfer learning, feature learning (including principal component and cluster analysis), multi-linear subspace learning, manifold learning, representation learning (including pre-dictionary learning), rule-based machine learning, anomaly detection, and variations or combinations thereof. After the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 890), and / or the machine learning model can be used by server 878 to remotely monitor the vehicle.
[0182] In some examples, server 878 can receive data from a vehicle and apply the data to a latest real-time neural network for real-time intelligent inference. Server 878 can include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 884, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 878 can include a deep learning infrastructure that uses only CPU-powered data centers.
[0183] The deep learning infrastructure of server 878 can have the ability of high-speed real-time inference and use that ability to evaluate and verify the condition of the processors, software, and / or related hardware within vehicle 800. For example, the deep learning infrastructure can receive periodic updates from vehicle 800, such as sequences of images and / or objects in which vehicle 800 is located within the sequence of images (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can execute its own neural network to identify the objects and compare them with those identified by vehicle 800. If the results do not match and the infrastructure concludes that the AI within vehicle 800 is not functioning properly, server 878 can send a signal to vehicle 800 that infers control, notifies the passengers, and commands the fail-safe computer of vehicle 800 to complete a safe parking operation.
[0184] For inference, server 878 can include GPU 884 and one or more programmable inference acceleration devices (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time responsiveness. In other examples, such as when less performance is required, servers powered by CPUs, FPGAs, and other processors can be used for inference.
[0185] Exemplary Computing Devices FIG. 9 is a block diagram of an example of a computing device 900 suitable for use in implementing some embodiments of the present disclosure. The computing device 900 may include an interconnect system 902 that indirectly or directly connects the following devices: a memory 904, one or more central processing units (CPUs) 906, one or more graphics processing units (GPUs) 908, a communication interface 910, an input / output (I / O) port 912, input / output components 914, a power supply 916, one or more presentation components 918 (e.g., a display), and one or more logic units 920. In at least one embodiment, the computing device 900 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). In a non-limiting example, one or more of the GPUs 908 may include one or more vGPUs, one or more of the CPUs 906 may include one or more vCPUs, and / or one or more of the logic units 920 may include one or more virtual logic units. As such, the computing device 900 may include individual components (e.g., a full GPU dedicated to the computing device 900), virtual components (e.g., a portion of a GPU dedicated to the computing device 900), or a combination thereof.
[0186] Although the various blocks of FIG. 9 are shown as being connected via an interconnect system 902 with lines, this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 918 such as a display device may be considered an I / O component 914 (e.g., if the display is a touch screen). As another example, the CPU 906 and / or the GPU 908 may include memory (e.g., the memory 904 may represent a storage device in addition to the memory of the GPU 908, the CPU 906, and / or other components). In other words, the computing device of FIG. 9 is merely exemplary. Categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types are all intended to be within the scope of the computing device of FIG. 9 and thus are not distinguished.
[0187] The interconnect system 902 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 902 may include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 906 may be directly connected to the memory 904. Further, the CPU 906 may be directly connected to the GPU 908. If there are direct or point-to-point connections between components, the interconnect system 902 may include a PCIe link for making the connections. In these examples, the PCI bus need not be included in the computing device 900.
[0188] The memory 904 may include any of a variety of computer-readable media. A computer-readable media may be any available media that can be accessed by the computing device 900. Computer-readable media may include both volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, computer-readable media may comprise computer storage media and communication media.
[0189] A computer storage medium can include both volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 904 can store computer-readable instructions (such as programs and / or program elements) such as an operating system. A computer storage medium can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 900. In this specification, a computer storage medium does not include a signal itself.
[0190] A computer storage medium can implement computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information distribution medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristic sets or that varies in such a way as to encode information in the signal. By way of example, and without limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.
[0191] The CPU 906 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to execute one or more of the methods and / or processes described herein. The CPU 906 may each include one or more (e.g., 1, 2, 4, 8, 28, 72, etc.) cores having the ability to process multiple software threads simultaneously. The CPU 906 may include any type of processor and, depending on the type of computing device 900 implemented, may include different types of processors (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 900, the processor may be an Advanced RISC Machines (ARM) processor implemented using reduced instruction set computing (RISC), or an x86 processor implemented using complex instruction set computing (CISC). The computing device 900 may include one or more CPUs 906 within one or more microprocessors or auxiliary coprocessors, such as a computing coprocessor.
[0192] In addition to or instead of the CPU 906, the GPU 908 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to execute one or more of the methods and / or processes described herein. One or more of the GPU 908 may be an integrated GPU (e.g., it may be with one or more of the CPU 906), and / or one or more of the GPU 908 may be a discrete GPU. In an embodiment, one or more of the GPU 908 may be one or more coprocessors of the CPU 906. The GPU 908 may be used by the computing device 900 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPU 908 may be used for general-purpose computing on the GPU (GPGPU). The GPU 908 may include hundreds or thousands of cores having the ability to process hundreds or thousands of software threads simultaneously. The GPU 908 can generate pixel data for an output image in response to rendering commands (e.g., rendering commands from the CPU 906 received via a host interface). The GPU 908 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory may be included as part of the memory 904. The GPU 908 may include two or more GPUs operating in parallel (e.g., via a link). The link can connect directly to the GPUs (e.g., using NVLINK), or can connect the GPUs via a switch (e.g., using NVSwitch). When coupled together, each GPU 908 can generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., the first GPU for the first image and the second GPU for the second image). Each GPU can include its own memory or can share memory with other GPUs.
[0193] In addition to and / or instead of the CPU 906 and / or the GPU 908, the logic unit 920 may be configured to execute at least some of the computer-readable instructions to control one or more of the computing devices 900 to execute one or more of the methods and / or processes described herein. In an example, the CPU 906, the GPU 908, and / or the logic unit 920 can execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more of the logic units 920 may be part of and / or integrated with one or more of the CPU 906 and / or the GPU 908, and / or one or more of the logic units 920 may be discrete components with respect to the CPU 906 and / or the GPU 908 or otherwise external to them. In an example, one or more of the logic units 920 may be one or more coprocessors of one or more of the CPU 906 and / or one or more of the GPU 908.
[0194] Examples of the logic unit 920 include one or more processing cores and / or their components, such as tensor cores (TC), tensor processing units (TPU), pixel visual cores (PVC), vision processing units (VPU), graphics processing clusters (GPC), texture processing clusters (TPC), streaming multiprocessors (SM), tree traversal units (TTU), artificial intelligence accelerators (AIA), deep learning accelerators (DLA), arithmetic logic units (ALU), application specific integrated circuits (ASIC), floating point units (FPU), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like.
[0195] The communication interface 910 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 900 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 910 can include components and functions for enabling communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating via Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0196] The I / O port 912 can enable the computing device 900 to be logically connected to other devices, including some of which may be built-in (e.g., integrated) into the computing device 900, such as I / O components 914, presentation components 918, and / or other components. Exemplary I / O components 914 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, etc. The I / O components 914 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be sent to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, face recognition, biometric recognition, gesture recognition on and adjacent to the screen, air gestures, head and gaze tracking, and touch recognition related to the display of the computing device 900 (as will be described in more detail later). The computing device 900 can include a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. Additionally, the computing device 900 can include an accelerometer or gyroscope (e.g., as part of an inertia measurement unit (IMU)) to enable motion detection. In some examples, the output of the accelerometer or gyroscope can be used by the computing device 900 to render immersive augmented reality or virtual reality.
[0197] The power supply device 916 can include a hard-wired power supply device, a battery power supply device, or a combination thereof. The power supply device 916 can provide power to the computing device 900 to enable the components of the computing device 900 to operate.
[0198] The presentation component 918 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 918 can receive data from other components (e.g., the GPU 908, the CPU 906, etc.) and output the data (e.g., as an image, a video, an audio, etc.). Sub-section heading
[0199] Exemplary data center FIG. 10 shows an exemplary data center 1000 that can be used in at least one embodiment of the present disclosure. The data center 1000 may include a data center infrastructure layer 1010, a framework layer 1020, a software layer 1030, and / or an application layer 1040.
[0200] As shown in FIG. 10, the data center infrastructure layer 1010 may include a resource orchestrator 1012, grouped computing resources 1014, and node computing resources (“node C.R.”) 1016(1) to 1016(N), where “N” represents any positive integer. In at least one embodiment, the node C.R. 1016(1) to 1016(N) may include any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors or graphics processing units (GPU), etc.), memory devices (e.g., dynamic random access memory), storage devices (e.g., solid state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and / or cooling modules, etc., but are not limited thereto. In some embodiments, one or more of the node C.R. 1016(1) to 1016(N) may correspond to a server having one or more of the computing resources described above. Additionally, in some embodiments, the node C.R. 1016(1) to 10161(N) may include one or more virtual components such as vGPU, vCPU, and / or the like, and / or one or more of the node C.R. 1016(1) to 1016(N) may correspond to a virtual machine (VM).
[0201] In at least one embodiment, the grouped computing resources 1014 can include a separate group of node C.R.s 1016 housed within one or more racks (not shown), or many racks (likewise not shown) housed in data centers at various geographical locations. The separate group of node C.R.s 1016 within the grouped computing resources 1014 can include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, some node C.R.s 1016 that include CPUs, GPUs, and / or other processors can be grouped within one or more racks to provide computing resources for supporting one or more workloads. The one or more racks can also include any number of power modules, cooling modules, and / or network switches in any combination.
[0202] The resource orchestrator 1022 can configure or otherwise control one or more node C.R.s 1016(1) - 1016(N) and / or the grouped computing resources 1014. In at least one embodiment, the resource orchestrator 1022 can include a software design infrastructure ("SDI") management entity for the data center 1000. The resource orchestrator 1022 can include hardware, software, or some combination thereof.
[0203] In at least one embodiment, as shown in FIG. 10, the framework layer 1020 may include a job scheduler 1032, a configuration manager 1034, a resource manager 1036, and / or a distributed file system 1038. The framework layer 1020 may include a framework for supporting software 1032 of the software layer 1030 and / or one or more applications 1042 of the application layer 1040. The software 1032 or the application 1042 may each include web-based service software or an application such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1020 may be a kind of free and open-source software web application framework such as Apache Spark (trademark) (hereinafter, Spark), which can utilize the distributed file system 1038 for large-scale data processing (e.g., "big data"), but is not limited thereto. In at least one embodiment, the job scheduler 1032 may include a Spark driver for facilitating the scheduling of workloads supported by various layers of the data center 1000. The configuration manager 1034 may have the ability to configure different layers such as the software layer 1030 and the framework layer 1020 including Spark and the distributed file system 1038 for supporting large-scale data processing. The resource manager 1036 may have the ability to manage the clustered or grouped computing resources mapped or assigned for the support of the distributed file system 1038 and the job scheduler 1032. In at least one embodiment, the clustered or grouped computing resources may include the grouped computing resources 1014 in the data center infrastructure layer 1010. The resource manager 1036 may manage these mapped or assigned computing resources in cooperation with the resource orchestrator 1012.
[0204] In at least one embodiment, the software 1032 included in the software layer 1030 may include software used by at least a portion of nodes C.R. 1016(1) to 1016(N), grouped computing resources 1014, and / or the distributed file system 1038 of the framework layer 1020. One or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.
[0205] In at least one embodiment, the application 1042 included in the application layer 1040 may include one or more types of applications used by at least a portion of nodes C.R. 1016(1) to 1016(N), grouped computing resources 1014, and / or the distributed file system 1038 of the framework layer 1020. One or more types of applications may include, but are not limited to, machine learning applications including any number of genomics applications, cognitive computing, and training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in combination with one or more embodiments.
[0206] In at least one embodiment, any of the configuration manager 1034, the resource manager 1036, and the resource orchestrator 1012 can implement any number and type of self-corrective actions based on any amount and type of data obtained in any technically feasible way. The self-corrective actions can free the data center operator of the data center 1000 from, in some cases, making bad configuration decisions and, in some cases, avoiding underutilized and / or underperforming parts of the data center.
[0207] The data center 1000 can include tools, services, software, or other resources for training one or more machine learning models according to one or more embodiments described herein or for predicting or inferring information using one or more machine learning models. For example, a machine learning model can be trained by calculating weight parameters according to a neural network architecture using the software and / or computing resources described above with respect to the data center 1000. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks can be used to infer or predict information using the resources described above with respect to the data center 1000 by using weight parameters calculated through one or more training techniques such as, but not limited to, those described herein.
[0208] In at least one embodiment, data center 1000 may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the above resources. Further, the one or more software and / or hardware resources may be configured as a service that enables a user to train or perform inference of information such as image recognition, voice recognition, or other artificial intelligence services.
[0209] Exemplary Network Environment A network environment suitable for use in implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be implemented on one or more instances of the computing device 900 of FIG. 9 - for example, each device may include similar components, features, and / or functions of the computing device 900. Additionally, when a backend device (e.g., a server, NAS, etc.) is implemented, the backend device may be included as part of the data center 1000, an example of which is described in more detail herein with respect to FIG. 10.
[0210] The components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. The network can include multiple networks, or a network of networks. By way of example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or the public switched telephone network (PSTN), and / or one or more private networks. Where the network includes a wireless communication network, components such as base stations, communication towers, and even access points (and other components) can provide a wireless connection.
[0211] A compatible network environment can include one or more peer-to-peer network environments - where no server is included in the network environment - and one or more client-server network environments - where one or more servers are included in the network environment. In a peer-to-peer network environment, the functions described herein with respect to a server can be implemented on any number of client devices.
[0212] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment may be implemented on one or more servers that may include one or more core network servers and / or edge servers, and may include a framework layer, a job scheduler, a resource manager, and a distributed file system. The framework layer may include a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or application may each include web-based service software or an application. In an embodiment, one or more of the client devices may use web-based service software or an application (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be a kind of free and open-source software web application framework such that a distributed file system can be used for large-scale data processing (e.g., "big data"), but is not limited thereto.
[0213] A cloud-based network environment may provide cloud computing and / or cloud storage that executes any combination of the computing functions and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions may be distributed to multiple locations from a central server or core server (e.g., one or more data centers that may be distributed across a state, region, country, the globe, etc.). When the connection to a user (e.g., a client device) is relatively close to an edge server, the core server may designate at least a portion of the functions to the edge server. The cloud-based network environment may be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0214] The client device may include at least some of the components, features, and functions of the exemplary computing device 900 described herein with respect to FIG. 9. By way of example and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smart watch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, game device or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, household electronic device, workstation, edge device, any combination of these detailed devices, or any other suitable device.
[0215] The present disclosure may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer or other machine, such as a portable information terminal or other handheld device. Generally, program modules include routines, programs, objects, components, data structures, etc., and refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be implemented in a variety of configurations, including handheld devices, household appliances, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.
[0216] As used herein, a description of "and / or" with respect to two or more elements should be construed to mean either only one element, or a combination of elements. For example, "element A, element B, and / or element C" can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Additionally, "at least one of element A or element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0217] The subject matter of this disclosure is described with particularity to meet statutory requirements. However, the description itself is not intended to limit the scope of the disclosure. Rather, the inventors intend that the claimed subject matter may be otherwise implemented, including in combination with other current or future technologies, in different steps or combinations of steps similar to those described in this document. Further, the terms "step" and / or "block" may be used herein to imply different elements of a method being used, but these terms should not be construed as implying any particular order among the various steps disclosed herein except where the order of individual steps is explicitly recited and when so recited.
Claims
1. applying sensor data representative of intersections within a field of view of at least one sensor to a neural network; calculating, using the neural network, data representative of a three-dimensional (3D) world space location corresponding to the intersection and one or more confidence values corresponding to semantic information based at least in part on the sensor data; decoding the data representing the 3D world space position to determine a position of one or more line segments; decoding the data representing the one or more confidence values to determine associated semantic information of the one or more line segments; performing, by an autonomous machine, one or more actions based at least in part on the positions of the one or more line segments and the associated semantic information; A method comprising:
2. 2. The method of claim 1 , wherein the data representing the 3D world space positions comprises 3D world space positions of key points corresponding to the line segments, the key points comprising, for one or more of the line segments, at least one of an endpoint or a center point of the one or more line segments.
3. 3. The method of claim 2, wherein the key points of a line segment of the one or more line segments include a center point, and the neural network further calculates a direction vector corresponding to the center point that indicates a direction of travel associated with the line segment.
4. 3. The method of claim 2, further comprising generating one or more potential paths between the two or more line segments based at least in part on directional vectors associated with two or more of the key points of the line segments, and wherein performing the one or more actions comprises selecting one of the one or more potential paths.
5. 2. The method of claim 1 , wherein the neural network is trained based at least in part on transforming a predicted output of the neural network into a 2D space and comparing the predicted output in the 2D space to 2D ground truth data using a loss function.
6. 2. The method of claim 1 , wherein the neural network is trained based at least in part on comparing a 3D output of the neural network to one or more geometric consistency constraints associated with intersections using a loss function.
7. 2. The method of claim 1 , wherein the 3D world space position is predicted in an object space having an origin corresponding to a position on the autonomous machine, and a size of the object space is determined based at least in part on an analysis of a plurality of intersections.
8. 2. The method of claim 1 , wherein the neural network includes an encoder portion that processes the sensor data to compute latent space vectors, and wherein the neural network includes a decoder portion that processes the latent space vectors to generate the data representing the 3D world space locations corresponding to the intersections and the confidence values corresponding to the semantic information.
9. 9. The method of claim 8, wherein the at least one sensor includes a first sensor that generates a first subset of the sensor data and a second sensor that generates a second subset of the sensor data, and the encoder portion of the neural network includes a first encoder instance that processes the first subset of the sensor data to generate a first portion of the latent space vector and a second encoder instance that processes the second subset of the sensor data to generate a second portion of the latent space vector, the latent space vector being generated at least in part based on concatenating the first and second portions.
10. The method of claim 9 , wherein the first encoder instance and the second encoder instance are executed in parallel using one or more parallel processing units.
11. applying data representing a two-dimensional (2D) space to a neural network, said data being generated using at least one sensor; using the neural network to compute a representation of a three-dimensional (3D) world space position based at least in part on the data; transforming the 3D world space position into the 2D space to generate a 2D space position; comparing the 2D spatial location to a 2D spatial ground truth location associated with the sensor data; updating one or more parameters of the neural network based at least in part on the comparing step; A method comprising:
12. 12. The method of claim 11 , wherein the comparing step is performed using a loss function, the method further comprising: comparing the 3D world space position to one or more geometric consistency constraints using another loss function, and the updating of the one or more parameters of the neural network is further based at least in part on the comparing of the 3D world space position to the one or more geometric consistency constraints.
13. 13. The method of claim 12, wherein the 3D world space location corresponds to a line segment of an intersection, and the one or more geometric consistency constraints include at least one of a smoothness constraint corresponding to the line segment, a straightness constraint corresponding to the line segment, or a lane width constraint corresponding to the line segment.
14. The method of claim 11 , wherein the converting step is based at least in part on at least one of intrinsic or extrinsic parameters of the at least one sensor.
15. 12. The method of claim 11 , wherein the neural network includes an encoder portion that processes the sensor data to compute latent space vectors, and wherein the neural network includes a decoder portion that processes the latent space vectors to generate the data representing the 3D world space positions.
16. applying sensor data representative of intersections within a field of view of at least one sensor to a neural network; using the neural network to calculate data representative of three-dimensional (3D) world space positions corresponding to the intersection line segments based at least in part on the sensor data; and performing one or more actions associated with controlling an autonomous machine based at least in part on the 3D world space position corresponding to the line segment.
23. A processor comprising: one or more circuits for:
17. 17. The processor of claim 16, wherein the neural network includes an encoder portion that processes the sensor data to compute latent space vectors, and wherein the neural network includes a decoder portion that processes the latent space vectors to generate the data representing the 3D world space positions corresponding to the intersection line segments.
18. 18. The processor of claim 17, wherein the at least one sensor includes a first sensor that generates a first subset of the sensor data and a second sensor that generates a second subset of the sensor data, and the encoder portion of the neural network includes a first encoder instance that processes the first subset of the sensor data to generate a first portion of the latent space vector and a second encoder instance that processes the second subset of the sensor data to generate a second portion of the latent space vector, the latent space vector being generated at least in part based on concatenating the first and second portions.
19. 17. The processor of claim 16, wherein the neural network is trained based at least in part on transforming a predicted output of the neural network into a 2D space and comparing the predicted output in the 2D space to 2D ground truth data using a loss function.
20. 17. The processor of claim 16, wherein the neural network is trained based at least in part on comparing a 3D output of the neural network to one or more geometric consistency constraints associated with intersections using a loss function.
21. 17. The processor of claim 16, wherein the processor is included in at least one of a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system for performing simulation operations, a system for performing deep learning operations, a system implemented using edge devices, a system implemented using a robot, a system incorporating one or more virtual machines (VMs), a system implemented at least partially in a data center, or a system implemented at least partially using cloud computing resources.
Citation Information
Patent Citations
Intersection region detection and classification for autonomous machine applications
US11436837B2
Intersection detection and classification in autonomous machine applications
US11648945B2
Intersection pose detection in autonomous machine applications
US12013244B2