DETECTING LINE MARKINGS FOR AUTONOMOUS AND SEMI-AUTOMOUS SYSTEMS AND APPLICATIONS
Patent Information
- Application Number
- DE102025110665
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-17
- Filing Date
- 2025-03-19
- Publication Date
- 2025-09-25
Smart Images

Figure 00000049_0000 
Figure 00000050_0000 
Figure 00000051_0000
Abstract
Description
BACKGROUND
[0001] High-resolution (HD), standard-definition (SD), navigation, and / or other map types serve a variety of functions in autonomous and semi-autonomous driving. These detailed maps can, for example, provide a precise localization reference, allowing autonomous or semi-autonomous vehicles to accurately determine their position in the environment by comparing real-time sensor data with existing map features. Furthermore, HD maps can contribute to path planning. For example, an autonomous vehicle can use map features such as road geometry, lane markings, and traffic signs to plan safe and effective trajectories.Furthermore, HD maps can provide a semantic understanding of the environment, encoding classifications of objects such as traffic lights and stop signs, and improving the vehicle's ability to interpret complex scenarios and make informed decisions based on contextual information. Furthermore, HD maps can provide a reliable reference point in situations where sensor data may be ambiguous or incomplete. Real-time or near-real-time map updates can enable autonomous vehicles to quickly adapt to changes in the environment, ensuring continuous accuracy and responsiveness to dynamic road conditions. Therefore, HD maps can provide spatial awareness to autonomous and semi-autonomous driving systems, enabling safe and efficient navigation in diverse and dynamic landscapes.
[0002] Conventional techniques for generating HD maps have a variety of disadvantages. For example, conventional techniques typically generate HD maps by projecting images generated using in-vehicle cameras for data collection onto the road surface. However, due to perspective distortion, visual features located far from the camera are often distorted in the map. Furthermore, visual features of interest based on scene changes over time are often occluded in images, which can lead to inaccuracies in the map. Some conventional techniques have aimed to detect features, such as lane lines or boundaries, from these projected images, but because visual features of interest are often distorted or occluded, the detected features have limited accuracy.Some techniques have aimed to apply semantic segmentation or line segment detection to these projected images, but this process requires significant computational effort during post-processing, for example, to connect pieces of the same line segment from different images. Therefore, there is a need for improved detection and map generation techniques. SUMMARY
[0003] The invention is defined by the claims. To illustrate the invention, aspects and embodiments are described herein, which may or may not be within the scope of the claims.
[0004] Various embodiments are disclosed in which sensor data representing a 3D environment may be collected using one or more ego machines while the ego machines navigate through the 3D environment. The sensor data may be projected onto a 2D representation of the ground or other surface, and this 2D representation may form a map representing a geographic region. The map may be divided into tiles, in which detected features (e.g., road lines, road markings, surface features, etc.) may be detected and used to detect demarcated regions, such as intersections, based on the geometry and proximity of the detected features. Therefore, new tiles may be centered around the detected regions, and the features may be detected from each resulting centered tile.The detected features can be aggregated, deduplicated and / or merged and used to label the map.
[0005] Embodiments of the present disclosure relate to the detection of navigation control lines for autonomous and semi-autonomous systems and applications. Systems and methods are disclosed that generate a labeled map (e.g., LiDAR, RADAR, ultrasound map, etc.) with detected navigation control lines (or other road or driving surface line types) for navigation, localization, and / or other applications in ego machines.
[0006] In contrast to conventional systems, navigation control lines can be detected and labeled in a map (e.g., LiDAR, RADAR, ultrasonic, image map, etc.). Instead of or in addition to cameras, data can be collected from ego machines using one or more LiDAR sensors and / or other sensor types, such as RADAR sensors, ultrasonic sensors, etc. For example, LiDAR sensor data can be collected from one or more ego machines, and the sensor data can be projected into a two-dimensional (2D) representation of the ground or other surface. In this example, the 2D representation of the ground can form a LiDAR map representing a geographic region observed by the one or more ego machines (e.g., over time). Continuing the example, the LiDAR map can be divided into tiles, and navigation control lines (e.g.,Road markings, lines on the road representing traffic signals, lines or other visual demarcations in an outdoor, indoor or warehouse environment, etc.) can be detected by individual tiles.
[0007] In some examples, detected navigation control lines may be used to detect intersections based on the geometry and proximity of the detected navigation control lines. Therefore, new tiles may be centered around the detected intersections, and navigation control lines may be jEach resulting intersection-centric tile can be detected. The detected navigation control lines from different tiles can be aggregated, deduplicated, and / or merged and used to label the LiDAR map. Therefore, the labeled LiDAR map can assist an autonomous or semi-autonomous vehicle or other ego-machine in navigating a physical environment, for example, by enabling the vehicle to accurately interpret and respond to traffic signals on the road, navigate safely through intersections, obey traffic laws, and / or it can assist the vehicle or machine in accurately determining its position and orientation within the road network.
[0008] The disclosure extends to all novel aspects or features described and / or illustrated herein.
[0009] Further features of the disclosure are characterized by the independent and dependent claims.
[0010] Any feature of one aspect of the disclosure may be applied to other aspects of the disclosure in any suitable combination. In particular, method aspects may be applied to device or system aspects, and vice versa.
[0011] Furthermore, features implemented in hardware may be implemented in software, and vice versa. Any reference to software and hardware features herein should be construed accordingly.
[0012] Any system or device feature described herein may also be provided as a method feature, and vice versa. System and / or device aspects described functionally (including means plus functional features) may alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.
[0013] It is also to be understood that certain combinations of the various features described and defined in each aspect of the disclosure may be implemented and / or provided and / or used independently of one another.
[0014] The disclosure also provides computer programs and computer program products comprising software code configured to, when executed on a data processing device, perform any of the methods described herein and / or embody any of the device and system features described herein, including any or all of the component steps of a method.
[0015] The disclosure also includes a computer or computer system (including networked or distributed systems) having an operating system that supports a computer program for performing the methods described herein and / or for embodying the device or system features described herein.
[0016] The disclosure also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0017] The disclosure also provides a signal carrying one or more of the aforementioned computer programs.
[0018] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0019] Aspects and embodiments of the disclosure will now be described, by way of example only, with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present systems and methods for detecting navigation control lines for autonomous and semi-autonomous systems and applications are described in detail below with reference to the accompanying drawings. In the drawings: Fig. 1 is a data flow diagram illustrating an example map generation pipeline, according to some embodiments of the present disclosure. Fig. 2 is a diagram illustrating an example intersection detector, according to some embodiments of the present disclosure; Fig. 3 illustrates an example map (e.g., LiDAR map) divided into tiles, according to some embodiments of the present disclosure; Fig. 4A-B illustrate exemplary detected navigation control lines, according to some embodiments of the present disclosure; Fig. 5A-B illustrate the generation of an exemplary intersection-centered tile, according to some embodiments of the present disclosure; Fig. 6 illustrates exemplary techniques for merging duplicates and grouping associated detected navigation control lines, according to some embodiments of the present disclosure; Fig. 7 is a flowchart illustrating a method for updating a map (e.g., LiDAR map) based on at least one or more navigation control lines, according to some embodiments of the present disclosure; Fig. 8 is a flowchart illustrating a method for detecting a refined set of navigation control lines, according to some embodiments of the present disclosure; Fig. 9A is an illustration of an exemplary autonomous vehicle, according to some embodiments of the present disclosure; Fig. 9B shows an example of camera locations and fields of view for the example autonomous vehicle from Fig. 9A, according to some embodiments of the present disclosure; Fig. 9C is a block diagram of an exemplary system architecture for the exemplary autonomous vehicle of Fig. 9A, according to some embodiments of the present disclosure; Fig. 9D is a system diagram for communication between one or more cloud-based servers and the example autonomous vehicle of Fig. 9A, according to some embodiments of the present disclosure; Fig. 10 is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure; and Fig. 11 is a block diagram of an exemplary data center suitable for use in implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0021] Systems and methods related to detecting lines, features, and / or road / surface markings for autonomous and semi-autonomous systems and applications are disclosed. For example, systems and methods are disclosed that, as a non-limiting example, project detected LiDAR intensity data onto a two-dimensional (2D) representation of a surface such as the ground (e.g., a LiDAR map), detect navigation control lines (e.g., traffic signal road lines) from individual tiles of the map, and label the map with detected lines and class labels. The present techniques can be used to create or update maps with navigation control lines for use by autonomous or semi-autonomous machines and other types of ego machines.
[0022] Although the present disclosure is made with respect to an exemplary autonomous or semi-autonomous vehicle or machine 900 (alternatively referred to herein as “vehicle 900” or “ego machine 900”), an example of which is disclosed with respect to Fig. 9A-9D), this is not limiting. The systems and methods described herein may be used, for example, without limitation, by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS)), autonomous vehicles or machines, guided and unguided robots or robotic platforms, warehouse vehicles, off-highway vehicles, vehicles coupled to one or more trailers, hydrofoils, boats, shuttle vehicles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, trains, underwater vehicles, remotely operated vehicles such as drones, and / or other types of vehicles. In addition, although the present disclosure is intended with respect to generating or updating maps (e.g.,a LiDAR map) with detected navigation control lines for use by road vehicles, this is not limiting, and the systems and methods described herein may be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technology area in which maps of 2D surfaces with navigation control lines are generated or updated.
[0023] In some embodiments, and taking as an example a use case where a map supports an autonomous or semi-autonomous vehicle navigating on roads, in some embodiments, sensor data (e.g., LiDAR data) may be collected using one or more ego machines (e.g., a fleet of data collection vehicles), and the sensor data may be projected onto a 2D representation of the ground or other surface. In an exemplary embodiment where the 2D representation of the ground is a LiDAR map representing a geographic region, the LiDAR map may be divided into tiles, and navigation control lines (e.g.,Traffic signal road lines or other road markings, lines or features on the road that represent traffic signals, such as crosswalk lines, stop lines, or yield lines, or lines or features or other demarcations in another environment, such as a warehouse, factory, building, park, plaza, etc., can be detected by individual tiles and used to detect intersections based on the geometry and proximity of the detected navigation control lines. Therefore, new tiles can be centered around the detected intersections, and navigation control lines can be detected by each resulting intersection-centered tile. The detected navigation control lines can be aggregated, deduplicated, and / or merged and used to label the map.Therefore, the labeled map can assist an autonomous vehicle in navigating a physical environment by enabling the vehicle to accurately interpret and respond to traffic signals, navigate safely through intersections, comply with traffic laws, and / or it can assist the vehicle in accurately determining its position and orientation within the road network, particularly at intersections and in areas controlled by traffic signals.
[0024] In some embodiments, LiDAR data collected from fleet vehicles may be sent to a map generation pipeline, which may be used to generate a map. LiDAR data (e.g., LiDAR intensity data) collected from any number of LiDAR sensors and / or any number of vehicles may be projected onto a 2D representation of a surface (e.g., the ground) of a three-dimensional (3D) space to form projected LiDAR intensity data (e.g., a top-down projection image). This projected LiDAR intensity data may take the form of a local representation of the 2D surface (e.g., ground) in a vehicle coordinate system (e.g., a projection image) of a corresponding data collection vehicle, a list of projected data points, and / or other forms.Therefore, the detected position of the vehicle in 3D space (or 3D world space) can be used to aggregate the projected LiDAR intensity data into a global representation of the surface (e.g., the ground) of the 3D space (e.g., a global LiDAR map, an HD map, etc.). For example, a global LiDAR map can represent projected LiDAR data in grayscale or color and can be regularly updated based on new data (e.g., changes to the physical road or to the navigation control lines on the road). Therefore, sensor data (e.g., LiDAR data) representing a geographic region can be used to create a map of the region. Although described primarily in relation to LiDAR, this is not limiting; other types or modalities of sensor data may also be used, such as radar, ultrasound, camera, etc.
[0025] In some embodiments, the map (e.g., global LiDAR map) may be divided into any suitable grid cells or tiles (e.g., of fixed size) to facilitate feature detection. These tiles may be either non-overlapping or overlapping with neighboring tiles. In some embodiments, a deep neural network (DNN) (e.g., LinE segment Transformers (LETR), TileNet V2, or any model that has the ability to detect lines and / or generate a list or other representation of the detected polylines) may be used to detect one or more classes of navigation control lines (e.g., line segments or polylines) from each tile of the map. In an exemplary embodiment, the output of the DNN may include multiple objects, where each object has two points (e.g.,two endpoints P1 and P2 that form a detected line segment), a classification label (e.g., a class indicating that the object does not correspond to a detected line; a supported class of a detected navigation control line, such as crosswalk lines, stop lines, yield lines; an indication that one of a variety of supported classes was detected; etc.), and a corresponding confidence. Therefore, a specified threshold confidence can be used to filter the DNN output and generate a list of detected lines and their corresponding class labels. In general, any known line detection technique can be applied to the individual tiles to detect and classify the navigation control lines present in each tile.
[0026] In some embodiments, the detected navigation control lines may be clustered to infer the locations of intersections in the map. For example, crosswalks typically have two parallel lines, and intersections may contain multiple crosswalks. Therefore, in some embodiments, detected crosswalk lines may be searched for corresponding crosswalk segments based on distance, orientation, and / or projected overlap. In some embodiments, detected crosswalk lines may be searched for other crosswalk lines in the same intersection based on distance and / or orientation. Therefore, if a detected navigation control line is classified as a crosswalk line (e.g., a type of navigation control line), a corresponding crosswalk line may be searched and grouped with nearby crosswalk lines to form an intersection.In some embodiments, each intersection may be searched for nearby detected stop and / or yield lines that may be clustered into the intersection. In some implementations, remaining detected stop and / or yield lines may be searched to identify the presence of other types of intersections (e.g., intersections without crosswalks, intersections other than four-way intersections) using appropriate distance and / or orientation thresholds. More generally, an intersection may be inferred from each of one or more detected navigation control lines that visually indicate the location of an intersection.
[0027] In some embodiments, a new intersection-centric tile may be created for each inferred intersection, the detection of navigation control lines may be repeated on each intersection-centric tile, and the detected navigation control lines from different tiles may be aggregated, deduplicated, and used to label or otherwise associate with the map (e.g., global LiDAR map). In some embodiments, to facilitate feature detection within an intersection-centric tile, the intersection-centric tile may be rotated (e.g., either clockwise or counterclockwise) to more closely align it with the boundaries of an intersection and maximize or increase the number of intersection features represented within the intersection-centric tile.In some embodiments, if the intersection is larger or smaller than the resolution of an intersection-centric tile, the intersection-centric tile may be resized (e.g., by applying a scaling factor such as 0.8 to 1.2). Additionally or alternatively, if an intersection is (e.g., significantly) larger than the resolution of an intersection-centric tile (e.g., a representation of an intersection containing 1500 pixels compared to an intersection-centric tile with a limit of 1000 pixels), the intersection-centric tile may be divided into smaller tiles, and navigation control lines may be detached from the intersection. jEach of the smaller tiles can be detected and aggregated to capture the navigation control lines located at the intersection. In some embodiments, crosswalk lines within the same crosswalk can be identified and paired to form corresponding polygons. Therefore, detected navigation control lines, detected regions bounded by detected navigation control lines, and / or corresponding class labels can be labeled on the map (e.g., global LiDAR map).
[0028] Therefore, the detected navigation control lines (e.g., traffic signal road lines) can be used to label a global LiDAR map, and the global LiDAR map can be distributed or otherwise accessed by any number of ego machines to facilitate navigation, localization, and / or other applications. The present techniques provide a variety of advantages over prior techniques. For example, by generating intersection-centric tiles, semantically meaningful features are effectively arranged in a single input representation, so that detecting navigation control lines from intersection-centric tiles focuses the DNN on more relevant information than prior techniques, resulting in improved detection accuracy and precision. Furthermore, various embodiments save significant computational effort.For example, using projected LiDAR data eliminates the computationally intensive backpropagation of image data. In another example, detecting navigation control lines from intersection-centered tiles should allow complete line segments to be detected from most intersections, eliminating the need for conventional techniques to connect separate pieces of the same line segment detected from different tiles. Furthermore, detecting navigation control lines from projected LiDAR data reduces and even eliminates many distortions and occlusions depicted in projected camera images, resulting in a more accurate representation of the immediate environment and thus more accurate line detections and more accurate downstream uses.Therefore, a labeled map generated using the present or similar techniques may improve the manner in which an autonomous or semi-autonomous vehicle or machine navigates and locates itself in a physical environment, particularly at intersections and in areas controlled by traffic signals.
[0029] Referring to Fig. 1 is Fig. 1 illustrates an exemplary map generation pipeline 100, according to some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are set forth as examples only. Other arrangements and elements (e.g., engines, interfaces, functions, arrangements, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as individual or distributed components, or in conjunction with other components, and in any suitable combination and location. Various functions described herein as being performed by entities may be performed by hardware, firmware, and / or software.For example, various functions may be performed by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein may be implemented using similar components, features, and / or functionality as those of the exemplary autonomous vehicle 900 of FIG. Fig. 9A-9D, the exemplary computing device 1000 of Fig. 10 and / or the exemplary data center 1100 from Fig. 11.
[0030] The map generation pipeline 100 may generate a map using data collected from one or more fleet vehicles (e.g., autonomous, semi-autonomous, non-autonomous vehicles) or other ego machines navigating on roads, driving surfaces, and / or other environments. For example, sensor data 110 (which in some embodiments may include LiDAR intensity data) may be collected using one or more ego machines (e.g., a fleet of data collection vehicles) and sent to the map generation pipeline 100 (which may, for example, be hosted at a remote location such as a data center), and a projection component 120 may project the sensor data 110 into a 2D representation of the ground or other surface.In some examples, a feature detector 130 (which may be referred to as a navigation control line detector 130 when used for line detection) may detect navigation control lines (e.g., traffic signal road lines or other road markings or lines on the road that represent traffic signals, such as crosswalk lines, stop lines, or yield lines) within the 2D representation of the ground or other surface. In an exemplary embodiment where the 2D representation of the ground is a LiDAR map representing a geographic region, an input generator 140 may divide the LiDAR map into tiles, and a line detector 150 may detect navigation control lines from the individual tiles.Continuing the example, an intersection detector 160 may use the detected navigation control lines to detect intersections based on the geometry and proximity of the detected navigation control lines. Therefore, a tile generator 170 may generate new tiles centered around the detected intersections and trigger the line detector 150 to detect navigation control lines from each resulting intersection-centered tile. A post-processing component 180 may aggregate, deduplicate, and / or merge detected navigation control lines, and a map labeling component 190 may use these detected navigation control lines to label the map. Therefore, the labeled map may be used by an autonomous or semi-autonomous vehicle (e.g.,the example autonomous or semi-autonomous vehicle or machine 900) or another ego machine to assist in navigating a physical environment, e.g., to enable the vehicle to accurately interpret and respond to traffic signals depicted on the map, safely navigate intersections, comply with traffic laws, and / or it may assist the vehicle in precisely determining its position and orientation within the road network or other navigable surface or environment, particularly at intersections and / or in areas controlled by traffic signals.
[0031] In the Fig. 1, the sensor data 110 may be collected using one or more fleet vehicles (e.g., ego machines). In some examples, sensor data 110 may include LiDAR data collected using any number of LiDAR sensors and / or any number of ego machines. However, this is only one example, and other types of sensor data may additionally or alternatively be used (e.g., data from RADAR sensors, ultrasonic sensors, inertial measurement units, GPS, GNSS or other position sensors, thermal sensors, etc.). For example, LiDAR intensity data may be projected onto a 2D surface and collected by one or more ego machines. In at least some examples, this projection may be performed by one or more ego machines, and a representation of the projected sensor data may be sent to the map generation pipeline 100 (e.g.,over any suitable network). Additionally or alternatively, another representation of the sensor data may be sent to the map generation pipeline 100, and the projection component 120 may operate at a remote location (e.g., in a data center, such as the data center 1100 that hosts the map generation pipeline 100). In some embodiments, the sensor data 110 may be received by the map generation pipeline 100 as a point cloud (e.g., a list of measured 3D points and corresponding reflectance properties), a projected representation, and / or another representation.
[0032] The projection component 120 may project the LiDAR or other sensor data (e.g., detected 3D points) from a 3D coordinate system (e.g., whether a global coordinate system such as a world map or a local coordinate system such as a vehicle-centered coordinate system) into a particular 2D view (e.g., a 2D map). For example, the projection component 120 may project into a 2D representation of a surface in the environment (e.g., a top-down view of the ground). In some embodiments, the projection component 120 may convert sensor data 110 into pixels of a projection image, and the pixels may be assigned values, such as grayscale or color values, that represent a corresponding measured value (e.g., LiDAR intensity) of the point projected onto each pixel. By projecting multiple observations (e.g.,LiDAR rotations) of the surface into the 2D representation, the 2D representation of the surface can represent the projected sensor data as a global map (e.g., LiDAR map) (e.g., a top-down view of the ground), which can have any number of channels storing all corresponding measurements (e.g., whether derived from LiDAR and / or other types of sensor data).
[0033] Therefore, in some embodiments, projection component 120 may use sensor data 110 (e.g., representing a geographic region) to create a map of the region. In some examples, feature detector 130 may detect navigation control lines within the 2D representation of the ground of a geographic region (e.g., the global LiDAR map representing it). For example, feature detector 130 may detect navigation control lines within the map (e.g., the global LiDAR map) created by projection component 120.
[0034] The feature detector 130 may divide the map (e.g., global LiDAR map) into tiles (e.g., fixed-size), use a deep neural network (DNN) to detect navigation control lines (e.g., traffic signal road lines, such as crosswalk lines, stop lines, road divider lines, bicycle lane lines, yield lines, painted signs or signals, etc.) and / or other features in each tile, cluster detected navigation control lines to infer the locations of intersections in the map, re-run the detection of navigation control lines on intersection-centered tiles, and / or apply deduplication to remove duplicates of detected navigation control lines (e.g., crosswalk lines).Therefore, to detect, classify, and label navigation control lines within the map constructed by projection component 120, feature detector 130 may use input generator 140, line detector 150, intersection detector 160, tile generator 170, and / or post-processing component 180.
[0035] The input generator 140 may divide the map (e.g., global LiDAR map) into any suitable grid cells or tiles (e.g., of fixed size) to facilitate the detection of navigation control lines. In some examples, these tiles may be non-overlapping or overlapping with neighboring tiles. Whether the tiles are overlapping or non-overlapping may be configurable (e.g., to make the tiles overlapping rather than non-overlapping, and vice versa). In some embodiments, the tiles may be of fixed size to facilitate the execution of a single DNN (e.g., or a single DNN architecture) to detect features of the navigation control line within each tile. Fig. Figure 3 illustrates an example LiDAR map divided into tiles, according to some embodiments of the present disclosure. In this example, the 2D representation of a geographic region (e.g., the global LiDAR map) may be divided into tiles (e.g., of a fixed size) to form a grid 310, and each tile may be applied to a DNN to detect and classify navigation control lines within the tile (e.g., tile 320).
[0036] Taking another look at Fig. 1, in some embodiments, the line detector 150 may be implemented using one or more neural networks, such as, but not limited to, a DNN or a convolutional neural network (CNN). For example, and without limitation, the line detector 150 may include any of a number of different networks or machine learning models, such as, for example, one or more machine learning models that include linear regression, logistic regression, decision trees, support vector machines (SVMs), the naive Bayes classifier, k-nearest neighbors (Knn), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g.,Autoencoders, convolutional, transformers, recurrent, perceptrons, long / short-term memory (LSTM), large language model (LLM), Hopfield, Boltzmann, deep belief, unfolding, generative adversarial, liquid state machine, etc.) and / or other types of machine learning models.
[0037] Therefore, the line detector 150 may use a machine learning model, such as a DNN, to detect and classify navigation control lines. In some implementations, any known line detection technique may be applied to the individual tiles to detect and classify the navigation control lines present in each tile. In some embodiments, a DNN (e.g., Transformer for LinE Segments (LETR), TileNet V2, or any model capable of detecting lines and / or generating a list or other representation of the detected polylines) may be used to detect one or more classes of navigation control lines (e.g., line segments or polylines) from each tile of the map (e.g., LiDAR map). Note that each tile may contain any number of channels of corresponding data, depending on the information encoded in the map.In an exemplary embodiment, the output of the DNN may be a plurality of objects, where each object encodes, represents, or otherwise identifies two points or pixels within a tile (e.g., two endpoints P1 and P2 that form a detected line segment), a classification label (e.g., a class indicating that the object does not correspond to a detected line; a supported class of a detected navigation control line, such as crosswalk lines, stop lines, yield lines; an indication that one of a plurality of supported classes was detected; etc.), and a corresponding confidence. In some implementations, the output of the DNN may be a fixed number of objects (e.g., 100), each having two points (e.g., pixels within the tile that represent corresponding endpoints), but not all objects may fall into a supported class of a detected navigation control line (e.g.,Crosswalk lines, stop lines, yield lines). Therefore, in at least some embodiments, threshold confidence may be applied to decode the DNN output (and to identify, for example, the detected navigation control lines).
[0038] Accordingly, the line detector 150 may decode the DNN output and generate a list of the detected lines and corresponding class labels (e.g., labeling the navigation control lines as crosswalk lines, stop lines, or yield lines). In some embodiments, the line detector 150 may apply a specified threshold confidence to the DNN output to identify the detected navigation control lines. In some embodiments, each object (e.g., each set of two points or pixels) is assigned a numerical value, and a threshold confidence may be used to separate all objects in the DNN output based on that numerical value. For example, a numerical value may be assigned to each classification label (e.g.,a class indicating that the object does not correspond to a detected line; a supported class of a detected navigation control line, such as crosswalk lines, stop lines, yield lines; an indication that one of a plurality of supported classes was detected; etc.), and a threshold confidence may be applied to filter out the numerical values that do not correspond to the detected navigation control lines. In this example, the line detector 150 can thus separate all objects that are not detected navigation control lines from all objects that are navigation control lines (e.g., crosswalk lines, stop lines, yield lines, etc.). Note that this is intended only as an example, and any suitable line detection technique may be used to detect lines of one or more specified classes.
[0039] The intersection detector 160 may cluster detected navigation control lines to infer the locations of intersections within each tile of the (e.g., LiDAR) map. Referring now to Fig. 2 is Fig. 2 is a diagram illustrating an example intersection detector 160, according to some embodiments of the present disclosure. In some examples, intersection detector 160 may receive a list of detected navigation control lines from line detector 150, and this list is represented by arrow 210. In some embodiments, intersection detector 160 may iterate through the detected crosswalk lines and search for other detected crosswalk lines that are close to (e.g., within one or more threshold distances) other crosswalk lines. For each pair of detected crosswalk lines within a threshold distance, intersection detector 160 may, for example, label and group the pairs of crosswalk lines as part of an inferred intersection. Additionally or alternatively, intersection detector 160 may, according to some embodiments,detected intersections) and iterate through one or more other classes of detected navigation control lines (e.g., stop lines, yield lines, etc.) to label the other nearby navigation control lines within the same intersection and cluster them into the inferred intersection.
[0040] In the Fig. 2, the intersection detector 160 includes a crosswalk identification component 220, a perpendicularity verification component 225, a parallel alignment verification component 230, a crosswalk line clustering component 230, an alternate crosswalk identification component 240, a stop line clustering component 250, and a yield line clustering component 260.
[0041] The crosswalk identification component 220 may iterate through the list of detected navigation control lines 210 (whether on a tile and / or aggregate basis), iterate through detected crosswalk lines, search for other detected crosswalk lines within a threshold distance, and cluster adjacent and / or nearby crosswalk lines. For each pair of detected crosswalk lines within a threshold distance, for example, the intersection detector 160 may determine whether the crosswalk lines are substantially perpendicular and substantially parallel to each other, and if so, the intersection detector 160 may label and group the pairs of crosswalk lines as part of an inferred intersection. In some embodiments, the intersection detector 160 may adjust its threshold angles to identify crosswalk lines at different types of intersections (e.g., 3-way intersections, 4-way intersections, 5-way intersections, etc.).) in different iterations. In the . Fig. 2, the crosswalk identification component 220 includes the perpendicularity verification component 225, the parallel alignment verification component 230, the crosswalk line clustering component 230, and the alternative crosswalk identification component 240.
[0042] The perpendicularity check component 225 and the parallel alignment check component 230 may be part of an overall alignment check that determines whether the relative orientation of detected navigation control lines conforms to a predetermined geometric pattern presented in an intersection. In at least some implementations, the perpendicularity check 225 may determine whether the crosswalk lines are adjacent to each other (e.g., meet at an angle within a certain threshold range) and / or are within a certain threshold distance (e.g., one or two feet) of each other, which may be used as an indication that these crosswalk lines are part of the same crosswalk. For example, the perpendicularity check 225 may determine that a pair of detected crosswalk lines is substantially perpendicular within a specified angle threshold.In at least some of these examples, the crosswalk clustering component 235 may label the pair as part of a (e.g., inferred) intersection. In some examples, crosswalks typically include two parallel lines, and intersections may include multiple crosswalks. In at least some embodiments, any pair of crosswalk lines that passes the alignment check (e.g., both the perpendicular alignment check and the parallel alignment check) may be labeled by the crosswalk clustering component 235 as part of (e.g., clustered into) a detected intersection and / or used to trigger the parallel alignment check 230 to check the pair for parallelism.
[0043] Additionally, in at least some implementations, parallel alignment check 230 may determine when a pair of crosswalk lines are substantially coplanar (i.e., do not intersect) within a threshold distance (e.g., a few feet) of each other, which may be used as an indication that these crosswalk lines are part of the same crosswalk. Additionally or alternatively, in at least some examples, parallel alignment check 230 may perform an overlap check to determine whether one crosswalk line (e.g., of a pair of crosswalk lines) projects to substantially the entire length of the other crosswalk line within a threshold distance (e.g., a few feet).In some embodiments, if the parallel alignment check 230 determines that a pair of crosswalk lines is substantially parallel and / or substantially project onto each other, the parallel alignment check 230 may determine that these parallel crosswalk lines are part of the same crosswalk. In some implementations, the alignment check component 225 (e.g., both the perpendicular alignment check and the parallel alignment check) and the parallel alignment check component 230 may be executed in any order. For example, any crosswalk line that satisfies both alignment checks (or either check in some embodiments) may be labeled by the crosswalk clustering component 235 as part of (e.g., clustered into) a detected intersection.Therefore, in at least some examples, the crosswalk clustering component 235 may label the pair as part of a (e.g., inferred) intersection.
[0044] Depending on the type of intersection, the alternative crosswalk identification component 240 may adjust the threshold angles to detect crosswalk lines at different types of intersections. For example, the threshold angles for a perpendicularity check may be adjusted depending on whether the inferred intersection is a 3-way intersection, a 4-way intersection, a 5-way intersection, or another type of intersection. In at least some embodiments, regardless of the type of intersection, the alternative crosswalk identification component 240 may adjust the threshold angles so that an alignment check (e.g., both the perpendicularity check and the parallel alignment check) and an overlap check (e.g., of the parallel alignment check component 230) can be performed on pairs of detected crosswalk lines to detect the pairs and convert them into an (e.g., inferred) intersection (e.g.,B. a detected intersection).
[0045] Additionally or alternatively, other navigation control lines may be detected and labeled as part of the detected intersection. For example, the stop line clustering component 250 may iterate through the detected intersections (e.g., for each inferred intersection) and search for detected stop lines within a threshold distance (e.g., from the inferred intersection) to cluster nearby stop lines into a corresponding (e.g., inferred) intersection. In some embodiments, an alignment check (e.g., both the perpendicular alignment check and the parallel alignment check) may be performed over detected crosswalk lines in the detected intersection to identify stop lines within a threshold distance from the crosswalk lines. Therefore, for example, the stop line clustering component 250 may select a reference line (e.g.,one of the detected crosswalk lines of the detected intersection), select an appropriate threshold distance (e.g., based on the type of reference line and how far away stop lines typically appear from that line type), apply the specified distance threshold to determine whether the stop line is part of the same intersection as the reference line, and / or perform an alignment check for stop lines (e.g., a stop line parallel to a reference crosswalk line of the detected intersection and / or a stop line that is perpendicular to the reference crosswalk line). In some implementations, the stop line clustering component 250 may detect stop lines that satisfy this alignment check and / or the distance threshold and cluster those stop lines into the detected intersection.
[0046] Additionally or alternatively, in some implementations, the yield line clustering component 260 may iterate through the detected intersections (e.g., for each inferred intersection) and search for detected yield lines within a threshold distance to cluster nearby yield lines into the intersection. For example, the same alignment check (e.g., as used to detect lines within a threshold angular alignment of each other) and / or a corresponding distance threshold (e.g., yield lines within a few feet of crosswalk lines) may be performed on all detected crosswalk lines in the detected intersection to identify yield lines within a threshold distance of the crosswalk lines.Accordingly, detected yield lines that are within a threshold distance from a reference crosswalk line and satisfy the alignment check may be clustered into the detected intersection by the yield line clustering component 260.
[0047] Now with brief reference to Fig. 4A and Fig. 4B illustrate Fig. 4A and Fig. 4B illustrates exemplary embodiments of detected navigation control lines determined at alternative crosswalk alignments. Intersections are constructed in different sizes and configurations with different types and quantities of navigation control lines. For example, Fig. 4A shows a four-way intersection where six crosswalk lines 410 and four stop lines 420 can be recognized. In another example, Fig. 4B depicts a three-way intersection where two crosswalk lines 430, a stop line 440, and a yield line 450 are detected. As shown in these exemplary embodiments, regardless of the type of intersection (e.g., two-way, three-way, four-way, five-way, etc.) and the number of navigation control lines located at that intersection, the intersection detector 160 can detect navigation control lines that are part of a (e.g., inferred) intersection and cluster those navigation control lines into a detected intersection.
[0048] Taking another look at Fig. 2, each detected intersection may be formed from the detected crosswalk lines (e.g., identified by crosswalk identification component 230 and clustered into the inferred intersection), the detected stop lines (e.g., identified by stop line clustering component 250 and clustered into the inferred intersection), and / or the detected yield lines (identified by yield line clustering component 260 and clustered into the inferred intersection). Therefore, intersection detector 160 may detect each intersection (e.g., whether within a tile or across tiles) by inferring intersections based on the detected navigation control lines (e.g., crosswalk lines, stop lines, yield lines, etc.). In at least some embodiments, tile generator 170 may receive a representation of each detected intersection (e.g., indicated by arrow 270) from intersection detector 160.
[0049] With further reference to Fig. 1, the tile generator 170 may generate a new intersection-centered tile for each detected (e.g., inferred) intersection, and the intersection-centered tile may be centered around the detected intersection. In at least some examples, the tile generator 170 may generate and / or orient an intersection-centered (e.g., intersection-centric) tile to match the boundaries of an inferred intersection, which should serve to maximize or increase the number of intersection features represented within the intersection-centered tile. In some embodiments, the tile generator 170 may rotate and / or resize (e.g., scale) an intersection-centered tile into larger or smaller tiles (e.g., up to some threshold). For example, to facilitate feature detection within an intersection-centered tile, the tile generator 170 may rotate (e.g.,either clockwise or counterclockwise) to align it more closely with the boundaries of the detected intersection. In some embodiments, if the intersection is larger or smaller than a target resolution for an intersection-centric tile, the tile generator 170 may resize the intersection-centric tile by applying a scaling factor (e.g., 0.8 to 1.2). Additionally or alternatively, if an intersection is (e.g., significantly) larger than a specified threshold resolution for an intersection-centric tile (e.g., a representation of an intersection containing 1500 pixels compared to an intersection-centric tile with a limit of 1000 pixels), the tile generator 170 may subdivide the intersection-centric tile into smaller tiles.
[0050] The line detector 150 can detect navigation control lines from each intersection-centered tile generated using the tile generator 170 (as described above). Therefore, once some or all of the tiles are generated, the line detector 150 can rerun the detection of navigation control lines on each new intersection-centered tile. Therefore, according to some embodiments of the present disclosure, the intersection-centered tile can be applied to the machine learning model of the line detector 150.
[0051] With reference to Fig. 5A and Fig. 5B illustrates, for example, Fig. 5B an exemplary intersection-centered tile 550. In at least some embodiments, the 2D representation represents the ground (e.g., the global LiDAR map) of a geographic region, as in Fig. 5A. In Fig. 5A, a detected intersection was inferred and localized covering (e.g., spanning) four tiles separated by the grid 510 (e.g., corresponding to Fig. 3 illustrated exemplary grid). Continuing the example, Fig. 5A eight crosswalk lines (e.g., 520) and four stop lines (e.g., 530). In at least some examples, the tile generator 170 may Fig. 1 generate the intersection-centered tile 540 as in Fig. 5B, which is centered around the detected intersection. Therefore, the intersection-centered tile 540 may be centered around the detected intersection. In at least some examples, the intersection-centered tile 540 may be applied to the line detector 150, which may detect and classify navigation control lines within the intersection-centered tile 540.
[0052] With further reference to Fig. 1, the post-processing component 180 may aggregate and / or deduplicate detected navigation control lines. In some embodiments, local sectors may overlap, which may result in duplicated detected navigation control lines (such as duplicates of navigation control lines detected by the line generator 150, for example, when the tiles overlap in a grid). In some examples, the post-processing component 180 may search for navigation control lines that are (e.g., substantially) close to each other (e.g., by determining the distance between the lines and applying a corresponding distance threshold) and perform an overlap check. In some embodiments, the post-processing component 180 (operating, for example, similarly to the intersection detector 160) may determine the threshold distances (e.g.,within a certain number of inches or centimeters) between the duplicated navigation control lines so that an alignment check (e.g., a parallel alignment check) and an overlap check (e.g., the parallel alignment check component 230) can be performed to detect duplicates of detected navigation control lines. If a pair of navigation control lines overlaps significantly more than a certain threshold (such as a 99 percent overlap), duplicate lines can be de-duplicated, for example, by the post-processing component 180. Additionally or alternatively, crosswalk lines at the same intersection can be identified (e.g.,based on an alignment check and within a certain distance threshold) and paired by the post-processing component 180 to form corresponding bounding boxes, polygons, or other connecting shapes.
[0053] Now with reference to Fig. 6 illustrates Fig. 6 shows an example of deduplicating detected navigation control lines (e.g., crosswalk lines and stop lines). As illustrated in the example left detected intersection 610 in a labeled map, multiple duplicated crosswalk lines 620 and duplicated stop lines 630 are present (e.g., depicted as thicker or overlapping lines). Therefore, the post-processing component 180 may apply deduplication to detected navigation control lines (e.g., represented by arrow 640). Additionally or alternatively, the post-processing component 180 may pair deduplicated crosswalk lines (e.g., represented by arrow 650) to form polygons 670, as illustrated in the example right detected intersection 660 in a labeled map. In at least some examples, these polygons may be used to label (and / or, e.g., update) a map.
[0054] With further reference to Fig. 1, the map labeling component 190 may use the detected polygons and / or the detected navigation control lines to label or otherwise update the map (e.g., global LiDAR map). For example, the detected navigation control lines may be composed of different tiles (e.g., those in Fig. 3, each tile (referred to as tile 320) and any other tile from grid 310 and / or any intersection-centered tile) may be aggregated and used to label a 2D representation of the ground or other surface (e.g., a LiDAR map). Therefore, detected navigation control lines, detected regions bounded by detected navigation control lines, and / or corresponding class labels on the map (e.g., global LiDAR map) may be labeled by the map labeling component 190.
[0055] With reference to Fig. 7 and Fig. 8, each block of the methods 700 and 800 described herein comprises a computational process that may be performed using any combination of hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. The methods may also be embodied as computer-usable instructions stored on computer storage media. The methods may be provided by a standalone application, a service or hosted service (standalone or in combination with another hosted service), or a plug-in for another product, to name a few. Additionally, the methods 700 and 800 are described by way of example with respect to the map generation pipeline 100 of Fig. 1. However, these methods may additionally or alternatively be performed by any system or combination of systems, including, but not limited to, those described herein.
[0056] Fig. 7 is a flowchart illustrating a method 700 for updating a map with detected navigation control lines, according to some embodiments of the present disclosure. The method 700 includes, at block B702, detecting, based at least on applying a representation of an intersection-centered tile of a map of a 2D surface to a neural network, one or more navigation control lines represented in the intersection-centered tile. The map represents, for example, projected LiDAR intensity data, and the line detector 150 may use a neural network (or other machine learning model) to detect and classify lines (e.g., navigation control lines) from intersection-centered tiles of the LiDAR map. Additionally or alternatively, the line detector 150 may decode a machine learning model output (e.g., DNN) and generate a list of the detected lines and corresponding class labels (e.g.,which label the navigation control lines as crosswalk lines, stop lines or yield lines).
[0057] The method 700 includes, at block B704, updating the map based on at least one or more navigation control lines. For example, the map labeling component 190 may label a map (e.g., global, LiDAR map) with detected navigation control lines, detected regions bounded by detected navigation control lines, and / or corresponding class labels. Therefore, the map labeling component 190 may use the detected navigation control lines to update the map (e.g., global, LiDAR map).
[0058] Fig. 8 is a flowchart illustrating a method 800 for detecting a refined set of navigation control lines, according to some embodiments of the present disclosure. The method 800 includes, at block B802, detecting an initial set of navigation control lines from one or more tiles of the map. For example, the line detector 150 may be configured with respect to the map generation pipeline 100 of Fig. 1 Use a machine learning model, such as a DNN, to detect and classify navigation control lines within one or more tiles of a LiDAR map.
[0059] The method 800 includes, at block B804, generating a representation of one or more detected intersections based at least on the clustering of the initial set of navigation control lines. For example, the tile generator 170 may generate a new intersection-centered tile for each detected (e.g., inferred) intersection, and the intersection-centered tile may be centered by the tile generator 170 around the detected intersection. In at least some examples, the tile generator 170 may generate and / or orient an intersection-centered tile to correspond to the boundaries of an inferred intersection, which should serve to maximize or increase the number of intersection features represented within the intersection-centered tile.
[0060] The method 800 includes, in block B806, detecting a refined set of navigation control lines from one or more intersection-centered tiles associated with one or more detected intersections. For example, the line detector 150 may detect navigation control lines from j detect any intersection-centered tile generated using the tile generator 170 and aggregate the detected navigation control lines to capture the navigation control lines located at the detected intersection (e.g., a refined set of navigation control lines).
[0061] The systems and methods described herein may be used by, without limitation, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), guided and unguided robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, hydrofoils, boats, shuttle vehicles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, trains, underwater vehicles, remotely operated vehicles such as drones, and / or other types of vehicles.Furthermore, the systems and methods described herein may be used for a variety of purposes, including, but not limited to, machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, generative AI, and / or other suitable applications.
[0062] The embodiments presented here may include a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented with a robot, aviation systems, media systems, boat systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems including one or more virtual machines (VMs), systems for performing operations for generating synthetic data, systems implemented at least partially in a data center, systems for performing operations with conversational AI, systems,that implement one or more language models - such as one or more large language models (LLMs), systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems that are implemented at least partially using cloud computing resources, and / or other types of systems. EXEMPLARY AUTONOMOUS VEHICLE
[0063] Fig. 9A illustrates an example autonomous vehicle 900, according to some embodiments of the present disclosure. The autonomous vehicle 900 (alternatively referred to herein as "vehicle 900") may include, without limitation, a passenger vehicle, such as a car, a truck, a bus, an emergency service vehicle, a shuttle, an electric or motorized bicycle, a motorcycle, a fire engine, a police vehicle, an ambulance, a boat, a construction vehicle, an underwater vehicle, a robotic vehicle, a drone, an aircraft, a vehicle coupled to a trailer (e.g., a semi-truck used to transport cargo), and / or another type of vehicle (e.g., that is unmanned and / or accommodates one or more passengers).Autonomous vehicles are generally described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and prior and future versions of this standard). The vehicle 900 may exhibit functionality consistent with one or more of the Level 3 through Level 5 autonomous driving levels.The vehicle 900 may exhibit functionality according to one or more of Level 1 through Level 5 of autonomous driving levels. For example, depending on the embodiment, the vehicle 900 may be capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomous" as used herein may include any and / or all types of autonomy for the vehicle 900 or other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, assisted autonomy, semi-autonomous, primarily autonomous, or another designation.
[0064] The vehicle 900 may include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. The vehicle 900 may include a propulsion system 950, such as an internal combustion engine, a hybrid electric power plant, a pure electric motor, and / or another type of propulsion system. The propulsion system 950 may be connected to a drivetrain of the vehicle 900, which may include a transmission to enable propulsion of the vehicle 900. The propulsion system 950 may be controlled in response to receiving signals from the throttle or accelerator 952.
[0065] A steering system 954, which may include a steering wheel, may be used to steer the vehicle 900 (e.g., along a desired path or route) when the propulsion system 950 is operating (e.g., when the vehicle is moving). The steering system 954 may receive signals from a steering actuator 956. The steering wheel may be optional for full automation (Level 5).
[0066] The brake sensor system 946 may be used to apply the vehicle brakes in response to receiving signals from the brake actuators 948 and / or the brake sensors.
[0067] The one or more controllers 936 that control one or more systems on chips (SoCs) 904 ( Fig. 9C) and / or GPUs may provide signals (e.g., representative of instructions) to one or more components and / or systems of the vehicle 900. For example, the one or more controllers may send signals to apply the vehicle brakes via one or more brake actuators 948, to apply the steering system 954 via one or more steering actuators 956, to apply the propulsion system 950 via one or more throttle / accelerator devices 952. The one or more controllers 936 may include one or more built-in (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and issue operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 900.The one or more controllers 936 may include a first controller 936 for autonomous driving functions, a second controller 936 for functional safety functions, a third controller 936 for artificial intelligence functions (e.g., computer vision), a fourth controller 936 for infotainment functions, a fifth controller 936 for emergency redundancy, and / or other controllers. In some examples, a single controller 936 may perform two or more of the above functionalities, two or more controllers 936 may perform a single functionality, and / or any combination thereof.
[0068] The one or more controllers 936 may provide the signals to control one or more components and / or systems of the vehicle 900 in response to sensor data received from one or more sensors (e.g., sensor inputs). The sensor data may, for example, B. and without limitation, received from: global navigation satellite systems (“GNSS”) sensor(s) 958 (e.g., global positioning system sensor(s)), RADAR sensor(s) 960, ultrasonic sensor(s) 962, LiDAR sensor(s) 964, inertial measurement unit (IMU) sensor(s) 966 (e.g., accelerometer(s), gyroscope(s), magnetic compass(es), magnetometer(s), etc.), microphone(s) 996, stereo camera(s) 968, wide-angle camera(s) 970 (e.g., fisheye cameras), infrared camera(s) 972, ambient camera(s) 974 (e.g.,360-degree cameras), long-range and / or medium-range camera(s) 998, speed sensor(s) 944 (e.g., for measuring the speed of the vehicle 900), vibration sensor(s) 942, steering sensor(s) 940, brake sensor(s) (e.g., as part of the brake sensor system 946), and / or one or more occupant monitoring system (OMS) sensors 901 (e.g., one or more interior cameras) and / or other sensor types.
[0069] One or more of the controllers 936 may receive inputs (e.g., in the form of input data) from an instrument cluster 932 of the vehicle 900 and provide outputs (e.g., in the form of output data, display data, etc.) via a human-machine interface (HMI) display 934, an audible annunciator, a speaker, and / or via other components of the vehicle 900. The outputs may include information such as vehicle speed, RPM, time, map data (e.g., the high-definition (“HD”) map 922 of Fig. 9C), location data (e.g., the location of the vehicle 900, e.g., on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and status of objects as perceived by the one or more controllers 936, etc. For example, information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about maneuvers that the vehicle has performed, is currently performing, or will perform (e.g., change lanes now, take exit 34B in two miles, etc.) may be displayed on the HMI display 934.
[0070] The vehicle 900 further includes a network interface 924 that may utilize one or more wireless antennas 926 and / or modems for communication over one or more networks. The network interface 924 may, for example, be capable of communication over Long-Term Evolution (LTE), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Global System for Mobile Communication (GSM), IMT-CDMA Multi-Carrier (CDMA2000), etc. The one or more wireless antennas 926 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.and / or Low Power Wide Area Networks (LPW ANs), such as LoRaWAN, SigFox, etc.
[0071] Fig. 9B is an example of camera locations and fields of view for the example autonomous vehicle 900 of Fig. 9A, according to some embodiments of the present disclosure. The cameras and respective fields of view are an example of one embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at various locations on the vehicle 900.
[0072] The camera types for the cameras may include, but are not limited to, digital cameras that may be configured for use with the components and / or systems of the vehicle 900. The one or more cameras may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera types may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The cameras may use rolling shutters, global shutters, another type of shutter, or a combination thereof.In some examples, the color filter array may include a red clear clear clear (RCCC) color filter array, a red clear clear blue (RCCB) color filter array, a red blue green clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor color filter array (RGGB), a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear-pixel cameras, such as cameras with an RCCC, an RCCB, and / or an RBGC color filter array, may be used to increase light sensitivity.
[0073] In some examples, one or more of the cameras can be used to run advanced driver assistance systems (ADAS) (e.g., as part of a redundant or fail-safe design). For example, a multifunction mono camera can be installed to provide features including lane departure warning, traffic sign assist, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0074] One or more of the cameras may be mounted in a bracket, such as a specially designed (three-dimensional ("3D") printed) bracket, to eliminate stray light and reflections from the vehicle interior (e.g., reflections from the dashboard reflected in the windshield mirrors) that could interfere with the camera's image data acquisition. With regard to the bracket for exterior mirrors, the exterior mirrors may be custom 3D printed so that the camera mounting plate is adapted to the shape of the exterior mirror. In some examples, the one or more cameras may be integrated into the exterior mirror. For side-mounted cameras, the one or more cameras may also be integrated into the four pillars at each corner of the cabin.
[0075] Cameras with a field of view that includes portions of the environment in front of the vehicle 900 (e.g., forward-facing cameras) can be used for the surrounding view to help identify forward paths and obstacles, and to provide information critical to establishing an occupancy grid and / or determining preferred vehicle paths with the assistance of one or more controllers 936 and / or control SoCs. Forward-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras can also be used for ADAS features and systems that include lane departure warnings (LDW), autonomous cruise control (ACC), and / or other features such as traffic sign detection.
[0076] A variety of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform that includes a complementary metal oxide semiconductor (CMOS) color imager. Another example is the wide-angle cameras 970, which can be used to capture objects that enter the field of view from the periphery (e.g., pedestrians, crossing vehicles, or bicycles). Although Fig. 9B illustrates only one wide-angle camera, the vehicle 900 may include any number (including zero) of wide-angle cameras 970. Furthermore, any number of remote cameras 998 (e.g., a remote stereo camera pair) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The one or more remote cameras 998 may also be used for object detection and classification, as well as basic object tracking.
[0077] Any number of stereo cameras 968 may also be included in a forward-facing configuration. In at least one embodiment, one or more of the stereo cameras 968 may include an integrated control unit comprising a scalable processing unit that can provide a programmable logic ("FPGA") and a multi-core microprocessor with an integrated controller area network ("CAN") or Ethernet interface on a single chip. Such a unit can be used to create a 3D map of the vehicle's surroundings that includes a distance estimate for all points in the image. One or more alternative stereo cameras 968 may include a compact stereo vision sensor that can include two camera lenses (one each on the left and right) and an image processing chip that can measure the distance between the vehicle and the target object and process the generated information (e.g.,metadata) to activate the autonomous emergency braking and lane keeping functions. Other types of stereo cameras 968 may be used in addition to or alternatively to those described here.
[0078] Cameras with a field of view that includes portions of the environment to the side of the vehicle 900 (e.g., side cameras) may be used for the environment view and provide information used to create and update the occupancy grid and to generate collision warnings in the event of a side impact. For example, the one or more environment cameras 974 (e.g., four environment cameras 974, as in Fig. 9B) may be positioned on the vehicle 900. The one or more surround cameras 974 may include one or more wide-angle cameras 970, one or more fisheye cameras, one or more 360-degree cameras, and / or the like. For example, four fisheye cameras may be mounted on the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may utilize three surround camera(s) 974 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround camera.
[0079] Cameras with a field of view that includes portions of the environment behind the vehicle 900 (e.g., rearview cameras) may be used for parking assistance, surround view, rear collision warnings, and occupancy grid creation and updating. A variety of cameras may be used, including, but not limited to, cameras that are also suitable as one or more forward-facing cameras (e.g., one or more long-range and / or medium-range cameras 998, one or more stereo cameras 968, one or more infrared cameras 972, etc.), as described herein.
[0080] Cameras with a field of view including portions of the interior environment within the cabin of the vehicle 900 (e.g., one or more OMS sensors 901) may be used as part of an occupant monitoring system (OMS), such as, but not limited to, a driver monitoring system (DMS). For example, OMS sensors (e.g., the one or more OMS sensors 901) may be used (e.g., by the one or more controllers 936) to track an occupant's and / or driver's gaze direction, head posture, and / or blinking. This gaze information may be used to determine the occupant's or driver's level of attention (e.g., to detect drowsiness, fatigue, and / or distraction) and / or to take appropriate action to prevent harm to the occupant or driver.In some embodiments, data from OMS sensors may be used to enable gaze-controlled operations initiated by the driver and / or other occupants, such as, but not limited to, adjusting cabin temperature and / or airflow, opening and closing windows, controlling cabin lighting, controlling entertainment systems, adjusting mirrors, adjusting seat positions, and / or other operations. In some embodiments, an OMS may be used for applications such as determining whether items and / or occupants have been left behind in a vehicle cabin (e.g., by detecting the presence of occupants after the driver has exited the vehicle).
[0081] Fig. 9C is a block diagram of an example system architecture for the example autonomous vehicle 900 of Fig. 9A, according to some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are set forth as examples only. Other arrangements and elements (e.g., engines, interfaces, functions, arrangements, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as individual or distributed components, or in conjunction with other components, and in any suitable combination and location. Various functions performed by entities described herein may be performed by hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory.
[0082] Each of the vehicle’s components, functions and systems 900 in Fig. 9C is illustrated as being connected via bus 902. Bus 902 may include a controller area network (CAN) data interface (alternatively referred to herein as a "CAN bus"). A CAN may be a network within vehicle 900 that serves to support the control of various features and functions of vehicle 900, such as the application of brakes, acceleration, braking, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, of which j Each has its own unique identifier (e.g., a CAN ID). The CAN bus can be read to determine the steering wheel angle, vehicle speed, engine speed (rpm), button positions, and / or other vehicle status indicators. The CAN bus can be ASIL B compliant.
[0083] Although bus 902 is described herein as a CAN bus, this is not intended to be a limitation. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to the CAN bus. Furthermore, while a single wire is used to represent bus 902, this is not intended to be a limitation. For example, there may be any number of buses 902, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses that use a different protocol. In some examples, two or more buses 902 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 902 may be used for collision avoidance functionality and a second bus 902 may be used for actuation control.In each example, each bus 902 may communicate with one of the components of the vehicle 900, and two or more buses 902 may communicate with the same components. In some examples, each SoC 904, each controller 936, and / or each computer within the vehicle may have access to the same input data (e.g., inputs from sensors of the vehicle 900) and may be connected to a common bus, such as the CAN bus.
[0084] The vehicle 900 may include one or more controllers 936 as described herein with respect to Fig. 9A. The one or more controllers 936 may be used for a variety of functions. The one or more controllers 936 may be coupled to any of the various other components and systems of the vehicle 900 and may be used for control of the vehicle 900, artificial intelligence of the vehicle 900, infotainment for the vehicle 900, and / or the like.
[0085] The vehicle 900 may include one or more systems on a chip (SoC) 904. The SoC 904 may include one or more CPUs 906, one or more GPUs 908, one or more processors 910, one or more caches 912, one or more accelerators 914, one or more data memories 916, and / or other unillustrated components and features. The one or more SoCs 904 may be used to control the vehicle 900 in a variety of platforms and systems. For example, the one or more SoCs 904 in a system (e.g., the system of the vehicle 900) may be combined with an HD card 922 that may be accessed via a network interface 924 from one or more servers (e.g., the one or more servers 978 of Fig. 9D) can receive map refreshes and / or updates.
[0086] The one or more CPUs 906 may include a CPU cluster or CPU complex (alternatively referred to herein as a "CCPLEX"). The one or more CPUs 906 may include multiple cores and / or L2 caches. For example, in some embodiments, the one or more CPUs 906 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the one or more CPUs 906 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The one or more CPUs 906 (e.g., the CCPLEX) may be configured to support concurrent operation of clusters, such that any combination of the clusters of the one or more CPUs 906 may be active at any given time.
[0087] The one or more CPUs 906 may implement power management functions that include one or more of the following features: individual hardware blocks may be automatically clocked when idle to conserve dynamic power; each core clock may be cycled through when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be cycled through independently; jEach core cluster may be independently clock-controlled if all cores are clock-controlled or power-controlled; and / or each core cluster may be independently power-controlled if all cores are power-controlled. The one or more CPUs 906 may further implement an enhanced power state management algorithm in which allowable power states and expected wake-up times are determined, and the hardware / microcode determines the best power state to enter for the core, cluster, and CCPLEX. The processing cores may support simplified power state entry sequences in software, offloading the work to the microcode.
[0088] The one or more GPUs 908 may include an integrated GPU (alternatively referred to herein as a "GPU"). The one or more GPUs 908 may be programmable and may be efficient for parallel workloads. The one or more GPUs 908 may, in some examples, utilize an enhanced Tensor instruction set. The one or more GPUs 908 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96 KB of memory capacity), and two or more of the streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512 KB of memory capacity). In some embodiments, the one or more GPUs 908 may include at least eight streaming microprocessors. The one or more GPUs 908 may utilize one or more application programming interfaces (APIs) for computations.In addition, the one or more GPUs 908 may use one or more parallel computing platforms and / or programming models (e.g., CUDA from NVIDIA).
[0089] The one or more GPUs 908 may be power-optimized for best performance in automotive and embedded use cases. For example, the one or more GPUs 908 may be fabricated on a fin field-effect transistor (FinFET). However, this is not a limitation, and the one or more GPUs 908 may also be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor may include a number of mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks. In such an example, each processing block can be assigned 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file.In addition, the streaming microprocessors can include independent parallel integer and floating-point datapaths to enable efficient execution of workloads with a mix of computations and addressing calculations. The streaming microprocessors can include an independent thread scheduling function to enable fine-grained synchronization and cooperation between parallel threads. The streaming microprocessors can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0090] The one or more GPUs 908 may include high-bandwidth memory (HBM) and / or a 16 GB HBM2 subsystem to provide, in some examples, a peak memory bandwidth of approximately 900 GB / second. In some examples, synchronous graphics random-access memory (SGRAM), such as double data rate synchronous memory type five (GDDR5), may be used in addition to or as an alternative to the HBM memory.
[0091] The one or more GPUs 908 may include unified memory technology that includes access counters to enable more accurate migration of memory pages to the processor that accesses them most frequently, thereby improving efficiency for processor-shared memory areas. In some examples, Address Translation Services (ATS) support may be used to allow the one or more GPUs 908 to directly access the page tables of the one or more CPUs 906. In such examples, if the memory management unit (MMU) of the one or more GPUs 908 fails, an address translation request may be sent to the one or more CPUs 906.In response, the one or more CPUs 906 may look up the virtual-to-physical mapping for the address in their page tables and send the translation back to the one or more GPUs 908. Thus, the unified memory technology may enable a single unified virtual address space for the memory of both the one or more CPUs 906 and the one or more GPUs 908, thereby simplifying the programming of the one or more GPUs 908 and the porting of applications to the one or more GPUs 908.
[0092] Additionally, the one or more GPUs 908 may include an access counter that can track the frequency of access by the one or more GPUs 908 to memory of other processors. The access counter can help move memory pages to the physical memory of the processor that accesses the pages most frequently.
[0093] The one or more SoCs 904 may include any number of caches 912, including those described herein. For example, the one or more caches 912 may include an L3 cache available to both the one or more CPUs 906 and the one or more GPUs 908 (e.g., connected to both the one or more CPUs 906 and the one or more GPUs 908). The one or more caches 912 may include a write-back cache that may track line states, e.g., by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include 4 MB or more, depending on the embodiment, although smaller cache sizes may also be used.
[0094] The one or more SoCs 904 may include one or more arithmetic logic units (ALUs) that may be utilized in performing processing related to any of the various tasks or operations of the vehicle 900, such as processing DNNs. Additionally, the one or more SoCs 904 may include one or more floating-point units (FPUs)—or other mathematical co-processors or numerical co-processors—for performing mathematical operations within the system. For example, the one or more SoCs 904 may include one or more FPUs integrated as execution units within one or more CPUs 906 and / or one or more GPUs 908.
[0095] The one or more SoCs 904 may include one or more accelerators 914 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the one or more SoCs 904 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. The large on-chip memory (e.g., 4MB SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster may be used to complement the one or more GPUs 908 and offload some of the tasks of the one or more GPUs 908 (e.g., to free up more cycles of the one or more GPUs 908 to perform other tasks). For example, the one or more accelerators 914 may be configured for targeted workloads (e.g.,Perception, convolutional neural networks (CNNs), etc.) that are robust enough to be suitable for acceleration can be used. The term "CNN" as used here can include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0096] The one or more accelerators 914 (e.g., the hardware acceleration cluster) may include a deep learning accelerator (DLA). The one or more DLAs may include one or more tensor processing units (TPUs) configured to provide an additional tens of trillion operations per second for deep learning applications and inferencing. The TPUs may be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The one or more DLAs may also be optimized for a specific set of neural network types and floating-point operations, as well as for inferencing. The design of the one or more DLAs may provide more performance per millimeter than a general-purpose GPU, far exceeding the performance of a CPU.The one or more TPUs can perform multiple functions, including a single-instance convolution function that supports, for example, INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions.
[0097] The one or more DLAs can quickly and efficiently execute neural networks, particularly CNNs, on processed or unprocessed data for a variety of functions, including, for example and without limitation: a CNN for object identification and detection using data from camera sensors; a CNN for distance estimation using data from camera sensors; a CNN for emergency vehicle detection and identification using data from microphones; a CNN for facial detection and vehicle owner identification using data from camera sensors; and / or a CNN for safety and / or security events.
[0098] The one or more DLAs can perform any function of the one or more GPUs 908, and by using an inference accelerator, a developer can, for example, dedicate either the one or more DLAs or the one or more GPUs 908 to each function. For example, the developer can focus the processing of CNNs and floating-point operations on the one or more DLAs and leave other functions to the one or more GPUs 908 and / or other accelerators 914.
[0099] The one or more accelerators 914 (e.g., the hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may also be referred to herein as a computer vision accelerator. The one or more PVAs may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The one or more PVAs may provide a balance between performance and flexibility. For example, and without limitation, each PVA may include any number of reduced instruction set computer (RISC) cores, direct memory access (DMA) cores, and / or any number of vector processors.
[0100] The RISC cores may interact with image sensors (e.g., the image sensors of any of the cameras described herein), image signal processors, and / or the like. Each of the RISC cores may include any amount of memory. The RISC cores may use any number of protocols, depending on the embodiment. In some examples, the RISC cores may execute a real-time operating system (RTOS). The RISC cores may be implemented with one or more integrated circuits, application-specific integrated circuits (ASICs), and / or memory devices. The RISC cores may include, for example, an instruction cache and / or tightly coupled RAM.
[0101] The DMA may enable components of the PVA(s) to access the system's memory independently of the one or more CPUs 906. The DMA may support any number of features designed to optimize the PVA, including, but not limited to, support for multi-dimensional addressing and / or circular addressing. In some examples, the DMA may support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0102] The vector processors may be programmable processors that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing functions. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or memory (e.g., VMEM). A VPU core may include a digital signal processor, such asA digital signal processor with single instruction multiple data (SIMD) and very long instruction words (VLIW). Combining SIMD and VLIW can increase throughput and speed.
[0103] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Therefore, in some examples, each of the vector processors may be configured to operate independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallelism. For example, in some embodiments, the multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may concurrently execute different computer vision algorithms on the same image, or even execute different algorithms on consecutive images or portions of an image.Among other things, any number of PVAs can be included in the hardware acceleration cluster, and any number of vector processors can be included in each of the PVAs. Furthermore, the one or more PVAs can contain additional memory for error correcting code (ECC) to increase the overall security of the system.
[0104] The one or more accelerators 914 (e.g., the hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM to the one or more accelerators 914. In some examples, the on-chip memory may include at least 4 MB of SRAM, consisting of, for example, and without limitation, eight field-configurable memory blocks accessible by both the PVA and the DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory through a backbone that provides high-speed access to the memory to the PVA and the DLA.The backbone may include an on-chip computer vision network connecting the PVA and DLA to the memory (e.g., using the APB).
[0105] The on-chip computer vision network can include an interface that determines that both the PVA and the DLA are delivering ready and valid signals before transmitting control signals / addresses / data. Such an interface can provide separate phases and channels for transmitting control signals / addresses / data, as well as bursty communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508, although other standards and protocols can also be used.
[0106] In some examples, the one or more SoCs 904 may include a real-time ray tracing hardware accelerator, as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model) to generate real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for simulation of sonar systems, for general wave propagation simulation, for comparison with LiDAR data for localization, and / or for other functions, and / or for other purposes. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more operations related to ray tracing.
[0107] The one or more accelerators 914 (e.g., the hardware accelerator cluster) have a wide range of uses for autonomous driving. The PVA can be a programmable vision accelerator that can be used for critical processing steps in ADAS and autonomous vehicles. The capabilities of the PVA are well suited to algorithmic areas that require predictable processing with low power and low latency. In itself, the PVA is well suited for semi-dense or dense regular computations, even on small datasets, that require predictable runtimes with low latency and low power. In the context of autonomous vehicle platforms, the PVAs are therefore designed to execute classical computer vision algorithms because they are efficient at object detection and operate on integer mathematics.
[0108] According to one embodiment of the technology, the PVA is used, for example, to perform computer stereo vision. In some examples, a semi-global matching-based algorithm may be used, although this is not intended as a limitation. Many applications for Level 3-5 autonomous driving require on-the-fly motion estimation or stereo matching (e.g., structure from motion, pedestrian detection, lane detection, etc.). The PVA can perform a computer stereo vision function on inputs from two monocular cameras.
[0109] In some examples, the PVA can be used to perform dense optical flow, such as processing raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar. In other examples, the PVA is used for depth-of-flight processing, such as processing raw flight data to provide processed flight time data.
[0110] The DLA can be used to power any type of network to improve control and driving safety; for example, this includes a neural network that outputs a confidence measure for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative "weight" to each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a confidence threshold and consider only those detections that exceed the threshold as true positives.In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most certain detections should be considered as triggers for AEB. The DLA may employ a neural network to regress the confidence value. The neural network may take as input at least a subset of parameters, such as, but not limited to, the bounding box dimensions, the ground plane estimate obtained (e.g., from another subsystem), the output of the inertial measurement unit (IMU) sensor 966 correlated with the orientation of the vehicle 900, distance, 3D position estimates of the object obtained by the neural network and / or other sensors (e.g., one or more LiDAR sensors 964 or one or more RADAR sensors 960).
[0111] The one or more SoCs 904 may include one or more data stores 916 (e.g., memory). The one or more data stores 916 may be on-chip memory on the one or more SoCs 904 in which neural networks to be executed on the GPU and / or the DLA may be stored. In some examples, the one or more data stores 916 may be large enough to store multiple instances of neural networks for redundancy and security. The one or more data stores 916 may include one or more L2 or L3 caches 912. The reference to the one or more data stores 916 may include a reference to the memory associated with the PVA, the DLA, and / or one or more other accelerators 914, as described herein.
[0112] The one or more SoCs 904 may include one or more processors 910 (e.g., embedded processors). The one or more processors 910 may include a boot and power management processor, which may be a dedicated processor and subsystem to handle boot power and management functions and associated security enforcement. The boot and power management processor may be part of the boot sequence of the one or more SoCs 904 and may provide runtime power management services. The boot and power management processor may provide clock and voltage programming, assisting with system transitions to a low-power state, managing the thermals and temperature sensors of the one or more SoCs 904, and / or managing the one or more SoCs 904 power states.Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the one or more SoCs 904 may use the ring oscillators to sense the temperatures of the one or more CPUs 906, the one or more GPUs 908, and / or the one or more accelerators 914. If it is determined that the temperatures exceed a threshold, the boot and power management processor may enter a temperature fault routine and place the one or more SoCs 904 into a lower power state and / or place the vehicle 900 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 900 to a safe stop).
[0113] The one or more processors 910 may also include a set of embedded processors that can serve as an audio processing engine. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0114] The one or more processors 910 may also include an always-on processor engine that can provide the necessary hardware functions to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0115] The one or more processors 910 may also include a safety cluster engine containing a dedicated processor subsystem for safety management of automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, an interrupt controller, etc.), and / or routing logic. In a safety mode, the two or more cores may operate in a lockstep mode, functioning as a single core with comparison logic that captures any differences between their operations.
[0116] The one or more processors 910 may also include a real-time camera engine, which may include a dedicated processor subsystem for managing the real-time camera.
[0117] The one or more processors 910 may further include a high dynamic range signal processor, which may include an image signal processor that is a hardware engine that is part of the camera processing pipeline.
[0118] The one or more processors 910 may include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the one or more wide-angle cameras 970, the one or more surround cameras 974, and / or the in-cabin surveillance camera sensors. The in-cabin surveillance camera sensor is preferably monitored by a neural network running on another instance of the enhanced SoC and configured to detect events in the cabin and respond accordingly.An in-cabin system can perform lip reading to activate cellular service and make a call, dictate emails, change the destination, activate or change the vehicle's infotainment system and settings, or enable voice-activated web browsing. Certain functions are available to the driver only when the vehicle is operating in autonomous mode and are disabled otherwise.
[0119] The video image compositor can incorporate enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, when motion occurs in a video, the noise reduction weights the spatial information accordingly, reducing the weight of information provided by neighboring frames. If a frame or section of a frame contains no motion, the temporal noise reduction performed by the video image compositor can use information from the previous frame to reduce noise in the current frame.
[0120] The video image compositor may also be configured to perform stereo distortion correction on the input stereo lens images. The video image compositor may also be used for user interface design when the operating system desktop is in use and the GPU(s) 908 do not need to constantly render new surfaces. Even when the GPU(s) 908 are powered on and actively performing 3D rendering, the video image compositor may be used to offload the GPU(s) 908, thereby improving performance and responsiveness.
[0121] The one or more SoCs 904 may also include a Mobile Industry Processor Interface (MIPI) serial camera interface for receiving video and input from cameras, a high-speed interface, and / or a video input block that may be used for camera and related pixel input functions. The one or more SoCs 904 may also include one or more software-controlled input / output controllers that may be used to receive I / O signals that are not associated with a specific role.
[0122] The one or more SoCs 904 may further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The one or more SoCs 904 may be used to process data from cameras (e.g., via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., one or more LiDAR sensors 964, one or more RADAR sensors 960, etc., which may be connected via Ethernet), data from bus 902 (e.g., speed of vehicle 900, steering wheel position, etc.), and data from one or more GNSS sensors 958 (e.g., connected via Ethernet or CAN bus).The one or more SoCs 904 may further include dedicated high-performance mass storage controllers, which may include their own DMA engines and which may be used to offload routine data management tasks from the one or more CPUs 906.
[0123] The one or more SoCs 904 may be an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture that supports and efficiently utilizes computer vision and ADAS techniques for diversity and redundancy, and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The one or more SoCs 904 may be faster, more reliable, and even more energy and space efficient than conventional systems. For example, the one or more accelerators 914, in combination with the one or more CPUs 906, the one or more GPUs 908, and the one or more data memories 916, may form a fast, efficient platform for Level 3-5 autonomous vehicles.
[0124] The technology thus offers capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs that can be configured using a high-level programming language, such as the C programming language, to execute a variety of processing algorithms on a wide variety of visual data. However, CPUs are often unable to meet the performance requirements of many image processing applications, such as execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a prerequisite for in-vehicle ADAS applications and a requirement for practical Level 3-5 autonomous vehicles.
[0125] In contrast to conventional systems, the technology described herein enables the simultaneous and / or sequential execution of multiple neural networks and the combination of the results to enable Level 3-5 autonomous driving functionality by providing a CPU complex, a GPU complex, and a hardware acceleration cluster. For example, a CNN running on the DLA or dGPU (e.g., the one or more GPUs 920) may include text and word recognition, allowing the supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA may further include a neural network capable of identifying, interpreting, and providing semantic understanding of the sign, and passing this semantic understanding to the path planning modules running on the CPU complex.
[0126] Another example is that multiple neural networks can run simultaneously, as required for Level 3, 4, or 5 driving. For example, a warning sign reading "Caution: Flashing lights indicate black ice," along with an electric light, can be interpreted independently or jointly by multiple neural networks. The sign itself can be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing lights indicate black ice" can be interpreted by a second deployed neural network, which informs the vehicle's path-planning software (preferably running on the CPU complex) that if flashing lights are detected, black ice is present.The turn signal can be identified across multiple images by a third neural network, which informs the vehicle's path planning software of the presence (or absence) of turn signals. All three neural networks can run simultaneously, e.g., within the DLA and / or on the one or more GPUs 908.
[0127] In some examples, a facial recognition and vehicle owner identification CNN may use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 900. The always-on sensor processing engine may be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner exits the vehicle. In this way, the one or more SoCs 904 provide security against theft and / or carjacking.
[0128] In another example, a CNN for detecting and identifying emergency vehicles may use data from microphones 996 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers to detect sirens and manually extract features, the one or more SoCs 904 use the CNN to classify environmental and urban sounds, as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to detect the relative approach speed of the emergency vehicle (e.g., by using the Doppler effect). The CNN may also be trained to identify emergency vehicles specific to the local area in which the vehicle operates, as identified by one or more GNSS sensors 958.For example, if the CNN is operating in Europe, it will attempt to detect European sirens, and if it is operating in the United States, the CNN will attempt to identify only North American sirens. Once an emergency vehicle is detected, a controller can be used to execute an emergency vehicle safety routine, slow the vehicle, pull over to the side of the road, park the vehicle, and / or idle the vehicle, using ultrasonic sensors 962, until the one or more emergency vehicles pass by.
[0129] The vehicle may include one or more CPUs 918 (e.g., one or more discrete CPUs or one or more dCPUs) that may be coupled to the one or more SoCs 904 via a high-speed connection (e.g., PCIe). The CPUs 918 may include, for example, an x86 processor. The CPUs 918 may be used, for example, to perform a variety of functions, including reconciling potentially inconsistent results between ADAS sensors and the one or more SoCs 904 and / or monitoring the status and health of the one or more controllers 936 and / or the infotainment SoC 930.
[0130] The vehicle 900 may include one or more GPUs 920 (e.g., one or more discrete GPUs or one or more dGPUs) that may be coupled to the one or more SoCs 904 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The one or more graphics processing units 920 may provide additional artificial intelligence capabilities, e.g., by executing redundant and / or distinct neural networks, and may be used to train and / or update neural networks based on inputs (e.g., sensor data) from sensors of the vehicle 900.
[0131] The vehicle 900 may further include the network interface 924, which may include one or more wireless antennas 926 (e.g., one or more wireless antennas for various communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 924 may be used to enable a wireless connection over the internet to the cloud (e.g., to the one or more servers 978 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct connection between the two vehicles and / or an indirect connection may be established (e.g., via networks and the internet). Direct connections may be established via vehicle-to-vehicle communication.Vehicle-to-vehicle communication may provide vehicle 900 with information about vehicles in the vicinity of vehicle 900 (e.g., vehicles in front of, beside, and / or behind vehicle 900). This functionality may be part of a cooperative adaptive cruise control feature of vehicle 900.
[0132] The network interface 924 may include an SoC that provides modulation and demodulation functions and enables the one and more controllers 936 to communicate over wireless networks. The network interface 924 may include a radio frequency front-end for upconverting from baseband to radio frequency and downconverting from radio frequency to baseband. The frequency conversions may be performed using known methods and / or super-heterodyne methods. In some examples, the radio frequency front-end functionality may be provided by a separate chip. The network interface may include wireless functionality for communicating over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0133] The vehicle 900 may further include one or more data stores 928 that may be stored off-chip (e.g., outside the SoCs 904). The one or more data stores 928 may include one or more memory elements, including RAM, SRAM, DRAM, VRAM, flash, hard drives, and / or other components and / or devices capable of storing at least one bit of data.
[0134] The vehicle 900 may further include one or more GNSS sensors 958. The one or more GNSS sensors 958 (e.g., GPS, assisted GPS sensors, differential GPS (DGPS) sensors, etc.) assist in mapping, sensing, occupancy grid creation, and / or path planning. Any number of GNSS sensors 958 may be used, including, for example, and without limitation, a GPS using a USB port with an Ethernet-to-serial (RS-232) bridge.
[0135] The vehicle 900 may further include one or more RADAR sensors 960. The one or more RADAR sensors 960 may be used by the vehicle 900 for long-range vehicle detection, even in darkness and / or adverse weather conditions. The functional safety level of the RADAR may be ASIL B. The one or more RADAR sensors 960 may utilize the CAN and / or bus 902 (e.g., to transmit the data generated by the one or more RADAR sensors 960) for control and access to object tracking data, with access to the raw data occurring over Ethernet in some examples. A variety of RADAR sensor types may be used. The one or more RADAR sensors 960 may be suitable for, for example, front, rear, and side RADAR, without limitation. In some examples, one or more Pulse Doppler RADAR sensors are used.
[0136] The one or more RADAR sensors 960 may include various configurations, such as long range with a narrow field of view, short range with a wide field of view, short range side coverage, etc. In some examples, long range RADAR may be used for adaptive cruise control functionality. Long range RADAR systems may provide a wide field of view realized through two or more independent scans, e.g., within a range of 250 m. The one or more RADAR sensors 960 may assist in distinguishing between static and moving objects and may be used by ADAS systems for emergency braking and forward collision warning. Long range RADAR sensors may include a monostatic multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and a high-speed CAN and FlexRay interface.In a six-antenna example, the middle four antennas can create a focused beam pattern designed to detect the surroundings of vehicle 900 at higher speeds with minimal interference from traffic in adjacent lanes. The other two antennas can expand the field of view so that vehicles entering or exiting the lane of vehicle 900 can be quickly detected.
[0137] Medium-range radar systems, for example, can have a range of up to 960 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 950 degrees (rear). Short-range radar systems can include, among other features, radar sensors designed for installation at both ends of the rear bumper. When installed at both ends of the rear bumper, such a radar sensor system can generate two beams that continuously monitor the blind spot area to the rear and side of the vehicle.
[0138] Short-range radar systems can be used in an ADAS system to detect blind spots and / or as lane change assistants.
[0139] The vehicle 900 may also include one or more ultrasonic sensors 962. The one or more ultrasonic sensors 962, which may be mounted on the front, rear, and / or sides of the vehicle 900, may be used for parking assistance and / or for creating and updating an occupancy grid. A plurality of ultrasonic sensors 962 may be used, and different ultrasonic sensors 962 may be used for different detection ranges (e.g., 2.5 m, 4 m). The one or more ultrasonic sensors 962 may operate at functional safety levels of ASIL B.
[0140] The vehicle 900 may include one or more LIDAR sensors 964. The one or more LIDAR sensors 964 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The one or more LIDAR sensors 964 may be ASIL B functional safety rated. In some examples, the vehicle 900 may include multiple LIDAR sensors 964 (e.g., two, four, six, etc.) that may use Ethernet (e.g., to deliver data to a Gigabit Ethernet switch).
[0141] In some examples, the one or more LiDAR sensors 964 may be capable of providing a list of objects and their distances for a 360-degree field of view. For example, commercially available LiDAR sensors 964 may have an advertised range of approximately 900 m, with an accuracy of 2 cm to 3 cm, and with support for a 900 Mbps Ethernet connection. In some examples, one or more non-prominent LiDAR sensors 964 may be used. In such examples, the one or more LiDAR sensors 964 may be implemented as a small device that may be embedded in the front, rear, sides, and / or corners of the vehicle 900. In such examples, the one or more LiDAR sensors 964 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees, with a range of 200 m, even for objects with low reflectivity.The one or more front-mounted LiDAR sensors 964 may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0142] In some examples, LiDAR technologies such as 3D Flash LiDAR may also be used. 3D Flash LiDAR uses a laser flash as the transmission source to illuminate the vehicle's surroundings up to approximately 200 m. A Flash LiDAR unit contains a receptor that records the time of flight of the laser pulse and the reflected light at each pixel, which in turn corresponds to the distance between the vehicle and the objects. Flash LiDAR can enable highly accurate and distortion-free images of the surroundings to be created with each laser flash. In some examples, four Flash LiDAR sensors may be deployed, one on each side of the vehicle. Available 3D Flash LiDAR systems incorporate a solid-state 3D focal plane array LIDAR camera that contains no moving parts other than a fan (e.g., a non-scanning LIDAR device).The flash LiDAR device can use a 5-nanosecond Class I (eye-safe) laser pulse per image and capture the reflected laser light as 3D range point clouds and co-registered intensity data. By using flash LiDAR, and because flash LiDAR is a solid-state device with no moving parts, the one or more LiDAR sensors 964 can be less susceptible to motion blur, vibration, and / or shock.
[0143] The vehicle may also include one or more IMU sensors 966. The one or more IMU sensors 966 may, in some examples, be located at the center of the rear axle of the vehicle 900. The one or more IMU sensors 966 may, for example and without limitation, include one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other types of sensors. In some examples, such as in six-axis applications, the one or more IMU sensors 966 may include accelerometers and gyroscopes, while in nine-axis applications, the one or more IMU sensors 966 may include accelerometers, gyroscopes, and magnetometers.
[0144] In some embodiments, the one or more IMU sensors 966 may be implemented as a miniaturized, high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical system (MEMS) inertial sensors, a high-sensitivity GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and attitude. Thus, in some examples, the one or more IMU sensors 966 may enable the vehicle 900 to estimate heading without requiring input from a magnetic sensor by directly observing GPS velocity changes and correlating them with the one or more IMU sensors 966. In some examples, the one or more IMU sensors 966 and the one or more GNSS sensors 958 may be combined into a single integrated unit.
[0145] The vehicle may include one or more microphones 996 mounted in and / or around the vehicle 900. The one or more microphones 996 may be used, among other things, to detect and identify emergency vehicles.
[0146] The vehicle may further include any number of camera types, including one or more stereo cameras 968, one or more wide-angle cameras 970, one or more infrared cameras 972, one or more surround cameras 974, one or more long-range and / or medium-range cameras 998, and / or other camera types. The cameras may be used to capture image data around the entire perimeter of the vehicle 900. The types of cameras used depend on the embodiments and requirements for the vehicle 900, and any combination of camera types may be used to provide the necessary coverage around the vehicle 900. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or a different number of cameras.The cameras may support, by way of example and without limitation, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the one or more cameras is referred to herein with reference to . Fig. 9A and Fig. 9B is described in more detail.
[0147] The vehicle 900 may further include one or more vibration sensors 942. The one or more vibration sensors 942 may measure vibrations from components of the vehicle, such as the one or more axles. For example, changes in vibrations may indicate a change in the road surface. In another example, when two or more vibration sensors 942 are used, the differences between the vibrations may be used to determine the friction or slippage of the road surface (e.g., when the difference in vibration is between a driven axle and a free-spinning axle).
[0148] The vehicle 900 may include an ADAS system 938. The ADAS system 938 may include an SoC in some examples. The ADAS system 938 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning systems (CWS), lane centering (LC), and / or other features and functions.
[0149] The ACC systems may use one or more radar sensors 960, one or more lidar sensors 964, and / or one or more cameras. The ACC systems may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 900 and automatically adjusts the vehicle speed to maintain a safe distance from preceding vehicles. Lateral ACC performs follow-through and advises the vehicle 900 to change lanes if necessary. Lateral ACC is related to other ADAS applications, such as LCA and CWS.
[0150] CACC utilizes information from other vehicles, which may be received via the network interface 924 and / or the one or more radio antennas 926 from other vehicles via a wireless connection or indirectly via a network connection (e.g., via the Internet). Direct connections may be provided via a vehicle-to-vehicle (V2V) communication link, while indirect connections may be an infrastructure-to-vehicle (I2V) communication link. In general, the V2V communication concept provides information about the immediately preceding vehicles (e.g., vehicles immediately in front of and in the same lane as vehicle 900), while the I2V communication concept provides information about traffic further ahead. CACC systems may include both I2V and V2V information sources.Given the information about the vehicles ahead of vehicle 900, CACC can be more reliable and has the potential to improve traffic flow and reduce congestion on the road.
[0151] FCW systems are designed to warn the driver of a hazard so they can take corrective action. FCW systems utilize a forward-facing camera and / or one or more radar sensors 960 coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to feedback to the driver, such as a display, speaker, and / or vibrating component. FCW systems can provide a warning, such as a sound, a visual warning, a vibration, and / or a rapid braking pulse.
[0152] AEB systems detect an impending frontal collision with another vehicle or object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. AEB systems may use one or more forward-facing cameras and / or one or more radar sensors 960 coupled with a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision; if the driver does not take corrective action, the AEB system can automatically apply the brakes to prevent or at least mitigate the effects of the predicted collision. AEB systems may incorporate techniques such as dynamic brake support and / or impending crash braking.
[0153] LDW systems provide visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle 900 crosses lane markings. An LDW system will not activate if the driver indicates an intentional lane departure by activating a turn signal. LDW systems may utilize forward-facing cameras coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to feedback to the driver, such as a display, speaker, and / or vibrating component.
[0154] LKA systems are a variant of LDW systems. LKA systems provide steering inputs or braking to correct the vehicle 900 when the vehicle 900 begins to depart from its lane.
[0155] BSW systems detect and warn the driver of vehicles in the car's blind spot. BSW systems may provide a visual, audible, and / or tactile warning signal to indicate that merging into or changing lanes is unsafe. The system may provide an additional warning when the driver activates a turn signal. BSW systems may utilize one or more rear-facing cameras and / or radar sensors 960 coupled with a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to feedback to the driver, such as a display, speaker, and / or vibrating component.
[0156] RCTW systems can provide visual, audible, and / or tactile notification when an object outside the range of the rearview camera is detected when the vehicle 900 is reversing. Some RCTW systems include AEB to ensure the vehicle brakes are applied to avoid a crash. RCTW systems can utilize one or more rear-facing RADAR sensors 960 coupled with a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to feedback to the driver, such as a display, speaker, and / or vibrating component.
[0157] Conventional ADAS systems can produce false positives, which can be annoying and distracting for the driver, but are typically not catastrophic because ADAS systems warn the driver and give them the opportunity to decide whether a safety issue truly exists and act accordingly. However, in an autonomous vehicle 900, in the event of conflicting results, the vehicle 900 must decide for itself whether to consider the result of a primary computer or a secondary computer (e.g., a first controller 936 or a second controller 936). In some embodiments, the ADAS system 938 may, for example, be a backup and / or secondary computer that provides perception information to a rationality module of the backup computer.The backup computer rationality monitor can run redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. The outputs of the ADAS system 938 can be provided to a supervising MCU. If the outputs of the primary computer and the secondary computer conflict, the supervising MCU must determine how to resolve the conflict to ensure safe operation.
[0158] In some examples, the primary computer may be configured to provide the monitoring MCU with a confidence value indicating the primary computer's confidence in the chosen outcome. If the confidence value exceeds a threshold, the monitoring MCU may follow the primary computer's instruction regardless of whether the secondary computer provides a conflicting or inconsistent result. If the confidence value does not meet the threshold and the primary and secondary computers indicate different results (e.g., a conflict), the monitoring MCU may arbitrate between the computers to determine the appropriate outcome.
[0159] The monitoring MCU may be configured to run one or more neural networks trained and configured to determine the conditions under which the secondary computer triggers false alarms based on the output from the primary and secondary computers. This allows the one or more neural networks in the monitoring MCU to learn when the output of the secondary computer can and cannot be trusted. For example, if the secondary computer is a radar-based FCW system, a neural network in the monitoring MCU can learn when the FCW system identifies metallic objects that are not actually hazardous, such as a drain grate or manhole cover, which triggers an alarm.Similarly, if the secondary computer is a camera-based LDW system, a neural network in the monitoring MCU can learn to override the LDW system when cyclists or pedestrians are present and lane departure is indeed the safest maneuver. In embodiments including one or more neural networks running on the monitoring MCU, the monitoring MCU can include at least one DLA or GPU suitable for executing the one or more neural networks with associated memory. In preferred embodiments, the monitoring MCU can comprise and / or be included as a component of the one or more SoCs 904.
[0160] In other examples, the ADAS system 938 may include a secondary computer that executes the ADAS functionality according to classical computer vision rules. Thus, the secondary computer may use classical computer vision rules (if-then), and the presence of one or more neural networks in the supervising MCU may improve reliability, safety, and performance. For example, the diverse implementation and intentional non-identity make the overall system more fault-tolerant, especially against errors caused by software (or software-hardware interfaces).For example, if a software bug or error occurs in the software on the primary computer and the non-identical software code on the secondary computer produces the same overall result, the monitoring MCU can have greater confidence that the overall result is correct and the bug in the software or hardware on the primary computer does not cause a significant error.
[0161] In some examples, the output of the ADAS system 938 may be fed into the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 938 displays a forward collision warning due to an object immediately in front of the vehicle, the perception block may use this information in identifying objects. In other examples, the secondary computer may have its own neural network trained to reduce the risk of false positives, as described herein.
[0162] The vehicle 900 may further include the infotainment SoC 930 (e.g., an in-vehicle infotainment (IVI) system). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 930 may include a combination of hardware and software that can be used to provide the vehicle 900 with audio (e.g., music, a personal digital assistant, navigation directions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assist, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, oil level, door open / close, air filter information, etc.).The infotainment SoC 930 may include, for example, radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free calling, a heads-up display (HUD), an HMI display 934, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. The infotainment SoC 930 may further be used to provide information (e.g., visual and / or audible) to one or more users of the vehicle, such as information from the ADAS system 938, autonomous driving information such as planned vehicle maneuvers, road layouts, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0163] The infotainment SoC 930 may include GPU functionality. The infotainment SoC 930 may communicate with other devices, systems, and / or components of the vehicle 900 via the bus 902 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 930 may be coupled to a supervisory MCU so that the GPU of the infotainment system may perform some self-driving functions in the event that the one or more primary controllers 936 (e.g., the primary and / or backup computers of the vehicle 900) fail. In such an example, the infotainment SoC 930 may place the vehicle 900 into a chauffeur-to-safe-stop mode, as described herein.
[0164] The vehicle 900 may further include an instrument cluster 932 (e.g., a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 932 may include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 932 may include a number of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift position indicator, seat belt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag system (SRS) information, lighting controls, safety system controls, navigation information, etc. In some examples, information from the infotainment SoC 930 and the instrument cluster 932 may be displayed and / or shared. Thus, the instrument cluster 932 may be included as part of the infotainment SoC 930, or vice versa.
[0165] Fig. 9D is a system diagram for communication between the one or more cloud-based servers and the example autonomous vehicle 900 of Fig. 9A, according to some embodiments of the present disclosure. The system 976 may include the one or more servers 978, the one or more networks 990, and the vehicles, including the vehicle 900. The server(s) 978 may include a plurality of GPUs 984(A)-984(H) (referred to herein as GPUs 984), PCIe switches 982(A)-982(D) (referred to herein as PCIe switches 982), and / or CPUs 980(A)-980(B) (referred to herein as CPUs 980). The GPUs 984, the CPUs 980, and the PCIe switches may be interconnected using high-speed interconnects such as, without limitation, the NVIDIA-developed NVLink interfaces 988 and / or PCIe interconnects 986. In some examples, the GPUs 984 are connected via NVLink and / or NVSwitch SoC, and the GPUs 984 and the PCIe switches 982 are connected via PCIe connections. Although eight GPUs 984, two CPUs 980, and two PCIe switches are illustrated, this is not to be construed as a limitation.Depending on the embodiment, each of the servers 978 may include any number of GPUs 984, CPUs 980, and / or PCIe switches. For example, the one or more servers 978 may each include eight, sixteen, thirty-two, and / or more GPUs 984.
[0166] The one or more servers 978 may receive, via the one or more networks 990 and from the vehicles, image data representative of images depicting unexpected or changed road conditions, such as recently commenced roadwork. The one or more servers 978 may transmit, via the one or more networks 990 and to the vehicles, neural networks 992, updated neural networks 992, and / or map information 994 containing information about traffic and road conditions. The updates to the map information 994 may include updates to the HD map 922, such as information about construction, potholes, detours, flooding, and / or other obstacles.In some examples, the neural networks 992, the updated neural networks 992, and / or the map information 994 may result from new training and / or experience represented in the data received from any number of vehicles in the environment and / or may be based on training performed in a data center (e.g., using the one or more servers 978 and / or other servers).
[0167] The one or more servers 978 may be used to train machine learning models (e.g., neural networks) based on training data. The training data may be generated using the vehicles and / or in a simulation (e.g., with a game engine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training may be performed using one or more classes of machine learning techniques, including, without limitation, supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including replacement dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning models are trained, the machine learning models may be used by the vehicles (e.g., transmitted to the vehicles via the one or more networks 990) and / or the machine learning models may be used by the one or more servers 978 to remotely monitor the vehicles.
[0168] In some examples, the one or more servers 978 may receive data from the vehicles and apply the data to real-time, real-time neural networks for intelligent inference. The one or more servers 978 may include deep learning supercomputers and / or dedicated AI computers powered by GPUs 984, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, the one or more servers 978 may include a deep learning infrastructure using only CPU-powered data centers.
[0169] The deep learning infrastructure of the one or more servers 978 may be capable of performing rapid inference in real time and may utilize this capability to evaluate and verify the state of the processors, software, and / or associated hardware in the vehicle 900. For example, the deep learning infrastructure may receive periodic updates from the vehicle 900, such as a sequence of images and / or objects that the vehicle 900 has located in that sequence of images (e.g., via computer vision and / or other machine learning object classification techniques).The deep learning infrastructure may run its own neural network to identify the objects and compare them to the objects identified by the vehicle 900, and if the results do not match and the infrastructure concludes that the AI in the vehicle 900 is not functioning properly, the one or more servers 978 may send a signal to the vehicle 900 instructing a fail-safe computer of the vehicle 900 to take control, notify passengers, and perform a safe parking maneuver.
[0170] For inferencing, the one or more servers 978 may include GPUs 984 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-driven servers and inference accelerators may enable real-time responsiveness. In other examples, such as when performance is less critical, servers powered by CPUs, FPGAs, and other processors may be used for inferencing. EXAMPLE CALCULATION DEVICE
[0171] Fig. 10 is a block diagram of an example computing device 1000 suitable for use in implementing some embodiments of the present disclosure. The computing device 1000 may include an interconnect system 1002 that directly or indirectly couples the following devices: memory 1004, one or more central processing units (CPUs) 1006, one or more graphics processing units (GPUs) 1008, a communications interface 1010, input / output (I / O) ports 1012, input / output components 1014, a power supply 1016, one or more presentation components 1018 (e.g., display(s)), and one or more logic units 1020. In at least one embodiment, the one or more computing devices 1000 may include one or more virtual machines (VMs), and / or each of the components thereof may include virtual components (e.g., virtual hardware components).As non-limiting examples, one or more of the GPUs 1008 may include one or more vGPUs, one or more of the CPUs 1006 may include one or more vCPUs, and / or one or more of the logic units 1020 may include one or more virtual logic units. Thus, a computing device 1000 may include discrete components (e.g., a full GPU associated with the computing device 1000), virtual components (e.g., a portion of a GPU associated with the computing device 1000), or a combination thereof.
[0172] Although the different blocks of Fig. 10 as being connected to wires via the interconnect system 1002, this is not intended as a limitation and is for clarity only. For example, in some embodiments, a presentation component 1018, such as a display device, may be considered an I / O component 1014 (e.g., if the display is a touchscreen). As another example, the CPUs 1006 and / or GPUs 1008 may include memory (e.g., memory 1004 may represent a storage device in addition to the memory of the GPUs 1008, the CPUs 1006, and / or other components). Thus, the computing device of Fig. 10 is merely illustrative. No distinction is made between categories such as “workstation”, “server”, “laptop”, “desktop”, “tablet”, “client device”, “mobile device”, “handheld device”, “game console”, “electronic control unit (ECU)”, “virtual reality system” and / or other device or system types, since all within the scope of the computing device of Fig. 10 come into consideration.
[0173] The interconnect system 1002 may represent one or more interconnects or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 1002 may include one or more bus or interconnect types, such as an Industry Standard Architecture (ISA) bus, an Extended ISA bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, and / or another type of bus or interconnect. In some embodiments, there are direct connections between components. For example, the CPU 1006 may be directly connected to the memory 1004. Further, the CPU 1006 may be directly connected to the GPU 1008.For a direct or point-to-point connection between components, interconnect system 1002 may include a PCIe link to establish the connection. In these examples, a PCI bus need not be included in computing device 1000.
[0174] Memory 1004 may include a variety of computer-readable media. The computer-readable media may be any available media accessible by the computing device 1000. The computer-readable media may include both volatile and non-volatile media, as well as removable and non-removable media. By way of example and without limitation, the computer-readable media may include computer storage media and communication media.
[0175] The computer storage media may include both volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1004 may store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system).Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, Digital Versatile Disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 1000. As used herein, computer storage media does not per se include signals.
[0176] The computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and may include any media for conveying information. The term "modulated data signal" may refer to a signal having one or more of its properties adjusted or altered to encode information in the signal. The computer storage media may include, for example, and is not limited to, wired media, such as a wired network or a direct-wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the above should also be included within the scope of computer-readable media.
[0177] The one or more CPUs 1006 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the methods and / or processes described herein. The CPUs 1006 may each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of concurrently executing a plurality of software threads. The CPUs 1006 may include any type of processor and may include different types of processors depending on the type of computing device 1000 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 1000, the processor may be, for example, an Advanced RISC Machines (ARM) processor implemented with Reduced Instruction Set Computing (RISC) or an x86 processor implemented with Complex Instruction Set Computing (CISC). Computing device 1000 may include one or more CPUs 1006, in addition to one or more microprocessors or additional coprocessors, such as math coprocessors.
[0178] In addition to or alternatively to the CPU(s) 1006, the GPU(s) 1008 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the methods and / or processes described herein. One or more of the GPUs 1008 may be an integrated GPU (e.g., with one or more of the CPUs 1006) and / or one or more of the GPUs 1008 may be a discrete GPU. In embodiments, one or more of the GPUs 1008 may be a co-processor of one or more of the CPUs 1006. The one or more GPUs 1008 may be used by the computing device 1000 to render graphics (e.g., 3D graphics) or to perform general-purpose computations. The GPUs 1008 may be used, for example, for general-purpose computing on GPUs (GPGPU).The one or more GPUs 1008 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The one or more GPUs 1008 may generate pixel data for output images in response to rendering commands (e.g., rendering commands from the one or more CPUs 1006 received via a host interface). The GPUs 1008 may include graphics memory, such as random access memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory may be included as part of the random access memory 1004. The GPUs 1008 may include two or more GPUs operating in parallel (e.g., via a link). The link may connect the GPUs directly (e.g., using NVLINK) or connect the GPUs via a switch (e.g., using NVSwitch).When combined, each GPU can generate 1008 pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can contain its own memory or share memory with other GPUs.
[0179] In addition to or alternatively to the one or more CPUs 1006 and / or the one or more GPUs 1008, the one or more logic units 1020 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the methods and / or processes described herein. In embodiments, the CPUs 1006, the GPUs 1008, and / or the one or more logic units 1020 may discretely or jointly execute any combination of the methods, processes, and / or portions thereof. One or more of the logic units 1020 may be part of one or more of the CPUs 1006 and / or one or more of the GPUs 1008, and / or one or more of the logic units 1020 may be discrete components or otherwise external to the CPUs 1006 and / or the GPUs 1008.In embodiments, one or more of the logic units 1020 may be a co-processor of one or more of the CPUs 1006 and / or one or more of the GPUs 1008.
[0180] Examples of the one or more logic units 1020 include one or more processing cores and / or components thereof, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel visual cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application-specific integrated circuits (Application-Specific Integrated Circuits, ASICs), floating point units (FPUs),Input / Output (I / O) elements, Peripheral Component Interconnect (PCI) or PCI Express (PCIe) elements, and / or similar.
[0181] The communication interface 1010 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1000 to communicate with other computing devices over an electronic network, including wired and / or wireless communication. The communication interface 1010 may include components and functions that enable communication over a variety of networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the one or more logic units 1020 and / or the communication interface 1010 may include one or more data processing units (DPUs) to transfer data received over a network and / or via the interconnect system 1002 directly to one or more GPUs 1008 (e.g., a memory thereof).
[0182] The I / O ports 1012 may enable the computing device 1000 to be logically coupled to other devices, including the I / O components 1014, the one or more presentation components 1018, and / or other components, some of which may be built into (e.g., integrated) the computing device 1000. Illustrative I / O components 1014 include a microphone, a mouse, a keyboard, a joystick, a gamepad, a game controller, a satellite dish, a scanner, a printer, a wireless device, etc. The I / O components 1014 may provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, the inputs may be communicated to a suitable network element for further processing.An NUI may implement any combination of speech capture, stylus capture, facial capture, biometric capture, both on-screen and off-screen gesture capture, air gestures, head and eye tracking, and touch capture (as further described below) associated with a display of the computing device 1000. The computing device 1000 may include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the computing device 1000 may include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion capture. In some examples, the output of the accelerometers or gyroscopes from the computing device 1000 may be used to present immersive augmented reality or virtual reality.
[0183] Power supply 1016 may include a hardwired power supply, a battery power supply, or a combination thereof. Power supply 1016 may supply power to computing device 1000 to enable operation of the components of computing device 1000.
[0184] The one or more presentation components 1018 may include a display (e.g., a monitor, a touchscreen, a television monitor, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The one or more presentation components 1018 may receive data from other components (e.g., the one or more GPUs 1008, the one or more CPUs 1006, DPUs, etc.) and output the data (e.g., as an image, video, audio, etc.). EXEMPLARY DATA CENTER
[0185] Fig. Figure 11 illustrates an example data center 1100 that may be used in at least one embodiment of the present disclosure. Data center 1100 may include a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and / or an application layer 1140.
[0186] As in Fig. 11, the data center infrastructure layer 1110 may include a resource orchestrator 1112, clustered computing resources 1114, and node computing resources (“node CRs”) 1116(1)-1116(N), where “N” represents any positive integer. In at least one embodiment, the node CRs 1116(1)-1116(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules and / or cooling modules, etc.In some embodiments, one or more Node CRs among Node CRs 1116(1)-1116(N) may correspond to a server having one or more of the computing resources mentioned above. Furthermore, in some embodiments, Node CRs 1116(1)-1116(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of Node CRs 1116(1)-1116(N) may correspond to a virtual machine (VM).
[0187] In at least one embodiment, the grouped computing resources 1114 may include separate groupings of node CRs 1116 housed in one or more racks (not shown) or in many racks in data centers in different geographic locations (also not shown). Separate groupings of node CRs 1116 within grouped computing resources 1114 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 1116, including CPUs, GPUs, DPUs, and / or other processors, may be grouped in one or more racks to provide computing resources to support one or more workloads.The one or more racks may also contain any number of power modules, cooling modules, and / or network switches in any combination.
[0188] Resource orchestrator 1112 may configure or otherwise control one or more node CRs 1116(1)-1116(N) and / or clustered computing resources 1114. In at least one embodiment, resource orchestrator 1112 may include an entity for managing the software design infrastructure (SDI) for data center 1100. Resource orchestrator 1112 may include hardware, software, or a combination thereof.
[0189] In at least one embodiment, as in Fig. 11, the framework layer 1120 may include a job scheduler 1133, a configuration manager 1134, a resource manager 1136, and / or a distributed file system 1138. The framework layer 1120 may include a framework that supports the software 1132 of the software layer 1130 and / or one or more applications 1142 of the application layer 1140. The software 1132 or the one or more applications 1142 may each include web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1120 may be some type of free and open source software web application framework, such as, but not limited to, Apache Spark™ (hereinafter "Spark"), which may utilize a distributed file system 1138 for processing large amounts of data (e.g., "Big Data").In at least one embodiment, job scheduler 1133 may include a Spark driver to facilitate scheduling workloads supported by different layers of data center 1100. Configuration manager 1134 may be capable of configuring different layers, such as software layer 1130 and framework layer 1120, which includes Spark and distributed file system 1138, to support processing large amounts of data. Resource manager 1136 may be capable of managing clustered or grouped computing resources allocated or assigned to support distributed file system 1138 and job scheduler 1133. In at least one embodiment, the clustered or grouped computing resources may include clustered computing resource 1114 at infrastructure layer 1110 of the data center.The resource manager 1136 may coordinate with the resource orchestrator 1112 to manage these allocated or assigned computing resources.
[0190] In at least one embodiment, the software 1132 included in software layer 1130 may include software used by at least portions of node CRs 1116(1)-1116(N), clustered computing resources 1114, and / or distributed file system 1138 of framework layer 1120. One or more types of software may include, but are not limited to, web page searching software, email virus scanning software, database software, and streaming video content software.
[0191] In at least one embodiment, the applications 1142 included in the application layer 1140 may include one or more types of applications used by at least portions of the node CRs 1116(1)-1116(N), the clustered computing resources 1114, and / or the distributed file system 1138 of the framework layer 1120. One or more types of applications may include, but are not limited to, any number of genomic applications, cognitive computation, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in connection with one or more embodiments.
[0192] In at least one embodiment, one of configuration manager 1134, resource manager 1136, and resource orchestrator 1112 may implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible manner. Self-modifying actions may relieve a data center operator of data center 1100 from making potentially poor configuration decisions and potentially avoid underutilized and / or poorly performing sections of a data center.
[0193] Data center 1100 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weighting parameters according to a neural network architecture using software and / or computing resources described above with respect to data center 1100.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with reference to data center 1100 using weighting parameters calculated by one or more training techniques, such as, but not limited to, those described herein.
[0194] In at least one embodiment, data center 1100 may utilize CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inferencing using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to enable users to train or infer information, such as image capture, speech capture, or other artificial intelligence services. EXAMPLE NETWORK ENVIRONMENTS
[0195] Network environments suitable for implementing embodiments of the disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented on one or more instances of the one or more computing devices 1000 of Fig. 10 - e.g., each device may include similar components, features, and / or functionality of the one or more computing devices 1000. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may also be included as part of a data center 1100, an example of which is described herein with reference to Fig. 11 is described in more detail.
[0196] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can contain multiple networks or a network of networks. For example, the network can contain one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network contains a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0197] Compatible network environments may include one or more peer-to-peer network environments—in which case, a server cannot be included in a network environment—and one or more client-server network environments—in which case, one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described herein with respect to one or more servers may be implemented on any number of client devices.
[0198] In at least one embodiment, a network environment may include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer may include a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or the one or more applications may each include web-based service software or applications. In embodiments, one or more of the client devices may utilize the web-based service software or applications (e.g.,by accessing the service software and / or applications through one or more application programming interfaces (APIs). The framework layer may be some type of free and open-source software web application framework, e.g., using, but not limited to, a distributed file system for processing large amounts of data (e.g., "Big Data").
[0199] A cloud-based network environment may provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Each of these various functions may be distributed from central or core servers (e.g., from one or more data centers that may be located across a state, region, country, globe, etc.) across multiple locations. When a connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers may offload at least a portion of the functionality to the one or more edge servers. A cloud-based network environment may be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0200] The one or more client devices may include at least some of the components, features, and functions of the one or more described herein with respect to Fig.10. By way of example, and not limitation, a client device may be embodied as a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smart watch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or global positioning device, a video player, a video camera, a surveillance device or surveillance system, a vehicle, a boat, a hydrofoil, a virtual machine, a drone, a robot, a portable communication device, a hospital device, a gaming device or gaming system, an entertainment system, a vehicle computing system, an embedded system controller, a remote control, an appliance, a consumer electronics device, a workstation, an edge device,any combination of these described devices or any other suitable device.
[0201] The disclosure of this application also includes the following numbered clauses: Clause 1 One or more processors comprising processing circuitry for: Detecting, based at least on applying a representation of an intersection-centered tile of a LiDAR map of a two-dimensional (2D) surface to a neural network, one or more Navigation control lines depicted in the intersection-centered tile; and updating the LiDAR map based on at least the one or more navigation control lines. Clause 2 The one or more processors of Clause 1, wherein the 2D surface represents a ground surface, and wherein the processing circuitry is further to generate the LiDAR map based at least on projecting LiDAR intensity data collected using one or more ego machines onto the 2D surface. Clause 3 The one or more processors of any preceding clause, wherein the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based on at least an initial set of navigation control lines detected from one or more tiles of the LiDAR map. Clause 4 The one or more processors of any preceding clause, wherein the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of navigation control lines detected from the LiDAR map for detected crosswalk lines forming a detected crosswalk. Clause 5 The one or more processors of any preceding clause, wherein the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of navigation control lines detected from the LiDAR map for detected lines forming different demarcated regions in a common intersection. Clause 6 The one or more processors of any preceding clause, wherein the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based on at least clustering one or more detected lines into the inferred intersection. Clause 7 The one or more processors according to any one of the preceding clauses, wherein the processing circuitry is further arranged to: Detecting an initial set of navigation control lines from one or more tiles of the LiDAR map; Generating a representation of one or more detected intersections based at least on clustering the initial set of navigation control lines; and Detecting a refined set of navigation control lines from one or more intersection-centered tiles associated with one or more detected intersections. Clause 8 The one or more processors according to any one of the preceding clauses, wherein the processing circuitry is further arranged to: Detecting an initial set of navigation control lines from one or more tiles of the LiDAR map; and Detecting one or more intersections based at least on the geometry and proximity of the initial set of navigation control lines. Clause 9 The one or more processors according to any one of the preceding clauses, wherein the processing circuitry comprises at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing operations using conversational AI; a system for implementing one or more language models; a system for implementing one or more large language models (LLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources. Clause 10 A system comprising one or more processors for detecting one or more lines represented in the intersection-centered tile based at least on processing a representation of an intersection-centered tile of a map of a two-dimensional (2D) surface using a neural network. Clause 11 The system of Clause 10, wherein the 2D surface represents a ground surface, and wherein the one or more processors are further operable to generate the map based at least on projecting intensity data collected using one or more ego machines onto the 2D surface. Clause 12 The system of any of clauses 10-11, wherein the one or more processors are further to generate the intersection-centered tile around an inferred intersection detected based on at least an initial set of lines detected from one or more tiles of the map. Clause 13 The system of any of clauses 10-12, wherein the one or more processors are further to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of lines detected from the map for detected crosswalk lines forming a detected crosswalk. Clause 14 The system of any of clauses 10-13, wherein the one or more processors are further to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of lines detected from the map for detected lines that form different demarcated regions in a common intersection. Clause 15 The system of any of clauses 10-14, wherein the one or more processors are further to generate the intersection-centered tile around an inferred intersection detected based on at least clustering one or more detected lines into the inferred intersection. Clause 16 A system according to any one of Clauses 10-15, wherein the one or more processors are further configured to: Detecting an initial set of lines from one or more tiles of the map; Generating a representation of one or more detected intersections based at least on clustering the initial set of lines; and detecting a refined set of lines from one or more intersection-centered tiles associated with one or more detected intersections. Clause 17 A system according to any one of Clauses 10-16, wherein the one or more processors are further configured to: Detecting an initial set of lines from one or more tiles of the map; and Detecting one or more intersections based at least on the geometry and proximity of the initial set of lines. Clause 18 A system as defined in any of Clauses 10-17, which system includes at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing operations using conversational AI; a system for implementing one or more language models; a system for implementing one or more large language models (LLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources. Clause 19 Procedure, comprising: Detecting, based at least on processing a representation of a tile of a map centered on a detected intersection using a neural network, one or more lines represented in the tile; and updating the map based at least on the one or more lines. Clause 20 Procedures under Clause 19, where the procedure is carried out by at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing operations using conversational AI; a system for implementing one or more language models; a system for implementing one or more large language models (LLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources.
[0202] The disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules include routines, programs, objects, components, data structures, etc., and refer to code that performs specific tasks or implements specific abstract data types. The disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc.The disclosure may also be practiced in distributed computing environments in which tasks are performed by remote processing devices that are interconnected via a network for communication.
[0203] As used herein, any reference to "and / or" in reference to two or more elements should be interpreted to mean only one element or a combination of elements. For example, "Element A, Element B, and / or Element C" may include only Element A, only Element B, only Element C, Element A and Element B, Element A and Element C, Element B and Element C, or Elements A, B, and C. Furthermore, "at least one of Element A or Element B" may include at least one of Element A, at least one of Element B, or at least one of Element A and at least one of Element B. Further, "at least one of Element A and Element B" may include at least one of Element A, at least one of Element B, or at least one of Element A and at least one of Element B.
[0204] The subject matter of the present disclosure is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of the present disclosure. Rather, the inventors contemplated that the claimed subject matter may be embodied in other ways to include various steps or combinations of steps similar to those described herein in conjunction with other present or future technologies. Moreover, while the terms "step" and / or "block" may be used herein to refer to various elements of the methods employed, the terms should not be interpreted to imply any particular ordering among or between the various steps disclosed herein, unless the order of each step is expressly described. EXEMPLARY VERBAL SUPPORT
[0205] In an example implementation, one or more processors include processing circuitry to: detect, based at least on applying a representation of an intersection-centered tile of a LiDAR map of a two-dimensional (2D) surface to a neural network, one or more navigation control lines represented in the intersection-centered tile; and update the LiDAR map based at least on the one or more navigation control lines.
[0206] In any combination of any of the elements of any of the above implementations of the one or more processors, the 2D surface represents a ground surface, and the processing circuitry is further operable to generate the LiDAR map based at least on projecting LiDAR intensity data collected using one or more ego machines onto the 2D surface.
[0207] In any combination of any of the elements of any of the above implementations of the one or more processors, the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based on at least an initial set of navigation control lines detected from one or more tiles of the LiDAR map. In some embodiments, the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based on at least searching an initial set of navigation control lines detected from the LiDAR map for detected crosswalk lines that form a detected crosswalk.
[0208] In any combination of any of the elements of any of the above implementations of the one or more processors, the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of navigation control lines detected from the LiDAR map for detected lines that form different demarcated regions in a common intersection.
[0209] In any combination of any of the elements of any of the above implementations of the one or more processors, the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based on at least clustering one or more detected lines into the inferred intersection.
[0210] In any combination of any of the elements of any of the above implementations of the one or more processors, the processing circuitry is further to: detect an initial set of navigation control lines from one or more tiles of the LiDAR map; generate a representation of one or more detected intersections based at least on clustering the initial set of navigation control lines; and detect a refined set of navigation control lines from one or more intersection-centered tiles associated with one or more detected intersections.
[0211] In any combination of any of the elements of any of the above implementations of the one and or more processors, the processing circuitry is further to: detect an initial set of navigation control lines from one or more tiles of the LiDAR map; and detect one or more intersections based at least on the geometry and proximity of the initial set of navigation control lines.
[0212] In any combination of any of the elements of any of the above implementations of the one or more processors, the processing circuitry comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device;a system implemented using a robot; a system for performing operations using conversational AI; a system implementing one or more language models; a system implementing one or more large-scale language models (LLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0213] In an example implementation, a system includes one or more processors to detect one or more lines represented in the intersection-centered tile based at least on processing a representation of an intersection-centered tile of a map of a two-dimensional (2D) surface using a neural network.
[0214] In any combination of any of the elements of any of the above implementations of the system, the 2D surface represents a ground surface, and the one or more processors are further operable to generate the map based at least on projecting intensity data collected using one or more ego machines onto the 2D surface.
[0215] In any combination of any of the elements of any of the above implementations of the system, the one or more processors are further to generate the intersection-centered tile around an inferred intersection detected based on at least an initial set of lines detected from one or more tiles of the map.
[0216] In any combination of any of the elements of any of the above implementations of the system, the one or more processors are further operable to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of lines detected from the map for detected crosswalk lines that form a detected crosswalk.
[0217] In any combination of any of the elements of any of the above implementations of the system, the one or more processors are further to generate the intersection-centered tile around an inferred intersection that is detected based at least on searching an initial set of lines detected from the map for detected lines that form different bounded regions in a common intersection.
[0218] In any combination of any of the elements of any of the above implementations of the system, the one or more processors are further to generate the intersection-centered tile around an inferred intersection detected based on at least clustering one or more detected lines into the inferred intersection.
[0219] In any combination of any of the elements of any of the above implementations of the system, the one or more processors are further operable to: detect an initial set of lines from one or more tiles of the map; generate a representation of one or more detected intersections based at least on clustering the initial set of lines; and detect a refined set of lines from one or more intersection-centered tiles associated with one or more detected intersections.
[0220] In any combination of any of the elements of any of the above implementations of the system, the one or more processors are further to: detect an initial set of lines from one or more tiles of the map; and detect one or more intersections based at least on the geometry and proximity of the initial set of lines.
[0221] In any combination of any of the elements of any of the above implementations of the system, the system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot;a system for performing operations with conversational AI; a system that implements one or more language models; a system that implements one or more large language models (LLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that includes one or more virtual machines (VMs); a system that is implemented at least in part in a data center; or a system that is implemented at least in part using cloud computing resources.
[0222] In an example implementation, a method comprises: detecting, based at least on processing a representation of a tile of a map centered around a detected intersection using a neural network, one or more lines represented in the tile; and updating the map based at least on the one or more lines.
[0223] In any combination of each of the elements of each of the above implementations of the method, the method is performed by at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device;a system implemented using a robot; a system for performing operations using conversational AI; a system implementing one or more language models; a system implementing one or more large-scale language models (LLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0224] It is to be understood that aspects and embodiments described above are purely exemplary and that modifications of details may be made within the scope of the claims.
[0225] Each device, method, and feature disclosed in the description, and (where appropriate) the claims and drawings may be provided independently or in any suitable combination.
[0226] Reference signs appearing in the claims are for illustrative purposes only and do not limit the scope of the claims. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] US 16 / 101,232
[0106] Cited non-patent literature
[0000] Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles) (Standard No. J3016-201806, published on 15 June 2018, Standard No. J3016-201609, published on 30 September 2016
[0063]
Claims
[1] One or more processors comprising processing circuitry for: Detecting, based at least on applying a representation of an intersection-centered tile of a LiDAR map of a two-dimensional (2D) surface to a neural network, one or more navigation control lines represented in the intersection-centered tile; and Updating the LiDAR map based on at least one or more navigation control lines. [2] The one or more processors of claim 1, wherein the 2D surface represents a ground surface, and wherein the processing circuitry is further to generate the LiDAR map based at least on projecting LiDAR intensity data collected using one or more ego machines onto the 2D surface. [3] The one or more processors of claim 1 or claim 2, wherein the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based on at least an initial set of navigation control lines detected from one or more tiles of the LiDAR map. [4] The one or more processors of any preceding claim, wherein the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of navigation control lines detected from the LiDAR map for detected crosswalk lines forming a detected crosswalk. [5] The one or more processors of any preceding claim, wherein the processing circuitry is further operable to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of navigation control lines detected from the LiDAR map for detected lines forming different demarcated regions in a common intersection. [6] The one or more processors of any preceding claim, wherein the processing circuitry is further to generate the intersection-centered tile around an inferred intersection detected based at least on clustering one or more detected lines into the inferred intersection. [7] The one or more processors according to any one of the preceding claims, wherein the processing circuitry is further arranged to: Detecting an initial set of navigation control lines from one or more tiles of the LiDAR map; Generating a representation of one or more detected intersections based at least on clustering the initial set of navigation control lines; and Detecting a refined set of navigation control lines from one or more intersection-centered tiles associated with one or more detected intersections. [8] The one or more processors according to any one of the preceding claims, wherein the processing circuitry is further arranged to: Detecting an initial set of navigation control lines from one or more tiles of the LiDAR map; and Detecting one or more intersections based at least on the geometry and proximity of the initial set of navigation control lines. [9] The one or more processors according to any one of the preceding claims, wherein the processing circuitry comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing operations using conversational AI; a system for implementing one or more language models; a system for implementing one or more large language models (LLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources. [10] A system comprising one or more processors for detecting one or more lines represented in the intersection-centered tile based at least on processing a representation of an intersection-centered tile of a map of a two-dimensional (2D) surface using a neural network. [11] The system of claim 10, wherein the 2D surface represents a ground surface, and wherein the one or more processors are further operable to generate the map based at least on projecting intensity data collected using one or more ego machines onto the 2D surface. [12] The system of claim 10 or claim 11, wherein the one or more processors are further to generate the intersection-centered tile around an inferred intersection detected based on at least an initial set of lines detected from one or more tiles of the map. [13] The system of any of claims 10-12, wherein the one or more processors are further operable to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of lines detected from the map for detected crosswalk lines forming a detected crosswalk. [14] The system of any of claims 10-13, wherein the one or more processors are further operable to generate the intersection-centered tile around an inferred intersection detected based at least on searching an initial set of lines detected from the map for detected lines forming different bounded regions in a common intersection. [15] The system of any of the preceding claims 10-14, wherein the one or more processors are further to generate the intersection-centered tile around an inferred intersection detected based at least on clustering one or more detected lines into the inferred intersection. [16] The system of any of the preceding claims 10-15, wherein the one or more processors further serve to: Detecting an initial set of lines from one or more tiles of the map; Generating a representation of one or more detected intersections based at least on clustering the initial set of lines; and Detecting a refined set of lines from one or more intersection-centered tiles associated with one or more detected intersections. [17] The system of any of the preceding claims 10-16, wherein the one or more processors further serve to: Detecting an initial set of lines from one or more tiles of the map; and Detecting one or more intersections based at least on the geometry and proximity of the initial set of lines. [18] A system according to any one of claims 10-17, wherein the system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing operations using conversational AI; a system for implementing one or more language models; a system for implementing one or more large language models (LLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources. [19] Method comprising: Detecting, based at least on processing a representation of a tile of a map centered on a detected intersection using a neural network, one or more lines represented in the tile; and Update the map based on at least one or more lines. [20] The method of claim 19, wherein the method is performed by at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing operations using conversational AI; a system for implementing one or more language models; a system for implementing one or more large language models (LLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources.
Citation Information
Patent Citations
US-PATENTANMELDUNGNR.16/101,232