DETECTION OF LINE SEGMENTS OF TRANSPORTATION MONUMENTS FOR AUTONOMOUS AND SEMI-AUTOMATIC SYSTEMS AND APPLICATIONS
Machine learning models for detecting road marking line segments enhance navigation precision and safety in autonomous vehicles by accurately determining vehicle locations and navigation rules based on road markings, addressing the limitations of conventional systems in environments with sparse traffic features.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-03-26
AI Technical Summary
Conventional systems fail to accurately determine the locations of road markings for precise longitudinal localization of vehicles, especially in environments with few traffic signs or signals, limiting the navigation capabilities of autonomous and semi-autonomous systems.
The use of machine learning models to process sensor data for detecting and determining the locations of line segments associated with road markings, enabling accurate localization and navigation by aligning these segments with map data, thereby enhancing the safety and precision of vehicle maneuvers.
Enables precise determination of vehicle locations along roads without relying on traffic signs, improving navigation accuracy and safety by accurately identifying permitted and restricted navigation areas based on road markings.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] For vehicles or machines (e.g., autonomous vehicles, semi-autonomous vehicles, robots, etc.) to operate safely in various environments, they must be capable of performing vehicle maneuvers such as lane keeping, lane changes, lane splitting, turning, stopping and starting at intersections, pedestrian crossings, and the like, and / or other vehicle or machine maneuvers. For a vehicle to navigate surface roads (e.g., side streets, neighborhood streets, etc.) and highways (e.g., multi-lane roads), it must be able to navigate within one or more subdivisions or boundaries (e.g., lanes, intersections, pedestrian crossings, barriers, etc.) of a roadway, which are often marked using road markings such as solid lines, dashed lines, and / or the like.In many situations, mapping and localization are therefore important processes for carrying out these autonomous and / or semi-autonomous functions.
[0002] As such, maps—such as navigation maps, standard definition maps (SD maps), and / or high definition maps (HD maps)—can be used to locate vehicles within environments. For example, vehicles can generate sensor data using various sensors and then align at least some of that sensor data with features on the map to perform localization. For instance, if a vehicle is navigating along a road (such as a highway), the vehicle can align some of the sensor data with specific features on the road represented by a map—such as lane markings, curbs, and / or road edges—to locate the vehicle laterally within a specific lane.The vehicle can also align another part of the sensor data with reference to additional features - such as traffic signs, traffic poles, trees, static structures and / or traffic signals - to align the vehicle longitudinally along the road (e.g. in the direction of travel of the vehicle).
[0003] However, in some situations, these additional features may not be located within the area of the environment in which the vehicle is navigating. For example, certain highways and / or other types of roads may have very few traffic signs, traffic poles, and / or traffic signals located along the road, making it difficult for the vehicle to perform longitudinal localization. Additionally, while maps may provide some information associated with road markings, such as the types of road markings, they may not include sufficient information to assist in the longitudinal localization of vehicles using such road markings.This is because conventional systems that generate maps may not include the functionality to determine specific details regarding road markings, which can then be used for accurate or precise localization. OVERVIEW
[0004] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not be within the scope of the claims are described herein.
[0005] Embodiments of the present disclosure relate to the detection of line segments of traffic features for autonomous and / or semi-autonomous systems and applications. The systems and methods described herein can determine the locations of line segments (e.g., line markings) associated with road markings within environments. For example, one or more machine learning models can process sensor data (e.g., image data, etc.) to determine points associated with the line segments as represented by the sensor data and / or directional information (e.g., direction vectors) associated with the points. As described herein, the points can be associated with the edges of the line segments, the midpoints of the line segments, and / or other locations of the line segments.Additionally, the points and / or direction information can be represented using various techniques, such as Cartesian coordinates and / or polar coordinates associated with the sensor data. The systems and procedures described herein may further include: performing operations based on the locations of these line segments, such as updating a map used for localization, performing localization, route planning, controlling, and / or determining trajectories to be navigated.
[0006] In contrast to conventional systems, the systems of the present disclosure are, in some embodiments, capable of automatically determining the locations of line segments within environments and / or updating a map to include information associated with the line segments. As such, and as described in more detail herein, machines using the map to perform localization can accurately determine and / or adjust locations along roads, such as longitudinal locations along roads that do not include other traffic features (e.g., traffic poles, traffic signs, etc.) over long distances. For example, the machines can align identified line segments represented by the sensor data with the line segments represented by the map to determine the locations of the machines longitudinally along the roads.
[0007] In addition, unlike conventional systems, the systems of this disclosure can increase the safety of machines navigating within the environment by precisely determining the locations of line segments for different types of road markings, as the machines are able to determine rules for navigating within the environment more accurately. For example, the machines can better determine locations where they are permitted to navigate over road markings (e.g., where line markings are present) and / or locations where they are not permitted to navigate over line markings (e.g., where double solid lines are present).
[0008] The revelation extends to any novel aspects or features described and / or illustrated herein.
[0009] Further features of the disclosure are characterized by the independent and dependent claims.
[0010] Any feature in one aspect of the disclosure can be applied to other aspects of the disclosure in any suitable combination. In particular, procedural aspects can be applied to device or system aspects and vice versa.
[0011] Furthermore, features implemented in hardware can also be implemented in software, and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.
[0012] Each system or device feature as described herein can also be provided as a process feature, and vice versa. Functionally described system and / or device aspects (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory.
[0013] It should also be apparent that specific combinations of the various features described and defined in any aspect of the disclosure can be implemented and / or provided and / or used independently of one another.
[0014] This disclosure also provides computer programs and computer program products comprising software code adapted, when executed on a data processing device, to perform any of the procedures described herein and / or to embody any of the device and system features described herein, including any or all of the partial steps of a procedure.
[0015] The disclosure also provides a computer or computing system (including networked or distributed systems) that includes an operating system supporting a computer program for performing any of the procedures described herein and / or embodying any of the device or system features described herein.
[0016] The revelation also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0017] The revelation also provides a signal that carries one or more of the aforementioned computer programs.
[0018] The disclosure extends to processes and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0019] Aspects and embodiments of the disclosure will now be described exclusively by way of example with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present systems and methods for detecting line segments of traffic features for autonomous and semi-autonomous systems and applications are described in detail below with reference to the attached drawings, wherein: Fig. 1A illustrates an exemplary data flow representation for a process of determining information associated with line segments of traffic features in accordance with some embodiments of the present disclosure; Fig. 1B An example of one or more machine learning models trained to determine information associated with line segments of traffic features is illustrated in accordance with some embodiments of the present disclosure; Fig. 2 an example of a sensor representation representing road markings with different line segments, in accordance with some embodiments of the present disclosure; Fig. 3A-3C Illustrate examples of determining information associated with line segments corresponding to road markings in accordance with some embodiments of the present disclosure; Fig. 4A-4B illustrate examples of different types of outputs in accordance with some embodiments of the present disclosure, which can generate one or more machine learning models, wherein the outputs include information associated with line segments; Fig. 5 a data flow diagram illustrating a process for training one or more machine learning models to determine information associated with line segments, in accordance with some embodiments of the present disclosure; Fig. 6 An example of information that may be included in ground-truth data for training one or more machine learning models is illustrated in accordance with some embodiments of the present disclosure; Fig. 7-8 Data flow diagrams illustrating methods for using one or more machine learning models to determine information associated with line segments, in accordance with some embodiments of the present disclosure; Fig. 9A is an illustration of an exemplary autonomous vehicle in accordance with some embodiments of the present disclosure; Fig. 9B is an example of camera locations and fields of view for the exemplary autonomous vehicle of Fig. 9A in accordance with some embodiments of the present disclosure; Fig. 9C a block diagram of an exemplary system architecture for the exemplary autonomous vehicle of Fig. 9A in accordance with some embodiments of the present disclosure; Fig. 9D a system representation in accordance with some embodiments of the present disclosure for communication between a cloud-based server(s) and the exemplary autonomous vehicle of Fig. 9A is; Fig. 10 is a block representation of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure; and Fig. 11 is a block representation of an exemplary data center suitable for use in implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0021] Systems and methods relating to the detection of line segments of traffic features for autonomous and / or semi-autonomous systems and applications are disclosed. Although the present disclosure relates to an exemplary autonomous or semi-autonomous vehicle or machine 900 (hereinafter alternatively referred to as "Vehicle 900", "Ego-Vehicle 900", "Ego-Machine 900" or "Machine 900"), an example of which is given with reference to Fig. The fact that the systems and procedures described herein may be described in sections 9A-9D is not intended to be limiting. For example, the systems and procedures described herein may be used without restriction by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, guided and unguided robots or robotic platforms, warehouse vehicles, all-terrain vehicles, vehicles coupled to one or more trailers, flying vehicles, boats, shuttle vehicles, emergency vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Additionally, although the present disclosure may be described in relation to the detection of line segments associated with traffic features and / or the performance of localization in autonomous or semi-autonomous systems and applications, this is not intended to be limiting, and the systems and methods described herein may be used for augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications and / or any other technology spaces where object perception and / or mapping may be used.
[0022] For example, a system can obtain sensor data generated by one or more sensors from one or more machines navigating an environment. As described herein, the sensor data can include, but are not limited to, image data generated by image sensor(s), LiDAR data generated by LiDAR sensor(s), radar data generated by radar sensor(s), and / or any other type of sensor data generated by any other type of sensor. Additionally, sensor representations provided by the sensor data—such as images, point clouds, and / or the like—can represent at least road markings located within an environment.For example, the sensor representations can represent solid road markings, dash markings, double road markings, mid-line road markings, two-way road markings, main road markings, arrows, stop lines, pedestrian crossing lines, and / or any other type of road marking. As such, at least some of the road markings can include various line segments, such as dash markings.
[0023] The system(s) can then process at least a portion of the sensor data using one or more machine learning models (the model(s)) trained to determine information associated with at least the line segments. For example, based at least on processing the sensor data, the model(s) can generate and / or output data representing points associated with the line segments, directional information (e.g., vectors) associated with the points, boundary shapes (e.g., bounding boxes, etc.) associated with the line segments, and / or any other information. As described herein, in some examples, the points may be associated with the edges of the line segments (e.g.,Starting points (associated with the front faces of line segments) and endpoints (associated with the back faces of line segments, along the direction of travel) can be used, and the directions can be directed towards the midpoints of the line segments. Additionally, or alternatively, in some examples, the points can be associated with the midpoints of the line segments, and the directions can be directed towards the edges of the line segments. However, these are only two examples of the types of points and / or directions that can be associated with line segments.
[0024] In some examples, the points can be associated with specific parts of the sensor representations, such as pixels of images represented by image data. Additionally, in some examples, the points and / or directions can be represented using different coordinate systems. For a first example, in a Cartesian coordinate system, the points can be represented using x-coordinates and y-coordinates, and the directions can be represented using components in the x-coordinate direction and components in the y-coordinate direction. For a second example, in a polar coordinate system, the points can be represented using distances and angles relative to reference points, and the directions can be represented using additional angles.For a third example, the points can be represented using three-dimensional coordinate locations (3D coordinate locations), such as when the points are projected from a two-dimensional space (2D space) onto a 3D space.
[0025] In some examples, the system(s) can perform various types of operations using the information associated with the line segments. For example, in some examples, the system can update a map of the environment using the information to specify, for example, the locations of the line segments, the number of line segments per unit length of a road marking, and / or any other information. In some examples, the system can use the information to locate a machine within the environment. For example, the system can align the locations of the line segments, as determined using the model(s), with the locations of the line segments as represented by the map to locate the machine.In some examples, the system(s) can determine one or more trajectories for the machine to navigate, such as based on rules associated with the road markings that enclose the line segments.
[0026] As described herein, the model(s) can be trained to determine the information associated with the line segments. For example, the system(s) (and / or one or more additional systems) can train the model(s) using training input data—such as image data, LiDAR data, radar data, and / or any other type of sensor data—along with appropriate ground truth data. For example, the ground truth data can represent at least boundary shapes associated with line segments, points associated with the line segments, and / or directional information associated with the points.As described in more detail herein, the system(s) may then use one or more training engines configured to determine one or more losses using the model(s) processing the training input data and the ground truth data. For example, the training engine(s) may determine the loss(s) at least based on comparing the outputs with the ground truth data. The training engine(s) may then update one or more parameters and / or one or more weights associated with the model(s) using the loss(s).
[0027] While the examples herein relate to determining information associated with line segments of road markings, similar processes can be used in other examples to determine information associated with segments of other types of features. For example, similar processes can be used to determine information associated with line segments of traffic signs, traffic signals, traffic masts, and / or any other type of traffic feature, and / or structures, machinery, and / or any other type of object.
[0028] In some examples, the model(s) can be packaged as a microservice—such as an inference microservice (e.g., NVIDIA NIMs)—which can enclose a container (e.g., an operating system (OS)-level virtualization package). This container can include an application programming interface (API) layer, a server layer, a runtime layer, and / or a model engine. For example, the inference microservice can include the container itself and the model(s) (e.g., weights and biases). In some cases, such as when the model(s) is / are small enough (e.g., has a sufficiently small number of parameters), the model(s) can be enclosed within the container itself.In some embodiments, the model(s) described herein can be used as an inference microservice to accelerate model development on a cloud, data center, or edge computing system while ensuring data security. The inference microservice may include, for example: one or more APIs, a pre-configured container for simplified deployment, an optimized inference engine (e.g., built using standardized AI model deployment and execution software, such as NVIDIA's Triton Inference Server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations that deliver low latency and high throughput for production applications—such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g.,(including identity, metrics, health checks, and / or monitoring). The model(s) described herein may be included as part of the microservice along with accelerated infrastructure, capable of being deployed and / or orchestrated with a single command and autoscaling with a container orchestration system on accelerated infrastructure (e.g., from a single device to data center scale). As such, the inference microservice may include the model(s) (optimized, for example, for high-performance inference), inference runtime software for executing the machine learning model(s) and providing outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identity, and / or other monitoring.In some embodiments, the inference microservice may include software for performing in-place replacement and / or updates of the machine learning model(s). During replacement or update, the software performing the replacement / update may maintain user configurations of the inference runtime software and enterprise management software.
[0029] In some embodiments of the systems described herein, procedures can be performed in a simulation environment (e.g., NVIDIA's DriveSIM) using simulated data (e.g., simulated sensor data from a virtual or simulated machine). For example, simulated sensor data can be used to perform various operations within the simulation environment, such as determining information associated with line segments and / or performing localization using that information. These simulated operations can be used to test the performance of the underlying algorithms, systems, and / or processes before their deployment in the real world. In some cases, the simulation can be used to generate synthetic training data—e.g., training data containing landmarks, features, objects, road markings, line segments, etc.include, so that the synthetic training data (in addition to or as an alternative to real-world data) can then be processed to perform one or more of the processes described herein.
[0030] In one example, when a simulation environment is used for testing, validation, training, etc., the simulation environment and / or associated training data can be rendered or otherwise generated using one or more light transport algorithms—such as ray tracing and / or path tracing algorithms. In some embodiments, the simulation environment and / or one or more objects, features, or components thereof can be generated or managed within a three-dimensional content collaboration platform (3D content collaboration platform) (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physical AI, and / or other use cases, applications, or services. For example, the content collaboration platform or system can be a system for using or deploying Universal Scene Descriptors (USDs) (e.g.,OpenUSD data can be used to manage objects, features, scenes, etc., within a simulated environment, digital environment, etc. The platform can incorporate real-world physics simulation, such as using NVIDIA's PhysX SDK, to simulate real-world physics and physical interactions with simulations hosted by the platform. The platform can integrate OpenUSD, along with ray tracing / path tracing / light transport simulation (e.g., NVIDIA's RTX rendering technologies), into software tools and simulation workflows for building, training, deploying, or testing AI systems—such as systems for testing, validating, training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automotive, robotics, machinery, or other applications.
[0031] In some embodiments, teleoperation or remote control of a vehicle or other machine can be performed using a remote control or teleoperation system. For example, the systems and methods described herein can be used to identify lane lines, road boundary lines, longitudinal features, etc., which may be included in a visualization or mapping of an environment, to assist a remote operator in controlling—or providing waypoints or other information for control or navigation—an autonomous or semi-autonomous machine in an environment.
[0032] The systems and procedures described herein may be used, without limitation, by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, guided and unguided robots or robotic platforms, warehouse vehicles, all-terrain vehicles, vehicles coupled to one or more trailers, flying vehicles, boats, shuttle vehicles, emergency vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Furthermore, the systems and methods described herein can be used for a number of purposes, for example, but not limited to, machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actuator simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.
[0033] Disclosed embodiments may be included in a number of different systems, such as automotive systems (e.g.a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, antenna systems, media systems, boat systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twinning operations, systems implemented using an edge device, systems implementing large language models (LLMs), systems implementing one or more multimodal models, systems using or employing one or more inference microservices, systems that integrate one or more machine learning models into a service or microservice along with an OS-level virtualization package (e.g.,include a container), systems that include one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems that are at least partially implemented in a data center, systems for performing conversational AI operations, systems for performing light transport simulation, systems for performing collaborative content creation for 3D assets, systems for performing generative AI operations, systems that are at least partially implemented using cloud computing resources, and / or other types of systems.
[0034] With reference to Fig. 1A illustrated Fig. Figure 1A is an exemplary data flow diagram for a process of determining information associated with line segments of traffic features in accordance with some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and other elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in conjunction with other components, and in any suitable combination and position.Various functions described herein as being performed by entities may be executed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory. In some embodiments, the systems described herein may employ methods and processes using components, features, and / or functionality similar to those of the exemplary autonomous vehicle 900 from [source missing]. Fig. 9A-9D, the exemplary computing device 1000 of Fig. 10 and / or the exemplary data center 1100 of Fig. 11 similar, will be executed.
[0035] For example, the process 100 can include one or more sensors 102 that generate sensor data 104 representing an environment. As described herein, the sensor data 104 can include, but are not limited to, image data 104 generated using image sensor(s) 102, LiDAR data 104 generated using LiDAR sensor(s) 102, RADAR data 104 generated using RADAR sensor(s) 102, and / or any other type of sensor data 104 generated using any other type of sensor 102. Additionally, sensor representations represented by the sensor data 104—such as images, point clouds, and / or the like—can represent at least road markings located within an environment.For example, the sensor representations can represent continuous road markings, dash road markings, double road markings, midway road markings, two-way road markings, main road markings, arrows, stop lines, pedestrian crossing lines and / or any other type of road marking.
[0036] As such, at least some road markings can include various line segments, such as dash markings. For example, a road marking can include a number of marked segments (e.g., the dash markings) and a number of road segments, with the road segments located between the marked segments. For example, a road marking can include a marked segment, followed by a road segment, followed by a marked segment, followed by a road segment, and / or so on. Additionally, a line segment can include any dimension, such as a rectangle that includes both a width and a length. Furthermore, line segments associated with a road marking can include similar dimensions and / or line segments associated with a road marking can include varying dimensions.
[0037] For example, illustrated Fig. 2 An example of a sensor representation representing road markings with different line segments, in accordance with some embodiments of the present disclosure. As shown, the sensor representation can include an image 202 illustrating at least one first road marking consisting of line segments 204(1)-(3) (also referred to singularly as "line segment 204" or plurally as "line segments 204") and a second road marking consisting of line segments 206(1)-(3) (also referred to singularly as "line segment 206" or plurally as "line segments 206"). Between the line segments 204 and the line segments 206, which can also be referred to as "marked segments", there are additional line segments associated with a road surface 208, which can also be referred to as "road segments".
[0038] Referring to the example of Fig. 1A The process 100 can include the following: applying the sensor data 104 to one or more machine learning models 106 (the model(s) 106) configured to process the sensor data 104, and, at a minimum, generating output data 108 based on the processing, representing information associated with the road markings. As shown, the information associated with the road markings (or other navigable surface markings) can include at least points 110, which mark the locations of line segments associated with the road markings, and direction information 112 associated with the points 110. In other examples, however, the output data 108 can represent additional information associated with the road markings, such as boundary shapes (e.g.,Boundary frames), which at least partially enclose the line segments, classifications that specify the types of road markings, 2D locations (e.g. image space) and / or 3D locations (e.g. space) of the road markings and / or so on.
[0039] As described herein, in some examples, the points 110 may be associated with the edges of the line segments (e.g., start points associated with the front faces of the line segments and end points associated with the back faces of the line segments, along the direction of travel), and the direction indicators 112 may be directed toward the midpoints of the line segments. Furthermore, or alternatively, in some examples, the points 110 may be associated with the midpoints of the line segments, and the direction indicators 112 may be directed toward the edges of the line segments. However, these are only two examples of the types of points 110 and / or direction indicators 112 that may be associated with line segments. For example, in other examples, the points 110 and / or the direction indicators 112 may be associated with any other locations corresponding to the line segments of the road markings.
[0040] For example, illustrated Fig. 1B An example of one or more machine learning models 114 (which may include and / or be similar to the model(s) 106) trained to determine information associated with line segments, in accordance with some embodiments of the present disclosure. The model(s) 114 may be a machine learning model that can be used to perform one or more of the processes described herein. The model(s) 114 may include or be described as a convolutional neural network, and may therefore be referred to herein alternatively as a convolutional neural network 114, convolutional network 114, or CNN (Convolutional Neural Network) 114.
[0041] As described herein, the model(s) 114 can use sensor data 116 (which may include and / or represent the sensor data 104) as input. The sensor data 114 can be inputted into one or more feature extractor layers 118 of the model(s) 114. The feature extractor layer(s) 118 can include any number of layers 118, such as layers 118A-118C. One or more of the layers 118 can include an input layer. The input layer can hold values associated with the sensor data 116. For example, if the sensor data 116 is an image / images, the input layer can hold values that are representative of the raw pixel values of the image(s) as volume (e.g., a width, W, a height, H, and color channels, C (e.g., RGB), such as 32x.32x.3), and / or a batch size B (e.g., if batching is used).
[0042] One or more layers can include convolutional layers. The convolutional layers can compute the output of neurons connected to local regions in an input layer (e.g., the input layer), with each neuron calculating a dot product between its weights and a small region to which it is connected in the input volume. As a result, a convolutional layer can be another volume, with one of the dimensions based on the number of filters applied (e.g., width, height, and the number of filters, such as 32x32x12 if 12 were the number of filters).
[0043] One or more of the layers can enclose a Rectified Linear Unit (ReLU) layer. The ReLU layer(s) can include an element-wise activation function, such as max. (0, x), thresholding at zero. The resulting volume of a ReLU layer can be the same as the volume of the input to the ReLU layer.
[0044] One or more of the layers 118 can include a pooling layer. The pooling layer can perform a downsampling operation along the spatial dimensions (e.g., height and width), which can result in a smaller volume than the input of the pooling layer (e.g., 16x16x12 from the 32x32x12 input volume). In some examples, the model model(s) 114 may not include any pooling layers. In other examples, other types of convolutional layers may be used instead of pooling layers. In some examples, the feature extractor layer(s) 118 may include alternating convolutional layers and pooling layers.
[0045] One or more of the layers 118 can include a fully connected layer. Each neuron in the fully connected layer(s) can be connected to each of the neurons in the preceding volume. The fully connected layer can compute class scores, and the resulting volume can be 1x1xN (where N is a number of classes). In some examples, the feature extractor layer(s) 118 can include a fully connected layer, while in other examples, the fully connected layer of the model(s) 114 can be the fully connected layer separate from the feature extractor layer(s) 118. In some examples, no fully connected layers are used by the feature extractor layer(s) 118 and / or the model(s) 114 collectively, in an effort to increase processing times and reduce computational resource requirements.In such examples where no fully connected layers are used, the model(s) 114 can be described as a fully convolutional network.
[0046] In some examples, one or more of the layers 118 may include a deconvolutional layer(s). However, the use of the term deconvolutional can be misleading and is not intended to be restrictive. For example, the deconvolutional layer(s) may alternatively be referred to as transposed convolutional layers or fractionally progressive convolutional layers. The deconvolutional layer(s) may be used to perform upsampling on the output of a preceding layer. For example, the deconvolutional layer(s) may be used to upsample to a spatial resolution equal to the spatial resolution of the input images (e.g., the sensor data 116) for the model(s) 114, or to upsample to the spatial input resolution of a subsequent layer.
[0047] Although input layers, convolutional layers, pooling layers, ReLU layers, deconvolutional layers, and fully connected layers are discussed herein with reference to the feature extractor layer(s) 118, this is not intended to be restrictive. For example, additional or alternative layers 118 may be used in the feature extractor layer(s) 118, such as normalization layers, SoftMax layers, and / or other layer types.
[0048] The output of feature extractor layer(s) 118 can be an input for segment layer(s) 120. Segment layer(s) 120A-C can use one or more of the layer types described herein with reference to feature extractor layer(s) 118. As described herein, in some examples, segment layer(s) 120 may not include fully connected layers to reduce processing speeds and computational resource requirements. In such examples, segment layers 120 may be referred to as fully convolutional layers.
[0049] Different sequences and numbers of layers 118 and 120 of model(s) 114 can be used, depending on the environment. For example, if two or more cameras or other sensor types are used to generate inputs, there may be a different sequence and number of layers 118 and 120 for one or more of the sensors. As another example, a different sequence and number of layers may be used depending on the type of sensor used to generate the sensor data 116 or the type of sensor data 116 (e.g., RGB YUV, etc.). As such, the sequence and number of layers 118 and 120 of model(s) 114 are not limited to a single architecture.
[0050] Furthermore, some layers 118 and 120—such as the feature extractor layer(s) 118 and / or the segment layer(s) 120—may include parameters (e.g., weights and / or biases), while others, such as the ReLU layers and pooling layers, may not. In some examples, the parameters can be learned by the model(s) 114 during training. Additionally, some layers 118 and 120—such as the convolutional layer(s), the deconvolutional layer(s), and the pooling layer(s)—may include additional hyperparameters (e.g., learning rate, progress, epochs, kernel size, number of filters, type of pooling for pooling layers, etc.), while others, such as the ReLU layer(s), may not.Various activation functions can be used, including but not limited to ReLU, Leaky-ReLU, Sigmoid, hyperbolic tangent (tan h), exponential linear unit (ELU), etc. The parameters, hyperparameters, and / or activation functions should not be restrictive and may vary depending on the implementation.
[0051] In any given example, the model(s) 114 can generate point data 122, representing the points associated with the line segments, and direction data 124, representing the directions associated with the points. For example, in some examples, the point data 122 and / or the direction data 124 can include and / or be similar to the output data 108.
[0052] Additionally, illustrate Fig. 3A-3C Examples of determining information associated with line segments of traffic features corresponding to road markings, in accordance with some embodiments of the present disclosure. As by the example of Fig. As illustrated in Figure 3A, the model(s) 106, at least based on processing image data representing an image 302 of a line segment 304 (which may include and / or be similar to one of the line segments 204 and 206), can determine information that includes at least a first point 306(1) associated with a first edge 308(1) of the line segment 304, and a second point 306(2) associated with a second edge 308(2) of the line segment 304, the edges 308(1)-(2) being located where the road surface transitions to the line segment 304. While the example of Fig. Figure 3A illustrates that the points 306(1)-(2) are located at the midpoints of the margins 308(1)-(2), but in other examples the points 306(1)-(2) may be located at other places along the margins 308(1)-(2).
[0053] The model(s) 106 can determine additional information, including at least one first direction 310(1) associated with the first point 306(1) and a second direction 310(2) associated with the second point 306(2). For example, the first direction 310(1) can correspond to a first direction vector starting at the first point 306(1) and directed towards the midpoint of line segment 304. Additionally, the second direction 310(2) can correspond to a second direction vector starting at the second point 306(2) and also directed towards the midpoint of line segment 304. While the example of Fig. 3A describes the determination of two points 306(1)-(2) and two directions 310(1)-(2) associated with both edges 308(1)-(2) of the line segment 304. In another example, the model(s) 106 may only determine one of the points 306(1)-(2) and / or one of the directions 310(1)-(2), such as when one of the edges 308(1)-(2) is blocked within the image.
[0054] As exemplified by Fig. As illustrated in Figure 3B, the model(s) 106, based at least on processing the image data representing image 302, can determine information that now includes at least one point 312 located approximately at the midpoint of line segment 304. Additionally, the model(s) 106 can determine further information that includes at least direction information 314(1)-(2) associated with point 312. For example, the first direction information 314(1) can include a first direction vector starting at point 312 and directed towards the midpoint of the first boundary 308(1), and the second direction information 314(2) can include a second direction vector starting at point 312 and directed towards the midpoint of the second boundary 308(2). While the example of Fig. Figure 3B illustrates that point 312 is located approximately at the midpoint of line segment 304; in other examples, point 312 may be located at any other location within line segment 304.
[0055] As exemplified by Fig. As illustrated in Figure 3C, an image 316 can represent a different type of road marking, including Botts Dots 318(1)-(5) (also referred to singularly as "Botts Dot 318" or plurally as "Botts Dots 318"). As such, in some examples, the model(s) 106 can be trained to group a number of the Botts Dots 318, such as three (and / or any other number), to generate the line segments 320 (although for clarity only one is labeled) associated with the road marking. Additionally, the model(s) 106 can, in turn, determine information associated with the line segments 320, such as using the technique of Fig. 3A and / or the technology of Fig. 3B.
[0056] For example, in the example of Fig. 3C can determine the model(s), at least based on processing image data representing image 316, information that includes at least one first point 322(1) associated with a first boundary 324(1) of line segment 320, and a second point 322(2) associated with a second boundary 324(2) of line segment 320. Additionally, the model(s) can determine information that includes at least one first direction indicator 326(1) associated with the first point 322(1), and a second direction indicator 326(2) associated with the second point 322(2). For example, the first direction indicator 326(1) can correspond to a first direction vector that starts at the first point 322(1) and points towards the midpoint of line segment 320.Additionally, the second direction specification 326(2) can correspond to a second direction vector that starts at the second point 322(2) and is also directed towards the midpoint of the line segment 320.
[0057] While the example of Fig. 3C the use of the technology of Fig. 3A illustrates how to determine points 322(1)-(2) and directions 326(1)-(2). In other examples, the model(s) can provide 106 points and / or directions for line segments 320 using the technique of Fig. 3B determine. Additionally, model(s) 106 will be used, while the example of Fig. 3C initially illustrates the determination of the line segments 320 associated with the Botts Dots 318; in other examples, the model(s) may not determine the line segments 320. Instead, the model(s) may determine 106 points and / or directions associated with the individual Botts Dots 318 (e.g., one point and / or one direction for each Botts Dot 318).
[0058] Referring again to the example of Fig. 1A, in some examples, points 110 can be associated with specific parts of the sensor representations, such as pixels of images represented by image data. Additionally, in some examples, points 110 and / or 112 can be represented as directions using different coordinate systems. For a first example, in a Cartesian coordinate system, points 110 can be represented using x-coordinate locations and y-coordinate locations, and directions 112 can be represented using components in the x-coordinate direction and components in the y-coordinate direction. For a second example, in a polar coordinate system, points 110 can be represented using distances and angles with respect to reference points, and directions 112 can be represented using additional angles.For a third example, the points 110 can be represented using three-dimensional coordinate locations (3D coordinate locations), such as when the points 110 are projected from a two-dimensional space (2D space) onto a 3D space.
[0059] Additionally, the model(s) 106 can output different types of output data 108 in some examples. For a first example, the model(s) 106 can be trained to generate output data 108 representing the locations of the points 110 and the direction information 112. For example, the output data 108 can represent the x-coordinate locations and the y-coordinate locations of the points 110, along with the components of the x-coordinate direction and the lower components of the y-coordinate direction, or the direction information 112.
[0060] For a second example, the model(s) 106 can be trained to generate output data 108 representing probabilities associated with the points 110 and / or the directions 112. For example, the output data 108 can represent probabilities that points 110 are located at different x-coordinate locations, probabilities that points 110 are located at different y-coordinate locations, probabilities that directions 112 include different values for components in the x-coordinate direction, and / or probabilities that directions 112 include different values for components in the y-coordinate direction. In such an example, the model(s) 106 and / or one or more processing components 126 can then process the output data 108 to determine the actual points 110 and direction information 112 associated with the line segments.For example, the model(s) 106 and / or the processing component(s) 126 can use the probabilities to satisfy a threshold probability to determine the x-coordinate locations and the y-coordinate locations of the points 110 and / or the components in the x-coordinate direction and the components in the y-coordinate direction for the direction specifications 112.
[0061] In a third example, the model(s) can be trained to generate output data 108 representing information for individual parts (e.g., pixels) associated with sensor representations. For example, the information can specify the locations of the nearest points 110 with respect to the individual parts. In such an example, the model(s) 106 and / or the processing component(s) 126 can then process the output data 108 to determine at least the actual points 110 associated with the line segments. While these are just a few examples of different outputs that can be generated by the model(s) 106, in other examples the model(s) 106 can generate additional and / or alternative outputs associated with the points 110 and / or the direction information 112.
[0062] For example, illustrate Fig. 4A-4B Examples of different types of outputs in accordance with some embodiments of the present disclosure that can generate the model(s) 106, wherein the outputs include information associated with line segments. As illustrated by the example of Fig. As illustrated in Figure 4A, the model(s) 106 can generate an initial output 402 that includes the locations of points 404(1)-(N) (also referred to singularly as "point 404" or plurally as "points 404"), such as pixels, associated with a sensor representation. Additionally, the initial output 402 includes probabilities 406(1)-(N) (also referred to singularly as "probability 406" or plurally as "probabilities 406") associated with the points 404. For example, the probabilities 406 can indicate a probability that the points 404 include actual points associated with the line segments. In some examples, the model(s) 106 can be trained to output a specific number of points 404, such as 200 points (and / or any other number of points).In some examples, the model(s) 106 can be trained to output a specific number of points 404 based on the sensor representation, such as one point 404 for each pixel of an image.
[0063] The model(s) 106 and / or the processing component(s) 126 can then use the probabilities 406 to determine which of the points 404 include the actual points associated with the line segments. For example, the model(s) 106 and / or the processing component(s) 126 can determine that the points 404 associated with probabilities 406 that satisfy a threshold probability include the actual points. In other examples, however, the model(s) 106 and / or the processing component(s) 126 can use additional and / or alternative techniques to identify the actual points using the probabilities 406.
[0064] As exemplified by Fig. As illustrated in Figure 4B, the model(s) 106 can generate a second output 408 that includes the locations of points 410(1)-(O) (also referred to singularly as "point 410" or plurally as "points 410"), such as pixels in images and / or points from LiDAR, associated with a sensor representation. In some examples, the points 410 can include any number of points associated with the sensor representation, such as a number of points 410 representing a subset of the pixels in the sensor representation, and / or a number of points 410 representing all of the pixels in the sensor representation. Additionally, the second output 408 includes location information 412(1)-(O) (also referred to as "location information 412") associated with the points 410. As described herein, and for a point 410, the location information 412 can be moved to the nearest point (e.g.Specify the nearest pixel associated with a line segment as represented by the sensor representation. For example, the location information 412 can specify a direction, a distance (e.g., a number of pixels), and / or any other type of location information that can be used to identify the nearest line segment point.
[0065] The model(s) 106 and / or the processing component(s) 126 can then use the locations of the points 410 together with the location information 412 to determine the actual points associated with the line segments. For example, for a cluster of points 410 that at least partially surrounds an actual point, the model(s) 106 and / or the processing component(s) 126 can use the location information 412 associated with the cluster of points 410 to identify the location of the actual point that lies within the cluster of points 410.
[0066] Referring again to the example of Fig. 1. The process 100 may then include the generation and / or output of line data 128, representing information associated with the road markings, by the model(s) 106 and / or the processing component(s) 126. For example, the line data 128 may represent the points 110, the direction information 112, and / or additional information such as the number of line segments detected per sensor representation, the number of line segments that are road markings, the dimensions of the line segments, and / or any other information.
[0067] In some examples, process 100 may include the use of line data 128 by one or more mapping components 130 to update a map associated with the environment, where the map is represented by map data 132. For example, the mapping component(s) 130 may update the map to indicate the locations of the road marking line segments, the orientation of the line segments, the number of line segments per length of the road markings, and / or any other information associated with the line segments. Additionally or alternatively, in some examples, process 100 may include providing the line data 128 and / or the map data 132 to one or more machines 134 that navigate the environment.In such examples, the machine(s) 134 can then use the information associated with the line segments and / or the map to perform one or more operations, such as locating the machine(s) 134 within its environment.
[0068] For example, in some examples, a machine 134 can align the line segments, as represented by the line data 128, with the line segments, as represented by the map, to determine at least one location of the machine 134 with respect to the road. As described herein, the location can include a longitudinal location with respect to the road, such as along the direction of travel associated with the machine 134. In this way, even if the environment does not include other types of traffic features, such as traffic poles, traffic signals, and / or road signs, the machine 134 is still able to perform accurate localization in both a lateral and a longitudinal direction.
[0069] As described herein, the model(s) 106 can be trained to determine at least the information associated with the line segments. For example, illustrates Fig. Figure 5 shows a data flow diagram illustrating a process 500 for training the model(s) 106 in accordance with some embodiments of the present disclosure to determine information associated with line segments. As shown, the model(s) 106 can be trained using the training data 502. In some examples, the training data 502 can be similar to the sensor data 104 that are subsequently trained by the model(s) 106, such as by including image data, LiDAR data, radar data, and / or any other type of sensor data. For example, the training data 502 can include image data representing pictures of different types of road markings.
[0070] The model(s) 106 can be trained using the training data 502 together with corresponding ground truth data 504. As shown, the ground truth data 504 can represent at least points 506 associated with line segments, direction information 508 associated with the line segments, and / or additional information 510 associated with the line segments, such as boundary shapes (e.g., bounding boxes, etc.) that specify the dimensions of the line segments. For example, in some examples, the points 506 can specify the pixel locations (and / or other types of sublocations) associated with the edges of the line segments, the midpoints of the line segments, and / or any other points associated with the line segments.Additionally, the direction information 508 can represent direction vectors associated with the points 506, directed in specific directions, such as towards the midpoints of the line segments and / or the edges of the line segments. Furthermore, the additional information 510 can include the boundary shapes that at least partially enclose the line segments. As described, the ground truth data 504 can be produced as follows: synthetically (e.g., generated from computer models or renderings), real (e.g., conceived and produced from real-world data), machine-automated (e.g., using feature analysis and learning to extract features from data and then generate labels), human-annotated (e.g., a labeler or annotation expert defines the location of the labels), and / or a combination thereof.In some examples, for instance, for each instance of training data 502, there may be corresponding ground truth data 504.
[0071] For example, illustrated Fig. 6 An example of information that may be included in ground-truth data for training the model(s) 106, in accordance with some embodiments of the present disclosure. As shown, the ground-truth data may represent at least the locations of points 602(1)-(6) associated with the line segments 204, direction information 604(1)-(6) associated with the line segments 204, the locations of points 606(1)-(6) associated with the line segments 206, and direction information 608(1)-(6) associated with the line segments 206. As described herein, the ground-truth data may represent points 602(1)-(6) and 606(1)-(6) using various techniques, such as using Cartesian coordinate locations, polar coordinate locations, and / or any other type of location information.Additionally, the Ground Truth data can represent direction information 604(1)-(6) and 608(1)-(6) using various techniques, such as information describing direction vectors.
[0072] As such, in some examples, the ground truth data may represent similar information associated with line segments 204 and 206, on which model(s) 106 is / are trained to determine, generate, and / or output them. For example, model(s) 106 may, at least based on processing by model(s) 106 of image data representing image 202, such as using process 100 of Fig. 1A, be configured to determine the points 602(1)-(6) associated with line segments 204, the direction information 604(1)-(6) associated with line segments 204, the points 606(1)-(6) associated with line segments 206, and / or the direction information 608(1)-(6) associated with line segments 206.
[0073] Referring again to the example of Fig. 5. One or more training engines 512 can use one or more loss functions that measure loss (e.g., errors) in outputs 514 compared to the ground truth data 504. As shown, the outputs 514 can include predicted points 516, predicted directions 518, and / or predicted additional information 520 (e.g., boundary shapes, etc.). Any type of loss function can be used, such as cross-entropy loss, mean squared error, mean absolute error, mean bias error, line segmentation loss, and / or other loss function types. In some examples, different outputs 514 can have different loss functions.For example, the predicted points 516 can be associated with a first loss function, the predicted directions 518 can be associated with a second loss function, and / or the predicted additional information 520 can be associated with a third loss function. In such examples, the loss functions can be combined to form a total loss, and the total loss can be used to train the model(s) 106 (e.g., to update their parameters). In one example, backward computations can be performed to recursively calculate gradients of the loss function(s) with respect to training parameters. In some examples, weights and biases of the model(s) 106 can be used to calculate these gradients.
[0074] In some examples, the model(s) 106 may include one or more new machine learning models that are specifically trained to determine the information associated with the line segments. However, in some examples, the model(s) 106 may have previously been trained to determine other information, such as other information associated with road markings. For example, the model(s) 106 may have previously been trained to determine types of road markings (e.g., dash markings, solid markings, etc.), locations of the road markings, and / or any other information. In such examples, the model(s) 106, by performing this further as in the example of Fig. 6 described trainings, then be trained to determine both the original information associated with the road markings and this additional information associated with the line segments of the road markings.
[0075] In some examples, the model(s) 106 can be packaged as a microservice—such as an inference microservice (e.g., NVIDIA NIMs)—which can enclose a container (e.g., an operating system (OS)-level virtualization package). This container can include an application programming interface (API) layer, a server layer, a runtime layer, and / or a model engine. For example, the inference microservice can include the container itself and the model (e.g., weights and biases). In some cases, such as when the model(s) 106 is / are small enough (e.g., has a sufficiently small number of parameters), the model can be enclosed within the container itself.In some embodiments, the model(s) described herein can be used as an inference microservice to accelerate model development on a cloud, data center, or edge computing system while ensuring data security. The inference microservice may include, for example: one or more APIs, a pre-configured container for simplified deployment, an optimized inference engine (e.g., built using standardized AI model deployment and execution software, such as NVIDIA's Triton Inference Server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations that deliver low latency and high throughput for production applications—such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g.,including identity, metrics, health checks, and / or monitoring). The model(s) described herein 106 may be included as part of the microservice along with accelerated infrastructure, capable of being deployed and / or orchestrated with a single command and autoscaling with a container orchestration system on accelerated infrastructure (e.g., from a single device to data center scale). As such, the inference microservice may include the machine learning model(s) (optimized, for example, for high-performance inference), inference runtime software for executing the machine learning model(s) 106 and providing outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identity, and / or other monitoring.In some embodiments, the inference microservice may include software for performing in-place replacement and / or updating of the machine learning model(s). During replacement or update, the software performing the replacement / update may maintain user configurations of the inference runtime software and enterprise management software.
[0076] Now, referring to Fig. 7-8 Each block of Methods 700 and 800 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. Methods 700 and 800 can also be embodied as computationally usable instructions stored in computer storage media. Methods 700 and 800 can be provided by a standalone application, a service or hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name only a few. Additionally, these Methods 700 and 800 are illustrated by way of example with reference to Fig. 1A described. However, these methods 700 and 800 may additionally or alternatively be performed by any system or any combination of systems, including, but not limited to, those described herein.
[0077] Fig. Figure 7 illustrates a data flow diagram showing a method 700 for using one or more machine learning models to determine information associated with line segments, in accordance with some embodiments of the present disclosure. The method 700, at block B702, may include: obtaining sensor data representative of one or more sensor representations, wherein the one or more sensor representations are representative of one or more line segments associated with one or more traffic features in an environment. For example, the sensor(s) 102 may generate the sensor data 104 representing the sensor representation(s), such as image data representing one or more images, LiDAR data representing one or more point clouds, and / or any other type of sensor representation.As described herein, the sensor representation(s) can represent the feature(s), such as a road marking, which includes the line segment(s), such as one or more marked segments and one or more road segments between the marked segment(s).
[0078] Method 700, at block B704, may include: Determining, using one or more machine learning models and based at least on the sensor data, one or more first points associated with one or more first edges of the one or more line segments, and one or more second points associated with one or more second edges of the one or more line segments. For example, the model(s) 106 may process the sensor data 104 and, based at least on the processing, generate the output data 108, which represent at least the first point(s) 110 associated with the first edge(s) of the line segment(s), and the second point(s) 110 associated with the first edge(s) of the line segment(s).Additionally, in some examples, the output data 108 can also represent the direction information 112, which is associated with the points 110.
[0079] Procedure 700, at block B706, may include: performing one or more operations based on at least one or more first points and one or more second points. For example, the mapping component(s) 130 may update the map using points 110 (and / or direction information 112), the machine(s) 134 may perform localization using the map and / or points 110, the machine(s) 134 may determine one or more trajectories to navigate using points 110, and any other process may be performed.
[0080] Fig. Figure 8 illustrates a data flow diagram showing another method 800 for using one or more machine learning models to determine information associated with line segments, in accordance with some embodiments of the present disclosure. Method 800, at block B802, may include: obtaining sensor data representative of one or more sensor representations, wherein the one or more sensor representations are representative of one or more line segments associated with one or more traffic features in an environment. For example, the sensor(s) 102 may generate the sensor data 104 representing the sensor representation(s), such as image data representing one or more images, LiDAR data representing one or more point clouds, and / or any other type of sensor representation.As described herein, the sensor representation(s) can represent the feature(s), such as a road marking, which includes the line segment(s), such as one or more marked segments and one or more road segments between the marked segment(s).
[0081] Method 800, in block B804, may include: Determining, using one or more machine learning models and based at least on the sensor data, one or more points associated with the one or more line segments and one or more direction indicators associated with the one or more points. For example, the model(s) 106 may process the sensor data 104 and, based at least on the processing, generate the output data 108, which represents at least the point(s) 110 associated with the line segment(s) and the direction indicator(s) 112 associated with the point(s) 110.As described herein, the point(s) 110 may be located on one or more edges of the line segment(s), one or more midpoints of the line segment(s) and / or any other location associated with the line segment(s).
[0082] Procedure 800, at block B806, may include: performing one or more operations based on at least one or more points and one or more directions. For example, the mapping component(s) 130 may update the map using point(s) 110 and / or direction(s) 112; the machine(s) 134 may perform localization using the map, point(s) 110 and / or direction(s) 112; the machine(s) 134 may determine one or more trajectories to navigate using point(s) 110 and / or direction(s) 112; and / or any other process may be performed. EXEMPLARY AUTONOMOUS VEHICLE
[0083] Fig. Figure 9A is an illustration of an exemplary autonomous vehicle 900 in accordance with some embodiments of the present disclosure. The autonomous vehicle 900 (alternatively referred to herein as the “vehicle 900”) may, without limitation, include a passenger vehicle, such as a car, truck, bus, rescue vehicle, shuttle, electric or motorized bicycle, motorcycle, fire engine, police vehicle, ambulance, boat, construction vehicle, underwater vehicle, robotic vehicle, drone, aircraft, a vehicle coupled with a trailer (e.g., a semi-trailer truck used for carrying cargo), and / or another type of vehicle (e.g., one that is unmanned and / or that carries one or more passengers).Autonomous vehicles are generally described in terms of automation levels, as defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in their "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 16, 2018; Standard No. J3016-201609, published September 30, 2016; and earlier and future versions of this standard). The Vehicle 900 may be capable of functionality corresponding to one or more of Levels 3 through 6 of autonomous driving levels. The Vehicle 900 may be capable of functionality corresponding to one or more of Levels 1 through 6 of autonomous driving levels.The Vehicle 900, for example, may be capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 6), depending on the configuration. The term "autonomous," as used herein, may include any and / or all types of autonomy for the Vehicle 900 or any other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assistive autonomy, semi-autonomous, mainly autonomous, or any other designation.
[0084] The vehicle 900 can include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. The vehicle 900 can include a drive system 960, such as an internal combustion engine, a hybrid-electric drive unit, a fully electric motor, and / or another type of drive system. The drive system 960 can be connected to a drivetrain of the vehicle 900, which may include a transmission to enable the propulsion of the vehicle 900. The drive system 960 can be controlled in response to signals received from the throttle / accelerator device 962.
[0085] A steering system 964, which may include a steering wheel, can be used to steer the vehicle 900 (e.g., along a desired path or route) when the drive system 960 is in operation (e.g., when the vehicle is in motion). The steering system 964 can receive signals from a steering actuator 966. The steering wheel may be optional for full automation functionality (Level 6).
[0086] The brake sensor system 946 can be used to operate the vehicle brakes in response to receiving signals from the brake actuators 948 and / or brake sensors.
[0087] One or more controllers 936, which include one or more CPU(s), system-on-chips (“SoCs”) 904 ( Fig. The controller(s) 936, which may include one or more GPUs (e.g., 9C), can provide signals (e.g., representing commands) for one or more components and / or systems of the vehicle 900. For example, the controller(s) can send signals to operate the vehicle brakes via one or more brake actuators 948, to operate the steering system 964 via one or more steering actuators 966, and / or to operate the propulsion system 960 via throttle / accelerator device 962. The controller(s) 936 can include one or more (e.g., integrated) onboard computing devices (e.g., supercomputers) that process sensor signals and issue operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 900.The controller(s) 936 may include a first controller 936 for autonomous driving functions, a second controller 936 for functional safety functions, a third controller 936 for artificial intelligence functionality (e.g., computer vision), a fourth controller 936 for infotainment functionality, a fifth controller 936 for redundancy in emergency situations, and / or other controllers. In some embodiments, a single controller 936 may handle two or more of the above functionalities, two or more controllers 936 may handle a single functionality, and / or any combination thereof.
[0088] The controller(s) 936 can provide signals to control one or more components and / or one or more systems of the vehicle 900 in response to sensor data received from one or more sensors (e.g. sensor inputs). The sensor data can be received, for example, without restriction, from the following: Global Navigation Satellite Systems (GNSS) sensors 968 (e.g., global positioning system sensors), radar sensors 960, ultrasonic sensors 962, lidar sensors 964, inertial measurement unit (IMU) sensors 966 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 996, stereo cameras 968, wide-angle cameras 970 (e.g., fisheye cameras), infrared cameras 972, surround-view cameras 974 (e.g., 360-degree cameras), long-range and / or medium-range cameras 998, velocity sensors 944 (e.g.,for measuring the speed of the vehicle 900), vibration sensors 942, steering sensors 940, brake sensors (e.g. as part of the brake sensor system 946) and / or other sensor types.
[0089] One or more of the controllers 936 can receive inputs (e.g., represented by input data) from an instrument cluster 932 of the vehicle 900 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 934, an acoustic alarm, a loudspeaker, and / or via other components of the vehicle 900. The outputs can include information such as vehicle speed, velocity, time, map data (e.g., the high-definition ("HD") map 922 from Fig. 9C), location data (e.g., the location data of vehicle 900, as shown on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the controller(s) 936, etc. The HMI display 934 can, for example, show information about the presence of one or more objects (e.g., a road sign, caution sign, traffic light sequence, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, exit 34B in two miles, etc.).
[0090] The vehicle 900 may further include a network interface 924, which can use one or more wireless antenna(s) 926 and / or one or more modem(s) to communicate over one or more networks. For example, the network interface 924 may be capable of communication via Long-Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”), etc. The wireless antenna(s) 926 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area network(s), such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or for a low-power wide-area network (“LPWANs”), such as LoRaWAN, SigFox, etc.
[0091] Fig. 9B is an example of camera locations and fields of view for the exemplary autonomous vehicle 900 from Fig. 9A in accordance with embodiments of the present disclosure. The cameras and their respective fields of view are an exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different locations on the vehicle 900.
[0092] The camera types may include, but are not limited to, digital cameras adapted for use with the components and / or systems of the Vehicle 900. The camera(s) may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. Depending on the configuration, the camera types may be capable of any frame rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The cameras may be capable of using rolling shutters, global shutters, another type of shutter, or a combination thereof.In some embodiments, the color filter array may include a red-clear-clear-clear (RCCC) color filter array, a red-clear-clear-blue (RCCB) color filter array, a red-blue-green clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear-pixel cameras, such as cameras with an RCCC, RCCB, and / or RBGC color filter array, may be used in an effort to increase light sensitivity.
[0093] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe system). For example, a multi-function monocular camera can be installed to provide functions such as lane departure warning, traffic sign recognition, and intelligent headlight control. One or more of the cameras (e.g., all of the cameras) can simultaneously record and provide image data (e.g., video).
[0094] One or more of the cameras can be mounted in a custom-designed (three-dimensionally printed) assembly to eliminate stray light and reflections from inside the car (e.g., reflections off the dashboard in the windshield mirrors) that could impair the camera's image capture capabilities. Regarding side mirror mounting assemblies, the assemblies can be custom 3D printed so that the camera mounting plate matches the shape of the side mirror. In some examples, the camera(s) can be integrated into the side mirror. For side-view cameras, the camera(s) can also be integrated within four pillars at each corner of the cabin.
[0095] Cameras with a field of view that includes parts of the environment in front of the vehicle 900 (e.g., forward-facing cameras) can be used for surround view to identify ahead routes and obstacles and, with the aid of one or more controllers 936 and / or control SoCs, to assist in providing information critical for generating an occupancy grid and / or determining preferred vehicle routes. Forward-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras can also be used for ADAS functions and systems, including lane departure warnings (LDW), autonomous cruise control (ACC), and / or other functions, such as traffic sign recognition.
[0096] Various cameras can be used in a forward-facing configuration, for example, including a monocular camera platform that incorporates a CMOS (Complementary Metal Oxide Semiconductor) color imager. Another example could be one or more wide-angle cameras (970) that can be used to detect objects entering the field of view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). Although in Fig. While Figure 9B illustrates only one wide-angle camera, there can be any number (including zero) of wide-angle cameras 970 on the vehicle 900. Additionally, any number of remote cameras 998 (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not yet been trained. The remote camera(s) 998 can also be used for object detection and classification, as well as simple object tracking.
[0097] Any number of stereo cameras 968 can also be included in a forward-facing configuration. In at least one embodiment, one or more stereo camera(s) 968 can include an integrated control unit comprising a scalable processing unit that can provide programmable logic (“FPGA”) and a multi-core microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle's surroundings, including a distance estimate for all points in the image. An alternative stereo camera(s) 968 can include a compact stereo vision sensor, which can include two camera lenses (one each on the left and right) and an image processing chip that measures the distance from the vehicle to the target object and processes the generated information (e.g.,Metadata) can be used to activate the autonomous emergency braking and lane departure warning functions. Other types of Stereo Camera 968 may be used in addition to or as an alternative to those described herein.
[0098] Cameras with a field of view that includes parts of the surroundings at the side of the vehicle 900 (e.g., side-view cameras) can be used for surround view, which provides information used to generate and update the occupancy grid and to generate side-impact collision warnings. For example, one or more surround camera(s) 974 (e.g., four surround cameras 974 as shown in Fig. (Figure 9B illustrates) is positioned on the vehicle 900. The surround view camera(s) 974 can include one or more wide-angle cameras 970, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround view cameras 974 (e.g., left, right, and rear) and can use one or more other cameras (e.g., a forward-facing camera) as a fourth surround view camera.
[0099] Cameras with a field of view that includes parts of the area behind the vehicle 900 (e.g., rear-view cameras) can be used for parking assistance, surround view, rear collision warnings, and generating and updating the occupancy grid. A wide range of cameras can be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., long-range and / or medium-range cameras 998, stereo cameras 968, infrared cameras 972, etc.), as described herein.
[0100] Fig. 9C is a block diagram of an exemplary system architecture for the exemplary autonomous vehicle 900 from Fig. 9A in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and other elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as discrete or distributed components, or in conjunction with other components, and in any suitable combination and position. Various functions described herein as being performed by entities may be executed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.
[0101] The components, features and systems of the 900 vehicle in Fig. 9C are each illustrated as connected via bus 902. Bus 902 may include a Controller Area Network (CAN) data interface (alternatively referred to herein as the CAN bus). A CAN can be a network within the vehicle 900 that is used to support the control of various features and functionality of the vehicle 900, such as actuation of the brakes, acceleration, braking, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to obtain steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL-B compliant.
[0102] Although described herein as a CAN bus, this is not intended to be restrictive. For example, FlexRay and / or Ethernet may be used in addition to or as an alternative to the CAN bus. Furthermore, although a single line is used to represent the 902 bus, this is not intended to be restrictive. For example, there may be any number of 902 buses, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more 902 buses may be used to perform different functions and / or for redundancy. For example, a first 902 bus may be used for collision avoidance functionality, and a second 902 bus may be used for actuation control.In any given example, the bus 902 can communicate with any of the components of the vehicle 900, and two or more buses 902 can communicate with the same components. In some examples, each SoC 904, each controller 936, and / or each computer within the vehicle can have access to the same input data (e.g., inputs from sensors of the vehicle 900) and be connected to a common bus, such as the CAN bus.
[0103] The vehicle 900 can include one or more controllers 936, such as those mentioned herein in relation to Fig. 9A are described. The 936 controller(s) can be used for a number of functions. The 936 controller(s) can be coupled with any of the various other components and systems of the 900 vehicle and used for controlling the 900 vehicle, the 900 vehicle's artificial intelligence, the 900 vehicle's infotainment system, and / or the like.
[0104] The vehicle 900 can include one or more system-on-a-chip (SoC) 904. The SoC 904 can include one or more CPUs 906, GPUs 908, one or more processors 910, caches 912, accelerators 914, data storage 916, and / or other components and features not illustrated. The SoC(s) 904 can be used to control the vehicle 900 in a variety of platforms and systems. For example, the SoC(s) 904 in a system (e.g., the system of the vehicle 900) can be combined with an HD card 922, enabling map refreshes and / or updates via a network interface 924 from one or more servers (e.g., the server(s) 978). Fig. 9D) can be obtained.
[0105] The CPU(s) 906 may include a CPU cluster or CPU complex (alternatively referred to herein as "CCPLEX"). The CPU(s) 906 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU(s) 906 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU(s) 906 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The CPU(s) 906 (e.g., the CCPLEX) may be configured to support concurrent cluster operation, allowing a combination of CPU(s) 906 clusters to be active at any given time.
[0106] The CPU(s) 906 can implement power management capabilities that include one or more of the following features: individual hardware blocks can be automatically clock-gated (automatically disconnected from the clock signal in a gate) when idle to save dynamic power; each core clock can be disconnected when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core can be independently power-gated; each core cluster can be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster can be independently power-gated when all cores are power-gated.The CPU(s) 906 can further implement an improved power state management algorithm, specifying the permissible power states and expected wake-up times, and the hardware / microcode for the core, cluster, and CCPLEX determining the best power state to enter. The processing cores can support simplified software sequences for entering power states, offloading the work to the microcode.
[0107] The GPU(s) 908 may include an integrated GPU (alternatively referred to herein as the "iGPU"). The GPU(s) 908 may be programmable and efficient for parallel workloads. The GPU(s) 908 may, in some examples, use an enhanced tensor instruction set. The GPU(s) 908 may include one or more streaming microprocessors, each of which may include an L1 cache (e.g., an L1 cache with a minimum storage capacity of 96 KB), and two or more of the streaming microprocessors may share an L2 cache (e.g., an L2 cache with a storage capacity of 612 KB). In some embodiments, the GPU(s) 908 may include at least eight streaming microprocessors. The GPU(s) 908 can use Compute Application Programming Interface(s) (API(s)). Furthermore, the GPU(s) 908 can utilize one or more parallel computing platforms and / or programming models (e.g.,NVIDIA's CUDA).
[0108] The GPU(s) 908 can be performance-optimized to achieve the best performance in automotive and embedded applications. For example, the GPU(s) 908 can be manufactured using a FinFET field-effect transistor. However, this is not intended to be a limitation, and the GPU(s) 908 can be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can contain a number of mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block could be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two NVIDIA TENSOR CORES for mixed-precision deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit and / or a 64 KB register file.Furthermore, streaming microprocessors can include independent parallel integer and floating-point data paths to enable efficient execution of workloads with a mix of computational and addressing operations. Streaming microprocessors can include independent thread scheduling functions to allow for finer-grained synchronization and collaboration between parallel threads. Streaming microprocessors can also include a combined L1 data cache and shared memory to improve performance while simplifying programming.
[0109] The GPU(s) 908 can include high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / second in some examples. In some examples, synchronous graphics random access memory (SGRAM), such as double-data-rate type five (GDDR6) synchronous graphics random access memory, can be used in addition to or as an alternative to the HBM memory.
[0110] The GPU(s) 908 can include unified memory technology, including access counters, to allow more accurate migration of memory pages to the processor that accesses them most frequently, thereby improving efficiency for memory areas shared by processors. In some examples, support for address translation services (ATS) can be used to allow the GPU(s) 908 to directly access the page tables of the CPU(s) 906. In such examples, an address translation request can be sent to the CPU(s) 906 if the memory management unit (MMU) of the GPU(s) 908 experiences an erroneous access. In response, the CPU(s) 906 can search its page tables for the virtual-to-physical mapping for the address and send the translation back to the GPU(s) 908.As such, Unified Memory technology can allow a single, unified virtual address space for the memory of both the CPU(s) 906 and the GPU(s) 908, thereby simplifying the programming of the GPU(s) 908 and the porting of applications to the GPU(s) 908.
[0111] Furthermore, GPU(s) 908 can include an access counter that tracks the frequency of GPU(s) 908 access to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently.
[0112] The SoC(s) 904 can include any number of Cache(s) 912, including those described herein. For example, the Cache(s) 912 can include an L3 cache available to both the CPU(s) 906 and the GPU(s) 908 (e.g., connected to both the CPU(s) 906 and the GPU(s) 908). The Cache(s) 912 can include a write-back cache capable of tracking row states, e.g., using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache can include 4 MB or more, although smaller cache sizes can be used.
[0113] The SoC(s) 904 may include one or more arithmetic logic units (ALUs) that can be used to perform processing with respect to any of the tasks or operations of the Vehicle 900, such as processing DNNs. Furthermore, the SoC(s) 904 may include one or more floating-point units (FPUs) – or other mathematical or numerical coprocessor types – for performing mathematical operations within the system. For example, the SoC(s) 104 may include one or more FPUs that are integrated as execution units within one or more CPU(s) 906 and / or GPU(s) 908.
[0114] The SoC(s) 904 can include one or more Accelerators 914 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC(s) 904 can include a hardware acceleration cluster, which may include optimized hardware accelerators and / or a large amount of on-chip memory. The large on-chip memory (e.g., 4 MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to complement the GPU(s) 908 and offload some tasks from the GPU(s) 908 (e.g., to free up more GPU cycles for other tasks). As an example, the Accelerator 914 can be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.).), which are stable enough to be amenable to acceleration. The term “CNN”, as used herein, can include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., such as those used for object detection).
[0115] The Accelerator 914 (e.g., the hardware acceleration cluster) can include one or more deep learning accelerators (DLAs). The DLA(s) can include one or more tensor processing units (TPUs), which can be configured to provide an additional ten trillion operations per second for deep learning applications and inference. The TPUs can be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The DLA(s) can further be optimized for a specific set of neural network types and floating-point operations, as well as for inference. The design of the DLA(s) can deliver more performance per millimeter than a general-purpose GPU and far surpasses the performance of a CPU.The TPU(s) can perform several functions, including a single-instance folding function that supports, for example, INT8, INT16 and FP16 data types for both features and weights, as well as post-processor functions.
[0116] The DLA(s) can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data for any variety of functions, including, but not limited to: a CNN for object identification and detection using camera sensor data; a CNN for distance estimation using camera sensor data; a CNN for emergency vehicle detection and identification using microphone data; a CNN for facial recognition and vehicle owner identification using camera sensor data; and / or a CNN for security-related events.
[0117] The DLA(s) can perform any function of the GPU(s) 908, and using an inference accelerator, a designer can, for example, use either the DLA(s) or the GPU(s) 908 for any given function. A designer can, for instance, focus the processing of CNNs and floating-point operations on the DLA(s) and leave other functions to the GPU(s) 908 and / or another accelerator 914.
[0118] The Accelerator 914 (e.g., the hardware acceleration cluster) can include one or more programmable vision accelerators (PVAs), which may alternatively be referred to herein as computer vision accelerators. The PVA(s) can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving applications, augmented reality (AR), and / or virtual reality (VR). The PVA(s) can provide a balance between performance and flexibility. For example, and without limitation, each PVA can include any number of reduced instruction set compute cores (RISC), direct memory access (DMA), and / or any number of vector processors.
[0119] The RISC cores can interact with image sensors (e.g., the image sensors of any of the cameras described herein), image signal processor(s), and / or the like. Each RISC core can include any memory location. The RISC cores can use any of a number of protocols, depending on the implementation. In some examples, the RISC cores can run a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.
[0120] The DMA can enable components of the PVA(s) to access system memory independently of the CPU(s) 906. The DMA can support any number of features provided to optimize the PVA, including, but not limited to, support for multidimensional addressing and / or circular addressing. In some examples, the DMA can support up to six or more addressing dimensions, which may include block width, block height, block depth, horizontal block gradation, vertical block gradation, and / or depth gradation.
[0121] Vector processors can be programmable processors designed to efficiently and flexibly execute computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, DMA engine(s) (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may act as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). A VPU core may include a digital signal processor, such as a single-instruction multiple data (SIMD) or very-long instruction word (VLIW) signal processor. The combination of SIMD and VLIW can improve throughput and speed.
[0122] Each vector processor can include an instruction cache and be coupled to dedicated memory. Consequently, in some examples, each vector processor can be configured to operate independently of the others. In other examples, vector processors enclosed in a given PVA can be configured to utilize data parallelism. For example, in some embodiments, the multitude of vector processors enclosed in a single PVA can execute the same computer vision algorithm, but in different regions of an image. In other examples, the vector processors enclosed in a given PVA can simultaneously execute different computer vision algorithms on the same image, or even different algorithms on sequential images or portions of an image.Among other things, any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each of the PVAs. Furthermore, the PVA(s) can include additional ECC (Error Correction Code) memory to improve overall system security.
[0123] The Accelerator 914 (e.g., the hardware acceleration cluster) can include a computer vision network on a single chip and SRAM to provide high-bandwidth, low-latency SRAM for the Accelerator 914. In some examples, the on-chip memory can include at least 4 MB of SRAM, consisting, for example, and without limitation, of eight field-configurable memory blocks accessible by both the PVA and the DLA. Each pair of memory blocks can include an Advanced Peripheral Bus (APB) interface, a configuration circuit, a controller, and a multiplexer. Any type of memory can be used. The PVA and DLA can access the memory via a backbone, providing high-speed memory access for both the PVA and DLA.The backbone can include a computer vision network on a chip that connects the PVA and DLA to the memory (e.g., using the APB).
[0124] The computer vision network can include an interface on a single chip that, prior to the transmission of control signals / addresses / data, ensures that both the PVA and the DLA provide ready-to-use and valid signals. Such an interface can provide separate phases and channels for the transmission of control signals / addresses / data, as well as burst communication for continuous data transfer. This type of interface can conform to ISO 26262 or IEC 61608 standards, although other standards and protocols can also be used.
[0125] In some examples, the SoC(s) 904 may include a real-time ray-tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray-tracing hardware accelerator can be used for fast and efficient determination of the positions and extents of objects (e.g., within a world model), for generating real-time visualization simulations, for interpreting radar signals, for synthesizing and / or analyzing sound propagation, for simulating SONAR systems, for general wave propagation simulation, for comparison with lidar data for localization purposes, and / or for other functions and / or other purposes. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray-tracing-related operations.
[0126] The Accelerator 914 (e.g., the hardware accelerator cluster) has a wide range of applications in autonomous driving. The PVA can be a programmable vision accelerator used for key processing stages in ADAS and autonomous vehicles. The PVA's capabilities are well-suited for algorithmic domains requiring predictable processing with low power consumption and low latency. In other words, the PVA performs well with semi-dense or dense regular computations, even with small datasets, that require predictable runtimes with low latency and low power consumption. In the context of autonomous vehicle platforms, PVAs are therefore designed to execute classical computer vision algorithms, as they are efficient at object detection and operate with integer mathematics.
[0127] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. A semi-global matching-based algorithm can be used in some examples, although this is not intended to be a limiting factor. Many Level 3-6 autonomous driving applications require motion estimation / stereo matching during operation (on-the-fly) (e.g., structure from motion, pedestrian detection, lane detection, etc.). The PVA can perform computer stereo vision functions with input from two monocular cameras.
[0128] In some examples, the PVA can be used to perform dense optical flow processing, such as processing raw radar data (e.g., using 4D Fast Fourier Transform) to produce processed radar data. In other examples, the PVA is used for time-of-flight depth processing, by processing the raw time-of-flight data to produce, for example, processed time-of-flight data.
[0129] The DLA can be used to operate any type of network to improve control and driving safety, for example, a neural network that outputs a confidence score for each object detection. Such a confidence score can be interpreted as a probability or as providing a relative "weight" to each detection compared to others. This confidence score allows the system to make further decisions about which detections should be considered true positives and not false positives. For example, the system can set a confidence threshold and only consider detections that exceed this threshold as true positives.In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically initiate emergency braking, which is obviously undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. The DLA can operate a neural network to regress the confidence level. The neural network can use as input at least a subset of parameters, such as the bounding frame dimensions, the ground plane estimate obtained (e.g. from another subsystem), the output of an inertial measurement unit (IMU) sensor 966 correlated with the vehicle's orientation 900, the distance, 3D position estimates of the object obtained from the neural network and / or other sensors (e.g., LIDAR sensor(s) 964 or RADAR sensor(s) 960).
[0130] The SoC(s) 904 may include one or more data stores 916 (e.g., memory). The data store 916 may be on-chip memory of the SoC(s) 904 capable of storing neural networks to be executed on the GPU and / or the DLA. In some examples, the data store 916 may be large enough to store multiple instances of neural networks for redundancy and safety. The data store 912 may include L2 or L3 cache(s) 912. Reference to the data store 916 may include reference to the memory associated with the PVA, DLA, and / or other accelerator(s) 914, as described herein.
[0131] The SoC(s) 904 can include one or more Processor(s) 910 (e.g., embedded processors). The Processor(s) 910 can include a Boot and Power Management Processor, which may be a dedicated processor and subsystem that handles boot power and management functions and associated security enforcement. The Boot and Power Management Processor can be part of the SoC(s) 904 boot sequence and provide runtime power management services. The Boot and Power Management Processor can provide clock and voltage programming, support for low-power system state transitions, management of the SoC(s) 904's thermals and temperature sensors, and / or management of the SoC(s) 904's power supply states.Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the SoC(s) 904 can use ring oscillators to detect temperatures of the CPU(s) 906, GPU(s) 908, and / or accelerator(s) 914. If it is determined that the temperatures exceed a threshold, the boot and power management processor can enter a temperature fault routine and put the SoC(s) 904 into a lower power consumption state and / or put the vehicle 900 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 900 to a safe stop).
[0132] The 910 processor(s) may further include a set of embedded processors that can serve as an audio processing engine. The audio processing engine can be an audio subsystem that provides full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0133] The 910 processor(s) may further include an always-on processor engine that can provide the necessary hardware features to support the management of low-power sensors and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0134] The 910 processor(s) can further include a security cluster engine, which comprises a dedicated processor subsystem for handling security management for automotive applications. The security cluster engine can include two or more processor cores, tightly coupled RAM, supporting peripherals (such as timers, an interrupt controller, etc.), and / or routing logic. In a security mode, the two or more cores can operate in lockstep mode, functioning as a single core with comparison logic to detect differences between their operations.
[0135] The 910 processor(s) may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0136] The 910 processor(s) may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0137] The 910 processor(s) can include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to produce the final image for the playback window. The video image compositor can perform lens distortion correction on the 970 wide-angle camera(s), the 974 surround-view camera(s), and / or the in-cabin surveillance camera sensors. An in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the Advanced SoC, configured to identify and respond to in-cabin events.An in-cabin system can perform lip reading to activate cellular service and make phone calls, dictate emails, change the vehicle's destination, activate or change the infotainment system and vehicle settings, or provide voice-activated web browsing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are otherwise disabled.
[0138] The video image compositor can include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if motion occurs in a video, the noise reduction can weight spatial information accordingly, thereby reducing the impact of information provided by adjacent frames. If an image or part of an image does not contain motion, the temporal noise reduction performed by the video image compositor can use information from the previous image to reduce noise in the current image.
[0139] The video image compositor can also be configured to perform stereo equalization on the input stereo lens frames. Furthermore, the video image compositor can be used for user interface composition when the operating system desktop is in use and the GPU(s) 908 are not required to continuously render new surfaces. The video image compositor can also be used to offload the GPU(s) 908, even when the GPU(s) 908 are powered on and actively performing 3D rendering, to improve performance and responsiveness.
[0140] The SoC(s) 904 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface for receiving video and camera input, a high-speed interface, and / or a video input block that can be used for camera and associated pixel input functions. The SoC(s) 904 may also include one or more input / output controllers that can be software-controlled and used to receive I / O signals not designated for a specific role.
[0141] The SoC(s) 904 can also include a wide selection of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC(s) 904 can be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., LiDAR sensor(s) 964, RADAR sensor(s) 960, etc., which may be connected via Ethernet), data from bus 902 (e.g., vehicle speed 900, steering wheel position, etc.), data from GNSS sensor(s) 968 (e.g., connected via an Ethernet bus or a CAN bus), etc. The SoC(s) 904 can also include dedicated high-performance mass storage controllers, which can include their own DMA engines and can be used to offload routine data management tasks from the CPU(s) 906.
[0142] The 904 SoC(s) can be an end-to-end platform with a flexible architecture encompassing automation levels 3-6, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS techniques for diversity and redundancy, and offers a platform for a flexible, reliable powertrain software stack, along with deep learning tools. The 904 SoC(s) can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, the 914 accelerator(s), in combination with the 906 CPU(s), 908 GPU(s), and 916 data storage(s), can provide a fast, efficient platform for autonomous vehicles operating at levels 3-6.
[0143] This technology thus provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be run on CPUs that can be configured using a higher-level programming language, such as C, to execute a wide range of processing algorithms on a wide range of visual data. However, CPUs are often unable to meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical Level 3-6 autonomous vehicles.
[0144] Unlike conventional systems, the technology described herein, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, allows multiple neural networks to run simultaneously and / or sequentially and to combine the results to enable autonomous driving functionality at levels 3-6. For example, a CNN running on the DLA or the dGPU (e.g., the GPU(s) 920) can include text and word recognition, allowing the supercomputer to read and understand traffic signs, including those for which the neural network has not been specifically trained. The DLA can further include a neural network capable of identifying and interpreting a sign, providing a semantic understanding of the sign, and passing this semantic understanding to the route planning modules running on the CPU complex.
[0145] As another example, multiple networks can operate simultaneously, as required for Level 3, 4, or 6 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate black ice" along with an electric light can be interpreted independently or jointly by several neural networks. The sign itself can be identified as a traffic sign by a first neural network (e.g., a trained neural network), while the text "Flashing lights indicate black ice" can be interpreted by a second neural network, which informs the vehicle's route planning software (preferably running on the CPU complex) that black ice is present when flashing lights are detected.The flashing light can be identified by a third, deployed neural network over several frames, thus informing the vehicle's route planning software of the presence (or absence) of flashing lights. All three neural networks can operate simultaneously, for example, within the DLA and / or on the GPU(s) 908.
[0146] In some examples, a CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or vehicle owner of the Vehicle 900. The always-on sensor processing engine can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner leaves. In this way, the SoC(s) 904 provides security against theft and / or forced vehicle removal.
[0147] In another example, a CNN for emergency vehicle detection and identification can use data from microphones 996 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers to detect sirens and manually extract features, the SoC(s) 904 can use the CNN to classify ambient and urban noise as well as visual data. In a preferred embodiment, the CNN, operating on the DLA, is trained to identify the relative approach speed of the emergency vehicle (e.g., using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle operates, as identified by the GNSS sensor(s) 968.For example, when operating in Europe, the CNN will attempt to detect European sirens, and when operating in the United States, the CNN will attempt to identify only North American sirens. A control program, with the assistance of ultrasonic sensors 962, can be used, once an emergency vehicle is detected, to execute a safety routine for the emergency vehicle, such as slowing down, pulling over to the side of the road, parking the vehicle, and / or holding the vehicle in neutral until the emergency vehicle(s) has / have passed.
[0148] The vehicle can include one or more CPU(s) 918 (e.g., discrete CPU(s) or dCPU(s)) that can be coupled to the SoC(s) 904 via a high-speed connection (e.g., PCIe). The CPU(s) 918 can include, for example, an x86 processor. The CPU(s) 918 can be used to perform any of a number of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC(s) 904 and / or monitoring the status and state of the Controller(s) 936 and / or Infotainment SoC 930.
[0149] The Vehicle 900 can include one or more GPU(s) 920 (e.g., discrete GPU(s) or dGPU(s)) that can be coupled to the SoC(s) 904 via a high-speed connection (e.g., NVIDIA's NVLINK). The GPU(s) 920 can provide additional artificial intelligence functionality, such as running redundant and / or different neural networks, and can be used to train and / or update neural networks based on input (e.g., sensor data) from sensors in the Vehicle 900.
[0150] The vehicle 900 can also include the network interface 924, which can include one or more wireless antennas 926 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 924 can be used to enable wireless connectivity via the internet to the cloud (e.g., to the server(s) 978 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). Communication with other vehicles can be established via a direct connection between the two vehicles and / or an indirect connection (e.g., via networks and the internet). Direct connections can be provided using a vehicle-to-vehicle communication link.The vehicle-to-vehicle communication link can provide the Vehicle 900 with information about vehicles in its vicinity (e.g., vehicles in front of, to the side of, and / or behind the Vehicle 900). This functionality can be part of a cooperative adaptive cruise control feature of the Vehicle 900.
[0151] The network interface 924 can include a SoC that provides modulation and demodulation functionality, enabling the controller(s) 936 to communicate over wireless networks. The network interface 924 can include a radio frequency front end for upconverting baseband to radio frequency and downconverting radio frequency to baseband. The frequency conversions can be performed using known processes and / or superheterodyne processes. In some examples, the radio frequency front-end functionality can be provided by a separate chip. The network interface can include wireless functionality for communication over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0152] The vehicle 900 may further include one or more data storage devices 928, which may include off-chip (e.g., off-SoC(s) 904) storage. The data storage device(s) 928 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash, hard disks, and / or other components and / or devices capable of storing at least one data bit.
[0153] The vehicle 900 can also include one or more GNSS sensor(s) 968. The GNSS sensor(s) 968 (e.g., GPS, supported GPS sensors, differential GPS sensors (DGPS), etc.) serve to support mapping, perception, occupancy grid generation, and / or route planning functions. Any number of GNSS sensor(s) 968 can be used, including, for example, and without limitation, one GPS using a USB port with an Ethernet-to-serial (e.g., RS-232) bridge.
[0154] The vehicle 900 can also include one or more RADAR sensor(s) 960. The RADAR sensor(s) 960 can be used by the vehicle 900 for long-range vehicle detection, even in darkness and / or adverse weather conditions. The functional RADAR safety levels can be ASIL B. In some examples, the RADAR sensor(s) 960 can use CAN and / or bus 902 (e.g., to transmit data generated by the RADAR sensor(s) 960) for control and access to object tracking data, with access to Ethernet for raw data. A wide range of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor(s) 960 can be suitable for front, rear, and side RADAR applications. In some examples, a pulse Doppler radar sensor(s) is / are used.
[0155] The 960 RADAR sensor(s) can include various configurations, such as long-range with a narrow field of view, short-range with a wide field of view, side coverage with short-range, etc. In some examples, long-range RADAR can be used for adaptive cruise control functionality. The long-range RADAR systems can provide a wide field of view, achieved through two or more independent scans, for example, within a range of 260 m. The 960 RADAR sensor(s) can assist in distinguishing between stationary and moving objects and can be used by ADAS systems for emergency braking assistance and forward collision warning. Long-range radar sensors can include monostatic multimodal radar with multiple (e.g., six or more) fixed radar antennas and a high-speed CAN and FlexRay interface.In an example with six antennas, the central four antennas can generate a focused beam pattern designed to record the area around vehicle 900 at higher speeds with minimal interference from traffic in adjacent lanes. The other two antennas can expand the field of view, allowing vehicles entering or leaving vehicle 900's lane to be detected quickly.
[0156] Mid-range radar systems, for example, can include a range of up to 960 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 960 degrees (rear). Short-range radar systems can, without limitation, include radar sensors designed for installation at both ends of the rear bumper. When such radar sensor systems are installed at both ends of the rear bumper, they can generate two beams that continuously monitor the blind spot in one direction, rearward and close to the vehicle.
[0157] Short-range radar systems can be used in the ADAS system to detect a blind spot and / or to assist with a lane change.
[0158] The vehicle 900 can also include one or more ultrasonic sensor(s) 962. The ultrasonic sensor(s) 962 can be positioned on the front, rear, and / or sides of the vehicle 900 and can be used for parking assistance and / or for generating and updating an occupancy grid. A wide range of ultrasonic sensor(s) 962 can be used, and different ultrasonic sensor(s) 962 can be used for different detection ranges (e.g., 2.6 m, 4 m). The ultrasonic sensor(s) 962 can operate according to the functional safety level ASIL B.
[0159] The vehicle 900 can include one or more LiDAR sensor(s) 964. The LiDAR sensor(s) 964 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor(s) 964 can operate at functional safety level ASIL B. In some examples, the vehicle 900 can include multiple LiDAR sensors 964 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0160] In some examples, the LiDAR sensor(s) 964 can provide a list of objects and their distances for a 360-degree field of view. A commercially available LiDAR sensor(s) 964, for example, may have a specified range of approximately 900 m with an accuracy of 2 cm to 3 cm and support for a 900 Mbps Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 964 may be used. In such examples, the LiDAR sensor(s) 964 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of the vehicle 900. The LIDAR sensor(s) 964 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 36 degrees in such examples, with a range of 200 m, even for objects with low reflectivity.One or more front-mounted LIDAR sensor(s) 964 can be configured for a horizontal field of view between 46 degrees and 136 degrees.
[0161] In some examples, LiDAR technologies, such as 3D flash LiDAR, can also be used. 3D flash LiDAR uses a laser flash as a transmission source to illuminate the vehicle's surroundings up to approximately 200 m. A flash LiDAR unit includes a sensor that records the laser pulse time-of-flight and the reflected light at each pixel, which corresponds to the vehicle's range to objects. With flash LiDAR, highly accurate and distortion-free images of the surroundings can be generated with each laser flash. In some examples, four flash LiDAR sensors can be used, one on each side of the vehicle. Available 3D flash LiDAR systems include a solid-state 3D LiDAR camera with a rigid array and no moving parts except for a fan (e.g., a non-scanning LiDAR device).The Blitz LIDAR device can use a 6-nanosecond Class I (eye-safe) laser pulse per frame and capture reflected laser light in the form of 3D distance point clouds and co-recorded intensity data. By using Blitz LIDAR, and because Blitz LIDAR is a solid-state device with no moving parts, the LIDAR sensor(s) 964 may be less susceptible to motion blur, vibration, and / or shock.
[0162] The vehicle may further include one or more IMU sensor(s) 966. The IMU sensor(s) 966 may be located in the center of the rear axle of the vehicle 900 in some examples. The IMU sensor(s) 966 may include, for example, without limitation, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In some examples, such as in six-axis applications, the IMU sensor(s) 966 may include accelerometers and gyroscopes, while in nine-axis applications, the IMU sensor(s) 966 may include accelerometers, gyroscopes, and magnetometers.
[0163] In some embodiments, the IMU sensor(s) 966 can be implemented as a miniature, high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical system (MEMS) inertial sensors, a high-sensitivity GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and attitude. As such, in some examples, the IMU sensor(s) 966 can enable the vehicle 900 to estimate its course without requiring input from a magnetic sensor by directly observing and correlating velocity changes from a GPS to the IMU sensor(s) 966. In some examples, the IMU sensor(s) 966 and the GNSS sensor(s) 968 can be combined in a single integrated unit.
[0164] The vehicle can include one or more microphone(s) 996, which are placed in and / or around the vehicle 900. The microphone(s) 996 can be used, among other things, for emergency vehicle detection and identification.
[0165] The vehicle may also include any number of camera types, including one or more stereo cameras 968, wide-angle cameras 970, infrared cameras 972, surround-view cameras 974, long-range cameras and / or medium-range cameras 998, and / or other camera types. The cameras may be used to capture image data around the entire periphery of the vehicle 900. The types of cameras used depend on the embodiments and requirements for the vehicle 900, and any combination of camera types may be used to provide the necessary coverage around the vehicle 900. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or any other number of cameras.The cameras can, for example, support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet communication without restriction. Each camera is described in more detail below. Fig. 9A and Fig. 9B described.
[0166] The vehicle 900 may also include one or more vibration sensors 942. The vibration sensor(s) 942 may measure vibrations of vehicle components, such as the axle(s). For example, changes in vibration may indicate a change in the road surface. In another example, if two or more vibration sensors 942 are used, the differences between the vibrations may be used to determine the friction or slipperiness of the road surface (e.g., if the difference in vibration is between a driven axle and a freely rotating axle).
[0167] The vehicle 900 may include an ADAS system 938. The ADAS system 938 may include a SoC in some examples. The ADAS system 938 may include systems for autonomous / adaptive / automatic distance control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warnings (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning systems (CWS), lane centering (LC), and / or other features and functions.
[0168] The ACC systems can use one or more radar sensors (960), lidar sensors (964), and / or one or more cameras. The ACC systems can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle (900) and automatically adjusts the vehicle's speed to maintain a safe distance from vehicles ahead. Lateral ACC maintains the distance and advises the vehicle (900) to change lanes if necessary. Lateral ACC is related to other ADAS applications, such as LCA and CWS.
[0169] CACC uses information from other vehicles, which can be received via the network interface 924 and / or the wireless antenna(s) 926 from other vehicles either wirelessly or indirectly via a network connection (e.g., the internet). Direct connections can be provided through a vehicle-to-vehicle (V2V) communication link, while indirect connections can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles directly in front of and in the same lane as vehicle 900), while the I2V communication concept provides information about traffic at a greater distance. CACC systems can incorporate one or both of the I2V and V2V information sources.Thanks to the information about the vehicles in front of the vehicle 900, the CACC can be more reliable and has the potential to improve traffic flow and reduce congestion on the road.
[0170] FCW systems are designed to alert the driver to a hazard, allowing the driver to take corrective action. FCW systems utilize a forward-facing camera and / or RADAR 960 sensors, coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component. FCW systems can provide a warning, for example, in the form of an audible signal, a visual warning, a vibration, and / or a rapid braking pulse.
[0171] AEB systems detect an impending forward collision with another vehicle or object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. AEB systems use one or more forward-facing cameras and / or one or more radar sensors coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it will typically first warn the driver to take corrective action to avoid the collision. If the driver does not take corrective action, the AEB system will automatically apply the brakes in an effort to avoid or at least mitigate the impact of the predicted collision. AEB systems may include techniques such as dynamic brake assist and / or anticipatory braking.
[0172] Lane Departure Warning (LDW) systems provide visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver if the vehicle crosses lane markings. An LDW system will not activate if the driver indicates an intention to leave the lane, for example, by activating a turn signal. LDW systems may use forward-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0173] LKA systems are a variant of LDW systems. LKA systems provide steering or braking input to correct vehicle 900 if it begins to leave its lane.
[0174] Blind Spot Warning (BSW) systems detect vehicles in a car's blind spot and warn the driver. BSW systems can provide a visual, audible, and / or tactile alert to indicate that merging or changing lanes is unsafe. The system can provide an additional warning if the driver uses a turn signal. BSW systems can use rear-facing camera(s) and / or radar sensor(s) coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0175] RCTW systems can provide visual, audible, and / or tactile alerts when an object outside the reversing camera's field of view is detected while the vehicle is reversing. Some RCTW systems include AEB to ensure the vehicle's brakes are applied to avoid a collision. RCTW systems can use one or more rear-facing radar sensors coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0176] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for a driver, but are generally not catastrophic because the ADAS systems warn the driver and allow them to decide whether a safety condition actually exists and act accordingly. In an autonomous vehicle 900, however, the vehicle 900 itself must decide, in the case of conflicting results, whether the result should be considered by a primary computer or a secondary computer (e.g., a first controller 936 or a second controller 936). For example, the ADAS system 938 in some embodiments may be a backup and / or secondary computer that provides perceptual information to a rationality module of the backup computer. The rationality monitor of the backup computer may run redundant software on hardware components to detect perceptual errors and dynamic driving tasks.Output from the ADAS system 938 can be provided to a higher-level MCU. If output from the primary and secondary computers conflict, the higher-level MCU must decide how to resolve the conflict to ensure safe operation.
[0177] In some examples, the primary computer can be configured to provide the higher-level MCU with a confidence score indicating its level of trust in the chosen result. If the confidence score exceeds a certain threshold, the higher-level MCU can follow the primary computer's guidance, regardless of whether the secondary computer provides a conflicting or inconsistent result. If the confidence score does not reach a certain threshold, and if the primary and secondary computers display different results (e.g., a conflict), the higher-level MCU can arbitrate between the computers to determine the appropriate result.
[0178] The higher-level MCU can be configured to run a neural network(s) trained and configured to determine, based on output from the primary and secondary computers, the conditions under which the secondary computer will generate false alarms. Thus, the neural network(s) in the higher-level MCU can learn when the secondary computer's output can be trusted and when it cannot. For example, if the secondary computer is a radar-based FCW system, the neural network(s) in the higher-level MCU can learn when the FCW system identifies metallic objects that are not actually hazards, such as a drainage grate or a manhole cover, triggering an alarm.Similarly, if the secondary computer is a camera-based lane departure warning (LDW) system, a neural network in the higher-level MCU can learn to override the LDW when cyclists or pedestrians are present and leaving the lane is indeed the safest maneuver. In embodiments that include one or more neural networks running on the higher-level MCU, the higher-level MCU can include at least one DLA or GPU suitable for running the neural network(s) with associated memory. In preferred embodiments, the higher-level MCU can include a component of the SoC(s) 904 and / or be included as such.
[0179] In other examples, the ADAS System 938 can include a secondary computer that performs ADAS functionality using traditional computer vision rules. As such, the secondary computer can use classic computer vision rules (if-then), and the presence of one or more neural networks in the higher-level MCU can improve reliability, safety, and performance. For example, the overall system becomes more fault-tolerant due to the different implementation and intentional non-identity, particularly with regard to errors caused by software functionality (or the software-hardware interface).For example, if there is a software bug or error in the software running on the primary computer, and the non-identical software code running on the secondary computer provides the same overall result, the higher-level MCU can have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer does not cause a significant error.
[0180] In some examples, the output of the ADAS system 938 can be fed into the perception block of the primary computer and / or into the dynamic driving task block of the primary computer. For example, if the ADAS system 938 displays a forward collision warning due to an object immediately in front of it, the perception block can use this information in object identification. In other examples, the secondary computer may have its own neural network that is trained, thus reducing the risk of false positives, as described herein.
[0181] The Vehicle 900 may further include the Infotainment SoC 930 (e.g., an In-Vehicle Infotainment System (IVI)). Although illustrated and described as an SoC, the Infotainment System may not be an SoC and may include two or more discrete components. The Infotainment SoC 930 may include a combination of hardware and software that can be used to provide the Vehicle 900 with audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, reversing parking assistance, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fluid level, oil level, door open / closed status, air filter information, etc.). The Infotainment SoC 930 may, for example,This includes radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, WiFi, steering wheel audio controls, hands-free voice control, a heads-up display (HUD), an HMI display 934, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. The Infotainment SoC 930 can also be used to provide information (e.g., visual and / or audible) to the vehicle's user(s), such as information from the ADAS system 938, autonomous driving information like planned vehicle maneuvers, trajectories, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0182] The Infotainment SoC 930 can include GPU functionality. The Infotainment SoC 930 can communicate with other devices, systems, and / or components of the vehicle 900 via bus 902 (e.g., CAN bus, Ethernet, etc.). In some examples, the Infotainment SoC 930 can be coupled with a higher-level MCU so that the infotainment system's GPU can perform some self-regulating functions if the primary controller(s) 936 (e.g., the primary and / or backup computer of the vehicle 900) fails. In such an example, the Infotainment SoC 930 can put the vehicle 900 into a chauffeur-to-safe-stop mode, as described herein.
[0183] The vehicle 900 may further include an instrument cluster 932 (e.g., a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 932 may include a controller and / or a supercomputer (e.g., a discrete controller or a discrete supercomputer). The instrument cluster 932 may include a set of instruments such as a speedometer, fuel gauge, oil pressure gauge, tachometer, odometer, turn signals, gear shift indicator, seatbelt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between the infotainment SoC 930 and the instrument cluster 932.In other words, the Instrument Cluster 932 can be included as part of the Infotainment SoC 930, or vice versa.
[0184] Fig. 9D is a system representation in accordance with some embodiments of the present disclosure for communication between one or more cloud-based server(s) and the exemplary autonomous vehicle 900 of Fig. 9A. The System 976 can include one or more Servers 978, one or more Network(s) 990, and vehicles, including the Vehicle 900. The Server(s) 978 can include a variety of GPUs 984(A)-984(H) (collectively referred to hereafter as GPUs 984), PCIe Switches 982(A)-982(H) (collectively referred to hereafter as PCIe Switches 982), and / or CPUs 980(A)-980(B) (collectively referred to hereafter as CPUs 980). The GPUs 984, CPUs 980, and PCIe Switches can be interconnected using high-speed links, such as, but not limited to, NVIDIA's NVLink 988 interfaces and / or PCIe 986 links. In some examples, the GPUs 984 are connected via an NVLink and / or NVSwitch SoC, and the GPUs 984 and PCIe switches 982 are connected via PCIe links. Although eight GPUs 984, two CPUs 980, and two PCIe switches are illustrated, this is not intended to be limiting.Depending on the configuration, each Server 978 can include any number of GPUs 984, CPUs 980, and / or PCIe switches. For example, the Server 978 can include eight, sixteen, thirty-two, and / or more GPUs 984.
[0185] The server(s) 978 can receive image data from the network(s) 990 and the vehicles, representing images showing unexpected or changed road conditions, such as recently started roadworks. The server(s) 978 can transmit neural networks 992, updated neural networks 992, and / or map information 994, including information regarding traffic and road conditions, to the vehicles via the network(s) 990 and the vehicles. Updates to the map information 994 may include updates to the HD map 922, such as information about construction sites, potholes, detours, floods, and / or other obstacles.In some examples, the neural networks 992, the updated neural networks 992 and / or the map information 994 may result from new training and / or new experiences represented in data received from any number of vehicles in the vicinity, and / or based on training performed at a data center (e.g. using server(s) 978 and / or other servers).
[0186] Server 978 can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by the vehicles and / or in a simulation (e.g., using a game engine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or preprocessed, while in other examples, the training data is untagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training can be performed according to one or more classes of machine learning techniques, including, but not limited to, classes such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representational learning (including sparse dictionary learning), rule-based machine learning, anomaly detection, and any other variants or combinations thereof. After the machine learning models have been trained, they can be used by the vehicles (e.g., transmitted to the vehicles via network(s) 990) and / or the machine learning models can be used by server(s) 978 for remote monitoring of the vehicles.
[0187] In other examples, the Server 978 can receive data from the vehicles and apply the data to current real-time neural networks for intelligent real-time inference. The Server 978 can include supercomputers for deep learning and / or dedicated AI computers powered by GPU(s) 984, such as NVIDIA's DGX and DGX Station machines. However, in some examples, the Server 978 can also include deep learning infrastructure that uses only CPU-powered data centers.
[0188] The deep learning infrastructure of the server(s) 978 can be capable of fast real-time inference and use this capability to assess and verify the state of the processors, software, and / or associated hardware in the vehicle 900. For example, the deep learning infrastructure can receive periodic updates from the vehicle 900, such as a sequence of images and / or objects that the vehicle 900 has located within that sequence of images (e.g., through computer vision and / or other machine learning object classification techniques).The deep learning infrastructure can operate its own neural network to identify the objects and compare them with the objects identified by the vehicle 900, and if the results do not match and the infrastructure concludes that the AI in the vehicle 900 is faulty, the server 978 can send a signal to the vehicle 900 and instruct a fail-safe computer of the vehicle 900 to take over control, notify the passengers and perform a safe parking maneuver.
[0189] The Server 978 can include GPU(s) 984 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-powered servers and inference acceleration can enable real-time responsiveness. In other examples, such as when performance is less critical, servers powered by CPUs, FPGAs, and other processors can be used for inference purposes. EXAMPLE CALCULATION DEVICE
[0190] Fig. 10 a block representation of an exemplary computing device(s) 1000 suitable for use in implementing some embodiments of the present disclosure. The computing device 1000 can include a connection system 1002 that directly or indirectly couples the following devices: memory 1004, one or more central processing units (CPUs) 1006, one or more graphics processing units (GPUs) 1008, a communication interface 1010, input / output (I / O) ports 1012, input / output components 1014, a power supply 1016, one or more presentation components 1018 (e.g., display(s)), and one or more logic units 1020. In at least one embodiment, the computing device(s) 1000 can include one or more virtual machines (VMs) and / or any of its components can include virtual components (e.g., virtual hardware components).For non-restrictive examples, one or more of the GPUs 1008 may comprise one or more vGPUs, one or more of the CPUs 1006 may comprise one or more vCPUs, and / or one or more of the logic units 1020 may comprise one or more virtual logic units. As such, a computing device (or devices) 1000 may include discrete components (e.g., a complete GPU dedicated to computing device 1000), virtual components (e.g., a portion of a GPU dedicated to computing device 1000), or a combination thereof.
[0191] Although the various blocks of Fig. Where components 10 are shown as connected via the connection system 1002 by lines, this is not intended to be restrictive and serves only for clarity. For example, in some embodiments, a presentation component 1018, such as a display device, can be considered an I / O component 1014 (e.g., if the display is a touchscreen). As another example, the CPUs 1006 and / or GPUs 1008 can include memory (e.g., the memory 1004 can be representative of a storage device, in addition to the memory of the GPUs 1008, the CPUs 1006, and / or other components). In other words, the computing device of Fig. Figure 10 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all fall under the category of computing device. Fig. 10 will be considered.
[0192] The connection system 1002 can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The connection system 1002 can include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standard Association (VESA) bus, a Peripheral Component Connection (PCI) bus, a Peripheral Component Connection Express (PCIe) bus, and / or other bus or link types. In some embodiments, there are direct connections between components. For example, the CPU 1006 can be directly connected to the memory 1004. Furthermore, the CPU 1006 can be directly connected to the GPU 1008. In the case of a direct or point-to-point connection between components, the connection system 1002 can include a PCIe link to implement the connection.In these examples, the computing device 1000 does not need to include a PCI bus.
[0193] Memory 1004 can include any of a range of computer-readable media. The computer-readable media can be any available media accessible to the computing device 1000. The computer-readable media can include both volatile and non-volatile media, and both removable and non-removable media. As an example, and not as a limitation, the computer-readable media can include computer storage media and communication media.
[0194] Computer storage media can include both volatile and non-volatile media and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory can store 1004 computer-readable instructions (e.g., representing a program and / or program element, such as an operating system).Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, Digital Versatile Discs (DVDs) or other optical disk storage, magnetic cartridges, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 1000. As used herein, computer storage media do not, per se, include signals.
[0195] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include information delivery media. The term "modulated data signal" can refer to a signal for which one or more of its characteristics are set or modified in a way that encodes information in the signal. By way of example, and not limited to this, computer storage media can include wired media, such as a wired network or a directly wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of any of the foregoing should also be included in the scope of computer-readable media.
[0196] The CPU(s) 1006 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 in order to perform one or more of the procedures and / or processes described herein. The CPU(s) 1006 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling a plurality of software threads concurrently. The CPU(s) 1006 can include any type of processor and can include different types of processors depending on the type of computing device 1000 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).For example, depending on the type of computing device 1000, the processor can be an Advanced RISC Machines (ARM) processor implemented using Reduced Instruction Set Computing (RISC), or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 1000 can include one or more CPUs 1006 in addition to one or more microprocessors or supplementary coprocessors, such as math coprocessors.
[0197] In addition to or as an alternative to the CPU(s) 1006, the GPU(s) 1008 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the procedures and / or processes described herein. One or more of the GPU(s) 1008 may be an integrated GPU (e.g., in one or more of the CPU(s) 1006) and / or one or more of the GPU(s) 1008 may be a discrete GPU. In embodiments, one or more of the GPU(s) 1008 may be a coprocessor of one or more of the CPU(s) 1006. The GPU(s) 1008 may be used by the computing device 1000 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, the GPU(s) 1008 can be used for general-purpose computing on GPUs (GPGPU).The GPU(s) 1008 can include hundreds or thousands of cores capable of handling hundreds or thousands of software threads simultaneously. The GPU(s) 1008 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 1006 received via a host interface). The GPU(s) 1008 can include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory can be included as part of the memory 1004. The GPU(s) 1008 can include two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using an NVLINK) or it can connect the GPUs via a switch (e.g., using an NVSwitch).When combined, each GPU can generate 1008 pixel data or GPGPU data for different dividers, or output for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or share memory with other GPUs.
[0198] In addition to or alternatively to the CPU(s) 1006 and / or the GPU(s) 1008, the logic unit(s) 1020 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 1006, the GPU(s) 1008, and / or the logic unit(s) 1020 may discretely or jointly perform any combination of the methods, processes, and / or parts thereof. One or more of the logic units 1020 may be part of and / or integrated into one or more of the CPU(s) 1006 and / or the GPU(s) 1008, and / or one or more of the logic units 1020 may be discrete components or otherwise external to the CPU(s) 1006 and / or GPU(s). the GPU(s) 1008.In embodiments, one or more of the logic units 1020 can be a coprocessor of one or more of the CPU(s) 1006 and / or one of the GPU(s) 1008.
[0199] Examples of logic unit(s) 1020 include: one or more processing cores and / or components thereof, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel visual cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating-point units (FPUs), and input / output (I / O) elements. Peripheral Component Connector (PCI) or Peripheral Component Connector Express (PCIe) components and / or the like.
[0200] The communication interface 1010 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 1000 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 1010 can include components and functionality to enable communication over any number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating via Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the logic unit(s) 1020 and / or communication interface 1010 may include one or more data processing units (DPUs) to transfer data received via a network and / or the connection system 1002 directly to (e.g., a memory of) one or more GPU(s) 1008.
[0201] The I / O ports 1012 allow the computing device 1000 to be logically coupled with other devices, including the I / O components 1014, the presentation component(s) 1018, and / or other components, some of which may be built into (e.g., integrated with) the computing device 1000. Illustrative I / O components 1014 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite table, scanner, printer, wireless device, etc. The I / O components 1014 can provide a natural user interface (NUI) that processes air gestures, speech, or other physical input generated by a user. In some cases, input can be transmitted to an appropriate network element for further processing.A NUI can implement any combination of speech recognition, pen recognition, facial recognition, biometric recognition, gesture recognition both on and adjacent to the screen, air gestures, head and eye tracking, and touch recognition (as described in more detail below), associated with a display of the Computing Device 1000. The Computing Device 1000 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. Additionally, the Computing Device 1000 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit, IMU) that enable motion detection.In some examples, the output of the accelerometers or gyroscopes can be used by the computing device 1000 to render immersive augmented reality or virtual reality.
[0202] The power supply 1016 can include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 1016 can provide power to the computing device 1000 to enable the components of the computing device 1000 to operate.
[0203] The presentation component(s) 1018 can include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 1018 can receive data from other components (e.g., the GPU(s) 1008, the CPU(s) 1006, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER
[0204] Fig. Figure 11 illustrates an exemplary data center 1100 that can be used in at least one embodiment of the present disclosure. The data center 1100 can include a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and / or an application layer 1140.
[0205] As in Fig. As shown in Figure 11, the data center infrastructure layer 1110 can include a resource orchestrator 1112, clustered compute resources 1114, and node compute resources (“node RR”) 1116(1)-1116(N), where “N” represents an integer, positive number. In at least one embodiment, Node-RR 1116(1)-1116(N) can include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), storage devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules and / or cooling modules, etc. In some embodiments, one or more Node-RRs can be controlled by the Node-RR1116(1)-1116(N) correspond to a server that has one or more of the computing resources mentioned above. Furthermore, in some embodiments, the node RR 1116(1)-1116(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or may correspond to one or more of the node RR 1116(1)-1116(N) of a virtual machine (VM).
[0206] In at least one embodiment, grouped compute resources 1114 can include separate groupings of node RR 1116 located in one or more racks (not shown), or many racks located in data centers at different geographic locations (also not shown). Separate groupings of node RR 1116 within grouped compute resources 1114 can include grouped compute, network, storage, or memory resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node RR 1116, including CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide compute resources to support one or more workloads.The one or more racks can also include any number of power modules, cooling modules and / or network switches in any combination.
[0207] The resource orchestrator 1112 can configure or otherwise control one or more node RR 1116(1)-1116(N) and / or grouped compute resources 1114. In at least one embodiment, the resource orchestrator 1112 can include a software design infrastructure (SDI) management entity for the data center 1100. The resource orchestrator 1112 can include hardware, software, or a combination thereof.
[0208] In at least one embodiment, as in Fig. As shown in Figure 11, a framework layer 1120 can include a job scheduler 1133, a configuration manager 1134, a resource manager 1136, and / or a distributed file system 1138. The framework layer 1120 can include a framework to support software 1132 of software layer 1130 and / or one or more application(s) 1142 of application layer 1140. The software 1132 or application(s) 1142 can each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1120 can be, but is not limited to, a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can use the distributed file system 1138 for large-volume data processing (e.g. "Big Data").In at least one embodiment, the job scheduler 1133 can include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 1100. The configuration manager 1134 can be capable of configuring various layers, such as the software layer 1130 and the framework layer 1120, including Spark and the distributed file system 1138, to support high-volume data processing. The resource manager 1136 can be capable of managing clustered or grouped compute resources allocated or assigned to support the distributed file system 1138 and the job scheduler 1133. In at least one embodiment, the clustered or grouped compute resources can include a grouped compute resource 1114 at the data center infrastructure layer 1110.The resource manager 1136 can coordinate with the resource orchestrator 1112 to manage these allocated or assigned computing resources.
[0209] In at least one embodiment, software 1132, which is enclosed in software layer 1130, can include software used by at least parts of the node RR 1116(1)-1116(N), grouped computing resources 1114, and / or distributed file system 1138 of framework layer 1120. One or more types of software can include, but are not limited to, internet web page search software, email virus scanning software, database software, and streaming video content software.
[0210] In at least one embodiment, one or more application(s) 1142 enclosed in the application layer 1140 may include one or more types of applications used by at least parts of the node RR 1116(1)-1116(N), grouped compute resources 1114, and / or distributed file system 1138 of the framework layer 1120. One or more types of applications may include, but are not limited to, any number of a genomics application, a cognitive computation application, and a machine learning application, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0211] In at least one embodiment, any configuration manager 1134, resource manager 1136, and resource orchestrator 1112 can implement any number and any type of self-modifying operations based on any set and any type of data captured in any technically feasible manner. Self-modifying operations can relieve a data center operator of data center 1100 of the burden of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly performing parts of a data center.
[0212] Data Center 1100 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weight parameters according to a neural network architecture using software and / or computing resources as described above in relation to Data Center 1100.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer and predict information regarding Data Center 1100 using the resources described above, by using weight parameters calculated via one or more training techniques, such as, but not limited to, those described herein.
[0213] In at least one embodiment, the data center can use 1100 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Additionally, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS
[0214] Network environments suitable for use in implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be connected to one or more instances of the computing device(s). Fig. 10. Each device can include similar components, features, and / or functionality to the computing device(s) 1000. Furthermore, if backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices can be included as part of a data center 1100; an example of this is described in more detail below with reference to… Fig. 11 described.
[0215] Components of a network environment can communicate with each other over a network, which can be wired, wireless, or both. The network can include multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the internet, and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0216] Compatible network environments can include one or more peer-to-peer network environments—in which case no server may be included in a network environment—and one or more client-server network environments—in which case one or more servers may be included in a network environment. In peer-to-peer network environments, functionality described herein with respect to one or more servers can be implemented on any number of client devices.
[0217] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework to support software of a software layer and / or one or more application(s) of an application layer. The software or application(s) can each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework, but is not limited to those that can use a distributed file system for large-volume data processing (e.g., "big data").
[0218] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination of computing and / or data storage functions as described herein (or one or more parts thereof). Any of these various functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers that may be distributed across a state, region, country, the Earth, etc.). If a connection to a user (e.g., a client device) is relatively close to an edge server, a core server may designate at least some of the functionality for that edge server(s). A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0219] The client device(s) may include at least some of the components, features, and functionality described herein with reference to Fig.The 10 exemplary computing device(s) described include 1000. By way of example, and not limited to, a client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or global positioning device, video player, video camera, surveillance device or system, vehicle, boat, flying vehicle, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, device, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.
[0220] The disclosure can generally be described in the context of computer code or machine-usable instructions, including computer-executable instructions such as program modules that are executed by a computer or other machine, such as a personal data assistant or a handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements specific abstract data types. The disclosure can be exercised in various system configurations, including handheld devices, consumer electronics, general-purpose computers, other specialized computing devices, etc. The disclosure can also be exercised in distributed computing environments where tasks are performed by remote processing devices linked via a communication network.
[0221] As used herein, any mention of "and / or" in relation to two or more elements shall be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0222] The subject matter of this disclosure is described herein with specificity to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have provided that the claimed subject matter may also be embodied in other ways to include different steps or combinations of steps similar to those described in this document, in conjunction with other current or future technologies. Furthermore, although the terms "step" and / or "block" may be used herein to denote different elements of methods employed, they should not be interpreted as implying a particular sequence of or between different steps disclosed herein, except where the sequence of individual steps is explicitly described. EXAMPLE PARAGRAPHS A: A method comprising: obtaining image data representative of at least one image illustrating one or more line markings associated with a road marking located within an environment; generating, using one or more machine learning models and based on the image data, output data specifying at least one or more first points associated with one or more first edges of the one or more line markings and one or more second points associated with one or more second edges of the one or more line markings, the one or more second edges being opposite the one or more first edges; comparing at least one of the one or more first edges or the one or more second edges with one or more edges encoded in a map;Performing a longitudinal localization of a machine with reference to the map, at least based on comparison; and performing one or more operations, at least based on the longitudinal localization. B: Method according to paragraph A, wherein the output data further specify one or more first direction information associated with the one or more first points and one or more second direction information associated with the one or more second points, and wherein the longitudinal localization is further based on at least one of the one or more first direction information or one or more second direction information. C: Method according to paragraph B, wherein: the one or more first directions include one or more first vectors directed from the one or more first points to one or more midpoints of the one or more line markings; and the one or more second directions include one or more second vectors directed from the one or more second points to one or more midpoints of the one or more line markings. D: Method according to any one of paragraphs AC, wherein: the one or more line markings include at least one first line marking associated with the road marking and a second line marking associated with the road marking; and at least one of the one or more first points and at least one of the one or more second points are associated with the first line marking; and at least one of the one or more first points and at least one of the one or more second points are associated with the second line marking. E: Method according to any of paragraphs AD, wherein at least one of: the one or more first points and the one or more second points are associated with first coordinate locations in a first coordinate direction associated with the image and second coordinate locations in a second coordinate direction associated with the image; or the one or more first points and the one or more second points are associated with distances and angles with respect to one or more reference points within the image. F: Method according to any of paragraphs AE, wherein: the output data represent a plurality of pixel locations associated with the image and a plurality of probabilities associated with the plurality of pixel locations; and the method further comprises: determining that at least a subset of the plurality of pixel locations is associated with at least a subset of the plurality of probabilities satisfying a threshold probability; and determining that the one or more first points and the one or more second points are located at the at least subset of the plurality of pixel locations. G: Method according to any of paragraphs AF, wherein: the output data represent points nearest to the pixels within the image; and the method further comprises determining that the one or more first points include a first part of the nearest points and the one or more second points include a second part of the nearest points. H: Method according to one of paragraphs AG, wherein, prior to deployment, the one or more machine learning models are evaluated within a simulation environment by processing at least simulated sensor data corresponding to virtual line markings. I: A system comprising: one or more processors for: obtaining image data representative of at least one or more images illustrating one or more line segments associated with one or more traffic features in an environment; determining, using one or more machine learning models and at least based on the image data, one or more points associated with one or more line segments and one or more direction information associated with the one or more points; and performing one or more operations at least based on the one or more points and the one or more direction information. J: System according to paragraph I, wherein: the one or more points include at least one or more first points associated with one or more first edges of the one or more line segments and one or more second points associated with one or more second edges of the one or more line segments; and the one or more direction indicators include at least one or more first direction indicators associated with the one or more first points and one or more second direction indicators associated with the one or more second points. K: System according to paragraph J, wherein: the one or more first directions include one or more first vectors starting at the one or more first points and directed towards one or more midpoints of the one or more line segments; and the one or more second directions include one or more second vectors starting at the one or more second points and directed towards one or more midpoints of the one or more line segments. L: System according to one of paragraphs IK, wherein: the one or more points are located at approximately one or more midpoints of the one or more line segments; and the one or more line direction indications start at the one or more points and are directed towards one or more edges of the one or more line segments. M: System according to paragraph L, wherein the one or more directions include: one or more first vectors starting at the one or more points and directed towards one or more first edges of the one or more edges; and one or more second vectors starting at the one or more points and directed towards one or more second edges of the one or more edges, the one or more second edges being opposite the one or more first edges. N: System according to one of paragraphs IM, wherein: one or more sections of the one or more line segments are occluded by one or more objects represented by the one or more images; and the one or more machine learning models fail to determine one or more second points associated with the one or more sections of the one or more line segments that are occluded. O: System according to one of paragraphs IN, wherein: the one or more traffic features include at least one road marking; the one or more line segments include at least one first line marking and one second line marking associated with the road marking; the one or more points include at least one first point associated with the first line marking and one second point associated with the second line marking; and the one or more directional indicators include at least one first directional indicator associated with the first point and one second directional indicator associated with the second point. P: System according to one of paragraphs IO, wherein: the one or more points are associated with one or more first coordinate locations in one or more first coordinate directions associated with the one or more images and one or more second coordinate locations in one or more second coordinate directions associated with the one or more images; and the one or more direction specifications are associated with one or more first values in the first coordinate direction and one or more second values in the second coordinate direction. Q: System according to any of paragraphs IP, wherein determining the one or more points comprises: generating, using the one or more machine learning models and at least based on the image data, an output that specifies one or more first coordinate locations associated with one or more pixels in a first coordinate direction and one or more second coordinate locations associated with the one or more pixels in a second coordinate direction; and determining the one or more points at least based on the one or more first coordinate locations and the one or more second coordinate locations. R: System according to any of paragraphs IQ, wherein the system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twinning operations; a system for performing a light transport simulation; a system for performing collaborative content creation for 3D assets; a system that provides one or more cloud gaming applications; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations;a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more vision language models (VLMs); a system for performing operations using one or more multimodal language models; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; systems that implement one or more multimodal language models; systems that use or employ one or more inference microservices; systems that incorporate one or more machine learning models in a service or microservice along with an OS-level virtualization package (e.g., a container);a system that involves one or more virtual machines (VMs); a system that is implemented at least partially in a data center; or a system that is implemented at least partially using cloud computing resources. S: One or more processors comprising: processing switching technology for performing longitudinal localization of a machine at least based on information associated with one or more line markings of one or more road markings within an environment, wherein the information is determined at least based on processing by one or more machine learning models of sensor data representative of the one or more road markings and includes at least one or more points associated with the one or more line markings and one or more directional information associated with the one or more points. T: The one or more processors according to paragraph S, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twinning operations; a system for performing a light transport simulation; a system for performing collaborative content creation for 3D assets; a system that provides one or more cloud gaming applications; a system for performing one or more deep learning operations; a system that is implemented using an edge device; a system that is implemented using a robot; a system for performing one or more generative AI operations;a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more vision language models (VLMs); a system for performing operations using one or more multimodal language models; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; systems that implement one or more multimodal language models; systems that use or employ one or more inference microservices; systems that incorporate one or more machine learning models in a service or microservice along with an OS-level virtualization package (e.g., a container);a system that involves one or more virtual machines (VMs); a system that is implemented at least partially in a data center; or a system that is implemented at least partially using cloud computing resources.
[0223] It is understood that the aspects and embodiments described above are merely examples and that modifications in detail may be made within the scope of the claims.
[0224] Each device, method and feature disclosed in the description and (if applicable) in the claims and drawings can be provided independently or in any suitable combination.
[0225] Reference numerals appearing in the claims serve only for illustration and are not intended to have any limiting effect on the scope of the claims. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 16 / 101,232
[0125] Cited non-patent literature
[0000] SAE) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard No. J3016-201806, published on June 16, 2018, Standard No. J3016-201609, published on September 30, 2016
[0083]
Citation Information
Patent Citations
US-PATENTANMELDUNGNR.16/101,232