PURSUITING A MULTI-DIMENSIONAL PATHGEOMETRY FOR AUTONOMOUS SYSTEMS AND APPLICATIONS
By employing Kalman filters to track and predict Bézier curves, the method addresses the inefficiencies of conventional path determination systems, enabling efficient 3D lane tracking for autonomous vehicles with reduced computational demands.
Patent Information
- Application Number
- DE102025133475
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-28
- Filing Date
- 2025-08-21
- Publication Date
- 2026-03-05
AI Technical Summary
Conventional systems for determining drivable paths for autonomous vehicles are suboptimal due to their reliance on computationally intensive data processing and inefficient methods like deep learning or particle filtering, which are not suitable for real-time applications in complex 3D environments.
The use of recursive models, specifically Kalman filters, to track and predict Bézier curves representing lane geometries, allowing for efficient 2D and 3D path geometry tracking with minimal input data, optimizing computational efficiency and accuracy.
The proposed method enables efficient and accurate 3D lane tracking with reduced computational resources, making it suitable for autonomous driving by relying on minimal input data and improving processing efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Designing a system to drive a vehicle or machine autonomously and without supervision at a level of safety required for practical acceptance is extremely challenging. For example, an autonomous or semi-autonomous vehicle should at least be capable of acting as the functional equivalent of an attentive driver, relying on a perception and action system with an incredible ability to identify and react to dynamic and static obstacles in a complex environment to avoid collisions with other objects or structures along its path. Therefore, determining navigable paths within complex environments that an autonomous vehicle might encounter is one of the most critical tasks an autonomous system must perform.
[0002] However, conventional systems for determining drivable paths for autonomous vehicles within an environment may be suboptimal due to the complexities associated with predicting and tracking these paths. For example, some conventional systems are only capable of detecting or tracking drivable paths in two-dimensional (2D) space, which is not ideal since autonomous systems must operate in real-world, three-dimensional (3D) environments. Furthermore, some conventional systems may rely on different types of data to make accurate path predictions. For instance, several different modalities, such as LiDAR data, radar data, and / or other types of sensor data, may be required to accurately predict or track paths within an environment.For this reason, conventional systems can be computationally intensive and / or inefficient, as they have to process such enormous data inputs.
[0003] Furthermore, some conventional systems can use deep learning or other techniques to generate a polyline representing a path for a vehicle. However, generating the polyline requires determining a large number of points—hundreds and / or even thousands—using deep learning. As a result, polyline representations can be inefficient to generate and, in some cases, over-parameterized and / or over-complex compared to standard road design, or require a large number of network parameters, which can reduce generalizability. In some cases, other conventional systems can use particle filtering to track drivable paths in an environment. However, particle filtering can be inefficient and unsuitable for autonomous driving and / or other real-time applications. BRIEF SUMMARY OF THE INVENTION
[0004] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not fall within the scope of protection of the claims are described herein.
[0005] Several examples are disclosed where geometries associated with one or more paths in an environment can be efficiently tracked and / or predicted using recursive models. For example, the disclosed systems and methods can use Kalman filters to track and predict control points corresponding to Bézier curves (e.g., 2D and / or 3D Bézier curves). The Bézier curves can represent geometries associated with one or more lanes of a road surface. In some cases, multiple Bézier curves can be used to represent a lane geometry, and multiple Kalman filters can be used to track and predict control points for each Bézier curve.For example, an edge of the lane can be represented using a first Bezier curve, and control points for the first Bezier curve can be tracked and predicted using multiple Kalman filters (e.g., for a 3D Bezier curve, one Kalman filter for each x, y, or z dimension).
[0006] Embodiments of the present disclosure relate to the tracking of a multidimensional path geometry for autonomous or semi-autonomous systems and applications. For example, the systems and methods described herein can use recursive models to efficiently track and / or predict geometries associated with paths in an environment. For example, the disclosed systems and methods can use Kalman filters and / or other recursive models to track and predict control points corresponding to Bézier curves (e.g., 2D and / or 3D Bézier curves), which can represent geometries associated with one or more lanes of a road surface.In some cases, the systems and methods of this disclosure can use multiple Bézier curves to represent the geometry of a lane, and they can also use multiple Kalman filters to track and predict control points for each Bézier curve. For example, an edge or centerline of the lane can be represented using a first Bézier curve, and control point coordinates for the first Bézier curve can be tracked and predicted using multiple Kalman filters (e.g., for a 3D Bézier curve, one Kalman filter for each x, y, or z dimension / coordinate). In some examples, the systems can compute a curve-shift matrix to shift previous states of the Kalman filters to an Ego Machine origin, and then predict new states of the Kalman filters by multiplying the previous states by the curve-shift matrix.Thus, the systems can implement a closed-form solution to apply the curve shift matrix and generate predictions (in the Kamal filtering processes) using the states from previous timestamps and the relative ego movement between the two timestamps.
[0007] In contrast to conventional systems, the systems of this disclosure, in some embodiments, are able to efficiently and accurately represent the geometry of the traversable path in 2D and / or 3D space using Bézier representations, and to further optimize the techniques by being able to track each dimension of a given Bézier representation independently. For example, by using Kalman filtering with efficient state representations and closed-form motion models, the systems of this disclosure are able to provide more optimized path geometry tracking than conventional systems, such as those described above, which can use particle filtering.Furthermore, in some embodiments, the systems of this disclosure are capable of performing 3D lane tracking, unlike conventional systems, while relying on minimal input data, such as image-based lane recognition data. By being able to rely on minimal input data and still perform 3D lane tracking, the systems of this disclosure may be more suitable for autonomous driving and other real-world applications, while simultaneously improving computational efficiency by processing less input data and other data.
[0008] The revelation extends to all novel aspects or features described and / or illustrated here.
[0009] Further features of the disclosure are characterized by the independent and dependent claims.
[0010] Any feature of one aspect of the disclosure can be applied in any suitable combination to other aspects of the disclosure. In particular, procedural aspects can be applied to apparatus or system aspects, and vice versa.
[0011] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features here should be interpreted accordingly.
[0012] Each system or device feature described here can also be provided as a process feature, and vice versa. System and / or device aspects that are functionally described (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated working memory.
[0013] It is also understood that certain combinations of the various features described and defined in each aspect of the revelation can be implemented and / or provided and / or used independently of one another.
[0014] The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods described herein when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all component steps of a method.
[0015] The disclosure also includes a computer or computer system (including networked or distributed systems) with an operating system that supports a computer program for carrying out the procedures described herein and / or for embodying the device or system features described herein.
[0016] The disclosure also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0017] The revelation also provides a signal that carries one or more of the aforementioned computer programs.
[0018] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0019] Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The systems and methods presented here for tracking a multidimensional path geometry for autonomous or semi-autonomous systems and applications are described in detail below with reference to the accompanying drawings. These show: Fig. 1 an exemplary data flow diagram for a process of tracing a multidimensional path geometry using Kalman filters, according to some embodiments of the present disclosure; Fig. 2 an exemplary image of drivable paths in an environment, according to some embodiments of the present disclosure; Fig. 3A an example of control points in a 3D space, according to some embodiments of the present disclosure; Fig. 3B is an example of a Bezier curve, which corresponds to the control points in the example of Fig. 3A corresponds, according to some embodiments of the present disclosure; Fig. 4. A data flow diagram illustrating an additional exemplary detail relating to certain operations of the example of Fig. 1 is assigned to the process described, according to some embodiments of the present disclosure; Fig. 5 An example of a system capable of performing one or more of the processes described herein, according to some embodiments of the present disclosure. Fig. 6 a flowchart illustrating an example of a method for tracing a multidimensional path geometry using Kalman filters, according to some embodiments of the present disclosure. Fig. 7 a flowchart illustrating an example of a method for using recursive models to predict the multidimensional path geometry according to some embodiments of the present disclosure. Fig. 8A an illustration of an exemplary autonomous vehicle, according to some embodiments of the present disclosure; Fig. 8B is an example of camera locations and fields of view for the exemplary autonomous vehicle from Fig. 8A, according to some embodiments of the present disclosure; Fig. 8C a block diagram of an exemplary system architecture for the exemplary autonomous vehicle from Fig. 8A, according to some embodiments of the present disclosure; Fig. 8D is a system diagram for the communication between one or more cloud-based servers and the example autonomous vehicle. Fig. 8A, according to some embodiments of the present disclosure; Fig. 9 a block diagram of an exemplary computing device suitable for use in the implementation of some embodiments of the present disclosure; and Fig. 10 a block diagram of an exemplary data center suitable for use in the implementation of some embodiments of the present disclosure. DETAILED DESCRIPTION
[0021] Systems and methods relating to the tracking of a three-dimensional (3D) path geometry for autonomous or semi-autonomous systems and applications are disclosed. Although the present disclosure relates to an exemplary autonomous or semi-autonomous vehicle or machine 800 (alternatively referred to herein as "Vehicle 800", "Ego-Vehicle 800", "Ego-Machine 800" or "Machine 800"), one example of which relates to Fig. The fact that the systems and procedures described here can be described (as described in sections 8A-8D) is not intended to be restrictive. For example, the systems and procedures described here can be used without restriction by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), steered and unsteered robots or robotic platforms, warehouse vehicles, all-terrain vehicles, vehicles coupled with one or more trailers, hydrofoils, boats, shuttles, emergency vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Although the present disclosure can be described in relation to predicting and tracking the geometry of a lane in autonomous or semi-autonomous systems and applications, this is not to be understood as a limitation beyond that, and the systems and methods described herein can also be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications and / or any other technology fields where prediction and tracking of the geometry of features is performed.
[0022] For example, one or more systems can receive sensor data generated using one or more sensors associated with a vehicle (e.g., a machine, an "ego vehicle," etc.), such as a semi-autonomous and / or autonomous vehicle. As described herein, in some examples, the sensor data may include image data generated using one or more image sensors, such as one or more of the vehicle's cameras, or it may include LiDAR data, radar data, ultrasonic data, and / or other data types generated using any number of sensor modalities.For example, when image data is used, the image data can be generated using one or more forward-facing image sensors on the vehicle, with the image data representing one or more images depicting an environment in front of the vehicle and / or in a direction in which the vehicle is essentially navigating. In some examples, the one or more systems can then process the sensor data using one or more processing techniques, which are described in more detail herein.
[0023] The one or more systems can then input the sensor data (and / or the processed sensor data) into one or more machine learning models configured to generate data associated with one or more paths (e.g., lanes, roads, etc.) along which the vehicle is to navigate. Based at least on the processing of the sensor data (and / or the processed sensor data), the one or more machine learning models can, for example, generate an initial output specifying points (e.g., Bézier points or control points) associated with one or more paths. As described herein, the one or more paths can be a primary path for the vehicle to navigate, one or more secondary paths adjacent to the primary path (e.g.,The output may contain one or more paths to the left of the first path, and / or one or more paths to the right of the first path, and / or one or more other paths, such as an exit path, a merge path, a lane split path, a contraflow path, and / or so forth, without limitation. Furthermore, the first output for a path may contain any number of points, such as two points, four points, eight points, twelve points, sixteen points, twenty points, fifty points, and / or any other number of points. As described herein, in some examples, such as...To reduce the computing resources required for generating one or more paths and / or to reduce the latency when generating one or more paths, the number of points for a path may be limited to a threshold number of points and / or a threshold number of points per distance of the path.
[0024] Furthermore, in some examples, based at least on the processing of the sensor data (and / or the processed sensor data), one or more machine learning models can generate a second output specifying one or more classifications associated with the points and / or the one or more paths. As described herein, a classification can include, but is not limited to, a vehicle or ego path (e.g., the path the vehicle is intended to navigate), an adjacent path, a left adjacent path, a right adjacent path, an edge path, a left-edge path, a right-edge path, a centerline path, a lane center path, an exit path, a merge path, a lane split path, a two-way traffic path, and / or any other type of path. In some examples, the second output can specify a separate classification for each point and / or path.In addition, or alternatively, in some examples the second output may specify one or more probabilities that are assigned to one or more classifications for each point and / or path.
[0025] The one or more systems (e.g., the one or more machine learning models, one or more post-processing components, etc.) can then use the initial output to determine one or more curves or other geometric representations associated with the one or more paths. For example, for a given path, the one or more systems can use one or more algorithms to generate a Bézier curve based on at least the points associated with the path. As described herein, in some examples, an algorithm may include a curve-fitting algorithm, such as…However, without restriction, a 2D Bézier curve-fitting algorithm, a 3D Bézier curve-fitting algorithm, a cubic Bézier curve-fitting algorithm, a higher-order Bézier curve-fitting algorithm, a splitting curve-fitting algorithm, any other type of Bézier algorithm, another type of curve-fitting algorithm, etc. Furthermore, or alternatively, in some examples, an algorithm may include another type of curve-fitting algorithm configured to generate a curve based on at least the points associated with the path. The one or more systems may then use similar processes to generate a respective curve (e.g., a Bézier curve) or other geometric representation associated with one or more (e.g., each) of the one or more other paths.
[0026] As described herein, the one or more systems can use one or more recursive models to track the curves and / or points output by the one or more machine learning models used by the one or more systems to determine the curves or other geometric representations. The recursive models can also be used to temporarily smooth the curves by making predictions based on previous states associated with the curves / points. For example, the one or more systems of this disclosure can use one or more Kalman filters to maintain (e.g., track, predict, etc.) various sets of points corresponding to the paths in the environment.Although many of the examples in the present disclosure are described in relation to the use of Kalman filters for tracking path geometry, these are only a few examples, and one or more of the systems in the present disclosure can track and predict path geometry using any other types of recursive models, algorithms, and techniques.
[0027] In some cases, the one or more systems can use multiple curves to represent the geometry of a given path. For example, the geometry of a left edge of a path can be represented using a first curve, the geometry of a right edge of the path can be represented using a second curve, the geometry of a center line or middle path of the lane can be represented using a third curve, and so on. Furthermore, since the one or more systems of the present disclosure can track and predict path geometry in a multidimensional space (e.g., 2D, 3D, etc.), the one or more systems can use multiple Kalman filters to maintain different state vectors for different dimensions of the curves.This means that for a single curve, the number of Kalman filters and state vectors used to maintain that curve can, in some examples, be equal to the number of dimensions (e.g., 2D, 3D, etc.) associated with the curve and / or the points.
[0028] As an example, control point locations in 3D space can be represented using three coordinate values, such as an x-coordinate, a y-coordinate, and a z-coordinate, and the geometry of a given curve can be defined using a number of 3D control points (e.g., 12 points per curve). Therefore, in the case of 3D points for a 3D curve, one or more systems can use three Kalman filters to maintain three distinct state vectors for each curve (e.g., one state vector for each x, y, and z coordinate dimension), and the state variables of the state vectors can correspond to the values of the 3D control points for that dimension. For example, for a first curve (of potentially multiple curves) corresponding to a first path (of potentially multiple paths), one or more systems can use a first Kalman filter to define first control point dimensions (e.g.,a second Kalman filter to maintain second control point dimensions (e.g., x-coordinate values) for the first curve (e.g., to track and predict), a second Kalman filter to maintain second control point dimensions (e.g., y-coordinate values) for the first curve, and a third Kalman filter to maintain third control point dimensions (e.g., z-coordinate values) for the first curve.
[0029] In some examples, one or more systems can track and predict geometries associated with multiple paths in an environment simultaneously. For example, one or more systems can maintain one or more first Kalman filters to track and predict the geometry for a first path in the environment, one or more second Kalman filters to track and predict the geometry for a second path in the environment, one or more third Kalman filters to track and predict the geometry for a third path in the environment, and so on.This means that one or more first Kalman filters can be used to track and maintain one or more first sets of 3D control points for one or more first Bézier curves representing the 3D geometry of the first path; one or more second Kalman filters can be used to track and maintain one or more second sets of 3D control points for one or more second Bézier curves representing the 3D geometry of the second path; one or more third Kalman filters can be used to track and maintain one or more third sets of 3D control points for one or more third Bézier curves representing the 3D geometry of the third path; and so on.
[0030] In some examples, the Kalman filters can include at least process models and measurement models. The process models (also called "prediction models") of the Kalman filters can be configured to predict current states (e.g., next states) of the Kalman filters based on at least previous states of the Kalman filters and the relative movement of a machine between those previous states and the current states. The measurement models of the Kalman filters can be configured to use various data sources to update or refine the predicted current states of the Kalman filters. For example, the measurement models can use sensor data, output from machine learning models, or any other information to update or refine the predicted state vectors.
[0031] As described herein, in some cases, the Kalman filter process models can use an optimized closed-form solution to generate predicted Bézier curve states using the relative motion of the machine. For example, if the machine is moving from a previous location to a current location, one or more systems can use the Kalman filter process models to predict the next states of the Bézier curves based on the previous states and the machine's motion. In some cases, the process models can update the tracked states (e.g., previous states) using data that specifies a relative motion of the machine between a current timestamp associated with the current states (e.g., state vectors) and a previous timestamp associated with the previous states, in order to generate one or more "transformed previous states."In some cases, the process models can then use one or more curve-shift matrices to shift the transformed previous states toward the center (e.g., the origin) of the machine and compare them to the one or more Bézier curves measured by the DNN. In other words, to predict the Bézier curve for the current or next time step, the one or more systems can multiply the transformed previous states by the curve-shift matrix.
[0032] To derive the process model and / or the curve-shift matrix, one or more systems can, for example, transform the one or more previous states of the Kalman filter (e.g., the control points of the previous time step) into "ego-movement-transformed" states by applying a rigid body transformation to each control point of the one or more previous states. However, since the DNN measurements begin approximately 2 meters in front of the machine (e.g., the ego vehicle) for each frame, the control points of the one or more previous states can be shifted backward relative to the DNN measurements for a current time step, depending on how far the ego machine has traveled forward.Therefore, the one or more systems can sample a multitude of polyline points using the one or more Bézier curves from the one or more previous states to determine the machine's current location relative to the polyline points / Bézier curve. That is, the one or more systems can identify a polyline point from the multitude of polyline points that corresponds to the machine's current location and / or immediately in front of the machine (e.g., approximately 2 meters in front of the machine), where the DNN measurements roughly begin. The one or more systems can use the polyline points in front of the machine's current location to determine one or more predicted sets of control points that correspond to the predicted Bézier curves for the current or next time step, starting and ending closer to the DNN measurements.
[0033] In various examples, the Kalman filter measurement models can obtain the predicted current state vectors from the process models and update or refine these current state vectors using, for example, sensor data representing actual path-associated measurements, outputs from machine learning models indicating predicted path-associated points based on the sensor data, or other data types. For instance, the measurement models can obtain the current state vectors containing first control point values corresponding to the first path-associated predictions. They can also obtain outputs from the machine learning models indicating second control point values corresponding to the second path-associated predictions. The measurement models can then fuse, combine, average, and otherwise manipulate the first and second control point values., to update the current state vectors.
[0034] The one or more systems can then use the updated current state vectors to determine one or more geometries associated with the paths. For example, using the updated current state vectors, the one or more systems can determine one or more Bézier curves associated with the paths, and the Bézier curves can indicate a 3D geometry associated with the various paths in the environment. The one or more systems can then initiate one or more operations associated with the machine, based at least on the geometries. For example, the one or more systems can select a path for the machine to follow, plan a trajectory for the machine to follow, or modify one or more behavior parameters of the machine based on the geometry of the path (e.g.,(increasing or decreasing the speed for paths with inclines or declines, increasing or decreasing the speed based on path curvature, etc.), or other operations.
[0035] In some embodiments, the systems and methods described herein can be performed within a simulation environment (e.g., NVIDIA DriveSIM) using simulated data (e.g., simulated sensor data from simulated sensors of a virtual or simulated machine). For example, simulated sensor data can be used (e.g., using one or more machine learning models, neural networks, etc.) to identify, detect, and / or map lane lines, road boundary lines, other lines, vertical structures / features, etc., within the simulation environment using points of a curve and / or one or more curve-fitting algorithms. This information can then be used to perform operations (e.g., control, navigation, planning operations, etc.) assigned to the virtual machine within the environment.These simulated operations can be used to test the performance of the underlying algorithms, systems, and / or processes before they are deployed in the real world. In some cases, the simulation can be used to generate synthetic training data, such as training data that includes regions of interest and / or subregions of interest within the simulation. In some embodiments, other methods can be used to generate synthetic training data in addition to or as an alternative to simulation. For example, the synthetic training data can be generated using neural rendering fields (NERFs), Gaussian splat techniques, diffusion models, electrostatic models (e.g., Poisson flow generative models (PFGMs), etc.).The synthetic training data (in addition to or as an alternative to real-world data) can then be processed to determine, for example, geometry, curvature, semantic information, classification information and / or other information related to features of interest, such as lines, longitudinal features (e.g., masts) and / or other features within a driving environment, a warehouse, etc.
[0036] In each example, such as when a simulation environment is used for testing, validation, training, etc., the simulation environment and / or the associated training data can be rendered or otherwise generated using one or more light transport algorithms, such as ray tracing and / or path tracing algorithms. In some embodiments, the simulation environment and / or one or more objects, features, or components thereof can be generated or managed within a three-dimensional (3D) content collaboration platform (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physical AI, and / or other use cases, applications, or services. For example, the content collaboration platform or system may include a system that uses Universal Scene Descriptor (USD) data (e.g., OpenUSD) for managing objects, features, scenes, etc.within a simulated environment, digital environment, etc. The platform can include real-world physics simulation, such as using NVIDIA's PhysX SDK, to simulate real-world physics and physical interactions with simulations hosted by the platform. The platform can integrate OpenUSD, along with ray tracing / path tracing / light transport simulation (e.g., NVIDIA's RTX rendering technologies), into software tools and simulation workflows to build, train, deploy, or test AI systems, such as systems for testing, validating, and training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automotive, robotics, machinery, or other applications.
[0037] In some embodiments, remote control of a vehicle and / or other machine can be performed using a remote control system. For example, the systems and methods described herein can be used to identify lane lines, road boundary lines, longitudinal features, etc., which may be included in a visualization or mapping of an environment, to assist a remote operator in controlling an autonomous or semi-autonomous machine through an environment—or to provide waypoints or other information for control or navigation.
[0038] The systems and procedures described here can be used without restriction by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g. in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, steered and unsteered robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled with one or more trailers, hydrofoils, boats, shuttles, emergency vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones and / or other vehicle types.Furthermore, the systems and methods described here can be used for a variety of purposes, including but not limited to machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulations (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.
[0039] The disclosed embodiments can include a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented with a robot, aviation systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems implementing language models, such as...Large Language Models (LLMs), Vision Language Models (VLMs) and / or multimodal language models, systems that include one or more virtual machines (VMs), systems for performing operations to generate synthetic data, systems that are at least partially implemented in a data center, systems for performing conversational AI operations, systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems for performing generative AI operations, systems that are at least partially implemented using cloud computing resources, and / or other types of systems.
[0040] With reference to Fig. 1 illustrates Fig. Figure 1 shows an exemplary data flow diagram for a process of tracing a multidimensional path geometry using Kalman filters, according to some embodiments of the present disclosure. It should be noted that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, arrays, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as single or distributed components, or in conjunction with other components, in any suitable combination and location. Various functions performed by entities described herein may be executed by hardware, firmware, and / or software.Various functions can be performed, for example, by a processor executing instructions stored in main memory. In some embodiments, the systems, methods, and processes described herein can be implemented using similar components, features, and / or functionality to that of the exemplary autonomous vehicle 800 from [reference missing]. Fig. 8A-8D, the exemplary computing device 900 from Fig. 9 and / or the exemplary data center 1000 from Fig. 10, will be executed.
[0041] The process 100 can be implemented using, with additional or alternative components, one or more sensors 102, one or more machine learning models 104, one or more Kalman filters 106, including at least one or more measurement models 108 and one or more process models 110, and a drive stack 112, which may contain one or more of a model component 114, a planning component 116, a control component 118, an avoidance component 120 and / or an actuation component 122.
[0042] In summary, the process 100 can include applying sensor data 124, generated using one or more sensors 102, to one or more machine learning models 104. The one or more machine learning models 104 can output point data 126 based on at least the sensor data 124. The one or more process models 110 of the one or more Kalman filters 106 can use state data 130 from a previous state of the one or more Kalman filters 106 and location data 132 associated with a machine to determine a predicted next state (or "current state") of the one or more Kalman filters 106, which can be represented using the predicted state data 128.The one or more measurement models 108 of the one or more Kalman filters 106 can use the point data 126 from the one or more machine learning models 104 to update or refine the predicted state data 128 and generate the state data 130, which represents the current state of the one or more Kalman filters 106. The path component 111 can then use the state data 130 to generate path data 134, which represents the geometry associated with a path in an environment. The drive stack 112, which can be associated with a machine, such as machine 800, can use the path data 134 to cause the machine to perform one or more operations.
[0043] The process 100 may involve one or more machine learning models 104 receiving one or more inputs, such as sensor data 124 generated using one or more sensors 102, and generating one or more outputs, such as point data 126 representing points associated with one or more paths. In some examples, the sensor data 124 may include image data generated using one or more image sensors (e.g., one or more cameras) of a machine. In some examples, the sensor data 124 may additionally or alternatively include other types of sensor data, such as LiDAR data generated using one or more LiDAR sensors, RADAR data generated using one or more RADAR sensors, and / or so on.
[0044] For example, it illustrates Fig. 2 An example of image data representing an image 202 associated with drivable paths in an environment 204, according to some embodiments of the present disclosure. As shown, the image 202 can depict one or more lanes of a driving surface 206 in the environment 204, as well as a machine 208 operating on the driving surface 206. The example of Fig. The depicted driving surface 206 contains two paths (e.g., lanes) separated by a double solid line and annotated as path labels – a left path 210 (e.g., path 1) and a right path 212 (e.g., path 2). Although the example of Fig. While it is stated in Section 2 that the driving surface 206 has two paths, in other examples the driving surface 206 can have any number of paths (e.g., 1, 2, 3, 4, 5, 6, 7, etc.), and the one or more systems disclosed herein can track and predict geometries for any number of the paths. The one or more path labels can contain edges or centerlines (hereafter also referred to as "middle paths") of the paths such that a left edge 210A, a centerline 210B, and a right edge 210C delimit path 210, and a left edge 212A, a centerline 212B, and a right edge 212C delimit path 212.
[0045] As described above and herein, the one or more systems can maintain multiple curves for each path of the driving surface 206. That is, the one or more machine learning models 104 can predict point data 126 for curves (e.g., Bezier curves) corresponding to one or more (e.g., each) of the left edges 210A and 212A, the centerlines 210B and 212B, and / or the right edges 210C and 212C of the paths 210 and 212. Furthermore, the one or more systems can use one or more of the one or more Kalman filters 106 to track and / or predict states (e.g. state vectors including the point data) for the curves corresponding to the left edges 210A and 212A, the center lines 210B and 212B and / or the right edges 210C and 212C of the paths 210 and 212.For example, the one or more systems can use three of the one or more Kalman filters 106 to maintain three state vectors for a curve corresponding to the centerline 210B of path 210 (e.g., path 1), such as one Kalman filter and one state vector for each control point dimension. That is, the one or more systems can use a first Kalman filter to maintain a first state vector for the x-coordinates of the control points for the curve, a second Kalman filter to maintain a second state vector for the y-coordinates of the control points for the curve, and a third Kalman filter to maintain a third state vector for the z-coordinates of the control points for the curve.
[0046] Referring again to the example of Fig. 1. In examples where the sensor data contains image data, the image data may include images (such as Figure 202 described above) of a field of view from one or more of the machine's image sensors, such as one or more stereo cameras, one or more wide-angle cameras (e.g., fisheye cameras), one or more infrared cameras, one or more surround-view cameras (e.g., 360-degree cameras), one or more long-range and / or medium-range cameras, and / or another type of camera on the machine. In some examples, the image data may be captured from a single image sensor with a forward-facing, substantially centered field of view with respect to a horizontal axis (e.g., from left to right) of the machine. The image data captured from this perspective may be useful for perception during navigation—e.g., within a lane, during a lane change, when turning, at an intersection, etc.-, because a forward-facing image sensor can contain a field of view that includes the machine's current lane, one or more adjacent lanes of the machine, and / or the boundaries of the driving surface. In some examples, more than one image sensor or other types of sensors (e.g., LiDAR sensor, radar sensor, etc.) may be used to include multiple fields of view.
[0047] In some examples, image data can be captured in one format (e.g., RCCB, RCCC, RBGC, etc.) and then converted to another format (e.g., during image data preprocessing). In some examples, the image data can be provided as input to a sensor data preprocessor (not shown) to generate preprocessed image data. Many types of images or formats can be used as input, for example, compressed images such as those in the Joint Photographic Experts Group (JPEG), Red-Green-Blue (RGB), or Luminance / Chrominance (YUV) formats; compressed images as single frames extracted from a compressed video format, such as... BH264 / Advanced Video Coding (AVC) or H.265 / High Efficiency Video Coding (HEVC), raw images such as those originating from Red Clear Blue (RCCB), Red Clear (RCCC) or other types of image sensors.In some examples, different formats and / or resolutions can be used for training the one or more machine learning models 104 than for inference (e.g., during the use of the one or more machine learning models 104 in the machine).
[0048] The sensor data preprocessor can use image data representing one or more images (or other data representations) and load the sensor data into memory as a multidimensional array / matrix (alternatively referred to as a tensor in some examples, or more specifically, an input tensor). The array size can be calculated and / or represented as W × H × C, where W is the image width in pixels, H is the height in pixels, and C is the number of color channels. Other types and arrangements of the input image components are possible without loss of generality. Furthermore, the batch size B can be used as a dimension (e.g., an additional fourth dimension) when batching is employed. Batching can be used for training and / or inference. Thus, the input tensor can represent an array of dimension W × H × C × B.Any arrangement of dimensions is possible, depending on the specific hardware and software used to implement the sensor data preprocessor. This arrangement can be chosen to maximize the training and / or inference performance of one or more machine learning models (104).
[0049] In some examples, a preprocessing image pipeline can be used by the sensor data preprocessor to process one or more raw images acquired by one or more sensors (e.g., one or more image sensors) and contained in the image data to produce preprocessed image data that can represent one or more input images for one or more input layers (e.g., one or more feature extraction layers) of the one or more machine learning models (104). An example of a suitable preprocessing image pipeline can take a raw RCCB Bayer type image (e.g., 1-channel) from the sensor and convert this image into a planar RCB image (e.g., 3-channel) stored in a fixed-precision format (e.g., 16 bits per channel).The preprocessing image pipeline may include decompanding, noise reduction, demosaicing, white balance, histogram calculation and / or adaptive global tonal assignment (e.g., in this order or in an alternative order).
[0050] If the sensor data preprocessor employs noise reduction, it may include bilateral denoising in the Bayer region. If the sensor data preprocessor employs demosaicing, it may include bilinear interpolation. If the sensor data preprocessor employs histogram calculation, it may be related to calculating a histogram for the C-channel and, in some examples, combined with decompanding or noise reduction. If the sensor data preprocessor employs adaptive global tone mapping, it may include performing an adaptive gamma-log transform. This may involve calculating a histogram, determining a midtone level, and / or estimating a maximum luminance using the midtone level.
[0051] The one or more machine learning models 104 can take as input one or more images or other data representations (e.g., LiDAR data, RADAR data, etc.) represented by the sensor data 124 to generate the one or more outputs (e.g., the point data 126). In some examples, the one or more machine learning models 104 can take as input one or more images represented by the sensor data 124 (e.g., after preprocessing) to generate the point data 126. Although examples relating to the use of neural networks, and in particular CNNs, as the one or more machine learning models 104 are described herein, this is not to be understood as a limitation. For example, and without limitation, the one or more machine learning models 104 described herein can be any type of machine learning model, such as...one or more machine learning models that use linear regression, logistic regression, decision trees, support vector machines (SVMs), the naive Bayes classifier, k-nearest neighbors (Knn), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., auto-encoder, convolutional, recurrent, perceptron, long / short-term / memory (LSTM), Hopfield, Boltzmann, deep-belief, unfolding, generative adversarial, liquid state machine, large language models, vision language models, multimodal language models, transformer, diffusion, etc.), and / or other types of machine learning models.
[0052] In some examples, the one or more machine learning models can be packaged as a microservice—such as an inference microservice (e.g., NVIDIA's NIMs)—which can contain a container (e.g., an operating system virtualization package) that may contain an application programming interface (API) layer, a server layer, a runtime layer, and / or a model engine. For example, the inference microservice can contain the container itself and the model (e.g., weights and biases). In some cases, such as when the one or more machine learning models are small enough (e.g., when they have a sufficiently small number of parameters), the model can be contained within the container itself.In some embodiments, one or more of the machine learning models described herein can be used as an inference microservice to accelerate the deployment of models in any cloud, data center, or edge computing system while ensuring data security. For example, the inference microservice may include one or more APIs, a pre-configured container for simplified deployment, an optimized inference engine (e.g., built using standardized software for deploying and running AI models, such as NVIDIA's Triton inference server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations that provide low latency and high throughput for production applications, such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g.,including identity, metrics, health checks, and / or monitoring). The one or more machine learning models described herein may be included as part of the microservice along with an accelerated infrastructure capable of being deployed with a single command and / or orchestrated and automatically scaled using a container orchestration system on an accelerated infrastructure (e.g., from a single device to the size of a data center). Therefore, the inference microservice may include the one or more machine learning models (e.g., optimized for high-performance inference), inference runtime software to execute the one or more machine learning models and provide outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise management software to provide health checks, identity, and other monitoring.In some implementations, the inference microservice may include software to perform an on-premises exchange and / or update of one or more machine learning models. During the exchange or update, the software performing the exchange / update may retain the user configurations of the inference runtime software and the enterprise management software.
[0053] The point data 126 can represent points (e.g., control or Bézier points) for one or more curves (e.g., Bézier curves) that represent geometries associated with the paths. In some examples, the one or more paths may include a primary path for the machine to navigate (e.g., an ego path), one or more secondary paths adjacent to the primary path (e.g., one or more paths to the left of the ego path and / or one or more paths to the right of the ego path), and / or one or more other paths, such as an exit path, a merge path, a lane split path, a contraflow path, and / or so on, without restriction. Furthermore, the primary output for a path can contain any number of points, such as...However, without limitation, two points, five points, eight points, twelve points, twenty points, fifty points, and / or any other number of points for any number of curves (e.g., 1 curve, 2 curves, 3 curves, etc.). As described herein, in some examples, such as to reduce the computational resources required to generate one or more paths and / or to reduce the latency in generating one or more paths, the number of points for a path may be limited by a threshold number of points (e.g., ten points, twelve points, fifteen points, twenty points, etc.) and / or a threshold number of points per distance associated with the path (e.g., twelve points per hundred meters of path).
[0054] The point data 126 can represent locations (e.g., coordinates) associated with the points, where the locations can include one or more different types of locations. In some examples, the locations associated with the points can include three-dimensional (3D) locations, such as x-coordinate locations, y-coordinate locations, and z-coordinate locations within the world coordinate system. In such examples, the 3D locations can refer to a given location, such as the machine 800 and / or the one or more sensors 102 used to generate the sensor data 124. Additionally, or alternatively, in some examples, the locations associated with the points can include two-dimensional (2D) locations, such as 2D locations relative to the machine and / or the one or more sensors 102 used to generate the sensor data 124.Furthermore, or alternatively, in some examples, the locations assigned to the points may correspond to the locations represented by the sensor representation associated with sensor data 124. For example, if sensor data 124 contains image data representing an image, then the locations may include pixel locations associated with that image.
[0055] In some examples, the locations associated with the points may still contain delta values. In such examples, the delta values may represent distances, such as pixel distances, in any direction (e.g., the x-direction, y-direction, z-direction, etc.) with respect to an anchor point (e.g., predetermined, fixed anchor points distributed at points in a single camera frame, LiDAR frame, etc.) or with respect to anchor points of an anchor line (e.g., a line with one or more anchor points along that line). For example, one or more machine learning models may be trained to predict delta values corresponding to locations (e.g., represented as distances from anchor points) of points associated with an edge or rail (e.g., a mean path, centerline, etc.) of a traversable path.For a given anchor point, one or more machine learning models 104 can output a series of delta values for one or more points (e.g., each point) associated with a path. Since the pixel coordinates or locations of the anchor points or anchor lines may be known by a path detection system (e.g., the path component 111), the delta values can be used to identify the pixel coordinates or locations corresponding to the points.
[0056] Fig. Figure 3A illustrates, for example, an example of control points 302(1)-302(12) (also collectively referred to as "control points 302") in a 3D space 304, according to some embodiments of the present disclosure. The 3D space 304 may correspond to the world coordinate system, and the control points 302 may define a curve representing a geometry associated with a path in an environment. For example, and with reference to Fig. 2. The control points 302 can correspond to a Bezier curve representing the geometry associated with the centerline 210B of path 210. As in the example of Fig. As illustrated in Figure 2, the driving surface curves from right to left and rises within the environment from the perspective or field of view of the image 202. The control points 302 also curve from right to left and rise within the 3D space 304. For example, control point 302(1) may correspond to a point in the environment 204 that is closest to the sensor (e.g., the camera) that generated the image 202, and control point 302(12) may correspond to another point in the environment 204 along the centerline 210B of the path 210 of the driving surface 206 (e.g., near or beyond the machine 208 in the example of Figure 2). Fig. 2) In some examples, the locations assigned to control points 302 can be defined using three-dimensional (3D) coordinates, such as x-coordinate locations, y-coordinate locations, and z-coordinate locations. Furthermore, although the example of Fig. Figure 3A illustrates that the number of control points 302 assigned to the path is equal to 12 control points, any number of control points 302 may be used in additional or alternative examples, and the number of control points 302 may depend on a number of factors, such as the line of sight in the environment, the range of one or more sensors, the complexity of the path, the curvature of the path, etc.
[0057] Referring again to the example of Fig. The process 100 can include the fact that the one or more Kalman filters 106 use the point data 126 to generate state data 130. For example, the one or more measurement models 108 of the one or more Kalman filters 106 can use the point data 126 to update or refine the predicted state data 128 determined by the one or more process models 110 of the one or more Kalman filters 106. As described above and herein, multiple curves can be used to represent the geometry of a given path. For example, the geometry of a left edge of a path can be represented using a first curve, the geometry of a right edge of the path can be represented using a second curve, the geometry of a centerline or middle path of the lane can be represented using a third curve, and so on.Furthermore, multiple Kalman filters 106 can be used to maintain different state vectors (e.g., predicted state data 128 and state data 130) for different dimensions of the curves. That is, for a single curve, the number of Kalman filters 106 and state vectors used to maintain that curve can, in some examples, be equal to the number of dimensions (e.g., 2D, 3D, etc.) associated with the curve and / or the points.
[0058] In some examples, multiple (e.g., three) Kalman filters 106 can be used to maintain each curve. For example, one or more systems can use a first Kalman filter 106 to maintain (e.g., track and predict) first control point dimensions (e.g., x-coordinate values) for a first curve, a second Kalman filter 106 to maintain second control point dimensions (e.g., y-coordinate values) for the first curve, and a third Kalman filter 106 to maintain third control point dimensions (e.g., z-coordinate values) for the first curve. Therefore, in at least one example, a total of nine Kalman filters 106 can be used to track and predict the geometry for a given path / lane. For example, three curves can be used to represent the geometry for each path (e.g.,Left edge curve, right edge curve and centerline curve), and 3 Kalman filters can be used to track and predict the control points for each curve (e.g. one Kalman filter for each 3D coordinate dimension for each curve).
[0059] Furthermore, one or more systems can track and predict geometries associated with multiple paths in an environment simultaneously. For example, one or more systems can use a first set of Kalman filters 106 to track and predict the geometry for a first path in the environment, a second set of Kalman filters 106 to track and predict the geometry for a second path in the environment, a third set of Kalman filters 106 to track and predict the geometry for a third path in the environment, and so on.In such an example, and continuing the above information, according to which a total of 9 Kalman filters can be used for each path, the one or more systems of the present disclosure can use 18 Kalman filters to track and predict 6 Bezier curves for 2 paths, 27 Kalman filters to track and predict 9 Bezier curves for 3 paths, 36 Kalman filters to track and predict 12 Bezier curves for 4 paths, and so on.
[0060] The one or more Kalman filters 106 can contain at least one or more process models 110 and one or more measurement models 108. The one or more process models 110 (in some cases also referred to as "prediction models") of the one or more Kalman filters 106 can be configured to predict current states (e.g., next states) of the one or more Kalman filters 106 based on at least previous states of the one or more Kalman filters 106 and the relative motion of a machine between the previous states and the current states. For example, the one or more process models 110 can generate or otherwise determine the predicted state data 128 based on at least the state data 130 (from the previous state) and the localization data 132, which can specify the relative motion of the machine.The one or more measurement models 108 of the one or more Kalman filters 106 can be configured to use different data sources to update or refine the predicted current states of the one or more Kalman filters 106. For example, the one or more measurement models 108 can use the point data 126 to update or refine the predicted state data 128, and the refined / updated predicted state data 128 can correspond to the state data 130 (of the current state).
[0061] As described herein, the one or more process models 110 of the one or more Kalman filters 106 can, in some cases, use an optimized closed-loop solution to generate predicted Bezier curve states using the relative motion of the machine. For example, if the machine is moving from a previous location to a current location, the one or more process models 110 of the one or more Kalman filters 106 can predict the next states of the Bezier curves based on the previous states and the machine motion. In some cases, the one or more process models 110 can update the tracked states (e.g., previous states) using data that shows a relative motion of the machine between a current, associated timestamp (e.g.,State vectors) and a previous timestamp associated with the previous states are specified to generate one or more "transformed previous states". The one or more process models 110 can then, in some cases, compute one or more curve-shift matrices to be used to shift the transformed previous states to the center of the machine (e.g., the origin), since the state values may still be from the perspective of the previous location.
[0062] To compute the curve displacement matrices, in some examples, one or more systems can determine one or more polyline representations of the curves using the points from the previous states. The polyline points can then be sampled to determine the current location of the machine relative to the polylines. That is, one or more systems can determine a polyline point from the polyline representation that corresponds to the current location of the machine relative to the polylines. The one or more systems can then use the polyline points prior to the current location of the machine to determine one or more new sets of control points to fit new Bézier curves to the polyline representations.In some examples, the new sets of control points can correspond to the curve-shift matrices, or the curve-shift matrices can be determined using these new sets of control points. The one or more process models 110 can then predict the current states of the Kalman filters 106 by, in some examples, multiplying the curve-shift matrices by the transformed previous states.
[0063] In various examples, the one or more measurement models 108 of the one or more Kalman filters 106 can obtain the predicted state data 128 from the process models 110 and update or refine the state vectors, for example, using the point data 126, the sensor data, the actual measurements associated with paths, or other types of data. For example, the one or more measurement models 108 can obtain the predicted state data 128, including state vectors with first control point values that correspond to the first predicted curves associated with the paths. The one or more measurement models 108 can also obtain the point data 126 from the one or more machine learning models 104 that specify second control point values that correspond to the second predicted curves associated with the paths.The one or more measurement models 108 can then fuse, combine, average, etc. the first control point values and the second control point values to update the predicted state data 128 and output the state data 130.
[0064] In some examples, the one or more measurement models 108 can weigh one or more values (e.g., point coordinate locations, etc.) in the predicted state data 128 against values in the point data 126 calculated by the one or more machine learning models 104. This enables the one or more Kalman filters 106 to output time-stable state data 130. In some examples, the one or more measurement models 108 can perform time smoothing of the different values using the following equation: final_value=a∗predicted_value+(1−a)∗measured_value
[0065] In equation (3) a can be a weighting factor, final_value can be the value of a point after smoothing, value vorhergesag , can be a point value calculated for a point of the predicted state data 128, and value gemessen can be a value calculated for a point by one or more machine learning models 104. However, this is only one example of how one or more measurement models 108 can fuse the point data 126 and the predicted state data 128 to generate the state data 130, and in additional or alternative examples, any techniques for combining, updating, or refining the predicted state data 128 based on the point data 126—or vice versa—can be used.
[0066] As in the example of Fig. As further illustrated in Figure 1, process 100 can include the path component 111 processing the state data 130 and generating path data 134 based on at least the processing, which represents one or more paths. Although the example of Fig. While the path component is illustrated as separate from the one or more machine learning models 104 in one example, in other examples the path component 111 may be included as part of the one or more machine learning models 104. For example, the path component 111 may correspond to one or more layers of the machine learning models 104 that are trained to perform one or more of the processes described herein with respect to the path component 111. In such examples, the one or more machine learning models 104 may be configured to output the path data 134 in addition to or as an alternative to the output of the point data 126.
[0067] The path component 111 can be configured to use the control points contained in the state data 130 to generate one or more curves representing one or more geometries associated with the one or more paths. For example, for a path, the path component 111 can use one or more algorithms to generate a curve, such as a Bézier curve, based on at least the points associated with the path. As described herein, in some examples, an algorithm may include a Bézier algorithm, such as, but without limitation, a two-dimensional Bézier curve-fitting algorithm, a three-dimensional Bézier curve-fitting algorithm, a cubic Bézier curve-fitting algorithm, a higher-order Bézier curve-fitting algorithm, a splitting Bézier curve-fitting algorithm, and / or any other type of Bézier algorithm.Additionally or alternatively, in some examples, an algorithm may include another type of algorithm configured to generate a curve based on at least the points associated with the path. The path component 111 may then use similar processes to generate a respective curve associated with one or more (e.g., each) of the other path(s).
[0068] As an example of generating a curve (or, for example, another geometric representation) associated with a path, and given a set of n + 1 points associated with the path, a curve can be generated using the following equation: B(t)=∑i=0nBin(t)Pi
[0069] In equation (1) t can be a value between 0 and 1, which determines the position along the curve, P i can be the i-th point and Am The Bernstein polynomial can be of degree n such that: Bin(t)=(ni)ti(1−t)n−i
[0070] For example, it illustrates Fig. 3B is an example of a Bézier curve 306, which corresponds to the control points 302 in the example of Fig. 3A corresponds, according to some embodiments of the present disclosure. As an example and with reference to Fig. 2. The curve 306 can be assigned to the center line 210B of the path 210 of the driving surface 206. In the example of Fig. 3B The state data 130 and / or a combination of several instances of the state data 130 can specify the respective locations of the points 302 within the 3D space 304. The path component 111 can then perform one or more of the processes described herein, such as using one or more Bézier algorithms to determine the curve 306 (e.g., a representation of the path geometry, such as a Bézier curve) associated with the path. As shown, using one or more Bézier algorithms, the curve 306 can start at a first point 302(1), end at a last point 302(12), and have a shape based on at least the other points 302(2)-(11). Additionally, by using one or more Bezier algorithms, curve 306 can contain a smooth, continuous curve that is intended to represent the actual path that the machine is to follow within an environment 204.Furthermore, the curve 306 can be represented using one or more Bézier algorithms with a smaller number of points 302 than with conventional methods, resulting in greater computational efficiency while maintaining - or even exceeding - the accuracy of the geometric representation of the path.
[0071] The process 100 can also include a drive stack 112 that uses the path data 134 and / or the one or more specified paths to cause the machine to perform one or more operations. As shown, the driving stack 112 can contain a perception component (e.g., corresponding to a perception layer of the driving stack 112), a model component 114, a planning component 116 (e.g., corresponding to a planning layer of the driving stack 112), a control component 118 (e.g., corresponding to a control layer of the driving stack 112), an avoidance component 120 (e.g., corresponding to an obstacle or collision avoidance layer of the driving stack 112), an actuation component 122 (e.g., corresponding to an actuation layer of the driving stack 112), and / or other components corresponding to additional and / or alternative layers of the driving stack 112.In some examples, the process 100 can be carried out by one or more perception components that can relay the layers of the driving stack 112 to the model component 114, as described in more detail herein.
[0072] Model component 114 can be used to generate, update, and / or define a world model. Model component 114 can use information generated by and received from one or more perception components of the driving stack 112 (e.g., the locations of rails or edges of drivable paths based on one or more path geometries, one or more path classifications, path data 134, etc.). The one or more perception components can include an obstacle perceiver, a path perceiver, a wait perceiver, a map perceiver, and / or one or more other perception components.For example, the world model can be defined, at least partially, based on the possibilities for obstacles, paths, and waiting conditions that can be perceived in real time or near real time by the obstacle perceiver, path perceiver, waiting perceiver, and / or map perceiver. Model component 114 can continuously update the world model based on recently generated and / or received inputs (e.g., data) from the obstacle perceiver, path perceiver, waiting perceiver, map perceiver, and / or other components of the autonomous machine's control system.
[0073] The world model can be used to help inform the planning component 116, the control component 118, the avoidance components 120, and / or the actuation component 122 of the travel stack 112. The obstacle detector can perform obstacle detection based on where the machine is allowed or able to travel (e.g., based on the location of the traversable paths defined by one or more path geometries) and how fast the machine can travel without colliding with an obstacle (e.g., an object such as a structure, entity, machine, etc.) detected by the machine's sensors.
[0074] The path perceiver can perform path perception, such as by perceiving nominal paths available in a given situation. In some examples, the path perceiver can further consider lane changes for path perception. A lane graph can represent the path or paths available to the machine and can be as simple as a single path on a highway on-ramp. In some examples, the lane graph can include paths to a desired lane and / or indicate available changes on the highway (or other road type), or include nearby lanes, lane changes, junctions, curves, highway interchanges, junctions, and / or other information. In at least some examples, the path perceiver can correspond to or include one or more machine learning models, one or more Kalman filters, and / or the path component.
[0075] The waiting senser may be responsible for determining restrictions for the machine as a result of rules, conventions, and / or practical considerations. For example, the rules, conventions, and / or practical considerations may relate to traffic lights, multi-way stops, right-of-way rules, merging, toll booths, barriers, police or other emergency personnel, road workers, stopped buses or other machinery, one-way bridge allocations, ferry landings, etc. Thus, the waiting senser can be used to identify potential obstacles and implement one or more controls (e.g., slowing down, stopping, etc.) that might not have been possible if one had relied solely on the obstacle senser.
[0076] The map perceiver may include a mechanism for determining behavior and, in some cases, for identifying specific examples of which conventions are applied at a particular location. For instance, based on data representing past trips or journeys, the map perceiver may determine that U-turns are not permitted at a certain intersection at certain times; that an electronic sign indicating the direction of lane changes depending on the time of day; that two traffic lights in close proximity (e.g., barely offset from each other) serve different streets; that in Rhode Island, the first car waiting to turn left at a traffic light is breaking the law if it initiates a turn before oncoming traffic when the light is green; and / or other information. The map perceiver may also inform the machine about static or stationary infrastructure objects and obstacles.The map sensing system can also generate information for the waiting sensing system and / or the path sensing system, for example to determine which traffic light at an intersection must be green for the machine to take a specific path.
[0077] In some examples, information from the map sitter can be sent, transmitted, and / or made available to one or more servers (e.g., to a map manager on one or more servers), and information from the one or more servers can be sent, transmitted, and / or made available to the map sitter and / or a localization manager of the machine. The map manager may include a cloud allocation application located remotely from the machine and accessible to the machine via one or more networks.For example, the machine's map perceiver and / or localization manager can communicate with the map manager and / or one or more other components or features of the one or more servers to inform the map perceiver and / or localization manager about the machine's past and present journeys or trips, as well as the past and present journeys or trips of other machines. The map manager can provide allocation outputs (e.g., map data) that can be localized by the localization manager based on a specific location of the machine, and the localized allocation outputs can be used by model component 114 to generate and / or update the world model.
[0078] Planning component 116 may include, among other components, features, and / or functionality, a route planner, a lane planner, a behavior planner, and a behavior selector. The route planner can use information from the map perceiver, the map manager, and / or the localization manager, among other sources, to generate a planned path. This path may consist of GNSS waypoints (e.g., GPS waypoints), 3D world coordinates (e.g., Cartesian, polar, etc.), coordinates relative to an origin point on the machine, and so on. The waypoints can represent a specific future distance for the machine, such as a number of city blocks, kilometers, feet, inches, miles, etc., which can be used as a destination for the lane planner.
[0079] The lane planner can use the lane graph (e.g., the lane graph from the path perceiver), poses of objects within the lane graph (e.g., according to the location manager), and / or a destination point and direction at a future distance from the route planner as inputs. The destination point and direction can be assigned to the best matching drivable point and direction in the lane graph (e.g., based on GNSS and / or compass direction). A graph search algorithm can then be executed on the lane graph from a current edge in the lane graph to find the shortest path to the destination point.
[0080] The behavior planner can determine the feasibility of basic machine behaviors, such as remaining in the lane or changing lanes left or right, so that the feasible behaviors can be aligned with the lane planner's output of the most desired behaviors. For example, if it is determined that the desired behavior is not safe and / or available, a default behavior can be selected instead (e.g., the default behavior might be to remain in the lane if the desired behavior or a lane change is not safe).
[0081] The control component 118 can follow a trajectory or path (lateral and longitudinal) received from the behavior selector of the planning component 116 (e.g., based on one or more path geometries and / or classifications) as closely as possible and within the machine's capabilities. The control component 118 can use close feedback to handle unplanned events or behaviors that are not modeled and / or anything that leads to deviations from the ideal (e.g., an unexpected delay). In some examples, the control component 118 can use a forward prediction model that takes the control as an input variable and generates predictions that can be compared to the desired state (e.g., the desired lateral and longitudinal path requested by the planning component 116). The one or more controls that minimize the deviation can then be determined.
[0082] Although planning component 116 and control component 118 are illustrated separately, this is not intended as a restriction. For example, in some examples, the boundary between planning component 116 and control component 118 may not be precisely defined. Therefore, at least some of the components, features, and / or functionality attributed to planning component 116 may also be assigned to control component 118, and vice versa. This may also apply to each of the separately illustrated components of the driving stack 112.
[0083] The avoidance component 120 can assist the machine in avoiding collisions with objects (e.g., dynamic and stationary objects). The avoidance component 120 can include a computational mechanism at a "primal level" of obstacle avoidance and act as a "survival brain" or "reptilian brain" for the machine. In some examples, the avoidance component 120 can be used independently of components, features, and / or functionality of the machine required to comply with traffic laws and drive considerately. In such examples, the avoidance component 120 can ignore traffic laws, rules of the road, and standards for considerate driving to ensure that no collisions occur between the machine and any objects.Therefore, the obstacle avoidance layer can be a separate layer from the traffic control layer, and the obstacle avoidance layer can ensure that the machine only performs safe actions from an obstacle avoidance perspective. The traffic control layer, on the other hand, can ensure that the machine obeys traffic rules and conventions and respects lawful and conventional right-of-way (as described herein).
[0084] In some examples, the traversable paths, as defined by the path geometries and / or the path data 134 corresponding to each of the path geometries, can be used by the avoidance component 120 to determine which controls or actions to take. For example, the traversable paths can provide the avoidance component 120 with information about where the machine can maneuver without colliding with objects, structures, and / or the like, or at least where no static structures can be found. In some examples, the avoidance component 120 can be implemented as a separate, independent feature of the machine. For example, the avoidance component 120 can operate separately (e.g., parallel to, before, and / or after) the planning layer, the control layer, the actuation layer, and / or other layers of the motion stack 112.
[0085] Now, with reference to Fig. 4 is Fig. 4. A data flow diagram illustrating an exemplary detail 400, which represents certain operations of the example of Fig. The process 100 described in Section 1 is associated with this disclosure, according to some embodiments. For example, the one or more process models 110 of the one or more Kalman filters 106 may include one or more state transition components 402 and one or more state prediction components 404, and the one or more measurement models 108 of the one or more Kalman filters 106 may include one or more mapping components 406 and one or more state update components 408.
[0086] The one or more state transition components 402 of the one or more process models 110 can receive previous state data 410 and the localization data 132 and determine one or more state transition matrices that represent how the state of the one or more Kalman filters 106 evolves over time. In some cases, the one or more state transition components 402 can update the previous state data 410 using the localization data 132, which can indicate a relative movement of the machine between a current timestamp and a previous timestamp associated with the previous state data 410, to generate one or more "transformed previous states".The one or more state transition components 402 can compute one or more curve shift matrices to be used to shift the transformed previous state data 410 to the center of the machine (e.g. the origin).
[0087] To compute the curve shift matrices, the one or more state transition components 402 can, in some examples, determine one or more polyline representations of the curves using the previous state data 410. The polyline points can then be sampled to determine a current location of the machine with respect to the polylines. That is, the one or more state transition components 402 can determine a polyline point of the polyline representation that corresponds to the current location of the machine with respect to the polylines. The one or more state transition components 402 can use the polyline points prior to the current location of the machine to determine one or more new sets of control points to fit new Bézier curves to the polyline representations.In some examples, the new sets of control points can correspond to the curve-shift matrices, or the curve-shift matrices can be determined using these new sets of control points. The one or more state prediction components 404 of the one or more process models 110 can then generate the predicted state data 128 by, in some examples, multiplying the curve-shift matrices with the transformed previous state data 410. Thus, the one or more process models 110 of the one or more Kalman filters 106 can use an optimized closed-form solution to generate predicted Bezier curve states using the relative motion of the machine.
[0088] In various examples, the one or more measurement models 108 of the one or more Kalman filters 106 can obtain the predicted state data 128 from the one or more process models 110. The one or more mapping components 406 of the one or more measurement models 108 can map the predicted state data 128 to the point data 126 obtained from the one or more machine learning models 104. That is, the one or more mapping components 406 can map predicted state data 128 for certain paths or curves to the point data 126 for those certain paths or curves. In some examples, the one or more mapping components 406 can map the predicted state data 128 for a path to the point data 126 based on the centerlines of the tracked / predicted paths.
[0089] For example, points for a first centerline curve can be assigned to a first confidence level, according to which the points are assigned to a first path, a second confidence level, according to which the points are assigned to a second path, a third confidence level, according to which the points are assigned to a third path, and so on. The one or more assignment components 406 can then determine that the points are assigned to the first path if the first confidence level contains a highest confidence level. Furthermore, the one or more assignment components 406 can perform similar processes for each of the other two paths in this example. The one or more assignment components 406 can then assign the points for the edge curves for the paths to their respective paths, based at least on the assignment of the centerline curve points to their respective paths.Using these or other techniques, the one or more mapping components 406 can determine which paths correspond to the point data 126 and / or the predicted state data 128.
[0090] The one or more state update components 408 can then update or refine the predicted state data 128 using the point data 126. The one or more state update components 408 of the one or more measurement models 108 can then output the updated / refined predicted state data 128 as the current state data 412. For example, the one or more measurement models 108 can fuse, combine, average, weight, etc., the values of the predicted state data 128 and the values of the point data 126 to generate or otherwise determine the current state data 412. In some examples, the one or more state update components 408 can output one or more values (e.g., point coordinate locations, etc.).) in the predicted state data 128 against values in the point data 126, which are calculated by the one or more machine learning models 104. This enables the one or more Kalman filters 106 to output temporally stable geometries. In some examples, the one or more state update components 408 can perform temporal smoothing of the various values, for example, using equation (3) described above and / or any other smoothing techniques.
[0091] Fig. Figure 5 illustrates an example of a system 502 that can perform one or more of the processes described herein, according to some embodiments of the present disclosure. As shown, the system 502 (which may represent and / or contain the one or more exemplary computing devices 900 and / or the exemplary data center 1000) can include one or more processors 504 (which may be similar to and / or contain the CPUs 906 and / or the GPUs 908), a memory 506 (which may be similar to and / or contain the memory 904), and the one or more sensors 102. For example, the memory 506 can store the one or more machine learning models 104, the one or more Kalman filters 106, the path component 111, and the driving stack 112.Furthermore, the one or more processors 504 can execute the one or more machine learning models 104, the one or more Kalman filters 106, the path component 111 and / or the drive stack 112 to perform one or more of the processes described herein.
[0092] For example, the System 502 can use one or more Kalman Filters 106 and / or other recursive models to track and predict control points corresponding to Bézier curves (e.g., 2D and / or 3D Bézier curves) that can represent geometries associated with one or more lanes of a road surface. In some cases, the System 502 can use multiple Bézier curves to represent the geometry of a lane, and it can also use multiple Kalman Filters 106 to track and predict control points for each Bézier curve. For example, the System 502 can represent an edge or centerline of the lane using a first Bézier curve, and control point coordinates for the first Bézier curve can be tracked and predicted using multiple Kalman Filters 106, as described herein.
[0093] Referring to Fig. 6 and Fig. 7. Each block of methods 600 and 700 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. Various functions can be performed, for example, by a processor executing instructions stored in main memory. The methods can also be embodied as computer-usable instructions stored on computer storage media. The methods can be provided by a standalone application, a service or hosted service (alone or in combination with another hosted service), or a plug-in for another product, to name just a few. Furthermore, methods 600 and 700 are illustrated by way of example with respect to Fig. 1 described. However, these procedures can additionally or alternatively be performed by any system or any combination of systems, including, but not limited to, the systems described herein.
[0094] Fig. Figure 6 is a flowchart illustrating an example of a method 600 for tracing a multidimensional path geometry using Kalman filters, according to some embodiments of the present disclosure. The method 600 may include, in block B602, determining one or more first state vectors associated with one or more first Kalman filters, wherein the one or more first state vectors contain one or more first values corresponding to one or more first Bezier representations associated with one or more first segments of one or more paths. For example, the one or more process models 110 may determine the predicted state data 128, which may contain the one or more state vectors.The one or more first state vectors can contain the one or more first values corresponding to the one or more first Bézier representations associated with the one or more first sections of the one or more paths.
[0095] Method 600 in block B604 may include determining one or more second state vectors associated with one or more second Kalman filters, wherein the one or more second state vectors contain one or more second values corresponding to one or more second Bézier representations associated with one or more second sections of the one or more paths. For example, the one or more process models 110 may determine a second instance of the predicted state data 128, which may contain the one or more second state vectors. The one or more second state vectors may contain the one or more second values corresponding to the one or more second Bézier representations associated with the one or more second sections of the one or more paths.In some examples, the one or more second values may be the same or different from the one or more first values, the one or more second Bézier representations may be the same or different from the one or more first Bézier representations, and / or the one or more second sections may be the same or different from the one or more first sections.
[0096] Method 600 in block B606 can include the computation of one or more geometries associated with the one or more paths, based on at least one or more first state vectors, one or more second state vectors, and data specifying at least one relative motion associated with a machine. For example, path component 111 can compute the one or more geometries associated with the one or more paths.In some examples, the one or more process models 110 can output the predicted state data 128 based at least on the one or more first state vectors, the one or more second state vectors, and the relative motion associated with the machine; then the one or more measurement models 108 can use the point data 126 and / or other outputs from machine learning models or perceptions to refine or update the predicted state data 128 and generate the state data 130; and the path component 111 can use the state data 130 to generate path data 134 representing the one or more geometries associated with the paths.
[0097] The procedure 600 can include in block B608 causing the machine to perform one or more operations based on at least one or more geometries associated with the one or more paths. For example, one or more components of the driving stack 112 can cause the machine to perform one or more operations based on at least one or more geometries associated with the one or more paths. For example, the planning component 116 can use the path data 134 to plan a trajectory for the machine to follow, or the avoidance component 120 can use the path data 134 to cause the machine to select a path that avoids one or more objects, etc.
[0098] Fig. Figure 7 is a flowchart illustrating an example of a method 700 for using recursive models to predict multidimensional path geometry, according to some embodiments of the present disclosure. The method 700 may include, in block B702, obtaining one or more first points corresponding to one or more first Bézier representations associated with one or more paths in an environment. For example, the one or more process models 110 may obtain the state data 130 containing the one or more first points corresponding to the one or more first Bézier representations associated with the one or more paths in the environment. In some examples, the one or more first points may correspond to the control point coordinate dimensions (e.g.,x-coordinate values, y-coordinate values or z-coordinate values) correspond to one or more of the first Bezier representations.
[0099] Method 700 in block B704 can include the calculation, based on at least one or more first points and a relative motion associated with a machine, of one or more second points corresponding to one or more second Bézier representations associated with the one or more paths. For example, the one or more process models 110 can calculate the predicted state data 128, which includes the one or more second points, based on at least the state data 130 and the localization data 132, which can specify the relative motion associated with the machine. The one or more second points of the predicted state data 128 can correspond to the one or more second Bézier representations associated with the one or more paths.In some examples, one or more second points can correspond to the control point coordinate dimensions (e.g., x-coordinate values, y-coordinate values, or z-coordinate values) of one or more second Bezier representations.
[0100] The procedure 700 in block B706 may include performing one or more operations assigned to the machine, based on at least one or more second Bézier representations. For example, the drive stack 112 may perform the one or more operations assigned to the machine 800, based on at least the path data 134, which specifies the one or more second Bézier representations. In some examples, the one or more measurement models 108 may update the one or more second Bézier representations to generate the state data 130, and the path component 111 may determine the path data 134 using the state data 130.In some examples, the planning component 116 of the driving stack 112 can use the path data 134 to plan a trajectory that the machine has to follow, or the avoidance component 120 can use the path data 134 to cause the machine to select a path that avoids one or more objects, etc. EXEMPLARY AUTONOMOUS VEHICLE
[0101] Fig. Figure 8A is an illustration of an exemplary autonomous vehicle 800, according to some embodiments of the present disclosure. The autonomous vehicle 800 (here alternatively referred to as "vehicle 800") may, without limitation, include: a passenger vehicle, such as a car, truck, bus, emergency service vehicle, shuttle, electric or motorized bicycle, motorcycle, fire engine, police vehicle, ambulance, boat, construction vehicle, underwater vehicle, robotic vehicle, drone, airplane, a vehicle coupled to a trailer (e.g., a semi-trailer truck used for transporting cargo), and / or another type of vehicle (e.g., one that is unmanned and / or carries one or more passengers).Autonomous vehicles are generally described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) standard "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and earlier and future versions of this standard). The Vehicle 800 can exhibit functionality corresponding to one or more of the Levels 3 through 5 of autonomous driving levels.The Vehicle 800 can exhibit functionality corresponding to one or more of the Levels 1 to 5 of autonomous driving. For example, depending on its configuration, the Vehicle 800 may be capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomous," as used herein, may encompass any and / or all types of autonomy for the Vehicle 800 or any other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, assistive autonomy, semi-autonomous, primary autonomous, or any other designation.
[0102] The vehicle 800 can include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. The vehicle 800 can include a propulsion system 850, such as an internal combustion engine, a hybrid electric power plant, a pure electric motor, and / or another type of propulsion. The propulsion system 850 can be connected to a drivetrain of the vehicle 800, which may include a transmission to enable the propulsion of the vehicle 800. The propulsion system 850 can be controlled in response to signals received from the throttle or accelerator device 852.
[0103] A steering system 854, which may include a steering wheel, can be used to steer the vehicle 800 (e.g., along a desired path or route) when the propulsion system 850 is in operation (e.g., when the vehicle is in motion). The steering system 854 can receive signals from a steering actuator 856. The steering wheel is optional for full automation (level 5).
[0104] The brake sensor system 846 can be used to actuate the vehicle brakes in response to receiving signals from the brake actuators 848 and / or the brake sensors.
[0105] The one or more 836 controllers, the one or more 804 systems-on-chips (SoCs) ( Fig. 8C) and / or GPUs, can supply signals (e.g., representing instructions) to one or more components and / or systems of the vehicle 800. For example, the one or more controllers can send signals to actuate the vehicle brakes via one or more brake actuators 848, to actuate the steering system 854 via one or more steering actuators 856, and to actuate the propulsion system 850 via one or more throttle / accelerator devices 852. The one or more controllers 836 can include one or more built-in (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and issue operating commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 800.The one or more Controller 836s can include a first Controller 836 for autonomous driving functions, a second Controller 836 for functional safety functions, a third Controller 836 for artificial intelligence functions (e.g., computer vision), a fourth Controller 836 for infotainment functions, a fifth Controller 836 for emergency redundancy, and / or other controllers. In some examples, a single Controller 836 can perform two or more of the above-mentioned functionalities, two or more Controller 836s can perform a single functionality, and / or any combination thereof.
[0106] The one or more controllers 836 can provide the signals for controlling one or more components and / or systems of the vehicle 800 in response to sensor data received from one or more sensors (e.g. sensor inputs). The sensor data can be received, for example, without restriction, by one or more of the following: Global Navigation Satellite Systems (GNSS) sensor(s) 858 (e.g., Global Positioning System sensor(s)), radar sensor(s) 860, ultrasonic sensor(s) 862, lidar sensor(s) 864, inertial measurement unit (IMU) sensor(s) 866 (e.g., accelerometer(s), gyroscope(s), magnetic compass(s), magnetometer(s), etc.), microphone(s) 896, stereo camera(s) 868, wide-angle camera(s) 870 (e.g., fisheye cameras), infrared camera(s) 872, ambient camera(s) 874 (e.g.,360-degree cameras), long-range and / or medium-range camera(s) 898, speed sensor(s) 844 (e.g. for measuring the speed of the vehicle 800), vibration sensor(s) 842, steering sensor(s) 840, brake sensor(s) (e.g. as part of the brake sensor system 846), and / or other sensor types.
[0107] One or more of the controllers 836 can receive inputs (e.g., in the form of input data) from an instrument cluster 832 of the vehicle 800 and provide outputs (e.g., in the form of output data, display data, etc.) via a human-machine interface (HMI) display 834, an audible alarm, a loudspeaker, and / or via other components of the vehicle 800. The outputs can include information such as vehicle speed, engine speed, time, map data (e.g., the high-definition (HD) map 822 from Fig. 8C), location data (e.g., the location of vehicle 800, e.g., on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the one or more controllers 836, etc. For example, the HMI display 834 can show information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about driving maneuvers that the vehicle has performed, is currently performing, or will perform (e.g., changing lanes now, taking exit 34B in two miles, etc.).
[0108] The vehicle 800 also includes a network interface 824, which can use one or more wireless antennas 826 and / or modems for communication over one or more networks. The network interface 824 can be suitable, for example, for communication over Long-Term Evolution (LTE), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Global System for Mobile Communication (GSM), IMT-CDMA Multi-Carrier (CDMA2000), etc. The one or more wireless antennas 826 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.and / or low power wide area networks (LPWANs), such as LoRaWAN, SigFox, etc.
[0109] Fig. 8B is an example of camera locations and fields of view for the exemplary autonomous vehicle 800. Fig. 8A, according to some embodiments of the present disclosure; The cameras and respective fields of view are an exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different locations on the vehicle 800.
[0110] The camera types may include, but are not limited to, digital cameras designed for use with the components and / or systems of the Vehicle 800. The one or more cameras may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. Depending on the configuration, the camera types may be capable of any frame rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The cameras may use roller shutters, global shutters, another type of shutter, or a combination thereof.In some examples, the color filter array may include a red-clear-clear-clear (RCCC) color filter array, a red-clear-clear-blue (RCCB) color filter array, a red-blue-green (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor color filter array (RGGB), a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, cameras with clear pixels, such as cameras with an RCCC, RCCB, and / or RBGC color filter array, may be used to increase light sensitivity.
[0111] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For instance, a multi-function monocular camera can be installed to provide features including lane departure warning, traffic sign recognition, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0112] One or more cameras can be mounted in a bracket, such as a specially designed (three-dimensional ("3D") printed bracket, to eliminate stray light and reflections from inside the vehicle (e.g., reflections of the dashboard in the windshield) that could interfere with the camera's image acquisition. Regarding the mounting of exterior mirrors, the mirrors can be individually 3D printed so that the camera mounting plate is shaped to fit the mirror. In some cases, the one or more cameras can be integrated into the exterior mirror. For side cameras, the one or more cameras can also be integrated into the four pillars at each corner of the cabin.
[0113] Cameras with a field of view that includes portions of the environment in front of the vehicle (e.g., forward-facing cameras) can be used for surround view to help identify forward paths and obstacles and to provide, with the aid of one or more Controller 836 and / or Control SoCs, information critical for creating an occupancy grid and / or determining preferred vehicle paths. Forward-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras can also be used for ADAS functions and systems that include lane departure warnings (LDW), autonomous cruise control (ACC), and / or other functions such as traffic sign recognition.
[0114] A variety of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform containing a complementary metal oxide semiconductor (CMOS) color imager. Another example is the 870 wide-angle camera series, which can be used to capture objects moving into view from the periphery (e.g., pedestrians, crossing vehicles, or bicycles). Although in Fig. While only one wide-angle camera is illustrated in Figure 8B, the vehicle 800 can contain any number (including zero) of wide-angle cameras 870. Furthermore, any number of long-range cameras 898 (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The one or more long-range cameras 898 can also be used for object detection and classification, as well as basic object tracking.
[0115] Any number of stereo cameras 868 can also be included in a forward-facing configuration. In at least one embodiment, one or more of the stereo cameras 868 can include an integrated control unit comprising a scalable processing unit that can provide programmable logic (“FPGA”) and a multicore microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to create a 3D map of the vehicle's surroundings that includes a distance estimate for all points in the image. Alternatively, one or more stereo cameras 868 can include a compact stereo vision sensor that may contain two camera lenses (one left and one right) and an image processing chip that measures the distance between the vehicle and the target object and processes the generated information (e.g.,Metadata) can be used to activate the autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 868 can be used in addition to or as an alternative to those described here.
[0116] Cameras with a field of view that includes sections of the environment to the sides of the vehicle 800 (e.g., side cameras) can be used for the surround view and provide information used to create and update the occupancy grid and to generate side-impact collision warnings. For example, one or more surround cameras 874 (e.g., four surround cameras 874, as in Fig. (8B illustrated) are positioned on the vehicle 800. The one or more surround-view cameras 874 can include one or more wide-angle cameras 870, one or more fisheye cameras, one or more 360-degree cameras, and / or the like. For example, four fisheye cameras can be mounted at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround-view cameras 874 (e.g., left, right, and rear) and one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.
[0117] Cameras with a field of view that includes sections of the area behind the vehicle 800 (e.g., reversing cameras) can be used for parking assistance, surround view, rear-impact warnings, and creating and updating the occupancy grid. A variety of cameras can be used, including cameras that are also suitable as one or more forward-facing cameras (e.g., one or more long-range and / or medium-range cameras 898, one or more stereo cameras 868, one or more infrared cameras 872, etc.), as described herein.
[0118] Fig. 8C a block diagram of an exemplary system architecture for the exemplary autonomous vehicle 800 from Fig. 8A, according to some embodiments of the present disclosure; It should be noted that these and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, arrays, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as single or distributed components, or in conjunction with other components, in any suitable combination and at any suitable location. Various functions described herein that are performed by entities may be executed by hardware, firmware, and / or software. Various functions may, for example, be performed by a processor executing instructions stored in a working memory.
[0119] Each of the components, features and systems of the 800 vehicle in Fig. 8C is illustrated as being connected via bus 802. Bus 802 may contain a Controller Area Network (CAN) data interface (here alternatively referred to as a "CAN bus"). A CAN can be a network within the vehicle 800 that serves to support the control of various features and functions of the vehicle 800, such as the operation of brakes, acceleration, braking, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to determine the steering wheel angle, vehicle speed, engine speed (rpm), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0120] Although the 802 bus is described here as a CAN bus, this is not intended as a limitation. For example, FlexRay and / or Ethernet can be used in addition to or as an alternative to the CAN bus. Furthermore, while a single line is used to represent the 802 bus, this is not meant as a restriction. There can be any number of 802 buses, which may contain one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more 802 buses can be used to perform different functions and / or for redundancy. For example, a first 802 bus can be used for collision avoidance functionality, and a second 802 bus can be used for actuation control.In each example, each 802 bus can communicate with one of the vehicle 800 components, and two or more 802 buses can communicate with the same components. In some examples, each 804 SoC, each 836 controller, and / or each computer within the vehicle can have access to the same input data (e.g., inputs from the vehicle 800 sensors) and be connected to a common bus, such as the CAN bus.
[0121] The vehicle 800 can contain one or more controllers 836, as shown here in relation to Fig. 8A are described. The one or more 836 controllers can be used for a variety of functions. The one or more 836 controllers can be coupled with one or more of the various other components and systems of the 800 vehicle and can be used for controlling the 800 vehicle, for the artificial intelligence of the 800 vehicle, for infotainment for the 800 vehicle, and / or the like.
[0122] The Vehicle 800 can contain one or more systems-on-a-chip (SoC) 804. The SoC 804 can contain one or more CPUs 806, one or more GPUs 808, one or more processors 810, one or more caches 812, one or more accelerators 814, one or more data storage devices 816, and / or other components and features not illustrated. The one or more SoCs 804 can be used to control the Vehicle 800 in a variety of platforms and systems. For example, the one or more SoCs 804 in a system (e.g., the Vehicle 800 system) can be combined with an HD card 822, which is accessed via a network interface 824 by one or more servers (e.g., the one or more servers 878). Fig. 8D) may receive map refreshes and / or updates.
[0123] The one or more CPUs 806 can contain a CPU cluster or CPU complex (hereafter referred to as "CCPLEX"). The one or more CPUs 806 can contain multiple cores and / or L2 caches. In some embodiments, the one or more CPUs 806 can, for example, contain eight cores in a coherent multiprocessor configuration. In some embodiments, the one or more CPUs 806 can contain four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The one or more CPUs 806 (e.g., the CCPLEX) can be configured to support the concurrent operation of clusters, so that any combination of clusters of the one or more CPUs 806 can be active at any given time.
[0124] The one or more CPUs 806 can implement power management features that include one or more of the following: individual hardware blocks can be automatically clocked when idle to dynamically conserve power; each core clock can be controlled when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core can be independently power-controlled; each core cluster can be independently clocked when all cores are clocked or power-controlled; and / or each core cluster can be independently power-controlled when all cores are power-controlled. The one or more CPUs 806 can also implement an enhanced power state management algorithm that establishes acceptable power states and expected wake-up times, and the hardware / microcode determines the best power state to input for the core, cluster, and CCPLEX.The processing kernels can support simplified sequences for inputting the energy state into the software, thereby offloading the work to the microcode.
[0125] The one or more GPUs 808 can include an integrated GPU (referred to herein alternatively as an "iGPU"). The one or more GPUs 808 can be programmable and can be efficient for parallel workloads. The one or more GPUs 808 can use an extended Tensor instruction set in some examples. The one or more GPUs 808 can include one or more streaming microprocessors, each of which can contain an L1 cache (for example, an L1 cache of at least 96 KB), and two or more of the streaming microprocessors can share an L2 cache (for example, an L2 cache of 512 KB). In some embodiments, the one or more GPUs 808 can contain at least eight streaming microprocessors. The one or more GPUs 808 can use one or more application programming interfaces (APIs) for computation.Furthermore, the one or more GPUs 808 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0126] The one or more GPUs 808 can be power-optimized for best performance in automotive and embedded applications. The one or more GPUs 808 can be manufactured, for example, on a FinFET field-effect transistor. However, this is not a limitation, and the one or more GPUs 808 can also be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can contain an array of mixed-precision processing cores, divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for Deep Learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit and / or a 64 KB register file.Furthermore, streaming microprocessors can include independent parallel integer and floating-point data paths to enable efficient execution of workloads with a mix of computations and addressing operations. Streaming microprocessors can include an independent thread scheduling function to allow fine-grained synchronization and cooperation between parallel threads. Streaming microprocessors can also include a combined L1 data cache and a shared memory unit to improve performance while simplifying programming.
[0127] The one or more GPUs 808 can include high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / second in some examples. In some examples, synchronous graphics random access memory (SGRAM), such as double-data-rate type five synchronous graphics random access memory (GDDR5), can be used in addition to or as an alternative to HBM memory.
[0128] The one or more GPUs 808 can incorporate a unified memory technology that includes access counters to enable more accurate migration of memory pages to the processor that accesses them most frequently, thereby improving the efficiency of memory areas shared by processors. In some examples, support for Address Translation Services (ATS) can be used so that the one or more GPUs 808 can directly access the page tables of the one or more CPUs 806. In such examples, if the Memory Management Unit (MMU) of the one or more GPUs 808 fails, an address translation request can be sent to the one or more CPUs 806.In response, the one or more 806 CPUs can search their page tables for the virtual-physical mapping for the address and send the translation back to the one or more 808 GPUs. This unified memory technology thus enables a single, unified virtual address space for the memory of both the one or more 806 CPUs and the one or more 808 GPUs, thereby simplifying the programming of the one or more 808 GPUs and the porting of applications to them.
[0129] Additionally, the one or more GPUs 808 can contain an access counter that tracks the frequency of accesses by the one or more GPUs 808 to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently.
[0130] The one or more 804 SoCs can contain any number of 812 caches, including those described here. The one or more 812 caches can, for example, contain an L3 cache available to both the one or more 806 CPUs and the one or more 808 GPUs (e.g., one connected to both the one or more 806 CPUs and the one or more 808 GPUs). The one or more 812 caches can contain a write-back cache capable of tracking row states, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache can be 4 MB or larger, depending on the implementation, although smaller cache sizes are also possible.
[0131] The one or more SoCs 804 can contain one or more Arithmetic Logic Units (ALUs) that can be used to perform processing related to one of the many tasks or operations of the Vehicle 800—such as DNN processing. Additionally, the one or more SoCs 804 can contain one or more Floating Point Units (FPUs)—or other mathematical or numerical coprocessors—for performing mathematical operations within the system. For example, the one or more SoCs 104 can contain one or more FPUs integrated as execution units into one or more CPUs 806 and / or one or more GPUs 808.
[0132] The one or more 804 SoCs can contain one or more 814 Accelerators (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the one or more 804 SoCs can contain a hardware acceleration cluster, which may include optimized hardware accelerators and / or a large amount of on-chip memory. The large on-chip memory (e.g., 4 MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used in conjunction with the one or more 808 GPUs, offloading some of the tasks performed by the one or more 808 GPUs (e.g., to free up more cycles of the one or more 808 GPUs for other tasks). The one or more 814 Accelerators can be used, for example, for specific workloads (e.g.,Perception, convolutional neural networks (CNNs), etc., are used that are stable enough to be suitable for acceleration. The term "CNN" as used here can include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0133] The one or more Accelerators 814 (e.g., the Hardware Acceleration Cluster) can include a Deep Learning Accelerator (DLA). The one or more DLAs can include one or more Tensor Processing Units (TPUs) configured to provide an additional ten trillion operations per second for deep learning applications and inference. The TPUs can be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The one or more DLAs can also be optimized for a specific set of neural network types and floating-point operations, as well as for inference. The design of the one or more DLAs can deliver more performance per millimeter than a general-purpose GPU and far surpasses the performance of a CPU.The one or more TPUs can perform multiple functions, including a convolution function for a single instance that supports, for example, INT8, INT16 and FP16 data types for both features and weights, as well as post-processor functions.
[0134] One or more DLAs can quickly and efficiently run neural networks, especially CNNs, on processed or unprocessed data for a variety of functions, including, but not limited to: a CNN for object identification and detection using camera sensor data; a CNN for distance estimation using camera sensor data; a CNN for emergency vehicle detection and identification using microphone data; a CNN for facial recognition and vehicle owner identification using camera sensor data; and / or a CNN for security and / or protection-related events.
[0135] The one or more DLAs can perform any function of the one or more GPUs 808, and by using an inference accelerator, a developer can, for example, allocate either the one or more DLAs or the one or more GPUs 808 to each function. For example, the developer can concentrate the processing of CNNs and floating-point operations on the one or more DLAs and leave other functions to the one or more GPUs 808 and / or other accelerators 814.
[0136] The one or more Accelerators 814 (e.g., the Hardware Acceleration Cluster) can contain a Programmable Vision Accelerator (PVA), also referred to here as a Computer Vision Accelerator. The one or more PVAs can be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The one or more PVAs can offer a balance between performance and flexibility. For example, each PVA can contain any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA) cores, and / or any number of vector processors, without limitation.
[0137] The RISC cores can interact with image sensors (e.g., the image sensors of one of the cameras described here), image signal processors, and / or the like. Each RISC core can contain any amount of memory. Depending on the implementation, the RISC cores can use any number of protocols. In some examples, the RISC cores can run a real-time operating system (RTOS). The RISC cores can be implemented with one or more integrated circuits, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.
[0138] The DMA allows components of the PVA(s) to access the system's main memory independently of the single or multiple 806 CPUs. The DMA can support any number of features that optimize the PVA, including, but not limited to, support for multidimensional and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0139] Vector processors can be programmable processors designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may contain a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA machines (e.g., two DMA machines), and / or other peripheral devices. The vector processing subsystem may operate as the primary processing machine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or working memory (e.g., VMEM).A VPU core can contain a digital signal processor, such as a single instruction, multiple data (SIMD) or a very long instruction word (VLIW). The combination of SIMD and VLIW can increase throughput and speed.
[0140] Each vector processor can contain an instruction cache and can be coupled to dedicated memory. Therefore, in some examples, each vector processor can be configured to operate independently of the others. In other examples, the vector processors contained in a particular PVA can be configured to use data parallelism. For example, in some embodiments, the multiple vector processors contained in a single PVA can execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors contained in a particular PVA can simultaneously execute different computer vision algorithms on the same image, or even different algorithms on successive images or sections of an image.Among other things, any number of PVAs can be included in the hardware acceleration cluster, and any number of vector processors can be contained in each of the PVAs. Furthermore, one or more PVAs can include additional memory for error-correcting code (ECC) to increase the overall security of the system.
[0141] The one or more Accelerator 814 units (e.g., the hardware acceleration cluster) can include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the Accelerator 814. In some examples, the on-chip memory can include at least 4 MB of SRAM, consisting, for example, and without limitation, of eight field-configurable memory blocks accessible to both the PVA and the DLA. Each pair of memory blocks can include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and the DLA can access the memory via a backbone, enabling high-speed memory access for both the PVA and the DLA.The backbone can include an on-chip computer vision network that connects the PVA and DLA to the main memory (e.g., using the APB).
[0142] The on-chip computer vision network can include an interface that, prior to the transmission of control signals / addresses / data, ensures that both the PVA and the DLA provide ready-to-use and valid signals. Such an interface can provide separate phases and channels for the transmission of control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, although other standards and protocols can also be used.
[0143] In some examples, one or more SoCs 804 can include a real-time ray tracing hardware accelerator as described in US patent application no. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model) for generating real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for simulating SONAR systems, for general wave propagation simulation, for comparison with lidar data for localization purposes, and / or for other functions and / or purposes. In some embodiments, one or more Tree Traversal Units (TTUs) can be used to perform one or more operations related to ray tracing.
[0144] The single or multiple Accelerator 814 (e.g., the hardware accelerator cluster) have a wide range of applications for autonomous driving. The PVA can be a programmable vision accelerator used for critical processing steps in ADAS and autonomous vehicles. The PVA's capabilities are well-suited to algorithmic domains that require predictable processing with low power consumption and low latency. In other words, the PVA is well-suited for semi-dense or dense regular computations, even with small datasets, that require predictable runtimes with low latency and low power consumption. In the context of autonomous vehicle platforms, PVAs are therefore designed to execute classic computer vision algorithms, as they are efficient at object detection and operate with integer mathematics.
[0145] According to one embodiment of the technology, the PVA is used, for example, to perform computer stereovision. In some examples, a semi-global matching-based algorithm can be used, although this is not intended as a limitation. Many applications for Level 3-5 autonomous driving require spontaneous motion estimation or stereo matching (e.g., structure of motion, pedestrian detection, lane detection, etc.). The PVA can perform computer stereovision on input from two monocular cameras.
[0146] In some examples, the PVA can be used to perform dense optical flow processing. This involves processing raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar data. In other examples, the PVA is used for time-of-flight depth processing, for example, by processing raw time-of-flight data to deliver processed time-of-flight data.
[0147] The DLA can be used to power any type of network to improve control and driving safety; this includes, for example, a neural network that outputs a confidence score for each object detection. Such a confidence score can be interpreted as a probability or as providing a relative "weighting" of each detection compared to other detections. This confidence score allows the system to make further decisions about which detections should be considered true positives and not false positives. For example, the system can set a confidence threshold and consider only those detections that exceed the threshold as true positives.In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically initiate emergency braking, which is obviously undesirable. Therefore, only the safest detections should be considered as triggers for AEB. The DLA can employ a neural network for confidence regression. The neural network can take as input at least a subset of parameters, such as the dimensions of the boundary frame, the ground plane estimate (obtained, for example, from another subsystem), the output of the inertial measurement unit (IMU) sensor 866 correlated with the vehicle's orientation 800, distance, and 3D position estimates of the object obtained from the neural network and / or other sensors (e.g., one or more LiDAR sensors 864 or one or more radar sensors 860).
[0148] The one or more SoCs 804 can contain the one or more datastores 816 (e.g., main memory). The one or more datastores 816 can be on-chip main memory on the one or more SoCs 804, where neural networks can be stored to run on the GPU and / or the DLA. In some examples, the one or more datastores 816 can be large enough to store multiple instances of neural networks for redundancy and security. The one or more datastores 812 can include one or more L2 or L3 caches 812. The reference to the one or more datastores 816 can include a reference to main memory allocated to the PVA, the DLA, and / or one or more other accelerators 814, as described here.
[0149] The one or more 804 SoCs can contain one or more 810 processors (e.g., embedded processors). The one or more 810 processors can contain a boot and power management processor, which can be a dedicated processor and subsystem to handle boot power and management functions and related security enforcement. The boot and power management processor can be part of the boot sequence of the one or more 804 SoCs and can provide runtime power management services. The boot and power management processor can provide clock and voltage programming, support for system transitions to a low-power state, management of the thermals and temperature sensors of the one or more 804 SoCs, and / or management of the one or more 804 SoC power states.Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the one or more SoCs 804 can use the ring oscillators to detect the temperatures of the one or more CPUs 806, the one or more GPUs 808, and / or the one or more accelerators 814. If it is determined that the temperatures exceed a threshold, the boot and power management processor can enter a temperature fault routine and put the one or more SoCs 804 into a lower power state and / or put the vehicle 800 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 800 to a safe stop).
[0150] The one or more 810 processors can also include a number of embedded processors that can serve as an audio processing engine. The audio processing engine can be an audio subsystem that provides full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0151] The one or more 810 processors can also include an always-on processor machine, which provides the necessary hardware functions to support low-power sensor management and wake-up from use cases. The always-on processor machine can include a processor core, tightly coupled RAM, supporting peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0152] The one or more 810 processors can also include a security cluster machine, which contains a dedicated processor subsystem for the security management of automotive applications. The security cluster machine can include two or more processor cores, tightly coupled RAM, supporting peripherals (such as timers, an interrupt controller, etc.), and / or routing logic. In a security mode, the two or more cores can operate in lockstep mode, functioning as a single core with comparison logic that detects any differences between their operations.
[0153] The one or more 810 processors can also contain a real-time camera machine, which may include a dedicated processor subsystem for managing the real-time camera.
[0154] The one or more 810 processors may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware machine that is part of the camera processing pipeline.
[0155] The one or more 810 processors can include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate the final image for the player window. The video image compositor can perform lens distortion correction on the one or more 870 wide-angle cameras, the one or more 874 surround-view cameras, and / or on the sensors of the in-cabin surveillance camera. The in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the extended SoC and configured to detect events in the cabin and respond accordingly.A system in the cabin can lip-read to activate mobile service and make a call, dictate emails, change the destination, activate or change the infotainment system and vehicle settings, or enable voice-controlled internet browsing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are otherwise deactivated.
[0156] The video image compositor can include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if there is motion in a video, the noise reduction weights the spatial information accordingly and reduces the impact of information provided by adjacent frames. If a frame or portion of a frame does not contain motion, the temporal noise reduction performed by the video image compositor can use information from the previous frame to reduce noise in the current frame.
[0157] The video image compositor can also be configured to perform stereo equalization of the input stereo lens images. Furthermore, the video image compositor can be used for user interface design when the operating system desktop is in use and the one or more GPUs 808 do not need to constantly render new surfaces. Even when the one or more GPUs 808 are powered on and actively performing 3D rendering, the video image compositor can be used to offload the workload from the GPUs 808, thus improving performance and responsiveness.
[0158] The one or more 804 SoCs can also include a serial camera interface with a Mobile Industry Processor Interface (MIPI) for receiving video and camera input, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions. The one or more 804 SoCs can also include one or more input / output controllers, one or more of which can be software-controlled and used for receiving I / O signals that are not assigned to a specific role.
[0159] The one or more 804 SoCs can also include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The one or more 804 SoCs can be used to process data from cameras (e.g., via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., one or more 864 LiDAR sensors, one or more 860 radar sensors, etc., which can be connected via Ethernet), data from the 802 bus (e.g., vehicle speed, steering wheel position, etc.), and data from one or more 858 GNSS sensors (e.g., connected via Ethernet or CAN bus).Furthermore, one or more 804 SoCs can contain dedicated high-performance mass storage controllers, which can include their own DMA machines and can be used to offload routine data management tasks from one or more 806 CPUs.
[0160] The single or multiple 804 SoCs can form an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture that supports and efficiently utilizes computer vision and ADAS techniques for diversity and redundancy, and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The single or multiple 804 SoCs can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, the single or multiple 814 accelerators, in combination with the single or multiple 806 CPUs, the single or multiple 808 GPUs, and the single or multiple 816 data stores, can form a fast, efficient platform for autonomous vehicles of levels 3-5.
[0161] This technology thus offers capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be run on CPUs that can be configured using a high-level programming language, such as C, to execute a variety of processing algorithms on a wide range of visual data. However, CPUs are often unable to meet the performance requirements of many computer vision applications, such as execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and a prerequisite for practical Level 3-5 autonomous vehicles.
[0162] Unlike conventional systems, the technology described herein, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, enables the simultaneous and / or sequential execution of multiple neural networks and the combination of their results to enable Level 3-5 autonomous driving functionality. For example, a CNN running on the DLA or the dGPU (e.g., one or more GPUs 820) can include text and word recognition, allowing the supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA can further include a neural network capable of identifying and interpreting the sign, providing a semantic understanding, and passing this semantic understanding to the path planning modules running on the CPU complex.
[0163] Another example is that multiple neural networks can run simultaneously, as required for driving at levels 3, 4, or 5. For instance, a warning sign reading "Caution: Flashing lights indicate black ice" accompanied by an electric light can be interpreted independently or jointly by several neural networks. The sign itself can be identified as a traffic sign by a first neural network (e.g., a trained neural network), while the text "Flashing lights indicate black ice" can be interpreted by a second neural network, which informs the vehicle's path planning software (preferably running on the CPU) that the presence of black ice indicates the presence of flashing lights.The turn signal can be identified across multiple images by a third neural network, which informs the vehicle's path planning software about the presence (or absence) of turn signals. All three neural networks can run simultaneously, for example, within the DLA and / or on one or more GPUs 808.
[0164] In some examples, a CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of the Vehicle 800. The always-on sensor processing unit can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner leaves. In this way, one or more SoCs 804 provide security against theft and / or carjacking.
[0165] In another example, a CNN for detecting and identifying emergency vehicles can use data from microphones 896 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers to detect sirens and manually extract features, the one or more SoCs 804 use the CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to detect the relative approach speed of the emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by one or more GNSS sensors 858.For example, the CNN will attempt to detect European sirens when operating in Europe, and when operating in the United States, the CNN will attempt to identify only North American sirens. Once an emergency vehicle is detected, a controller can be used to execute an emergency vehicle safety routine, slowing the vehicle down, pulling over to the side of the road, parking the vehicle, and / or letting the vehicle idle, using the 862 ultrasonic sensors, until one or more emergency vehicles pass.
[0166] The vehicle can contain one or more CPUs 818 (e.g., one or more discrete CPUs or one or more dCPUs) that can be coupled to the one or more SoCs 804 via a high-speed connection (e.g., PCIe). The CPUs 818 can, for example, contain an x86 processor. The CPUs 818 can be used, for example, to perform a variety of functions, including reconciling potentially inconsistent results between ADAS sensors and the one or more SoCs 804 and / or monitoring the status and health of the one or more Controllers 836 and / or the Infotainment SoC 830.
[0167] The Vehicle 800 can contain one or more GPUs 820 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to the one or more SoCs 804 via a high-speed connection (e.g., NVIDIA's NVLINK). The one or more GPUs 820 can provide additional artificial intelligence capabilities, such as running redundant and / or distinct neural networks, and can be used to train and / or update neural networks based on input (e.g., sensor data) from sensors in the Vehicle 800.
[0168] The vehicle 800 may also include the network interface 824, which may contain one or more wireless antennas 826 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 824 can be used to establish a wireless connection over the internet to the cloud (e.g., to the one or more servers 878 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct connection between the two vehicles and / or an indirect connection (e.g., via networks and the internet) can be established. Direct connections can be established via vehicle-to-vehicle communication.Vehicle-to-vehicle communication can provide the Vehicle 800 with information about vehicles in its vicinity (e.g., vehicles in front of, beside, and / or behind the Vehicle 800). This functionality can be part of a cooperative adaptive cruise control function of the Vehicle 800.
[0169] The 824 network interface can include a system-on-a-chip (SoC) that provides modulation and demodulation capabilities, enabling one or more 836 controllers to communicate over wireless networks. The 824 network interface can include a high-frequency (RF) front end for upconversion from baseband to RF and downconversion from RF to baseband. Frequency conversions can be performed using known methods and / or superheterodyne techniques. In some examples, the RF front-end functionality can be provided by a separate chip. The network interface can include wireless functionality for communication over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0170] The Vehicle 800 may further include one or more Data Stores 828, which may be located outside the chip (e.g., outside the SoCs 804). The one or more Data Stores 828 may contain one or more memory elements, including RAM, SRAM, DRAM, VRAM, Flash, hard disks, and / or other components and / or devices capable of storing at least one bit of data.
[0171] The Vehicle 800 can also include one or more GNSS Sensors 858. The one or more GNSS Sensors 858 (e.g., GPS, supported GPS sensors, differential GPS (DGPS) sensors, etc.) assist with mapping, perception, occupancy grid creation, and / or path planning. Any number of GNSS Sensors 858 can be used, including, for example, a single GPS unit that uses a USB connection with an Ethernet-to-serial (RS-232) bridge.
[0172] The vehicle 800 can also include one or more RADAR sensors 860. The vehicle 800 can use the one or more RADAR sensors 860 to detect vehicles at long range, even in darkness and / or adverse weather conditions. The functional safety level of the RADAR can be ASIL B. The one or more RADAR sensors 860 can use the CAN bus and / or the 802 bus (e.g., for transmitting the data generated by the one or more RADAR sensors 860) for control and access to object tracking data, with some examples using Ethernet for access to the raw data. A variety of RADAR sensor types can be used. The one or more RADAR sensors 860 can be suitable for front, rear, and side RADAR applications without restriction. In some examples, one or more pulse-Doppler RADAR sensors are used.
[0173] The single or multiple RADAR 860 sensors can incorporate various configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range with side coverage, and so on. In some cases, long-range RADAR can be used for adaptive cruise control. Long-range RADAR systems can provide a wide field of view, achieved through two or more independent scans, for example, within a range of 250 m. The single or multiple RADAR 860 sensors can assist in distinguishing between stationary and moving objects and can be used by ADAS systems for emergency braking and forward collision warning. Long-range RADAR sensors can incorporate a monostatic multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and a high-speed CAN and FlexRay interface.In an example with six antennas, the four central antennas can generate a focused beam pattern designed to detect the area around vehicle 800 at higher speeds with minimal interference from traffic in adjacent lanes. The other two antennas can expand the field of view, enabling the rapid detection of vehicles entering or exiting vehicle 800's lane.
[0174] Medium-range radar systems, for example, can have a range of up to 860 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 850 degrees (rear). Short-range radar systems can include, among other things, radar sensors designed for installation at both ends of the rear bumper. When such a radar sensor system is installed at both ends of the rear bumper, it can generate two beams that continuously monitor the blind spot behind and to the sides of the vehicle.
[0175] Short-range radar systems can be used in an ADAS system for blind spot detection and / or as a lane change assistant.
[0176] The vehicle 800 can also include one or more ultrasonic sensors 862. The one or more ultrasonic sensors 862, which can be mounted on the front, rear, and / or sides of the vehicle 800, can be used for parking assistance and / or for creating and updating an occupancy grid. A variety of ultrasonic sensors 862 can be used, and different ultrasonic sensors 862 can be used for different detection ranges (e.g., 2.5 m, 4 m). The one or more ultrasonic sensors 862 can operate with functional safety levels of ASIL B.
[0177] The vehicle 800 can contain one or more LiDAR sensors 864. The one or more LiDAR sensors 864 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The one or more LiDAR sensors 864 can meet the functional safety level ASIL B. In some examples, the vehicle 800 can contain multiple LiDAR sensors 864 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to deliver data to a Gigabit Ethernet switch).
[0178] In some examples, one or more LiDAR sensors 864 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LiDAR sensors 864 may, for example, have a specified range of approximately 800 m, with an accuracy of 2 cm to 3 cm and support for an 800 Mbit / s Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 864 may be used. In such examples, the one or more LiDAR sensors 864 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of the vehicle 800. In such examples, one or more LIDAR 864 sensors can provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees, with a range of 200 m, even with objects of low reflectivity.The one or more front-mounted LIDAR 864 sensors can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0179] In some examples, LiDAR technologies, such as 3D flash LiDAR, can also be used. 3D flash LiDAR uses a laser pulse as a transmission source to illuminate the vehicle's surroundings up to approximately 200 m. A flash LiDAR unit contains a sensor that records the travel time of the laser pulse and the reflected light at each pixel, which in turn corresponds to the distance between the vehicle and objects. Flash LiDAR can generate highly accurate and distortion-free images of the surroundings with each laser pulse. In some examples, four flash LiDAR sensors can be used, one on each side of the vehicle. Available 3D flash LiDAR systems include a solid-state 3D focal plane array LiDAR camera that contains no moving parts other than a fan (e.g., a non-scanning LiDAR device).The flash LIDAR device can use a 5-nanosecond pulse of a Class I (eye-safe) laser per frame and capture the reflected laser light in the form of 3D distance point clouds and co-registered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the single or multiple LIDAR 864 sensors can be less susceptible to motion blur, vibration, and / or shock.
[0180] The vehicle may also contain one or more IMU sensors 866. In some examples, the one or more IMU sensors 866 may be located in the center of the rear axle of the vehicle 800. The one or more IMU sensors 866 may, for example, and without limitation, contain one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In some examples, such as six-axis applications, the one or more IMU sensors 866 may contain accelerometers and gyroscopes, while in nine-axis applications, the one or more IMU sensors 866 may contain accelerometers, gyroscopes, and magnetometers.
[0181] In some embodiments, the one or more IMU sensors 866 can be implemented as a miniaturized, high-performance GPS-Aided Inertial Navigation System (GPS / INS) that combines inertial sensors of a microelectromechanical system (MEMS), a highly sensitive GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and orientation. Thus, in some examples, the one or more IMU sensors 866 can enable the vehicle 800 to estimate its course without requiring input from a magnetic sensor by directly observing and correlating velocity changes from the GPS with the one or more IMU sensors 866. In some examples, the one or more IMU sensors 866 and the one or more GNSS sensors 858 can be combined in a single integrated unit.
[0182] The vehicle may contain one or more microphones 896, which are mounted in and / or around the vehicle 800. The one or more microphones 896 may be used, among other things, for the detection and identification of emergency vehicles.
[0183] The vehicle may also include any number of camera types, including one or more stereo cameras 868, one or more wide-angle cameras 870, one or more infrared cameras 872, one or more surround-view cameras 874, one or more long-range and / or medium-range cameras 898, and / or other camera types. The cameras may be used to capture image data around the entire periphery of the vehicle 800. The types of cameras used depend on the embodiment and requirements of the vehicle 800, and any combination of camera types may be used to ensure the necessary coverage around the vehicle 800. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or any other number of cameras.The cameras can, for example and without limitation, support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the one or more cameras is described here with reference to... Fig. 8A and Fig. 8B is described in more detail.
[0184] The vehicle 800 may also include one or more vibration sensors 842. The one or more vibration sensors 842 can measure vibrations of vehicle components, such as one or more axles. For example, changes in vibration may indicate a change in the road surface. In another example, if two or more vibration sensors 842 are used, the differences between the vibrations can be used to determine friction or slippage on the road surface (e.g., if the difference in vibration is between a driven axle and a freely rotating axle).
[0185] The vehicle 800 may include an ADAS system 838. The ADAS system 838 may include a SoC in some examples. The ADAS system 838 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning systems (CWS), lane centering (LC), and / or other features and functions.
[0186] The ACC systems can use one or more radar sensors (860), one or more lidar sensors (864), and / or one or more cameras. The ACC systems can include longitudinal and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle (800) and automatically adjusts the vehicle speed to maintain a safe distance from vehicles ahead. Lateral ACC performs distance control and advises the vehicle (800) to change lanes if necessary. Lateral ACC interacts with other ADAS applications, such as LCA and CWS.
[0187] The CACC uses information from other vehicles, which can be received via the network interface 824 and / or the one or more wireless antennas 826 from other vehicles via a wireless connection or indirectly via a network connection (e.g., via the internet). Direct connections can be provided via a vehicle-to-vehicle (V2V) communication link, while indirect connections can be an infrastructure-to-vehicle (I2V) communication link. In general, the V2V communication concept provides information about the vehicles immediately ahead (e.g., vehicles directly in front of the vehicle 800 and in the same lane), while the I2V communication concept provides information about traffic further ahead. CACC systems can incorporate one or both of the I2V and V2V information sources.Given the information about the vehicles ahead of vehicle 800, CACC can be more reliable and has the potential to improve traffic flow and reduce congestion on the road.
[0188] FCW systems are designed to warn the driver of a hazard, allowing them to take corrective action. FCW systems use a forward-facing camera and / or one or more RADAR 860 sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver feedback system, such as a display, speaker, and / or vibrating component. FCW systems can provide a warning in the form of an audible signal, a visual warning, a vibration, and / or a rapid braking pulse.
[0189] AEB systems detect an impending head-on collision with another vehicle or object and can automatically apply the brakes if the driver does not take corrective action within a specific time or distance parameter. AEB systems can use one or more forward-facing cameras and / or one or more RADAR 860 sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first warns the driver so they can take corrective action to avoid the collision; if the driver does not take corrective action, the AEB system can automatically apply the brakes to prevent or at least mitigate the effects of the predicted collision. AEB systems may incorporate techniques such as dynamic brake assist and / or emergency braking for an impending collision.
[0190] Lane Departure Warning (LDW) systems provide visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver if the vehicle crosses lane markings. An LDW system will not activate if the driver indicates an intentional lane departure by using a turn signal. LDW systems may utilize forward-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the feedback system for the driver, such as a display, speaker, and / or vibrating component.
[0191] LKA systems are a variant of LDW systems. LKA systems provide steering or braking inputs to correct vehicle 800 if the vehicle 800 begins to leave its lane.
[0192] Blind Spot Warning (BSW) systems detect and warn the driver of vehicles in the car's blind spot. BSW systems can provide a visual, audible, and / or tactile warning to indicate that merging into or changing lanes is unsafe. The system can issue an additional warning if the driver activates a turn signal. BSW systems can use one or more rear-facing cameras and / or radar sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver for feedback, such as a display, speaker, and / or vibrating component.
[0193] RCTW systems can provide visual, audible, and / or tactile alerts when an object is detected outside the reversing camera's field of view while the vehicle is in reverse. Some RCTW systems incorporate AEB to ensure the vehicle's brakes are applied to prevent a collision. RCTW systems can utilize one or more rear-facing radar sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver for feedback, such as a display, speaker, and / or vibrating component.
[0194] Conventional ADAS systems can produce false positives, which, while annoying and distracting for the driver, are generally not catastrophic because the ADAS systems warn the driver and give them the opportunity to decide whether a safety issue truly exists and to act accordingly. However, in an autonomous vehicle 800, the vehicle 800 itself must decide, in the event of conflicting results, whether to follow the result from a primary computer or a secondary computer (e.g., a first controller 836 or a second controller 836). In some embodiments, the ADAS system 838 can, for example, be a backup and / or secondary computer that provides information about perception to a rationality module of the backup computer.The backup computer rationality monitor can run redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. The outputs of the ADAS system 838 can be provided to a monitoring MCU. If the outputs of the primary and secondary computers conflict, the monitoring MCU must determine how to resolve the conflict to ensure safe operation.
[0195] In some examples, the primary computer can be configured to provide the monitoring MCU with a confidence score indicating its confidence in the chosen outcome. If the confidence score exceeds a threshold, the monitoring MCU can follow the primary computer's instruction, regardless of whether the secondary computer returns a conflicting or inconsistent result. If the confidence score does not reach the threshold and the primary and secondary computers display different results (e.g., conflicting results), the monitoring MCU can mediate between the computers to determine the appropriate outcome.
[0196] The monitoring MCU can be configured to run one or more neural networks trained and configured to determine, based on the outputs of the primary and secondary computers, the conditions under which the secondary computer will trigger false alarms. This allows the one or more neural networks in the monitoring MCU to learn when the secondary computer's output can be trusted and when it cannot. For example, if the secondary computer is a radar-based FCW system, a neural network in the monitoring MCU can learn to trigger an alarm when the FCW system identifies metallic objects that do not actually pose a threat, such as a drain grate or manhole cover.Similarly, if the secondary computer is a camera-based lane departure warning (LDW) system, a neural network in the supervising MCU can learn to override the LDW system when cyclists or pedestrians are present and leaving the lane is indeed the safest maneuver. In embodiments containing one or more neural networks running on the supervising MCU, the supervising MCU can include at least one DLA or GPU suitable for executing the one or more neural networks with associated memory. In preferred embodiments, the supervising MCU can include and / or be contained as a component of one or more SoCs 804.
[0197] In other examples, the ADAS System 838 can include a secondary computer that executes the ADAS functionality according to the classical rules of computer vision. Thus, the secondary computer can use classical computer vision rules (if-then), and the presence of one or more neural networks in the monitoring MCU can improve reliability, safety, and performance. For example, the diverse implementation and intentional non-identity make the overall system more fault-tolerant, especially to errors caused by software (or software-hardware interfaces).For example, if a software bug or error occurs in the software on the primary computer and the non-identical software code on the secondary computer produces the same overall result, the monitoring MCU can have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer does not cause a significant error.
[0198] In some examples, the output of the ADAS system 838 can be fed into the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 838 displays a frontal collision warning due to an object directly in front of the vehicle, the perception block can use this information in object identification. In other examples, the secondary computer may have its own trained neural network, thus reducing the risk of false positives, as described herein.
[0199] The Vehicle 800 may also include the Infotainment SoC 830 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not actually be an SoC and may contain two or more discrete components. The Infotainment SoC 830 may include a combination of hardware and software that can be used to provide the Vehicle 800 with audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking sensors, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fluid level, oil level, door open / close status, air filter information, etc.).The Infotainment SoC 830 can include, for example, radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free calling, a head-up display (HUD), an HMI display 834, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. The Infotainment SoC 830 can also be used to provide information (e.g., visual and / or audible) to one or more vehicle users, such as information from the ADAS system 838, autonomous driving information such as planned vehicle maneuvers, road layouts, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0200] The Infotainment SoC 830 can include GPU functionality. The Infotainment SoC 830 can communicate with other devices, systems, and / or components of the Vehicle 800 via the 802 bus (e.g., CAN bus, Ethernet, etc.). In some examples, the Infotainment SoC 830 can be coupled with a monitoring MCU so that the Infotainment System's GPU can perform some self-driving functions if one or more of the Primary Controllers 836 (e.g., the Vehicle 800's primary and / or backup computers) fail. In such an example, the Infotainment SoC 830 can put the Vehicle 800 into a chauffeur-to-safe-stop mode, as described here.
[0201] The Vehicle 800 may also include an Instrument Cluster 832 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument panel, etc.). The Instrument Cluster 832 may include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). The Instrument Cluster 832 may include a number of instruments, such as a speedometer, fuel gauge, oil pressure gauge, tachometer, odometer, turn signals, shift position indicator, seat belt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag system (SRS) information, lighting controls, safety system controls, navigation information, etc. In some examples, information from the Infotainment SoC 830 and the Instrument Cluster 832 may be displayed and / or shared. In other words, the Instrument Cluster 832 may be included as part of the Infotainment SoC 830, or vice versa.
[0202] Fig. 8D is a system diagram for the communication between one or more cloud-based servers and the exemplary autonomous vehicle 800. Fig. 8A, according to some embodiments of the present disclosure; The system 876 may include the one or more servers 878, the one or more networks 890, and the vehicles, including the vehicle 800. The server(s) 878 may include multiple GPUs 884(A)-884(H) (here collectively referred to as GPUs 884), PCIe switches 882(A)-882(H) (here collectively referred to as PCIe switches 882), and / or CPUs 880(A)-880(B) (here collectively referred to as CPUs 880). The GPUs 884, the CPUs 880, and the PCIe switches may be interconnected by high-speed links, such as, but not limited to, NVIDIA's NVLink interfaces 888 and / or PCIe links 886. In some examples, the GPUs 884 are connected via NVLink and / or NVSwitch SoC, and the GPUs 884 and the PCIe switches 882 are connected via PCIe links.Although eight GPUs 884, two CPUs 880, and two PCIe switches are illustrated, this should not be interpreted as a limitation. Depending on the configuration, each Server 878 can contain any number of GPUs 884, CPUs 880, and / or PCIe switches. For example, one or more Server 878s can each contain eight, sixteen, thirty-two, and / or more GPUs 884.
[0203] The one or more servers 878 can receive image data from the vehicles via the one or more networks 890. This image data is representative of images showing unexpected or changed road conditions, such as recently started roadworks. The one or more servers 878 can transmit neural networks 892, updated neural networks 892, and / or map information 894 to the vehicles via the one or more networks 890. This map information contains information about traffic and road conditions. The map information updates 894 can include updates to the HD map 822, such as information about construction sites, potholes, detours, flooding, and / or other obstacles.In some examples, the neural networks 892, the updated neural networks 892 and / or the map information 894 may result from new training and / or experience represented in the data received from any number of vehicles in the environment, and / or may be based on training performed in a data center (e.g. using one or more servers 878 and / or other servers).
[0204] One or more Server 878s can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by the vehicles and / or in a simulation (e.g., using a game machine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or subjected to other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training can be performed using one or more classes of machine learning techniques, including, but not limited to, classes such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, diverse learning, representational learning (including substitute dictionary learning), rule-based machine learning, anomaly detection, and all variants or combinations thereof. Once the machine learning models are trained, they can be used by the vehicles (e.g., transmitted to the vehicles via one or more networks 890) and / or the machine learning models can be used by one or more servers 878 for remote monitoring of the vehicles.
[0205] In some examples, one or more Server 878s can receive data from the vehicles and apply that data to advanced neural networks in real time for intelligent, real-time inference. The one or more Server 878s can include deep learning supercomputers and / or dedicated AI computers powered by GPUs 884, such as NVIDIA's DGX and DGX Station machines. However, in some examples, the one or more Server 878s can include a deep learning infrastructure that utilizes only CPU-powered data centers.
[0206] The deep learning infrastructure of one or more Server 878 systems can perform fast, real-time inference and can use this capability to assess and verify the state of the processors, software, and / or associated hardware in the Vehicle 800. For example, the deep learning infrastructure can receive periodic updates from the Vehicle 800, such as a sequence of images and / or objects that the Vehicle 800 has located within that sequence (e.g., via computer vision and / or other machine learning object classification techniques).The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by the vehicle 800, and if the results do not match and the infrastructure concludes that the AI in the vehicle 800 is not working correctly, one or more servers 878 can send a signal to the vehicle 800, instructing a fail-safe computer in the vehicle 800 to take control, notify the passengers and perform a safe parking maneuver.
[0207] For inference, one or more Server 878s can include GPUs 884s and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-driven servers and inference accelerators can enable real-time responsiveness. In other scenarios, such as when performance is less critical, servers powered by CPUs, FPGAs, and other processors can be used for inference. EXAMPLE CALCULATION DEVICE
[0208] Fig. Figure 9 is a block diagram of an exemplary computing device 900 suitable for use in implementing some embodiments of the present disclosure. The computing device 900 may include a connection system 902 that directly or indirectly couples the following devices: main memory 904, one or more central processing units (CPUs) 906, one or more graphics processing units (GPUs) 908, a communication interface 910, input / output (I / O) ports 912, input / output components 914, a power supply 916, one or more presentation components 918 (e.g., display(s)), and one or more logic units 920. In at least one embodiment, the one or more computing devices 900 may comprise one or more virtual machines (VMs), and / or each of the components thereof may comprise virtual components (e.g., virtual hardware components).As non-restrictive examples, one or more of the GPUs 908 can comprise one or more vGPUs, one or more of the CPUs 906 can comprise one or more vCPUs, and / or one or more of the logic units 920 can comprise one or more virtual logic units. Thus, a compute device 900 can contain discrete components (e.g., a complete GPU allocated to the compute device 900), virtual components (e.g., a portion of a GPU allocated to the compute device 900), or a combination thereof.
[0209] Although the various blocks of Fig. Where components 9 are shown as connected via the connection system 902, this is not intended as a limitation and is for clarity only. In some embodiments, for example, a presentation component 918, such as a display device, may be considered an I / O component 914 (e.g., if the display is a touchscreen). As another example, the CPUs 906 and / or GPUs 908 may contain memory (e.g., the memory 904 may represent a storage device in addition to the memory of the GPUs 908, the CPUs 906, and / or other components). In other words, the computing device of Fig. Section 9 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all are within the scope of protection of the computing device of Fig. 9 are being considered.
[0210] The 902 connection system can represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 902 connection system can include one or more bus or connection types, such as an Industry Standard Architecture (ISA) bus, an Extended ISA bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, and / or another type of bus or connection. In some embodiments, there are direct connections between components. For example, the CPU 906 can be directly connected to the main memory 904. Furthermore, the CPU 906 can be directly connected to the GPU 908.In a direct or point-to-point connection between components, the 902 connection system can include a PCIe link to establish the connection. In these examples, a PCI bus does not need to be included in the 900 computing device.
[0211] The 904 main memory can contain a variety of computer-readable media. Computer-readable media can be any available media to which the 900 computing device can access. Computer-readable media can include both volatile and non-volatile media, as well as removable and non-removable media. For example, and without limitation, computer-readable media can include computer storage media and communication media.
[0212] Computer storage media can include both volatile and non-volatile media, and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, main memory can store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other storage technologies; CD-ROM, Digital Versatile Discs (DVDs), or other optical disk storage; magnetic cartridges, magnetic tapes, magnetic disk storage, or other magnetic storage devices; or any other medium that can be used to store the desired information and that the computer can access.As used here, computer storage media do not inherently contain signals.
[0213] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include any media for transmitting information. The term "modulated data signal" can refer to a signal in which one or more of its properties are set or modified to encode information within the signal. Computer storage media can include, but are not limited to, wired media, such as a wired network or a direct-wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the above should also be included in the scope of protection of the computer-readable media.
[0214] The one or more CPUs 906 can be configured to execute at least some of the computer-readable instructions to control one or more components of the Computing Device 900 to perform one or more of the procedures and / or processes described herein. The one or more CPUs 906 can each contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a multitude of software threads simultaneously. The one or more CPUs 906 can contain any type of processor and can contain different types of processors depending on the type of Computing Device 900 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 900, the processor can be, for example, an Advanced RISC Machine (ARM) processor implemented with Reduced Instruction Set Computing (RISC), or an x86 processor implemented with Complex Instruction Set Computing (CISC). The computing device 900 can contain one or more CPUs 906, in addition to one or more microprocessors or additional coprocessors, such as mathematical coprocessors.
[0215] In addition to or as an alternative to the one or more CPUs 906, the one or more GPUs 908 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the procedures and / or processes described herein. One or more of the GPUs 908 may be an integrated GPU (e.g., with one or more of the CPUs 906) and / or one or more of the GPUs 908 may be a discrete GPU. In embodiments, one or more of the GPUs 908 may be a coprocessor of one or more of the CPUs 906. The one or more GPUs 908 may be used by the computing device 900 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, the one or more GPUs 908 may be used for general-purpose computing on GPUs (GPGPU).The one or more GPUs 908 can contain hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The one or more GPUs 908 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the one or more CPUs 906 received via a host interface). The one or more GPUs 908 can include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory can be included as part of the 904 main memory. The one or more GPUs 908 can contain two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).When combined, each GPU can generate 908 pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can have its own dedicated memory or share memory with other GPUs.
[0216] In addition to or as an alternative to the one or more CPUs 906 and / or the one or more GPUs 908, the one or more logic units 920 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the methods and / or processes described herein. In embodiments, the one or more CPUs 906, the one or more GPUs 908, and / or the one or more logic units 920 may discretely or jointly execute any combination of the methods, processes, and / or sections thereof. One or more of the logic units 920 may be part of and / or integrated into one or more of the CPUs 906 and / or one or more of the GPUs 908, and / or one or more of the logic units 920 may be discrete components or otherwise separate from the CPUs 906 and / or the GPUs 908.In embodiments, one or more of the logic units 920 can be a co-processor of one or more of the CPUs 906 and / or one or more of the GPUs 908.
[0217] Examples of one or more Logic Units 920 include one or more processing cores and / or components thereof, such as Data Processing Units (DPUs), Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), Arithmetic Logic Units (ALUs), and application-specific integrated circuits. (Application-Specific Integrated Circuits, ASICs), Floating Point Units (FPUs),Input / output (I / O) elements, peripheral component interconnect (PCI) or PCI Express (PCIe) elements, and / or similar.
[0218] The Communication Interface 910 can include one or more receivers, transmitters, and / or transceivers that enable the Computing Device 900 to communicate with other computers over an electronic network, including wired and / or wireless communication. The Communication Interface 910 can include components and functions that enable communication over a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., Ethernet or InfiniBand communication), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the one or more logic units 920 and / or the communication interface 910 may contain one or more data processing units (DPUs) to transfer data received via a network and / or via the connection system 902 directly to one or more GPUs 908 (e.g., a working memory thereof).
[0219] The I / O ports 912 enable the Computing Device 900 to be logically coupled with other devices, including the I / O Components 914, one or more Presentation Components 918, and / or other components, some of which may be built into (e.g., integrated with) the Computing Device 900. Illustrative I / O Components 914 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O Components 914 can provide a Natural User Interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, the inputs can be transmitted to a suitable network element for further processing.A NUI can implement any combination of speech capture, stylus capture, face capture, biometric capture, gesture capture (both on-screen and off-screen), air gestures, head and eye tracking, and touch capture (as further described below) associated with a display of the Computing Device 900. The Computing Device 900 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the Computing Device 900 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion detection. In some examples, the output from the accelerometers or gyroscopes can be used by the Computing Device 900 to render immersive augmented reality or virtual reality.
[0220] The power supply 916 can include a hardwired power supply, a battery power supply, or a combination of both. The power supply 916 can supply power to the computing device 900 to enable the operation of the computing device 900's components.
[0221] The one or more presentation components 918 can include a display (e.g., a monitor, a touchscreen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The one or more presentation components 918 can receive data from other components (e.g., the one or more GPUs 908, the one or more CPUs 906, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER
[0222] Fig. Figure 10 illustrates an exemplary data center 1000 that can be used in at least one embodiment of the present disclosure. The data center 1000 can include an infrastructure layer 1010, a framework layer 1020, a software layer 1030, and / or an application layer 1040.
[0223] As in Fig. As shown in Figure 10, the infrastructure layer 1010 of the data center can contain a resource orchestrator 1012, clustered computing resources 1014 and node computing resources (“node CRs”) 1016(1)-1016(N), where “N” is any positive integer. In at least one embodiment, the node CRs 1016(1)-1016(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic solid-state memory), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power supply modules and / or cooling modules, etc.In some embodiments, one or more node CRs among node CRs 1016(1)-1016(N) may correspond to a server that has one or more of the computer resources mentioned above. Furthermore, in some embodiments, node CRs 1016(1)-1016(N) may contain one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of node CRs 1016(1)-1016(N) may correspond to a virtual machine (VM).
[0224] In at least one embodiment, the grouped computing resources 1014 can contain separate groupings of node CRs 1016, which are housed in one or more racks (not shown) or in many racks in data centers at different geographic locations (also not shown). Separate groupings of node CRs 1016 within grouped computing resources 1014 can contain grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 1016, including CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide computing resources to support one or more workloads.The one or more racks can also contain any number of power supply modules, cooling modules and / or network switches in any combination.
[0225] The resource orchestrator 1012 can configure or otherwise control one or more node CRs 1016(1)-1016(N) and / or grouped computing resources 1014. In at least one embodiment, the resource orchestrator 1012 can include a software design infrastructure (SDI) management entity for the data center 1000. The resource orchestrator 1012 can include hardware, software, or a combination thereof.
[0226] In at least one embodiment, as in Fig. As shown in Figure 10, the framework layer 1020 can contain a job scheduler 1033, a configuration manager 1034, a resource manager 1036, and / or a distributed file system 1038. The framework layer 1020 can contain a framework that supports the software 1032 of the software layer 1030 and / or one or more applications 1042 of the application layer 1040. The software 1032 or the one or more applications 1042 can each contain web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1020 can be a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can use a distributed file system 1038 for processing large amounts of data (e.g., "Big Data"), without being limited to it.In at least one embodiment, the job scheduler 1033 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 1000. The configuration manager 1034 can be able to configure different layers, such as the software layer 1030 and the framework layer 1020, which contains Spark and the distributed file system 1038, to support the processing of large amounts of data. The resource manager 1036 can be able to manage clustered or grouped computing resources allocated or assigned to support the distributed file system 1038 and the job scheduler 1033. In at least one embodiment, the clustered or grouped computing resources can include the grouped computing resource 1014 on the infrastructure layer 1010 of the data center.The Resource Manager 1036 can coordinate with the Resource Orchestrator 1012 to manage these allocated or assigned computer resources.
[0227] In at least one embodiment, the software contained in software layer 1030 may include software 1032 that is used by at least sections of the node CRs 1016(1)-1016(N), the grouped computing resources 1014, and / or the distributed file system 1038 of framework layer 1020. One or more types of software may include, among others, web search software, email virus scanning software, database software, and streaming video content software.
[0228] In at least one embodiment, the applications 1042 contained in the application layer 1040 may include one or more types of applications used by at least sections of the node CRs 1016(1)-1016(N), the grouped computing resources 1014, and / or the distributed file system 1038 of the framework layer 1020. One or more types of applications may include, but are not limited to, any number of genome applications, cognitive computations, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0229] In at least one embodiment, the configuration manager 1034, the resource manager 1036, and / or the resource orchestrator 1012 can implement any number and type of self-modifying actions based on any set and type of data acquired in any technically feasible way. Self-modifying actions can relieve a data center operator of data center 1000 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning sections of a data center.
[0230] The Data Center 1000 may contain tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weighting parameters according to a neural network architecture, using software and / or computing resources described above in relation to the Data Center 1000.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer or predict information using the resources described above with reference to the data center 1000 by using weighting parameters calculated by one or more training techniques such as those described herein, without limitation.
[0231] In at least one embodiment, the data center can use 1000 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or infer information, such as image capture, speech capture, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS
[0232] Network environments suitable for implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may run on one or more instances of the one or more computing devices 900 of Fig. 9 can be implemented - for example, each device can contain similar components, features, and / or functionality to one or more computing devices 900. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices can also be included as part of a data center 1000, an example of which is given herein with reference to Fig. 10 is described in more detail.
[0233] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can contain multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0234] Compatible network environments can contain one or more peer-to-peer network environments—in which case a server cannot be included in a network environment—and one or more client-server network environments—in which case one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described here can be implemented on any number of client devices with reference to one or more servers.
[0235] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework to support software of a software layer and / or one or more applications of an application layer. The software or the one or more applications may each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework that uses, for example, a distributed file system for processing large amounts of data (e.g., "Big Data"), but is not limited to this.
[0236] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or one or more parts) of the computing and / or data storage functions described herein. Each of these different functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers, which may be distributed across a state, region, country, the globe, etc.). If a connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers can offload at least some functionality to the one or more edge servers. A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0237] The one or more client devices can incorporate at least some of the components, features, and functions of the one or more devices referred to here. Fig.The exemplary computing devices described in Section 9 include 900. By way of example, and not as a limitation, a client device may be a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smartwatch, a portable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or global positioning device, a video player, a video camera, a surveillance device or surveillance system, a vehicle, a boat, a hydrofoil, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or gaming system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, a device, a consumer electronics device, a workstation, an edge device,any combination of these described devices or any other suitable device may be embodied.
[0238] The disclosure can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules that are executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules, which contain routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements certain abstract data types. The disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc.The revelation can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are connected to each other via a network for communication.
[0239] As used herein, any mention of "and / or" in relation to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0240] The subject matter of this disclosure is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of protection afforded by this disclosure. Rather, the inventors have considered that the claimed subject matter may also be embodied in other ways to include various steps or combinations of steps similar to those described in this document, in conjunction with other present or future technologies. Although the terms “step” and / or “block” may be used herein to denote various elements of the methods employed, these terms should not be interpreted as implying any particular sequence among or between the various steps disclosed herein, except where the sequence of each step is expressly described.
[0241] The disclosure of this application also includes the following numbered clauses: Clause 1 Method, comprising: Determining one or more first state vectors associated with one or more first Kalman filters, wherein the one or more first state vectors contain one or more first values corresponding to one or more first Bézier representations associated with one or more first sections of one or more paths; Determining one or more second state vectors associated with one or more second Kalman filters, wherein the one or more second state vectors contain one or more second values corresponding to one or more second Bézier representations associated with one or more second sections of the one or more paths;Computation of one or more geometries associated with the one or more paths, based on at least the one or more first state vectors, the one or more second state vectors, and data specifying at least one relative motion associated with a machine; and causing the machine to perform one or more operations, based on at least the one or more geometries associated with the one or more paths. Clause 2 Method according to Clause 1, wherein the one or more paths include at least one first path and one or more second paths, wherein the first path corresponds to a first lane of a driving surface occupied by the machine, and the one or more second paths correspond to one or more second lanes on the driving surface. Clause 3 Procedure according to one of Clauses 1-2, wherein: the one or more first sections of the one or more paths correspond to at least one or more first edges assigned to the first lane and one or more second edges assigned to the one or more second lanes, and the one or more second sections of the one or more paths correspond to at least one first center line assigned to the first lane and one or more second center lines assigned to the one or more second lanes. Clause 4 Method according to one of Clauses 1-3, wherein the one or more first Bézier representations and the one or more second Bézier representations contain at least one three-dimensional (3D) Bézier curve representing a geometry that is associated with a section of a path of the one or more paths. Clause 5 Procedure according to any of Clauses 1-4, wherein: the one or more first values contained in the one or more first state vectors correspond to a first dimension associated with one or more 3D control points for the 3D Bézier curve, and the one or more second values contained in the one or more second state vectors correspond to a second dimension associated with the one or more 3D control points for the 3D Bézier curve. Clause 6: Method according to any of Clauses 1-5, further comprising: applying, to one or more machine learning models, sensor data generated using one or more sensors associated with the machine, wherein the sensor data specify at least the relative motion associated with the machine; determining one or more first updated state vectors, based at least on updating the one or more first values using one or more first outputs of the one or more machine learning models;and determining one or more second updated state vectors based on at least updating the one or more second values using one or more second outputs of the one or more machine learning models, wherein the computation of the one or more geometries associated with the one or more paths is based on at least the one or more first updated state vectors and the one or more second updated state vectors. Clause 7 System, comprising: one or more processors, for: obtaining one or more first points corresponding to one or more first Bézier representations associated with one or more paths in an environment; calculating, based on at least the one or more first points and a relative motion associated with a machine, one or more second points corresponding to one or more second Bézier representations associated with the one or more paths; and performing one or more operations associated with the machine, based on at least the one or more second Bézier representations. Clause 8 System according to Clause 7, wherein the one or more first Bézier representations and the one or more second Bézier representations correspond to one or more three-dimensional (3D) Bézier curves representing one or more 3D geometries associated with the one or more paths. Clause 9 System according to one of Clauses 7-8, wherein the one or the several first Bézier representations contain at least one first Bézier curve and one or more second Bézier curves, wherein the first Bézier curve corresponds to a first section of at least one path of the one or the several paths and the one or the several second Bézier curves correspond to one or more second sections of the path. Clause 10 system according to one of clauses 7-9, wherein the first section of the path is a center line associated with the path, and the one or more second sections are one or more edges associated with the path. Clause 11 System according to one of Clauses 7-10, wherein the one or more paths correspond to one or more lanes assigned to a driving surface in the surrounding area, wherein the one or more lanes include at least one first lane and one or more second lanes. Clause 12 System according to one of Clauses 7-11, wherein the one or more first Bézier representations are assigned to one or more first sections of the one or more paths and the one or more second Bézier representations are assigned to one or more second sections of the one or more paths. Clause 13 System according to one of Clauses 7-12, wherein obtaining the one or more first points corresponding to the one or more first Bezier representations comprises: obtaining one or more first state vectors containing one or more first values representing one or more first coordinate locations of the one or more first points with respect to a first dimension of a multidimensional space; and obtaining one or more second state vectors containing one or more second values representing one or more second coordinate locations of the one or more first points with respect to a second dimension of the multidimensional space. Clause 14 System according to any of Clauses 7-13, wherein the one or more processors further serve to: apply, to one or more machine learning models, sensor data generated using one or more sensors associated with the machine; and compute, based on at least one or more outputs of the one or more machine learning models, one or more updated versions of the one or more second points corresponding to one or more updated Bezier representations associated with the one or more paths, wherein the execution of the one or more operations associated with the machine is further based on at least one or more updated Bezier representations. Clause 15 System according to one of Clauses 7-14, wherein the one or more first Bezier representations include at least one first group of multidimensional Bezier curves corresponding to one or more first sections of a first path in the environment, and one or more second groups of multidimensional Bezier curves corresponding to one or more second sections of one or more second paths in the environment. Clause 16 System according to one of Clauses 7-15, wherein the one or more processors further serve to: apply, to one or more values associated with the one or more first points, one or more displacement matrices determined on the basis of at least one or more polyline points associated with the one or more first Bezier representations; and wherein the calculation of the one or more second points is further based on at least the application of the one or more displacement matrices. Clause 17 System according to any of Clauses 7-16, wherein the system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models;a system for performing operations using one or more vision language models (VLMs); a system for performing operations using one or more multimodal language models; a system for performing one or more operations using conversational AI; a system for generating synthetic data; a system for presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources. Clause 18 At least one processor comprising: processing circuitry for performing one or more operations associated with a machine, based on at least one or more curves associated with one or more segments of one or more paths in an environment, wherein the one or more curves are determined using one or more Kalman filters to perform at least one of tracking or predicting one or more three-dimensional (3D) control coordinates corresponding to the one or more curves. Clause 19 Processor according to Clause 18, wherein the one or more curves comprise at least the following: one or more first 3D Bézier curves representing one or more first 3D geometries associated with a first path of the one or more paths; and one or more second 3D Bézier curves representing one or more second 3D geometries associated with one or more second paths of the one or more paths. Clause 20 Processor according to any of Clauses 18-19, wherein the processor comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models;a system for performing operations using one or more vision language models (VLMs); a system for performing operations using one or more multimodal language models; a system for performing one or more operations using conversational AI; a system for generating synthetic data; a system for presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources.
[0242] It is understood that the aspects and embodiments described above are purely exemplary and that modifications of details may be made within the scope of protection of the claims.
[0243] Each device, each method and each feature disclosed in the description, and (where applicable) the claims and drawings, may be provided independently or in any suitable combination.
[0244] Reference numerals appearing in the claims are for illustrative purposes only and do not restrict the scope of protection of the claims. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 16 / 101,232
[0143] Cited non-patent literature
[0000] Society of Automotive Engineers, SAE) (Standard No. J3016-201806, published on June 15, 2018, Standard No. J3016-201609, published on September 30, 2016
[0101]
Claims
[1] Procedure, encompassing: Determining one or more first state vectors associated with one or more first Kalman filters, wherein the one or more first state vectors contain one or more first values corresponding to one or more first Bezier representations associated with one or more first sections of one or more paths; Determining one or more second state vectors associated with one or more second Kalman filters, wherein the one or more second state vectors contain one or more second values corresponding to one or more second Bézier representations associated with one or more second sections of the one or more paths; Computation of one or more geometries associated with one or more paths, based on at least one or more first state vectors, one or more second state vectors, and data specifying at least one relative motion associated with a machine; and Causing the machine to perform one or more operations based on at least the one or more geometries associated with the one or more paths. [2] Method according to claim 1, wherein the one or more paths comprise at least one first path and one or more second paths, wherein the first path corresponds to a first lane of a driving surface occupied by the machine, and the one or more second paths correspond to one or more second lanes on the driving surface. [3] Method according to claim 2, wherein: the one or more first sections of the one or more paths correspond to at least one or more first edges assigned to the first lane and one or more second edges assigned to the one or more second lanes, and the one or more second sections of the one or more paths correspond to at least one first center line assigned to the first lane and one or more second center lines assigned to the one or more second lanes. [4] Method according to one of the preceding claims, wherein the one or more first Bézier representations and the one or more second Bézier representations include at least one three-dimensional (3D) Bezier curve representing a geometry that is associated with a section of a path of the one or more paths. [5] Method according to claim 4, wherein: the one or more first values contained in the one or more first state vectors correspond to a first dimension that is associated with one or more 3D control points for the 3D Bezier curve, and the one or more second values contained in the one or more second state vectors correspond to a second dimension that is assigned to the one or more 3D control points for the 3D Bézier curve. [6] Method according to any one of the preceding claims, further comprising: Applying, to one or more machine learning models, sensor data generated using one or more sensors associated with the machine, wherein the sensor data specify at least the relative motion associated with the machine; Determining one or more first updated state vectors, based at least on updating one or more first values using one or more first outputs of one or more machine learning models; and Determining one or more second updated state vectors based on at least updating one or more second values using one or more second outputs of one or more machine learning models, wherein the calculation of the one or more geometries associated with the one or more paths is based at least on the one or more first updated state vectors and the one or more second updated state vectors. [7] System encompassing: one or more processors, for example: Obtaining one or more first points corresponding to one or more first Bezier representations associated with one or more paths in an environment; Calculate, based on at least one or more first points and a relative motion associated with a machine, one or more second points corresponding to one or more second Bézier representations associated with the one or more paths; and Performing one or more operations assigned to the machine, based on at least one or more second Bézier representations. [8] System according to claim 7, wherein the one or more first Bézier representations and the one or more second Bézier representations correspond to one or more three-dimensional (3D) Bezier curves representing one or more 3D geometries associated with the one or more paths. [9] System according to claim 7 or claim 8, wherein the one or more first Bezier representations include at least one first Bezier curve and one or more second Bezier curves, wherein the first Bezier curve corresponds to a first section of at least one path of the one or more paths and the one or more second Bezier curves correspond to one or more second sections of the path. [10] System according to claim 9, wherein the first section of the path is a center line associated with the path, and the one or more second sections are one or more edges associated with the path. [11] System according to one of claims 7-10, wherein the one or more paths correspond to one or more lanes associated with a driving surface in the surrounding area, wherein the one or more lanes include at least one first lane and one or more second lanes. [12] System according to one of claims 7-11, wherein the one or more first Bézier representations are assigned to one or more first sections of the one or more paths and the one or more second Bézier representations are assigned to one or more second sections of the one or more paths. [13] System according to any of claims 7-12, wherein obtaining one or more first points corresponding to one or more first Bézier representations comprises: Obtaining one or more first state vectors containing one or more first values representing one or more first coordinate locations of the one or more first points with respect to a first dimension of a multidimensional space; and Obtain one or more second state vectors containing one or more second values representing one or more second coordinate locations of the one or more first points with respect to a second dimension of the multidimensional space. [14] System according to any one of claims 7-13, wherein the one or more processors further serve to: Applying sensor data generated using one or more sensors assigned to the machine to one or more machine learning models; and Calculate, based on at least one or more outputs of the one or more machine learning models, one or more updated versions of the one or more second points corresponding to one or more updated Bézier plots associated with the one or more paths, wherein the execution of the one or more operations assigned to the machine is furthermore based on at least one or more updated Bézier representations. [15] System according to one of claims 7-14, wherein the one or more first Bezier representations include at least one first group of multidimensional Bezier curves corresponding to one or more first sections of a first path in the environment, and one or more second groups of multidimensional Bezier curves corresponding to one or more second sections of one or more second paths in the environment. [16] System according to any one of claims 7-15, wherein the one or more processors further serve to: Applying to one or more values associated with the first one or more points, one or more displacement matrices determined based on at least one or more polyline points associated with the first one or more Bezier representations; and wherein the calculation of one or more second points is furthermore based on at least the application of one or more displacement matrices. [17] System according to any one of claims 7 to 16, wherein the system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system that is implemented using an edge device; a system that is implemented using a robot; a system for performing one or more operations using generative AI; a system for performing operations using one or more large language models; a system for performing operations using one or more Vision Language Models (VLMs); a system for performing operations using one or more multimodal language models; a system for performing one or more operations using conversational AI; a system for generating synthetic data; a system for presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [18] At least one processor, comprising: Processing circuits to perform one or more operations associated with a machine, based on at least one or more curves associated with one or more sections of one or more paths in an environment, wherein the one or more curves are determined using one or more Kalman filters to perform at least one of tracking or predicting one or more three-dimensional (3D) control coordinates corresponding to the one or more curves. [19] Processor according to claim 18, wherein the one or more curves comprise at least the following: one or more first 3D Bézier curves representing one or more first 3D geometries associated with a first path of one or more paths; and one or more second 3D Bézier curves representing one or more second 3D geometries associated with one or more second paths of the one or more paths. [20] Processor according to claim 18 or claim 19, wherein the processor comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system that is implemented using an edge device; a system that is implemented using a robot; a system for performing one or more operations using generative AI; a system for performing operations using one or more large language models; a system for performing operations using one or more Vision Language Models (VLMs); a system for performing operations using one or more multimodal language models; a system for performing one or more operations using conversational AI; a system for generating synthetic data; a system for presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources.
Citation Information
Patent Citations
Method for programmable timeouts of tree traversal mechanisms in hardware
US10885698B2
16/101,232