Differential and modular end-to-end stacks for autonomous systems and applications

By employing a differentiable and modular end-to-end stack in an autonomous machine system, and utilizing machine learning model sequences to achieve interpretability of perception and motion prediction and gradient backpropagation, the problem of reduced interpretability and versatility in modular architectures is solved, thereby improving the decision reliability of safety-critical applications.

CN121189413APending Publication Date: 2025-12-23NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510818010.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-21
Filing Date
2025-06-18
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

In autonomous machine systems, modular architectures suffer from inter-module composite errors, information bottlenecks, and integration challenges, leading to reduced interpretability and versatility, especially affecting the reliability of safety decisions in safety-critical applications.

Method used

Employing a differentiable and modular end-to-end stack, this system achieves interpretability of perception and motion prediction and gradient backpropagation through a sequence of machine learning models. This allows for interpretability of the output and gradient backpropagation, enabling downstream decisions to learn from upstream predictions.

Benefits of technology

It achieves interpretability and reusability of modular architecture in autonomous systems, while ensuring security and performance reliability, especially providing security assurance in security decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121189413A_ABST
    Figure CN121189413A_ABST
Patent Text Reader

Abstract

The invention relates to differentiable and modular end-to-end stacks for autonomous systems and applications. In various examples, the control stack may include a machine learning model (MLM) sequence that predicts differentiable output sequences, respectively, to determine one or more control sequences. The disclosed method can be used to implement a differentiable and modular end-to-end AV stack that allows interpretability of the output and back propagation of the gradient to learn upstream predictions with respect to downstream decisions. Various methods are provided for interfacing perception with motion prediction in a differentiatable manner, and interfacing motion prediction with motion planning and motion control in a differentiatable manner.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Autonomous machines, such as intelligent robotic systems or autonomous vehicles (AVs), are typically architectured in a modular fashion, for example, to perform detection, tracking, prediction, planning, and control. Modular architectures can offer high levels of reusability, interpretability, and versatility. However, they can also be prone to inter-module composite errors, information bottlenecks, and integration challenges. To overcome these challenges, the AV stack has been transformed into an end-to-end neural network. This approach benefits from eliminating information bottlenecks and scales performance as the dataset size increases. However, reusability, interpretability, and versatility are significantly reduced compared to modular architectures. In particular, interpretability is crucial for debugging and verification, and is essential for providing assurance for safety-critical applications such as AVs, where assurance is paramount for reliably making safety decisions using the AV stack. Summary of the Invention

[0002] Embodiments of this disclosure relate to differentiable and modular end-to-end stacks for autonomous and semi-autonomous systems and applications. The disclosed systems and methods can be used to implement autonomous driving stacks with modular architectures that have interpretable outputs while allowing upstream perception and prediction to be trained for downstream control objectives.

[0003] Compared to conventional systems, aspects of this disclosure provide control stacks for machines, such as autonomous vehicles (AVs), with sequences of machine learning models (MLMs) that predict differentiable output sequences to determine one or more control sequences. The disclosed methods can be used to implement differentiable and modular end-to-end AV stacks—allowing for interpretability of outputs and backpropagation of gradients, enabling upstream predictions to be learned from downstream decisions. Therefore, the systems and methods described herein provide various approaches for differentiably interfaced perception with motion prediction, and for differentiably interfaced motion prediction with motion planning and motion control. Attached Figure Description

[0004] The present system and method, with reference to the accompanying drawings, provide a detailed description of a differentiable and modular end-to-end stack for autonomous and semi-autonomous systems and applications, wherein:

[0005] Figure 1A Data flow diagrams of examples of processes for a differentiable and modular end-to-end stack for an autonomous or semi-autonomous machine, according to some embodiments of this disclosure.

[0006] Figure 1BIncludes example data flow diagrams of a process for training a differentiable and modular end-to-end stack for autonomous or semi-autonomous machines, according to some embodiments of this disclosure.

[0007] Figure 2 Includes data flow diagrams of examples of processes for performing object detection and tracking using a differentiable and modular end-to-end stack for autonomous and semi-autonomous machines, according to some embodiments of this disclosure;

[0008] Figure 3A Includes a data flow diagram of an example of a process for interfacing a tracker with a motion predictor using a differentiable and modular end-to-end stack for autonomous and semi-autonomous machines, according to some embodiments of this disclosure.

[0009] Figure 3B Includes a data flow diagram of an example of a process for interfacing a tracker with a motion predictor using a differentiable and modular end-to-end stack for autonomous and semi-autonomous machines, according to some embodiments of this disclosure.

[0010] Figure 4 Example illustrations of an environment including a predicted location distribution pattern with an agent and a corresponding planned trajectory of a machine, according to some embodiments of this disclosure;

[0011] Figure 5 This is a flowchart illustrating a method for controlling a machine using a differentiable end-to-end autonomous or semi-autonomous vehicle stack, according to some embodiments of this disclosure.

[0012] Figure 6 This is a flowchart illustrating a method for controlling a machine using a sequence of machine learning models (MLMs) that predict differentiable output sequences, according to some embodiments of the present disclosure.

[0013] Figure 7A These are illustrations of example autonomous vehicles according to some embodiments of the present disclosure;

[0014] Figure 7B According to some embodiments of this disclosure Figure 7A Examples of camera positions and fields of view for autonomous vehicles;

[0015] Figure 7C According to some embodiments of this disclosure Figure 7A A block diagram of an example system architecture for an example autonomous vehicle;

[0016] Figure 7D Cloud-based servers and according to some embodiments of this disclosure Figure 7A A schematic diagram of a system for communication between autonomous vehicles;

[0017] Figure 8 This is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and

[0018] Figure 9 This is a block diagram of an example data center applicable to implementing some embodiments of this disclosure. Detailed Implementation

[0019] Systems and methods relating to differentiable and modular end-to-end stacks for autonomous or semi-autonomous systems and applications are disclosed. Although this disclosure may relate to an example autonomous or semi-autonomous vehicle or machine 700 (which may alternatively be referred to herein as "vehicle 700," "self-vehicle 700," "machine 700," or "self-machine 700"), examples of which relate to... Figures 7A-7D The description herein is intended to be limiting. For example, the systems and methods described herein may be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Furthermore, while this disclosure may describe the control operations of determining machines (such as autonomous vehicles), this is not intended to be limiting, and the systems and methods described herein can be used in augmented reality, virtual reality, mixed reality, robotics, security and supervision, simulated environments (e.g., NVIDIA's DriveSIM), autonomous or semi-autonomous machine applications, and / or any other technological space that may use physical movement for evaluation.

[0020] Compared to conventional systems such as those described above, aspects of this disclosure provide a control stack for machines such as autonomous vehicles (AVs) having a sequence of machine learning models (MLMs) that each predict a differentiable sequence of outputs to determine one or more control sequences. The disclosed methods can be used to implement differentiable and modular end-to-end AV stacks—allowing for interpretability of outputs and backpropagation of gradients, enabling upstream predictions to be learned with respect to downstream decision-making.

[0021] In at least one embodiment, for perception, one or more first MLMs can be used to determine correspondence data between one or more object detections and one or more object tracks, wherein the correspondence data is differentiable with respect to one or more object detections and one or more object tracks. This correspondence data can be used to update one or more object tracks. Various methods can be used to interface perception with one or more second MLMs for motion prediction in a differentiable manner. In at least one embodiment, a combined solver uses the correspondence data to associate object detections with object tracks (e.g., corresponding to one or more previous frames) to determine updated object tracks. The updated object tracks can be applied to one or more second MLMs for motion prediction. When the combined solver associates object detections with object tracks in a non-differentiable manner, a portion of the computation graph may be non-differentiable, but still provides an overall differentiable end-to-end stack. In at least one embodiment, to increase the differentiability of the computation graph, a differentiable combinatorial solver can be used to associate object detection with object traces, and / or object detection and / or correspondence data can be applied to one or more second MLMs for motion prediction (e.g., instead of updated object traces).

[0022] In another aspect, at least one MLM of the motion planner may include at least one analytical function for: determining candidate trajectories for the machine based at least on future movements predicted using motion predictions, calculating cost values ​​for the candidate trajectories, and selecting a reference trajectory from the candidate trajectories. To provide differentiability, a term of the analytical function may correspond to the predicted future movements. The reference trajectory may be used by the motion controller to compute a control sequence for the machine using at least one analytical function of at least one MLM trained to generate predictions corresponding to control sequences, wherein these predictions are differentiable with respect to at least one parameter of the MLM used in each module of the control stack.

[0023] The systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles connected to one or more trailers, aircraft, ships, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a variety of purposes, including, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0024] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, marine systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems for implementing language models (such as large language models (LLM), visual language models (VLM), multimodal language models, etc.), systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0025] refer to Figure 1A , Figure 1AThis includes an example data flow diagram of process 100 for a differentiable and modular end-to-end stack for an autonomous or semi-autonomous machine according to some embodiments of this disclosure. It should be understood that such and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used in addition to or as alternatives to the arrangements and elements shown, and certain elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and can be implemented in any suitable combination and location. The various functions performed by the entities described herein can be performed by hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can use... Figures 7A-7D Examples of autonomous or semi-autonomous vehicles or machines 700 Figure 8 Example computing devices 800 and / or Figure 9 The example data center 900 uses components, features, and / or functions similar to those of other components, features, and / or functions to perform the same task.

[0026] In summary, the process may include: detector 104 receiving sensor data 102, which is used to generate and / or determine predictions corresponding to one or more object detections of one or more objects in the environment. One or more trackers 106 may use one or more object detections to generate and / or determine predictions corresponding to one or more object tracks and / or tracked objects or trajectories in the environment. Motion predictor 108 may use one or more object detections, one or more object tracks, and / or other data (e.g., correspondence data between object detections and object tracks) to generate and / or determine predictions (e.g., extended and / or future trajectory and / or location distribution) corresponding to future movements associated with one or more object detections. Motion planner 110 may use future movements to generate and / or determine at least one trajectory and / or future movement of the self-machine. Motion controller 112 may use at least one trajectory to generate a fourth prediction for one or more control sequences for the self-machine. Control component 114 may use one or more control sequences to perform one or more control operations on the machine.

[0027] In the various examples, each component or module 104-112 may include at least one machine learning model with at least one learned parameter, so that the component or module can be trained in an end-to-end manner. Therefore, the outputs of components 104-112 can be interpreted individually and provide formal guarantees to the stack, while allowing upstream perception and prediction to be trained with respect to downstream control objectives.

[0028] Process 100 may include generating and / or receiving sensor data 102 obtained using one or more sensors. In one or more embodiments, the sensors may include at least one of one or more physical sensors in a physical environment or one or more virtual sensors in a simulated environment. For example, one or more sensors may correspond to a physical or simulated version of vehicle 700, as described herein.

[0029] Sensor data 102 may include, but is not limited to, sensor data 102 from any sensor of vehicle 700 (and / or other vehicles or objects, in some examples such as robotic devices, VR systems, AR systems, etc.). For example, refer to Figures 7A-7C Sensor data 102 may include data generated by or using the following sensors (but is not limited to): Global Navigation Satellite System (GNSS) sensor 758 (e.g., Global Positioning System sensor, Differential GPS (DGPS) etc.), RADAR sensor 760, ultrasonic sensor 762, LIDAR sensor 764, Inertial Measurement Unit (IMU) sensor 766 (e.g., accelerometer, gyroscope, magnetic compass, magnetometer, etc.), microphone 796, stereo camera 768, wide-angle camera 770 (e.g., fisheye camera), infrared camera 772, surround camera 774 (e.g., 360-degree camera), long-range and / or medium-range camera 798, speed sensor 744 (e.g., for measuring the speed and / or distance traveled by vehicle 700) and / or other sensor types.

[0030] In some examples, sensor data 102 may include sensor data generated using one or more forward-facing sensors, lateral sensors, and / or rearward-facing sensors. This sensor data 102 can be used to identify, detect, classify, and / or track movement of objects around vehicle 700 in the environment. In embodiments, any number of sensors can be used to encompass multiple fields of view (e.g., Figure 7B The field of view of the remote camera 798, the forward stereo camera 768 and / or the forward wide-angle camera 770) and / or the sensing field (e.g., the sensing field of the LiDAR sensor 764, the RADAR sensor 760 and so on).

[0031] Sensor data 102 may include image data representing images, image data representing video (e.g., a snapshot of video), data representing the sensor's sensing field (e.g., a depth map of a LiDAR sensor, a value map of an ultrasonic sensor, etc.), and / or data representing the sensor's measurement results. When sensor data 102 includes image data, any type of image data format may be used, such as (but not limited to) compressed images (e.g., Joint Image Experts Group (JPEG) or Luminosity / Chromatography (YUV) formats), compressed images as frames derived from compressed video formats (e.g., H.264 / Advanced Video Coding (AVC) or H.265 / High-Efficiency Video Coding (HEVC)), raw images (e.g., raw images derived from Red-to-Blue (RCCB), Red-to-Blue (RCCC), or other types of imaging sensors), and / or other formats. Furthermore, in some examples, sensor data 102 can be used in process 100 without any preprocessing (e.g., in raw or captured format), while in other examples, sensor data 102 may undergo preprocessing (e.g., noise balancing, depigmentation, scaling, cropping, enhancement, white balance, tone curve adjustment, etc., such as using a sensor data preprocessor (not shown)). As used herein, sensor data 102 may refer to unprocessed sensor data, preprocessed sensor data, or a combination thereof.

[0032] Sensor data 102 may be used at least in part by detector 104 to generate and / or determine one or more detection results for one or more entities (such as self-actualizing actors and / or other actors or entities (objects) or characteristics of the environment). Detection results may correspond to one or more states of the environment, where the states of the environment may correspond to one or more specific times or time steps. For example, detector 104 may be trained to predict or detect one or more state parameters of actors in the environment (e.g., vehicle 700 and other static or dynamic objects). In at least one embodiment, detector 104 may determine one or more control actions taken by one or more actors based at least on one or more states of the actors.

[0033] For example, the control actions of an actor may include one or more parameters corresponding to steering and / or acceleration. In at least one embodiment, the control actions may include one or more control variables corresponding to heading rate and / or longitudinal acceleration. The state of each entity or actor typically includes one or more of the following: position, velocity, direction or heading (e.g., direction of travel), rate, acceleration (e.g., scalar, rotation, etc.), attitude (e.g., orientation), and / or other information about the state of the actor or object. For example, the state may encode or represent the actor's position (e.g., (x,y) coordinates), the actor's unit orientation, and / or the actor's scalar velocity at a given time in two-dimensional space. In some examples, the state may encode or represent additional or alternative information, such as rotational velocity (e.g., yaw) and / or scalar acceleration in any direction, and / or any other abstract information associated with the entity, such as appearance, category, associated object, associated intent, state, etc. In at least one embodiment, the distance between states may be measured at least based on the 2D Euclidean distance between states. In at least one embodiment, the distance between trajectories may be based at least on the root mean square state distance over time.

[0034] Detector 104 may use any combination of sensors (such as GNSS sensor 758, IMU sensor 766, speed sensor 744, steering sensor 740, etc.) to determine one or more parameters of the state and / or one or more corresponding control actions. In at least one embodiment, detector 104 may use any combination of stereo camera 768, wide-angle camera 770, infrared camera 772, surround camera 774, long-range and / or mid-range camera 798, LiDAR sensor 764, RADAR sensor 760, microphone 796, ultrasonic sensor 762 and / or other sensors of vehicle 700 to determine and / or infer the state of objects in the environment (e.g., other than vehicle 700). In some examples, the state of an object (e.g., when one or more objects are another vehicle or a person using a client device capable of wireless communication) may be determined using wireless communication (e.g., vehicle-to-vehicle communication or device-to-vehicle communication) over one or more networks (such as, but not limited to, the networks described herein).

[0035] In at least one embodiment, detector 104 may be trained to detect and / or determine one or more features of the environmental state, for example, to provide context (e.g., semantic information) to the state of an entity. Examples of one or more features include road geometry features, road characteristic features (e.g., signs, road types, road markings, road conditions, etc.), weather features, visibility features, and / or other external features that may affect the control actions of at least one entity. In at least one embodiment, one or more detections may correspond to one or more driving maneuvers and / or the type of driving maneuver performed by one or more actors, such as lane changing maneuvers, overtaking maneuvers, following maneuvers, parking maneuvers, etc. In at least one embodiment, one or more detections may be assigned to one or more scenarios and / or associated with one or more scenarios. For example, a scenario may be defined using one or more parameters indicating one or more environmental features and / or driving maneuvers.

[0036] In at least one embodiment, a partially observable Markov decision process (POMDP) ​​can be used to model the dynamics of vehicle 700 and the scene surrounding vehicle 700. The POMDP can use tuples... To define it. As described in this paper, the state space S can include, for example, self-agents s. e Non-self-agent or entity s ne The states of other variables or parameters, such as those related to the environment map. m Corresponding variables or parameters. Observation space. This can refer to the space of observations received by vehicle 700 from detector 104 (e.g., corresponding to sensor data 102). Additionally, it refers to the control input space. It can refer to the space of the control input u of vehicle 700, the function f(s) t |s t-1 ,u t-1 () can refer to a random state transition function for a time instance or time step t.

[0037] In some examples, machine learning models (such as neural networks, e.g., convolutional neural networks) can be used to determine or detect parameters of the state of control actions and / or actors and / or the environment. For example, sensor data from sensors of vehicle 700 can be applied to one or more machine learning models to determine the state of objects and / or the environment. Neural networks can perform various functions on processed and / or unprocessed data. For example, but not limited to, convolutional neural networks can be used for object detection and recognition (e.g., using sensor data from cameras of vehicle 700), one or more convolutional neural networks can be used for distance estimation, object detection, object position detection and / or object pose detection or determination (e.g., using sensor data from cameras of vehicle 700), one or more convolutional neural networks can be used for emergency vehicle detection and recognition (e.g., using sensor data from microphones of vehicle 700), one or more convolutional neural networks can be used to identify and process safety events and / or safety-related events, and / or other machine learning models (MLMs) can be used. In examples using convolutional neural networks, any type of convolutional neural network can be used, including region-based convolutional neural networks (R-CNN), fast R-CNN, and / or other types. In addition to CNNs, or as an alternative to CNNs, it can be used to implement any other type of machine learning model.

[0038] For example, but not limited to, any of the various MLMs described herein may include one or more of any type of machine learning model, such as machine learning models using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), K-means clustering, control barrier functions, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., one or more autoencoders, convolutions, recurrent neural networks, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid machines, transformers, large language models, visual language models, multimodal language models, etc.), and / or other types of machine learning models and / or computer vision algorithms.

[0039] In embodiments where sensor data 102 at least partially corresponds to simulated sensor data, the simulated sensor data can be generated using one or more simulators. For example, the simulated sensor data may correspond to simulated data generated using a simulation application, such as an autonomous vehicle driving simulator, like NVIDIA's DriveSIM. In embodiments, the simulation may be generated or instantiated in an OMNIVERSE or METAVERSE environment, and / or one or more ray tracing or light transport algorithms may be used to generate more realistic lighting and shadows in the simulated environment.

[0040] Simulated data may include snapshots, images, samples, and / or other data about the state of the simulated or virtual world in each frame. For example, simulated sensor data may include information about an actor's position in the world, their velocity, acceleration, attitude, etc., information about the state of traffic lights or signals, information about the positions of traffic signs, stop lines, etc. The world state may be perceived by vehicle 700, other vehicles, and / or other systems.

[0041] In one or more embodiments, simulation data can be generated and / or analyzed based on one or more scenarios as described herein. For example, one or more scenarios of interest can be hard-coded, manually created, programmatically generated, sporadic, or otherwise represented in a virtual and / or computational environment. One or more observations can be determined from simulation data corresponding to one or more scenarios. For example, one or more scenarios can be assigned to one or more observations for learning driving behavior for those scenarios.

[0042] In at least one embodiment, at least a portion of detector 104 may be included in (e.g., of vehicle 700) a sensing component, module, system, and / or block. For example, detector 104 may provide one or more outputs of a sensing module and / or data for generating one or more outputs of the sensing module. In at least one embodiment, the sensing module may in the future include tracker 106 to generate one or more outputs.

[0043] Now for reference Figure 2 , Figure 2 This includes example data flow diagrams of a process 200 for performing object detection and tracking using a differentiable and modular end-to-end stack for autonomous machines, according to some embodiments of this disclosure. For example, but not limited to, detector 104 may include one or more encoders 204 and one or more decoders 206. In at least one embodiment, detector 104 may receive sensor measurements or observations corresponding to sensor data 102 and apply the sensor data 102 to encoder 204 to encode environmental features 210. The encoded environmental features 210 may be applied to decoder 206 to decode at least object detection results 212.

[0044] In at least one embodiment, detector 104 uses one or more backbone networks to extract environment features 210 and / or one or more portions thereof. For example, a LiDAR backbone may be used in conjunction with a sparse voxel encoder and a feature pyramid network to transform the point cloud into a bird's-eye view (BEV) feature space for environment features 210. For image data, multi-view images may be encoded into a perspective feature space for environment features 210 using a CNN such as a residual network (e.g., ResNet-101). In at least one embodiment, environment features 210 from multiple modalities and / or feature spaces may be fused to generate object detection result 212. For example, object features 214 may collect information from multiple feature spaces (e.g., perspective feature space and BEV feature space), and object features 214 may be processed by decoder 206 (e.g., a transformer decoder) to generate object detection result 212. In at least one embodiment, object detection result 212 includes 3D bounding boxes, but more generally may include information indicating one or more locations (e.g., 2D, 3D, etc.) of one or more objects in the environment.

[0045] Now for reference Figure 1B and Figure 2 , Figure 1B This includes an example data flow diagram of process 150 for training an autonomous machine using a differentiable and modular end-to-end stack according to some embodiments of this disclosure. Figure 2 As shown, a detection loss function can be used. To train one or more MLM 154s (such as encoder 204 and decoder 206), the detection loss function For example, it can be defined according to equation (1):

[0046]

[0047] Where θ l1 and θ focal It can indicate the weights of the corresponding loss function. This can refer to the focus loss used for classification, while Regression loss can be applied between the detection location and the ground truth bounding shape (e.g., box) location for each object.

[0048] Tracker 106 may be configured to generate and / or determine predictions corresponding to one or more object traces (e.g., object tracklets) and / or tracked objects or trajectories across multiple frames, timestamps, and / or observations or detections in the environment. In at least one embodiment, tracker 106 may perform data association based at least on object detection results from detector 104 linked across frames and / or timestamps. Furthermore, tracker 106 may determine and / or refine parameters of the state of the tracked object or entity.

[0049] Now for reference Figure 2 Tracker 106 may include, for example, a correspondence determiner 220 and a track updater 222. The correspondence determiner 220 may use track features 230 corresponding to one or more object tracks and object features 214 corresponding to one or more object detection results (e.g., object embeddings from decoder 206) to generate and / or determine correspondence data 232 between one or more object detection results and one or more object tracks.

[0050] In at least one embodiment, correspondence data 232 may indicate the association between one or more object detection results and one or more object traces. For example, correspondence data 232 may include or represent one or more values ​​(e.g., correspondence scores) that indicate the probability that one or more particular object detection results correspond to one or more particular object traces. In at least one embodiment, correspondence data 232 includes pairwise correspondence scores between objects (e.g., object detection results) and traces (e.g., trace fragments). In at least one embodiment, the correspondence scores are calculated based at least on the similarity of appearance and / or motion. In at least one embodiment, a set of correspondence scores may be calculated across frames (e.g., consecutive frames and / or pairwise frames), and these sets may be aggregated to determine correspondence data 232 as an aggregated set of correspondence scores (e.g., corresponding to all frames or additional frames). In at least one embodiment, the set of correspondence scores may be stored in one or more matrices.

[0051] In at least one embodiment, the correspondence determiner 220 computes correspondence data 232 in a differentiable manner for one or more object detection results and one or more object traces. For example, the correspondence determiner 220 may use one or more MLMs (e.g., neural networks) to determine predictions of the correspondence data 232 between one or more object detection results and one or more object traces. The MLM may receive data corresponding to one or more object detection results (e.g., object embeddings) and one or more object traces (e.g., trace features 230) and use this data to predict the correspondence data 232.

[0052] Now for reference Figure 1B and Figure 2 The process 150 for training one or more MLMs 156 to predict correspondence data 232 corresponding to trace fragments 152 can use supervised learning of correspondence scores, which utilizes, for example, an intermediate cross-entropy loss function for the estimated correspondence scores. In at least one embodiment, the intermediate cross-entropy loss function can be based at least on each tracked object having at most one match detection result. For example, the ground truth correspondence score matrix A g Each row and each column can be either a one-hot vector or a vector of all zeros. For A g All rows and columns containing one-hot vectors can be used to calculate cross-entropy loss. The corresponding rows and columns are applied to the estimated correspondence score matrix A. The appropriate object tracking loss function is shown using equation (2). An example, where columns in the truth correspondence score matrix It can be a one-hot vector, and the cross-entropy loss is applied to the j-th column. It can be defined as:

[0053]

[0054] Where M can represent the row number of the correspondence score matrix and the number of trace segments.

[0055] In at least one embodiment, the trace updater 222 updates the trace features 230 using object detection results 212 and / or object embeddings from the decoder 206 and correspondence data 232 for iteration in subsequent frames and / or process 200. For example, the trace updater 222 may update the trace features 230 based at least on the associated object detection results and motion and / or appearance information of the object traces indicated by the correspondence data 232. In at least one embodiment, the trace updater 222 may use... Figure 3A The combined solver 333 determines these associations. The trace updater 222 can also use object detection results 212 and / or object embeddings from the decoder 206 to update object features 214 for subsequent frames and / or iterations of process 200.

[0056] Now for reference Figure 3A , Figure 3AThis includes an example data flow diagram of a process 300A for interfacing a tracker with a motion predictor using a differentiable and modular end-to-end stack for autonomous machines, according to some embodiments of this disclosure. In at least one embodiment, a combined solver 333 (e.g., used by a trace updater 222) receives object detection results 212, correspondence data 232, object traces 330 (e.g., corresponding to trace features 230), and / or other data to associate object detection results 212 with object traces 330 (e.g., corresponding to one or more previous frames) to determine object traces 332 (e.g., corresponding to one or more current frames).

[0057] In at least one embodiment, the combined solver 333 uses a non-differentiable Hungarian algorithm to correlate object detection results with object tracks. Using a non-differentiable method to correlate object detection results with object tracks results in parts of the computation graph being non-differentiable. However, the process 150 for training a differentiable and modular end-to-end stack for autonomous machines can still be differentiable overall because of the prediction loss function from the motion predictor 108. This can be backpropagated to detector 104 and tracker 106 because object features 214 and corresponding matching trace features 230 are propagated in time and used to generate object traces 332 as input to motion predictor 108 (e.g., as...). Figure 3A (As shown). The motion predictor 108 can determine predictions of one or more future movements associated with the object detection results, based at least on applying object traces 332 to one or more MLMs.

[0058] In at least one embodiment, instead of using a non-differentiable Hungarian algorithm to correlate object detection results with object traces, the combined solver 333 can construct object traces 332 in a differentiable manner using a differentiable combined solver. Using a differentiable combined solver makes the entire computational graph differentiable. For example, the combined solver can be made differentiable at least based on treating the solver as a negative identity during backpropagation, where gradients are passed through the combined solver without any change in magnitude, but the signal is reversed. In at least one embodiment, the combined solver is implemented using a continuous function. For example, the combined solver 333 can implement a linear cost solver differentiability. In one or more embodiments, the combined solver 333 can be made differentiable using one or more graph neural networks, softmax relaxation, one or more differentiable sorting networks, and / or one or more differentiable assignment algorithms.

[0059] Besides about Figure 3A The process described in 300A is either outside of or a substitute for the method described therein. Figure 3B Process 300B can be used to provide a differentiable computational graph. Now refer to... Figure 3B , Figure 3B This includes an example data flow diagram of a process 300B for interfacing a tracker with a motion predictor using a differentiable and modular end-to-end stack for autonomous machines, according to some embodiments of this disclosure. Figure 3B In this method, the motion predictor 108 determines predictions of one or more future movements associated with one or more object detection results, based at least on applying correspondence data 232 and object detection results 212 to one or more MLMs. Therefore, in at least one embodiment, it is not necessary to construct object tracks 332 to input into the motion planner 110. Furthermore, when correspondence data 232 is applied to one or more MLMs of the motion planner 110, the one or more MLMs can account for uncertainties in the association between object detections and object tracks. In at least one embodiment, correspondence data 232 can be similarly used in other methods, such as those concerning... Figure 3A The method described.

[0060] In at least one embodiment, motion predictor 108 may use one or more observations and / or parameters of environmental states (e.g., current and / or historical) provided by detector 104 and / or tracker 106 to determine or generate one or more predicted future movements (e.g., locations) for one or more entities or actors in the environment. For example, motion predictor 108 may generate data indicating one or more predicted locations of one or more entities at one or more specific times or time steps. For example, motion predictor 108 may determine one or more predicted trajectories or tracks of one or more entities.

[0061] The motion predictor 108 can be implemented using various methods. For example, but not limited to, the motion predictor 108 can be implemented using one or more MLMs, such as... Figure 1BAt least one neural network 158 is shown. One or more MLMs can be trained to predict one or more predicted locations of one or more entities or actors in an environment, such as data representing and / or indicating one or more parameters of one or more future or predicted world states within one or more specific times or time steps. In at least one embodiment, one or more MLMs include a graph-structured recurrent neural network that predicts the future location distribution of an agent given its past trajectory history and the past trajectories of one or more neighboring agents (e.g., object trace 332). In at least one embodiment, one or more MLMs can use at least one neural network (such as a conditional variational autoencoder (CVAE)) to model the probabilities of multiple future trajectories.

[0062] In at least one embodiment, the motion predictor 108 takes the H-second state history of one or more agents as input and outputs a multimodal trajectory prediction for agent a∈A according to equation (3).

[0063]

[0064] Where k∈K is the mode of the output distribution, s refers to the state, and θ refers to the training parameters of the motion predictor 108. For the sake of brevity, this paper can use In at least one embodiment, the encoder of the motion planner 110 (e.g., CVAE) may use a recurrent long short-term memory (LSTM) network to process the agent state history and use graph-based attention to model inter-agent interactions. The decoder of the motion planner 110 (e.g., CVAE) may include a gated recurrent unit (GRU) that outputs a Gaussian mixture model (GMM) for each future time step. In at least one embodiment, the GMM pattern may correspond to one or more discrete latent states of the motion planner (e.g., a CVAE with K = 25 discrete latent states).

[0065] In at least one embodiment, to ensure that the prediction is dynamically feasible, the GMM can be defined by overcontrol and then integrated through a differentiable dynamic function to generate at least a portion of the trajectory. In at least one embodiment, the input state of at least one neural network 158 of the motion predictor 108 can be an input state enhanced with one or more variables (e.g., ego indicator variables), ego states, and / or other state-related information as described herein, such as for ego-agent relation reasoning.

[0066] Now for reference Figure 4 , Figure 4The illustrations include example environments according to some embodiments of the present disclosure, where the environment has patterns of predicted location distributions of agents and corresponding planned trajectories of machines. For example, Figure 4 An environment 400 including agent 410 is shown. Figure 4 Pattern 420 is also shown, which can correspond to the pattern of the output distribution described with respect to motion predictor 108.

[0067] Refer again Figure 1B The original prediction training objective of motion predictor 108 can be defined according to equation (4):

[0068]

[0069] The prediction loss function For example, the loss function could include the Information Maximization Variational Autoencoder (InfoVAE).

[0070] In at least one embodiment, motion planner 110 uses one or more observations of environmental states (e.g., current and / or historical) and / or observed or predicted parameters provided by detector 104, tracker 106, and / or motion predictor 108 to determine or generate one or more motion plans for a machine (e.g., vehicle 700). Each motion plan may indicate one or more potential locations for the machine at one or more specific times or time steps. For example, a motion plan may indicate a potential trajectory, a reference trajectory, and / or a planned trajectory for the machine. In at least one embodiment, a motion plan may include a reference trajectory or an actual trajectory that the machine will follow.

[0071] The motion planner 110 can be implemented using various methods. For example, but not limited to, the motion planner 110 can generate or determine one or more (e.g., multiple) candidate movements, such as Figure 1B The candidate movements 160 shown are (e.g., candidate trajectories and / or paths, one or more future positions or maneuvers, etc.). Candidate movements 160 may be generated or determined based on at least one or more of the positions, trajectories, and / or states of other entities predicted using motion predictor 108. Motion planner 110 can use candidate movements 160 to determine (e.g., select) the machine's final, reference, or planned path, trajectory, or movement, such as... Figure 1B The reference movement 162 shown is Figure 4 The reference movement 430 is shown in the diagram. For example, each path can be generated using multiple time steps or timestamps (e.g., 0, 1, 2, and 3), and each time step and / or the entire path can be assigned a weight or score, where the final value of each path is used to select the final path from the paths.

[0072] In at least one embodiment, motion planner 110 uses observations, predicted motion, and / or sensor data 102 to generate one or more suggested paths for vehicle 700. For example, motion planner 110 can generate any number of paths for vehicle 700 and can analyze these suggested paths (e.g., for safety or collision avoidance considerations, comfort considerations, consistency considerations, power / fuel consumption considerations, compliance with road rules, etc.) to determine which path to select as the actual or reference path / trajectory for vehicle 700. Motion controller 112 can generate control sequence 164 to follow the selected reference trajectory while avoiding collisions with other agents on the road, ensuring comfort, and meeting the dynamic constraints of vehicle 700.

[0073] In at least one embodiment, the motion planner 110 generates a set of N dynamically feasible candidate trajectories P = {s} for the machine. n ,u n} n∈N And select a reference trajectory or planned trajectory based on at least one or more corresponding cost values ​​(e.g., select the trajectory with the lowest cost). For example, a reference trajectory can be selected according to equation (5):

[0074]

[0075] Where C refers to the cost function. This refers to multi-model predictions from one or more agents of the motion predictor 108, where g is the given target and m is the lane map.

[0076] In at least one embodiment, the motion planner 110 uses one or more analytical functions to calculate one or more cost values ​​for one or more candidate moves 160 of the machine. The one or more analytical functions may include at least one parameter, and the one or more analytical functions are trained to generate predictions corresponding to the one or more cost values, wherein these predictions are differentiable with respect to the parameters of at least one neural network 158 of the motion predictor 108.

[0077] In at least one embodiment, the motion planner 110 may sample the state space to determine one or more candidate movements 160 of the machine, at least based on the machine's state (e.g., current state or initial state). The candidate movements may then be filtered, at least based on control constraints associated with the machine. For example, to generate candidate trajectories, the motion planner 110 may sample a set of terminal states and generate a trajectory using the machine's current state and that set of terminal states. In at least one embodiment, the motion planner 110 fits a spline function (e.g., a cubic spline function) from the current state to the terminal state to generate candidate trajectories. Dynamically infeasible trajectories may be filtered out and / or rejected, not considered as reference trajectories or movements for use by the motion controller 112.

[0078] In at least one embodiment, the motion planner 110 uses a lane map m to generate one or more candidate movements 160. The lane map m may, for example, represent lanes available to the vehicle 700 on a route (e.g., a road). For example, lanes may be modeled as directed edges in a graph. Each edge in the lane map may correspond to a specific lane segment, and nodes in the graph may represent lane endpoints or intersections. Each lane segment may be assigned various attributes, such as its location, width, and speed limits, as well as any traffic signs or road markings that may be present. Furthermore, the lane map may include information about traffic flow and congestion, which can be used for planning and control. The lane map may be used by the motion planner 110 to filter and / or limit the lanes under consideration, and / or determine lanes within a threshold distance (e.g., 4.5 meters) from the target state, where the heading difference between the lane and its self-target state is less than a threshold amount (e.g., 90 degrees).

[0079] In at least one embodiment, for each lane considered, the motion planner 110 generates a set of terminal states for a fixed set of acceleration values, based at least on the distance traveled along the lane, such as, but not limited to, a ∈ {-3, -2, -1, -0.5, 0, 0.5, 1, 2} m / s. 2 Lateral lane departure l ⊥ In the range {-0.5, 0, 0.5}m, the motion planner 110 can then fit a cubic spline function from the current state to the terminal state to minimize the mean control force and can discard dynamically infeasible trajectories given the control constraints of the vehicle 700.

[0080] As described herein, the cost function of the motion planner 110 may include one or more analytical functions. Because these functions are analytical, they are highly interpretable and can be configured to provide various safety assurances to motion planning. The cost function may include any of a variety of factors. For example, the cost function may be formulated to generate one or more cost values ​​based on at least one or more of the following: one or more predicted collisions between the machine and at least one of one or more agents; one or more distances between at least one state of the machine and at least one target state of the machine; one or more lane lateral deviations; one or more lane heading deviations; or one or more control forces. In at least one embodiment, the cost function may include momentum-shaped spatial distances with other agent functions, a collision time cost function (e.g., calculated assuming the agent continues to move at a constant speed along its current heading), a speed limit cost function, a distance cost function to a reference trajectory, a ego comfort cost function, and / or a reversing cost function that penalizes the vehicle for traveling against the heading direction of the current lane.

[0081] As described herein, each time step of a path (or trajectory) and / or the path as a whole can be assigned a score, where the final value for each path is used to select a final path, a reference path, or a planned path from a plurality of paths. For example, in at least one embodiment, the score can correspond to a cost function C, which may include a weighted ω x The weighted sum of x handcrafted terms, an unrestricted example of which can be represented by equation (6):

[0082]

[0083] These items respectively penalize collision, distance from the target, lateral lane departure, lane heading deviation, and control strength.

[0084] In at least one embodiment, one or more analytical functions include one or more terms that calculate a cost corresponding to the location of the agent predicted by at least one neural network 158 using motion predictor 108. For example, collision term C coll It may include one or more predictions from at least one neural network 158. In at least one embodiment, the collision term corresponds to equation (7):

[0085]

[0086] in, This refers to the k-th pattern of the predicted trajectory distribution of agent a, π. k It refers to the probability of the k-th pattern. "|" refers to the Gaussian radial basis function, while "|·|" refers to the Euclidean norm.

[0087] As described herein, one or more analytical functions used to compute cost values ​​may include one or more parameters that are trained to generate differentiable predictions about at least one neural network 158 of motion predictor 108 (e.g., parameters trained to predict the position and / or state of an actor). Therefore, motion planner 110 may be jointly trained with motion predictor 108, tracker 106, and detector 104. Various methods can be used to provide a differentiable motion planner 110. In at least one embodiment, motion planner 110 may classify a trajectory or movement as a target, reference, or planned trajectory or movement of the machine, at least based on a category distribution corresponding to one or more candidate trajectories or movements being evaluated. For example, the arm min operation in equation (5) may be relaxed to sample from the category distribution, with motion planner 110 acting as a classifier.

[0088] Refer again Figure 1B The original prediction training objective of the motion planner 110 can be defined according to equation (8):

[0089]

[0090] Among them, the motion planning loss function This could include, for example, a cross-entropy loss function, while s * This refers to the movement or trajectory of the target. The cross-entropy loss function can be used to effectively treat the motion planner 110 as a classifier during training, where the probability p n It can be defined according to equation (9):

[0091] p n = exp(-βC(s) n ,u n ; ·; ω)) / Z, (9)

[0092] Where β refers to the temperature parameter, and Z refers to the normalization constant.

[0093] Using the chain rule, the gradients with respect to the cost function C and the motion predictor 108 parameters θ can be defined according to equation (10):

[0094]

[0095] All items exist.

[0096] The motion controller 112 can follow a selected reference movement 162 while avoiding collisions with other agents on the road, ensuring comfort and meeting the dynamic constraints of the vehicle 700. In at least one embodiment, the motion controller 112 uses the reference movement 162 to compute a control sequence 164 (e.g., a continuous and smooth ego trajectory) for the machine using one or more analytical functions having one or more parameters, which are trained to generate predictions corresponding to the control sequence, and these predictions are differentiable with respect to at least one parameter of the one or more analytical functions used by the motion planner 110 to compute one or more cost values.

[0097] In at least one embodiment, the motion controller 112 performs model predictive control (MPC) within a finite horizon using an iterative box-constrained linear quadratic regulator (LQR) algorithm. In at least one embodiment, the motion controller 112 is implemented according to equation (11):

[0098]

[0099]

[0100] Where C refers to the cost function, f d This refers to dynamics, s init This refers to the current or initial machine state. u and This refers to the control limits. The cost function C can be the same as or different from the cost function used by the motion planner 110. The dynamically extended unicycle model can be used for dynamics f d The trajectory can be generated using u from the motion planner 110. plan Initialize. The motion controller 112 can use the dynamics f for each iteration i. d And the first and second Taylor approximations of the cost function C, around the current solution s (i) ,u (i) The quadratic LQR approximation of equation (11) is iteratively formed and solved. The trajectory can be updated to approximate the optimal LQR control while also reducing the original non-quadratic cost. The iteration can terminate at least based on reaching convergence and / or a threshold constraint.

[0101] In at least one embodiment, in addition to or in lieu of one or more analytical functions, the motion controller 112 and / or motion planner 110 also employ at least one neural network. By using one or more analytical functions, the algorithm is highly interpretable and can be configured to provide various safeguards to the control sequence 164. Furthermore, because the objectives of the motion planner 110, motion predictor 108, tracker 106, and detector 104 are differentiable, these components can be jointly trained, which can improve performance. For example, as... Figure 1B As shown, the gradient (e.g., with the motion control loss function) The corresponding gradient can be backpropagated from the output of motion controller 112 to detector 104, tracker 106, and / or sensing module, thereby allowing downstream control objectives of motion controller 112 to learn upstream predictions. In other words, process 150 may include parameters θ with respect to upstream module I. i calculate That is, the training objective of downstream component j The gradient is used to jointly train the parameters of each component for the overall objective using a gradient-based optimizer.

[0102] Therefore, for example, even with fewer parameters used in the MLM, the MLM of detector 104, tracker 106, motion predictor 108, and / or motion planner 110 can provide improved predictions, thereby improving the control sequence. In at least one embodiment, end-to-end training can allow for easier and / or more extensive extrapolation of the Pareto frontier to determine the machine's control sequence, thereby reducing the computational and / or storage requirements for training and / or deploying the MLM.

[0103] Various methods can be used to provide a differentiable motion controller 112. In at least one embodiment, in order for the motion controller 112 to be differentiable, it can be differentiated, at least based on the underlying Karush-Kuhn-Tucker (KKT) conditions of the last LQR approximation, with respect to the dynamics f. d and cost function C for control sequence trajectory s ctr Perform differentiation. The gradient can be analytically computed using one or more additional backpropagations of a modified iterative LQR solver. If the iLQR fails to converge, the gradient may not be propagated.

[0104] The gradient can be calculated based on equation (12) regarding the cost parameters of the motion planner 110:

[0105]

[0106] Furthermore, according to equation (13), the parameter θ of the motion predictor 108 is calculated:

[0107]

[0108] in This refers to the motion control loss function.

[0109] Various methods can be used to train the detector 104, tracker 106, motion predictor 108, motion planner 110, and motion controller 112. In at least one embodiment, open-loop training, such as reinforcement learning (RL) or imitation learning (IL), can be performed using human driving data. For RL, process 150 can be designed to optimize the hindsight cost, which is a post-hoc measurement of the quality of the ego trajectory (e.g., after observing the future trajectories of other agents in the scene). The hindsight cost can be considered a negative reward for RL. In at least one embodiment, RL can be designed to minimize the output ego trajectory s ctr ,u ctr The ex-post cost C on the training examples in the dataset H , Post-event cost C H It is possible to understand the future trajectory of non-self-agents or entities. The quality of the captured trajectory afterwards. In at least one embodiment, the post-hoc cost C... H The cost function C is the same as that used by motion planner 110, but it has a true future trajectory input instead of a prediction, and has a fixed ω parameters. In at least one embodiment, the recorded trajectory can be regarded as the true future, used to calculate the trajectory given a ground truth future. The prediction model is trained under the condition that both terms exist given a differentiable motion controller 112. In at least one embodiment, the loss of the motion predictor 108 may correspond to the observed future trajectory. The loss of the motion planner 110 may correspond to the candidate movement 160 with the lowest ex-post cost.

[0110] For IL, process 150 can be designed to minimize the mean squared error (MSE) between the output ego trajectory and the true value in the dataset. The goal of motion planner 110 can be s gt The total loss during training can be a linear combination of component-wise objectives.

[0111] Refer again Figure 4The diagram illustrates a scenario where agent 410 is traveling behind vehicle 700. Conventional methods might incorrectly predict that agent 410 is likely to collide with vehicle 700 from behind, causing vehicle 700 to drift into the intersection, which would be dangerous. In contrast, the disclosed method assigns a higher probability of agent 410 slowing down, thus causing vehicle 700 to slow down before entering the intersection. Therefore, the predictions made using process 100 are more realistic and accurate, resulting in more reasonable downstream planning.

[0112] The motion controller 112 of the autonomous or semi-autonomous driving software stack can transmit information indicating one or more control operations (e.g., corresponding to control sequence 164) to the control component 114 of the autonomous driving software stack, and the control component 114 can determine the control of the vehicle 700 based on the control operations to initiate the vehicle 700. One or more parts of process 100 can be completed as needed, for example at each time instance, interval, and / or for each environmental state, or when a new motion plan or control sequence is needed or expected, thereby generating, analyzing, and selecting new control operations and / or controls, and the vehicle 700 follows the control operations along the way.

[0113] Now for reference Figure 5 and Figure 6 Each block of methods 500 and 600, as well as other methods described herein, includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. These methods can also be embodied as computer-usable instructions stored on a computer storage medium. These methods can be provided by standalone applications, services, or managed services (independently or in combination with another managed service), or plug-ins to other products, to name a few. Furthermore, regarding Figure 1- Figure 4 These methods are described by way of example. However, these methods may be implemented additionally or alternatively by any system or any combination of systems, including but not limited to the systems described herein.

[0114] Now for reference Figure 5 , Figure 5This is a flowchart illustrating a method 500 for controlling a machine using a differentiable end-to-end autonomous vehicle stack, according to some embodiments of the present disclosure. At block B502, method 500 includes: determining a first prediction of one or more correspondence scores that are differentiable with respect to one or more object detection results and one or more object tracks. For example, at least the correspondence determiner 220 of tracker 106 can be used to determine a first prediction of correspondence data 232 between object detection results 212 and object tracks 330 using one or more first machine learning models (MLMs) and sensor data 102 obtained using one or more sensors associated with the machine, wherein the correspondence data 232 is differentiable with respect to object detection results 212 and object tracks 330.

[0115] At box B504, method 500 includes: determining a second prediction of one or more future movements that are differentiable with respect to one or more correspondence scores. For example, motion predictor 108 may use one or more second MLMs to determine a second prediction of one or more future movements associated with object detection result 212, wherein one or more future movements are differentiable with respect to correspondence data 232.

[0116] In box B506, method 500 includes: determining a third prediction of at least one trajectory, wherein at least one trajectory is differentiable with respect to future movements. For example, motion planner 110 may use one or more third MLMs to determine a third prediction of candidate movements 160 of the machine, wherein candidate movements 160 are differentiable with respect to one or more future movements.

[0117] In block B508, method 500 includes determining a fourth prediction of one or more control sequences that are differentiable with respect to at least one trajectory. For example, motion controller 112 may use one or more fourth MLMs to determine a fourth prediction of one or more control sequences 164 for a machine, wherein one or more control sequences 164 are differentiable with respect to candidate movement 160.

[0118] In box B510, method 500 includes performing one or more control operations. For example, control component 114 may perform one or more control operations on the machine based on at least one or more control sequences 164.

[0119] Now for reference Figure 6 , Figure 6This is a flowchart illustrating a method 600 for controlling a machine using a sequence of machine learning models (MLMs) predicting differentiable output sequences, according to some embodiments of the present disclosure. At block B602, method 600 includes: using the sequence of machine learning models (MLMs) predicting differentiable output sequences to determine one or more control sequences for the machine, the differentiable output sequences including: correspondence data between one or more object detections and one or more object tracks, one or more object motion predictions corresponding to the correspondence data, one or more motion plans corresponding to the one or more object motion predictions, and one or more control sequences corresponding to the one or more motion plans. For example, motion controller 112 may use the sequence of machine learning models (MLMs) corresponding to detector 104, tracker 106, motion predictor 108, and motion planner 110 and predicting a sequence of differentiable output sequences including object detection results 212, future moves, candidate moves 160, and control sequence 164 to determine control sequence 164 for the machine.

[0120] At box B604, method 600 includes performing one or more control operations. For example, control component 114 may perform one or more control operations on the machine based on at least one or more control sequences 164.

[0121] Example autonomous vehicles

[0122] Figure 7AThis is an illustration of an example autonomous vehicle 700 according to some embodiments of the present disclosure. The autonomous vehicle 700 (or, alternatively, referred to herein as “vehicle 700”) may include, but is not limited to, passenger vehicles such as cars, trucks, buses, ambulances, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, engineering vehicles, underwater vessels, robotic vehicles, drones, aircraft, vehicles attached to trailers (e.g., semi-trailers for transporting goods) and / or other types of vehicles (e.g., driverless and / or capable of accommodating one or more passengers). Autonomous vehicles are typically described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in its "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201506, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 700 is capable of performing one or more functions that meet Level 3 through Level 5 of autonomous driving. Vehicle 700 is capable of performing one or more functions that meet Level 1 through Level 5 of automated driving. For example, depending on the embodiment, vehicle 700 is capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term “autonomy” as used herein may include any and / or all types of autonomy for the 700 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, providing auxiliary autonomy, semi-autonomy, primary autonomy, or other names.

[0123] Vehicle 700 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 700 may include a propulsion system 750, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 750 may be connected to the drivetrain of vehicle 700, which may include a transmission, to enable propulsion of vehicle 700. Propulsion system 750 may be controlled in response to receiving a signal from throttle / accelerator 752.

[0124] A steering system 754, which may include a steering wheel, can be used to steer the vehicle 700 (e.g., along a desired path or route) when the propulsion system 750 is operating (e.g., when the vehicle is in motion). The steering system 754 may receive signals from the steering actuator 756. For fully automatic (level 5) functionality, the steering wheel may be optional.

[0125] The brake sensor system 746 can be used to operate the vehicle brakes in response to receiving signals from the brake actuator 748 and / or the brake sensor.

[0126] It can include one or more System-on-Chip (SoC) 704 ( Figure 7C One or more controllers 736, including and / or one or more GPUs, can provide (e.g., signals representing commands) to one or more components and / or systems of vehicle 700. For example, one or more controllers can send signals to operate vehicle brakes via one or more brake actuators 748, to operate steering system 754 via one or more steering actuators 756, and / or to operate propulsion system 750 via one or more throttles / accelerators 752. One or more controllers 736 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 700. One or more controllers 736 may include a first controller 736 for autonomous driving functions, a second controller 736 for functional safety functions, a third controller 736 for artificial intelligence functions (e.g., computer vision), a fourth controller 736 for infotainment functions, a fifth controller 736 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 736 can handle two or more of the functions described above, and two or more controllers 736 can handle a single function, and / or any combination thereof.

[0127] One or more controllers 736 may provide signals for controlling one or more components and / or systems of vehicle 700 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, but not limited to, a Global Navigation Satellite System (“GNSS”) sensor 758 (e.g., a Global Positioning System sensor), a RADAR sensor 760, an ultrasonic sensor 762, a LIDAR sensor 764, an Inertial Measurement Unit (IMU) sensor 766 (e.g., an accelerometer, gyroscope, magnetic compass, magnetometer, etc.), a microphone 796, a stereo camera 768, a wide-angle camera 770 (e.g., a fisheye camera), an infrared camera 772, a surround camera 774 (e.g., a 360-degree camera), a long-range and / or medium-range camera 798, a speed sensor 744 (e.g., for measuring the rate of vehicle 700), a vibration sensor 742, a steering sensor 740, a braking sensor (e.g., as part of a braking sensor system 746), and / or other sensor types.

[0128] One or more of the controllers 736 may receive inputs (e.g., represented by input data) from the instrument cluster 732 of the vehicle 700 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 734, an auditory signaling device, a speaker, and / or via other components of the vehicle 700. These outputs may include information such as vehicle speed, rate, time, map data (e.g., ...). Figure 7C Information such as high-definition (“HD”) maps 722, location data (e.g., the location of vehicle 700 on the map), orientation, the location of other vehicles (e.g., grid occupancy), and information about objects and their states perceived by controller 736, etc. For example, HMI display 734 may display information about the existence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).

[0129] The vehicle 700 further includes a network interface 724, which can communicate via one or more networks using one or more wireless antennas 726 and / or a modem. For example, the network interface 724 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. One or more wireless antennas 726 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or one or more low-power wide area networks (“LPWAN”) such as LoRaWAN, SigFox, etc.

[0130] Figure 7B For use in accordance with some embodiments of this disclosure Figure 7A This is an example of the camera position and field of view of an example autonomous vehicle 700. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included, and / or these cameras may be located at different positions on the vehicle 700.

[0131] The camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 700. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-white-white-white (RCCC) color filter array, a red-white-white-blue (RCCB) color filter array, a red-blue-green-white (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a sharp-pixel camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to improve light sensitivity.

[0132] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0133] One or more of the cameras can be mounted in mounting components such as custom-designed (3D-printed) parts to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard in the windshield mirror) that may interfere with the camera's image data capture capabilities. Regarding the wing mirror mounting components, the wing mirror components can be custom-3D-printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.

[0134] A camera with a field of view that includes the environment in front of the vehicle 700 (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 736 and / or control SoCs, to provide information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used in ADAS functions and systems, including Lane Departure Warning (“LDW”), Autonomous Cruise Control (“ACC”), and / or other functions such as traffic sign recognition.

[0135] A variety of cameras can be used in front-facing configurations, including, for example, monocular camera platforms including complementary metal-oxide-semiconductor (“CMOS”) color imagers. Another example could be a wide-angle camera 770, which can be used to perceive objects entering the field of view from the periphery (e.g., pedestrians, traffic at intersections, or bicycles). Although Figure 7B The middle image shows only one wide-angle camera, but any number (including zero) of wide-angle cameras 770 can be present on vehicle 700. Furthermore, any number of remote cameras 798 (e.g., long-view stereo camera pairs) can be used for depth-based object detection, especially for objects for which neural networks have not yet been trained. Remote cameras 798 can also be used for object detection and classification, as well as basic object tracking.

[0136] Any number of stereo cameras 768 may also be included in the front-mounted configuration. In at least one embodiment, one or more stereo cameras 768 may include an integrated control unit comprising a scalable processing unit that can provide a multi-core microprocessor and programmable logic (“FPGA”) with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 768 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip capable of measuring the distance from the vehicle to a target object and using the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 768 may be used in addition to those described herein or alternatively.

[0137] A camera (e.g., a side-view camera) having a field of view that includes the side of the vehicle 700 can be used for surround view, providing information for creating and updating occupancy grids and generating side-impact collision warnings. For example, a surround camera 774 (e.g., such as...) Figure 7B The four surround cameras 774 shown can be mounted on the vehicle 700. The surround cameras 774 can include wide-angle cameras 770, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 774 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.

[0138] A camera having a field of view that includes the environment behind the vehicle 700 (e.g., a rear-view camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy grids. A wide variety of cameras can be used, including but not limited to those also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range camera 798, stereo camera 768, infrared camera 772, etc.).

[0139] Figure 7C For use in accordance with some embodiments of this disclosure Figure 7AThe example autonomous vehicle 700 is illustrated in the block diagram of an example system architecture. It should be understood that this arrangement, and other arrangements described herein, are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities, which may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by these entities can be implemented via hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.

[0140] Figure 7C Each component, feature, and system in vehicle 700 is illustrated as being connected via bus 702. Bus 702 may include a Controller Area Network (CAN) data interface (or, alternatively, referred to herein as the "CAN bus"). CAN may be a network within vehicle 700 used to assist in the control of various features and functions of vehicle 700, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0141] Although bus 702 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to a CAN bus. Furthermore, although bus 702 is represented by a single line, this is not intended to be limiting. For example, any number of buses 702 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 702 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 702 may be used for a collision avoidance function, and a second bus 702 may be used for drive control. In any example, each bus 702 may communicate with any component of vehicle 700, and two or more buses 702 may communicate with the same component. In some examples, each SoC 704, each controller 736, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 700) and may be connected to a common bus such as a CAN bus.

[0142] Vehicle 700 may include one or more controllers 736, such as those described herein. Figure 7A The controllers described herein. Controller 736 can be used for a wide variety of functions. Controller 736 can be coupled to any other different components and systems of vehicle 700 and can be used for the control of vehicle 700, artificial intelligence of vehicle 700, infotainment and / or the like for vehicle 700.

[0143] Vehicle 700 may include one or more System-on-Chip (SoC) 704. SoC 704 may include CPU 706, GPU 708, processor 710, cache 712, accelerator 714, data storage 716, and / or other components and features not shown. SoC 704 can be used to control vehicle 700 across a wide variety of platforms and systems. For example, one or more SoCs 704 may be combined with an HD map 722 in a system (e.g., the system of vehicle 700), the HD map being accessible from one or more servers (e.g., via a network interface 724). Figure 7D One or more servers (778) receive map refresh and / or updates.

[0144] The CPU 706 may include a CPU cluster or a CPU complex (or, alternatively, referred to herein as "CCPLEX"). The CPU 706 may include multiple cores and / or L2 cache. For example, in some embodiments, the CPU 706 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 706 may include four dual-core clusters, each with a dedicated L2 cache (e.g., 2MB L2 cache). The CPU 706 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of the CPU 706 can be active at any given time.

[0145] The CPU 706 can implement power management capabilities including one or more of the following features: automatic clock gating of hardware blocks when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to the execution of WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. The CPU 706 can further implement enhanced algorithms for managing power states, specifying allowed power states and desired wake-up times, and the hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core can support simplified power state entry sequences in software, with this work offloaded to the microcode.

[0146] The GPU 708 may include an integrated GPU (or, alternatively, referred to herein as an "iGPU"). The GPU 708 may be programmable and efficient for parallel workloads. In some examples, the GPU 708 may use an enhanced tensor instruction set. The GPU 708 may include one or more streaming microprocessors, each of which may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 708 may include at least eight streaming microprocessors. The GPU 708 may use a computation application programming interface (API). Furthermore, the GPU 708 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0147] In automotive and embedded applications, the GPU 708 can be power-optimized for optimal performance. For example, the GPU 708 can be fabricated on FinFETs. However, this is not intended to be limiting, and the GPU 708 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can combine several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, dispatch units, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to leverage the mixture of computation and addressing computations for efficient execution of workloads. The streaming microprocessor can include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors can include a combination of L1 data cache and shared memory units to improve performance while simplifying programming.

[0148] The GPU 708 may include, in some examples, a High Bandwidth Memory (HBM) and / or a 16GB HBM2 memory subsystem providing a peak memory bandwidth of approximately 900GB / s. In some examples, in addition to HBM memory or alternatively, Synchronous Graphics Random Access Memory (SGRAM), such as Generation 5 Graphics Double Data Rate Synchronous Random Access Memory (GDDR5), may be used.

[0149] The GPU 708 may include unified memory technology, which includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support can be used to allow the GPU 708 to directly access the CPU 706 page tables. In such examples, when the GPU 708 Memory Management Unit (MMU) experiences a miss, the address translation request can be transferred to the CPU 706. In response, the CPU 706 can look up the virtual-physical mapping for the address in its page tables and transfer the translation back to the GPU 708. Thus, unified memory technology can allow a single unified virtual address space for the memory of both the CPU 706 and GPU 708, simplifying GPU 708 programming and porting applications to the GPU 708.

[0150] In addition, the GPU 708 may include access counters that track how frequently the GPU 708 accesses the memory of other processors. These access counters help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.

[0151] SoC 704 may include any number of caches 712, including those described herein. For example, cache 712 may include an L3 cache available to both CPU 706 and GPU 708 (e.g., it is connected to both CPU 706 and GPU 708). Cache 712 may include a write-back cache, which may track the state of rows, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but a smaller cache size may also be used.

[0152] SoC 704 may include one or more arithmetic logic units (ALUs) that can be used to perform processing of any of a variety of tasks or operations relating to vehicle 700, such as processing a DNN. Furthermore, SoC 704 may include a floating-point unit (FPU) or other mathematical coprocessor or digital coprocessor type for performing mathematical operations within the system. For example, SoC 704 may include one or more FPUs integrated as execution units within CPU 706 and / or GPU 708.

[0153] SoC 704 may include one or more accelerators 714 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, SoC 704 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement GPU 708 and offload some tasks from GPU 708 (e.g., freeing up more cycles of GPU 708 to perform other tasks). As an example, accelerator 714 can be used for targeted workloads (e.g., perceptrons, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. When used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0154] Accelerator 714 (e.g., a hardware acceleration cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations as well as inference. The DLA is designed to provide higher performance per millimeter than a general-purpose GPU and significantly outperform CPUs. The TPU can perform several functions, including single-instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.

[0155] DLA can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any function across a wide variety of applications, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.

[0156] The DLA can perform any function of the GPU 708, and by using inference accelerators, for example, a designer can make either the DLA or the GPU 708 target any function. For example, a designer can focus the CNN processing and floating-point operations on the DLA and leave other functions to the GPU 708 and / or other accelerators 714.

[0157] Accelerator 714 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0158] RISC cores can interact with image sensors (e.g., the image sensor of any camera described herein), image signal processors, and / or the like. Each of these RISC cores may include any amount of memory. Depending on the embodiment, the RISC core may use any of several protocols. In some examples, the RISC core may execute a real-time operating system (RTOS). RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.

[0159] DMA enables PVA components to access system memory independently of the CPU 706. DMA can support any number of features to provide optimizations to the PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0160] A vector processor can be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as, for example, a Single Instruction Multiple Data (SIMD) or Very Long Instruction Word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and speed.

[0161] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Consequently, in some examples, each of the vector processors may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each of these PVAs. Furthermore, the PVA may include additional error-correcting code (ECC) memory to enhance overall system security.

[0162] Accelerator 714 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 714. In some examples, on-chip memory may include at least 4MB of SRAM consisting of, for example, but not limited to, eight field-configurable memory blocks, accessible by both PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, controllers, and multiplexers. Any type of memory may be used. PVA and DLA may access memory via a backbone that provides high-speed memory access to PVA and DLA. The backbone may include (e.g., using an APB) an on-chip computer vision network that interconnects PVA and DLA to memory.

[0163] On-chip computer vision networks can include interfaces that ensure both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such interfaces can provide separate phases and channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, but other standards and protocols can also be used.

[0164] In some examples, the SoC 704 may include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison with LiDAR data for localization and / or other functional purposes, and / or for other uses. In some embodiments, one or more Tree Traversal Units (TTUs) may be used to perform one or more ray tracing-related operations.

[0165] Accelerator 714 (e.g., a hardware accelerator cluster) has broad applications in autonomous driving. PVAs can be programmable vision accelerators used in critical processing stages of ADAS and autonomous vehicles. The capabilities of PVAs are a good match for algorithmic domains requiring predictable processing, low power, and low latency. In other words, PVAs perform well in semi-dense or dense rule computation, even on small datasets requiring predictable runtimes with low latency and low power. Therefore, in the context of platforms for autonomous vehicles, PVAs are designed to run classical computer vision algorithms because they are efficient in object detection and integer mathematical operations.

[0166] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching-based algorithms may be used, but this is not intended to be limiting. Many applications for Level 3–5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras.

[0167] In some examples, PVA can be used to perform intensive optical flow, processing raw RADAR data (e.g., using 4D Fast Fourier Transform) to provide processed RADAR. In other examples, PVA is used for time-of-flight depth processing, which, for example, involves processing raw time-of-flight data to provide processed time-of-flight data.

[0168] DLA can be used to run any type of network to enhance control and driving safety, including, for example, neural networks that output a confidence metric for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative “weight” for each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a threshold for the confidence and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network to regress the confidence value. This neural network can take at least a subset of parameters as input, such as bounding box dimensions, ground plane estimates (e.g., from another subsystem), inertial measurement unit (IMU) sensor 766 outputs related to vehicle orientation and distance, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LiDAR sensor 764 or RADAR sensor 760), etc.

[0169] SoC 704 may include one or more data storage units 716 (e.g., memory). The data storage unit 716 may be on-chip memory of the SoC 704, which may store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and security, the data storage unit 716 may be large enough to store multiple instances of the neural network. The data storage unit 716 may include L2 or L3 cache 712. References to the data storage unit 716 may include references to memory associated with PVA, DLA, and / or other accelerators 714 as described herein.

[0170] SoC 704 may include one or more processors 710 (e.g., embedded processors). Processor 710 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as safety implementation. The startup and power management processor may be part of the SoC 704 startup sequence and may provide runtime power management services. The startup power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 704 thermal and temperature sensor management, and / or SoC 704 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 704 may use the ring oscillator to detect the temperature of CPU 706, GPU 708, and / or accelerator 714. If it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place SoC 704 into a lower power state and / or place vehicle 700 into a driver-safe parking mode (e.g., safely parking vehicle 700).

[0171] The processor 710 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio via multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.

[0172] The processor 710 may further include an always-on-processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. This always-on-processor engine may include a processor core, tightly coupled RAM, peripheral support (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0173] The processor 710 may further include a secure cluster engine, which includes a dedicated processor subsystem for handling security management for automotive applications. The secure cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.

[0174] The processor 710 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0175] The processor 710 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0176] Processor 710 may include a video image compositer, which may be (e.g., implemented on a microprocessor) a processing block, implementing video post-processing functions required by the video playback application to generate the final image for the player window. The video image compositer may perform lens distortion correction on the wide-angle camera 770, the surround camera 774, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate mobile phone services and make calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other situations.

[0177] Video image compositers can include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in the case of motion in the video, denoising appropriately weights spatial information, reducing the weight of information provided by neighboring frames. In cases where the image or part of the image does not contain motion, the temporal denoising performed by the video image compositer can use information from previous images to reduce noise in the current image.

[0178] The video image compositer can also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 708 does not need to continuously render new surfaces, the video image compositer can be further used for user interface components. Even when the GPU 708 is powered on and activated, performing 3D rendering, the video image compositer can be used to offload the GPU 708 to improve performance and responsiveness.

[0179] The SoC 704 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions for receiving video and input from a camera. The SoC 704 may further include an input / output controller that can be software-controlled and can be used to receive I / O signals not assigned to a specific role.

[0180] SoC 704 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. SoC 704 can be used to process data from cameras and sensors (e.g., LIDAR sensor 764, RADAR sensor 760, etc., connected via Gigabit Multimedia Serial Link and Ethernet), data from bus 702 (e.g., vehicle 700 speed, steering wheel position, etc.), and data from GNSS sensor 758 (connected via Ethernet or CAN bus). SoC 704 may further include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine, and which can be used to free up CPU 706 from routine data management tasks.

[0181] The SoC 704 can be an end-to-end platform with a flexible architecture spanning Automation Levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS technologies for diversity and redundancy, along with deep learning tools to deliver a flexible and reliable driving software stack. The SoC 704 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, when combined with the CPU 706, GPU 708, and data storage 716, the accelerator 714 can provide a fast and efficient platform for Level 3-5 autonomous vehicles.

[0182] Therefore, this technology offers capabilities and functionalities that cannot be achieved through conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages ​​such as C to execute a wide variety of processing algorithms across a diverse range of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.

[0183] In contrast to conventional systems, the techniques described herein, by providing CPU complexes, GPU complexes, and hardware-accelerated clusters, allow multiple neural networks to be executed simultaneously and / or sequentially, and the results combined to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executed on a DLA or dGPU (e.g., GPU 720) could include text and word recognition, allowing a supercomputer to read and understand traffic signs, including those for which neural networks have not yet been specifically trained. The DLA could further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on the CPU complex.

[0184] As another example, multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with a light can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a deployed second neural network that informs the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when the flashing lights are detected. The flashing lights can be identified by a deployed third neural network operating across multiple frames, informing the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within a DLA and / or on a GPU 708.

[0185] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 700. A processing engine always on the sensors can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in safe mode, to disable the vehicle when the owner leaves. In this way, SoC 704 provides security against theft and / or carjacking.

[0186] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 796 to detect and identify emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect siren and manually extract features, SoC 704 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to recognize the relative shut-off rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the localized area in which the vehicle operates, as identified by GNSS sensor 758. Thus, for example, when operating in Europe, the CNN will seek to detect European siren, and when operating in the United States, the CNN will seek to identify siren only in North America. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 762, the control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, stop, and / or idle until the emergency vehicle passes.

[0187] The vehicle may include a CPU 718 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 704 via a high-speed interconnect (e.g., PCIe). The CPU 718 may include, for example, an x86 processor. The CPU 718 can be used to perform any of a wide variety of functions, including, for example, arbitrating the results of potential inconsistencies between ADAS sensors and the SoC 704, and / or monitoring the status and health of the controller 736 and / or the infotainment SoC 730.

[0188] Vehicle 700 may include a GPU 720 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 704 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 720 may provide additional artificial intelligence capabilities, for example by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs from sensors of vehicle 700 (e.g., sensor data).

[0189] Vehicle 700 may further include a network interface 724, which may include one or more wireless antennas 726 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 724 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with server 778 and / or other network devices), with other vehicles, and / or with computing devices (e.g., passenger client devices). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across networks and via the Internet). A direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 700 with information about vehicles approaching vehicle 700 (e.g., vehicles in front, to the side, and / or behind vehicle 700). This functionality can be part of vehicle 700's cooperative adaptive cruise control function.

[0190] Network interface 724 may include a SoC that provides modulation and demodulation functions and enables controller 736 to communicate via a wireless network. Network interface 724 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed by known processes and / or using a superheterodyne process. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0191] Vehicle 700 may further include data storage 728, which may include off-chip (e.g., off-chip SoC 704) storage devices. Data storage 728 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.

[0192] Vehicle 700 may further include a GNSS sensor 758. The GNSS sensor 758 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used for assisted mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 758 can be used, including, for example, but not limited to, GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.

[0193] Vehicle 700 may further include a RADAR sensor 760. The RADAR sensor 760 can be used by vehicle 700 for remote vehicle detection even in dark and / or inclement weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 760 can use CAN and / or bus 702 (e.g., to transmit data generated by the RADAR sensor 760) for control and access to object tracking data, and in some examples, Ethernet access for access to raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 760 can be adapted for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.

[0194] The RADAR sensor 760 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 760 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assist and forward collision warning. The long-range RADAR sensor can include a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the vehicle 700's surroundings at higher rates with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 700's lane.

[0195] As an example, a mid-range RADAR system can include a range of up to 760m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 750 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor blind spots behind and beside the vehicle.

[0196] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.

[0197] Vehicle 700 may further include ultrasonic sensors 762. Ultrasonic sensors 762, which may be positioned at the front, rear, and / or sides of vehicle 700, can be used for parking assistance and / or creating and updating occupancy grids. A wide variety of ultrasonic sensors 762 can be used, and different ultrasonic sensors 762 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 762 can operate at functional safety level ASIL B.

[0198] Vehicle 700 may include a LIDAR sensor 764. The LIDAR sensor 764 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 764 may be of functional safety level ASIL B. In some examples, vehicle 700 may include multiple LIDAR sensors 764 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0199] In some examples, the LiDAR sensor 764 may be able to provide a list of objects and their distances within a 360-degree field of view. A commercially available LiDAR sensor 764 may have an advertising range of, for example, approximately 700m, with an accuracy of 2cm-3cm, and support for 700Mbps Ethernet connectivity. In some examples, one or more non-protruding LiDAR sensors 764 may be used. In such examples, the LiDAR sensor 764 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 700. In such examples, the LiDAR sensor 764 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, with a range of 200m. A front-mounted LiDAR sensor 764 may be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0200] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses a flash of laser light as the emission source to illuminate the vehicle's surroundings up to approximately 200 meters. A flash LiDAR unit includes a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-browsing LiDAR devices) without moving parts other than a fan. Flash LiDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using a flash LiDAR, and because a flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 764 is less susceptible to motion blur, vibration, and / or shock.

[0201] The vehicle may further include an IMU sensor 766. In some examples, the IMU sensor 766 may be located at the center of the rear axle of the vehicle 700. The IMU sensor 766 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 766 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 766 may include an accelerometer, a gyroscope, and a magnetometer.

[0202] However, in some embodiments of this disclosure, the IMU sensor 766 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 766 can enable the vehicle 700 to estimate heading by directly observing and correlating velocity changes from GPS to the IMU sensor 766 without input from a magnetic sensor. In some examples, the IMU sensor 766 and the GNSS sensor 758 can be combined into a single integrated unit.

[0203] The vehicle may include a microphone 796 placed in and / or around the vehicle 700. Among other things, the microphone 796 may be used for emergency vehicle detection and identification.

[0204] The vehicle may further include any number of camera types, including stereo camera 768, wide-angle camera 770, infrared camera 772, surround camera 774, long-range and / or mid-range camera 798, and / or other camera types. These cameras can be used to capture image data around the entire perimeter of the vehicle 700. The camera types used depend on the embodiment and the requirements of the vehicle 700, and any combination of camera types can be used to provide the necessary coverage around the vehicle 700. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with respect to... Figure 7A and Figure 7B It was described in more detail.

[0205] Vehicle 700 may further include vibration sensor 742. Vibration sensor 742 can measure vibrations of vehicle components such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 742 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when there is a vibration difference between a power drive shaft and a free-rotating shaft).

[0206] Vehicle 700 may include ADAS system 738. In some examples, ADAS system 738 may include SoC. ADAS system 738 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.

[0207] The ACC system can use a RADAR sensor 760, a LIDAR sensor 764, and / or a camera. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to vehicles immediately in front of vehicle 700 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance holding and, if necessary, advises vehicle 700 to change lanes. Lateral ACC is associated with other ADAS applications such as LCA and CWS.

[0208] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or through a network connection (e.g., via the Internet) through network interface 724 and / or wireless antenna 726. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Typically, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately in front of vehicle 700 and in the same lane), while the I2V communication concept provides information about traffic further ahead. A CACC system can include either or both of these I2V and V2V information sources. Given information about vehicles ahead of vehicle 700, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.

[0209] The Forward-Looking Warning (FCW) system is designed to alert the driver to hazards, enabling the driver to take corrective action. The FCW system uses a front-facing camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components. The FCW system can provide warnings in the form of, for example, audible, visual, haptic, and / or rapid braking pulses.

[0210] An AEB (Autonomous Emergency Braking) system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front-facing camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision approach braking.

[0211] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses a lane marking. When the driver indicates intentional lane departure, the LDW system is deactivated by activating a turn signal. The LDW system can utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.

[0212] The LKA system is a variation of the LDW system. If vehicle 700 begins to leave the lane, the LKA system provides corrective steering input or braking to vehicle 700.

[0213] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses turn signals. The BSW system may use one or more rear-facing cameras and / or one or more RADAR sensors 760, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically coupled to drive feedback, such as displays, speakers, and / or vibration components.

[0214] The RCTW system can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle is reversing. Some RCTW systems include AEB to ensure the application of the vehicle's brakes to avoid a collision. The RCTW system may use one or more rear-mounted RADAR sensors 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.

[0215] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for the driver, but typically not catastrophic, as ADAS systems alert the driver and allow them to determine whether a safe condition truly exists and take appropriate action. However, in an autonomous vehicle 700, in the event of conflicting results, the vehicle 700 itself must decide whether to heed the results from the main computer or auxiliary computer (e.g., the first controller 736 or the second controller 736). For example, in some embodiments, the ADAS system 738 may be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor may run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. Outputs from the ADAS system 738 may be provided to a supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0216] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold and the master and auxiliary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine the appropriate result.

[0217] The supervisory MCU can be configured to run a neural network trained and configured to determine the conditions under which the auxiliary computer provides a false alarm based on outputs from both the host and auxiliary computers. Thus, the neural network in the supervisory MCU can learn when the output of the auxiliary computer can be trusted and when it cannot. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not actually dangerous, such as a drain grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network using associated memory. In a preferred embodiment, the supervisory MCU may include components of and / or be included as components of the SoC 704.

[0218] In other examples, ADAS system 738 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. This allows the auxiliary computer to use classic computer vision rules (if-then), and the presence of neural networks in the supervising MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For instance, if a software vulnerability or bug exists in the software running on the host computer and non-identical software code running on the auxiliary computer provides the same overall result, the supervising MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.

[0219] In some examples, the output of ADAS system 738 can be fed to the perception block and / or the dynamic driving task block of the main computer. For example, if ADAS system 738 issues a forward collision warning because an object is immediately in front, the perception block can use this information when identifying the object. In other examples, the assistance computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.

[0220] Vehicle 700 may further include an infotainment SoC 730 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 730 may include a combination of hardware and software that can be used to provide vehicle 700 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.) and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 730 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 734, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems) and / or other components. The infotainment SoC 730 may further be used to provide information (e.g., visual and / or auditory) to the vehicle's users, such as information from the ADAS system 738, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0221] The infotainment SoC 730 may include GPU functionality. The infotainment SoC 730 can communicate with other devices, systems, and / or components of the vehicle 700 via bus 702 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 730 may be coupled to a supervisory MCU, allowing the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 736 (e.g., the primary and / or backup computer of the vehicle 700). In such an example, the infotainment SoC 730 may place the vehicle 700 into a driver-safe parking mode as described herein.

[0222] Vehicle 700 may further include an instrument cluster 732 (e.g., a digital instrument panel, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 732 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 732 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 730 and the instrument cluster 732. In other words, the instrument cluster 732 may be included as part of the infotainment SoC 730, or vice versa.

[0223] Figure 7D For cloud-based servers and according to some embodiments of this disclosure Figure 7A This is a system diagram illustrating communication between example autonomous vehicles 700. System 776 may include server 778, network 790, and vehicles including vehicle 700. Server 778 may include multiple GPUs 784(A)-784(H) (collectively referred to herein as GPU 784), PCIe switches 782(A)-782(H) (collectively referred to herein as PCIe switch 782), and / or CPUs 780(A)-780(B) (collectively referred to herein as CPU 780). GPUs 784, CPUs 780, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 786, such as, but not limited to, NVLink interface 788 developed by NVIDIA. In some examples, GPUs 784 are connected via NVLink and / or NVSwitch SoCs, and GPUs 784 and PCIe switches 782 are connected via PCIe interconnects. Although eight GPUs 784, two CPUs 780, and two PCIe switches are shown in the diagram, this is not intended to be limiting. Depending on the embodiment, each of the servers 778 may include any number of GPUs 784, CPUs 780, and / or PCIe switches. For example, each of the servers 778 may include eight, sixteen, thirty-two, and / or more GPUs 784.

[0224] Server 778 can receive image data from vehicles via network 790, representing images of unexpected or changed road conditions such as recently commenced roadworks. Server 778 can also transmit neural network 792, updated neural network 792, and / or map information 794, including information about traffic and road conditions, to vehicles via network 790. Updates to map information 794 may include updates to HD map 722, such as information about construction sites, potholes, bends, floods, or other obstacles. In some examples, neural network 792, updated neural network 792, and / or map information 794 may have been generated from new training and / or data received from any number of vehicles in the environment, and / or based on experience gained from training performed at a data center (e.g., using server 778 and / or other servers).

[0225] Server 778 can be used to train machine learning models (e.g., neural networks) based on training data. Training data can be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., where the neural network does not require supervised learning). Training can be performed according to any one or more classes of machine learning techniques, including but not limited to: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component analysis and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to the vehicle via network 790), and / or the machine learning model can be used by server 778 to remotely monitor the vehicle.

[0226] In some examples, server 778 can receive data from vehicles and apply that data to state-of-the-art real-time neural networks for real-time intelligent inference. Server 778 may include a deep learning supercomputer powered by GPU 784 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 778 may include a deep learning infrastructure in a data center that uses only CPU power.

[0227] The deep learning infrastructure of server 778 may be capable of rapid, real-time inference and can be used to assess and verify the health status of the processor, software, and / or associated hardware in vehicle 700. For example, the deep learning infrastructure may receive periodic updates from vehicle 700, such as image sequences and / or objects located in those image sequences that vehicle 700 has already located (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 700. If the results do not match and the infrastructure concludes that the AI ​​in vehicle 700 has malfunctioned, then server 778 may transmit a signal to vehicle 700 instructing its fail-safe computer to take control, notify passengers, and complete a safe stopping operation.

[0228] For inference, server 778 may include GPU 784 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration enables real-time response. In other examples, such as where performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.

[0229] Example computing device

[0230] Figure 8 The block diagram is provided for an example computing device 800 suitable for implementing some embodiments of the present disclosure. The computing device 800 may include an interconnect system 802 directly or indirectly coupled to the following devices: memory 804, one or more central processing units (CPUs) 806, one or more graphics processing units (GPUs) 808, a communication interface 810, input / output (I / O) ports 812, input / output components 814, a power supply 816, one or more presentation components 818 (e.g., displays), and one or more logic units 820. In at least one embodiment, the computing device 800 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 808 may include one or more vGPUs, one or more CPUs 806 may include one or more vCPUs, and / or one or more logic units 820 may include one or more virtual logic units. Therefore, computing device 800 may include discrete components (e.g., a complete GPU dedicated to computing device 800), virtual components (e.g., a portion of the GPU dedicated to computing device 800), or a combination thereof.

[0231] although Figure 8The various boxes are shown connected via an interconnect system 802 with wiring, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 818, such as a display device, may be considered an I / O component 814 (e.g., if the display is a touchscreen). As another example, CPU 806 and / or GPU 808 may include memory (e.g., memory 804 may represent a storage device other than the memory of GPU 808, CPU 806, and / or other components). In other words, Figure 8 The computing devices mentioned are merely illustrative. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all of these are considered within the same category. Figure 8 Within the scope of computing devices.

[0232] Interconnect system 802 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 802 may include one or more link or bus types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Fast (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, CPU 806 may be directly connected to memory 804. Furthermore, CPU 806 may be directly connected to GPU 808. In cases where there is a direct or point-to-point connection between components, interconnect system 802 may include a PCIe link to perform that connection. In these examples, a PCI bus is not required in computing device 800.

[0233] Memory 804 may include any medium of a wide variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by computing device 800. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. For example and without limitation, computer-readable media may include computer storage media and communication media.

[0234] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media, implemented in any way or by any method or technique for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 804 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 800. As used herein, computer storage media does not include the signal itself.

[0235] Computer storage media may contain computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and include any information transport medium. The term "modulated data signal" can refer to a signal whose characteristics are set or altered in a manner that encodes information into that signal. For example and without limitation, computer storage media may include wired media such as wired networks or direct wired connections, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.

[0236] CPU 806 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 800 to perform one or more of the methods and / or processes described herein. Each of CPU 806 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a large number of software threads simultaneously. CPU 806 may include any type of processor and may include different types of processors depending on the type of computing device 800 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 800, the processor may be an advanced RISC mechanism (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as math coprocessors, computing device 800 may also include one or more CPUs 806.

[0237] In addition to or replacing CPU 806, GPU 808 may also be configured to execute at least some computer-readable instructions to control one or more components of computing device 800 to perform one or more of the methods and / or processes described herein. One or more GPUs 808 may be integrated GPUs (e.g., having one or more CPUs 806) and / or one or more GPUs 808 may be discrete GPUs. In embodiments, one or more GPUs 808 may be coprocessors of one or more CPUs 806. Computing device 800 may use GPU 808 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU 808 may be used for general-purpose computing on a GPU (GPGPU). GPU 808 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. GPU 808 may generate pixel data for outputting an image in response to rendering commands (e.g., rendering commands received via a host interface from CPU 806). GPU 808 may include graphics memory, such as display memory, for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory may be included as part of memory 804. GPU 808 may include two or more GPUs operating in parallel (e.g., via links). The links may connect the GPUs directly (e.g., using NVLINK) or via a switch (e.g., using NVSwitch). When combined, each GPU 808 may generate different portions of pixel data or GPGPU data for different outputs (e.g., the first GPU for the first image, the second GPU for the second image). Each GPU may include its own memory or may share memory with other GPUs.

[0238] In addition to or replacing CPU 806 and / or GPU 808, logic unit 820 may be configured to execute at least some computer-readable instructions to control one or more components of computing device 800 to perform one or more methods and / or processes described herein. In embodiments, CPU 806, GPU 808, and / or logic unit 820 may perform any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 820 may be part of and / or integrated into one or more CPUs 806 and / or one or more GPUs 808, and / or one or more logic units 820 may be discrete components of CPU 806 and / or GPU 808 or otherwise external thereto. In embodiments, one or more logic units 820 may be processors of one or more CPUs 806 and / or one or more GPUs 808.

[0239] Examples of logic unit 820 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree traversal unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), arithmetic logic unit (ALU), application-specific integrated circuit (ASIC), floating-point unit (FPU), input / output (I / O) element, peripheral component interconnect (PCI) or peripheral component interconnect fast (PCIe) element, etc.

[0240] The communication interface 810 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 800 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. The communication interface 810 may include components and functions that enable communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 820 and / or the communication interface 810 may include one or more data processing units (DPUs) to directly transmit data received via a network and / or via interconnect system 802 to one or more GPUs 808 (e.g., memory within GPU 808).

[0241] I / O port 812 enables computing device 800 to be logically coupled to other devices, including I / O component 814, presentation component 818, and / or other components, some of which may be built into (e.g., integrated into) computing device 800. Illustrative I / O component 814 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, browsers, printers, wireless devices, and so on. I / O component 814 can provide a Natural User Interface (NUI) for processing user-generated air gestures, voice, or other physiological input. In some instances, the input may be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of computing device 800 (described in more detail below). Computing device 800 may include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. In addition, computing device 800 may include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by computing device 800 to render immersive augmented reality or virtual reality.

[0242] Power supply 816 may include hardwired power supply, battery power supply, or a combination thereof. Power supply 816 may supply power to computing device 800 so that components of computing device 800 can operate.

[0243] The presentation component 818 may include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 818 may receive data from other components (e.g., GPU 808, CPU 806, DPU, etc.) and output that data (e.g., as an image, video, sound, etc.).

[0244] Example Data Center

[0245] Figure 9 An example data center 900 is illustrated, which can be used in at least one embodiment of this disclosure. The data center 900 may include a data center infrastructure layer 910, a framework layer 920, a software layer 930, and an application layer 940.

[0246] like Figure 9As shown, the data center infrastructure layer 910 may include a resource coordinator 912, grouped computing resources 914, and node computing resources (“nodes CR”) 916(1)-916(N), where “N” represents any complete positive integer. In at least one embodiment, nodes CR916(1)-916(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules and cooling modules, etc. In some embodiments, one or more nodes CR916(1)-916(N) may correspond to servers having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CR916(1)-916(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of nodes CR916(1)-916(N) may correspond to virtual machines (VMs).

[0247] In at least one embodiment, the grouped computing resources 914 may include individual groups (not shown) of nodes CR916 housed in one or more racks, or a plurality of racks (also not shown) housed in data centers in various geographic locations. Individual groups of nodes CR916 within the grouped computing resources 914 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several nodes CR916, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.

[0248] Resource coordinator 912 may be configured or otherwise control one or more nodes CR916(1)-916(N) and / or grouped computing resources 914. In at least one embodiment, resource coordinator 912 may include a Software Design Infrastructure (SDI) management entity for data center 900. Resource coordinator 912 may include hardware, software, or some combination thereof.

[0249] In at least one embodiment, such as Figure 9As shown, framework layer 920 may include a job scheduler 933, a configuration manager 934, a resource manager 936, and a distributed file system 938. Framework layer 920 may include a framework of software 932 supporting software layer 930 and / or one or more applications 942 of application layer 940. Software 932 or application 942 may respectively include web-based service software or applications, such as service software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 920 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 938 for large-scale data processing (e.g., "big data"). TM (Hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 933 may include a Spark driver for facilitating the scheduling of workloads supported by various layers of the data center 900. In at least one embodiment, the configuration manager 934 may be able to configure different layers, such as the software layer 930 and the framework layer 920, which includes Spark and a distributed file system 938 for supporting large-scale data processing. The resource manager 936 is able to manage cluster or grouped computing resources mapped to or allocated to support the distributed file system 938 and the job scheduler 933. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 914 at the data center infrastructure layer 910. The resource manager 936 may coordinate with the resource coordinator 912 to manage these mapped or allocated computing resources.

[0250] In at least one embodiment, the software 932 included in the software layer 930 may include software used by at least a portion of the nodes CR916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of software may include, but are not limited to, Internet web page search software, email virus browsing software, database software, and streaming video content software.

[0251] In at least one embodiment, the application layer 940 may include one or more applications 942 that can be used by at least a portion of the nodes CR916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0252] In at least one embodiment, any of the configuration manager 934, resource manager 936, and resource coordinator 912 can perform any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can alleviate the risk of data center operators of data center 900 making potentially poor configuration decisions and can prevent underutilization and / or skewed portions of the data center.

[0253] Data center 900 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above regarding data center 900. In at least one embodiment, by using weight parameters calculated through one or more training techniques, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks, such as, but not limited to, those described herein, using the resources described above regarding data center 900.

[0254] In at least one embodiment, the data center 900 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0255] Example network environment

[0256] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may... Figure 8 This is implemented on one or more instances of computing device 800—for example, each device may include similar components, features, and / or functions of computing device 800. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of data center 900, examples of which are described herein. Figure 9 To describe in more detail.

[0257] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks, or networks within multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (e.g., the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.

[0258] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the server functionality described herein can be implemented on any number of client devices.

[0259] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting software at the software layer and / or one or more applications at the application layer. The software or applications may respectively include network-based service software or applications. In embodiments, one or more client devices may use the network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open-source software network application framework, such as one that can use a distributed file system for large-scale data processing (e.g., "big data").

[0260] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions can be distributed across multiple locations from a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0261] Client devices may include those described in this article. Figure 8 The example computing device 800 described includes at least some components, features, and functions. By way of example and not limitation, the client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, aircraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.

[0262] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a personal digital assistant or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be practiced in a wide variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. This disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices linked via a communication network.

[0263] As used herein, the phrase "and / or" relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" could include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0264] This document describes in detail the subject matter of this disclosure to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways to include steps different from or similar combinations of steps described herein in conjunction with other current or future techniques. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.

[0265] Example paragraph

[0266] 1. A method comprising: determining, using one or more first machine learning models (MLMs) and sensor data obtained using one or more sensors associated with a machine, a first prediction of one or more correspondence scores between one or more object detection results and one or more object tracks, wherein the one or more correspondence scores are differentiable with respect to the one or more object detection results and the one or more object tracks; determining, using one or more second MLMs, a second prediction of one or more future movements associated with the one or more object detection results, wherein the one or more future movements are differentiable with respect to the one or more correspondence scores; determining, using one or more third MLMs, a third prediction of at least one trajectory of the machine, wherein the at least one trajectory is differentiable with respect to the one or more future movements; determining, using one or more fourth MLMs, a fourth prediction of one or more control sequences for the machine, wherein the one or more control sequences are differentiable with respect to the at least one trajectory; and performing one or more control operations on the machine based at least on the one or more control sequences.

[0267] 2. The method according to 1, wherein the one or more object traces are applied to the one or more second MLMs to generate the second prediction, and the one or more object traces are generated based at least on: matching the one or more object detection results with the one or more object traces using one or more differentiable combinatorial solvers and the one or more correspondence scores; and updating one or more previous versions of the one or more object traces based at least on the matching.

[0268] 3. The method according to any one of 1-2, wherein the one or more object detection results and the one or more correspondence scores are applied to the one or more second MLMs to generate the second prediction.

[0269] 4. The method according to any one of 1-3, wherein the one or more first MLMs, the one or more second MLMs, the one or more third MLMs, and the one or more fourth MLMs are trained at least based on backpropagating the loss corresponding to the one or more control sequences through the one or more fourth MLMs, the one or more third MLMs, the one or more second MLMs, and the one or more first MLMs.

[0270] 5. The method according to any one of 1-4, wherein the one or more third MLMs comprise one or more analytical functions having at least one parameter, the one or more analytical functions being trained to generate the third prediction of the at least one trajectory of the machine based at least on the one or more future movements.

[0271] 6. The method according to any one of 1-5, wherein the one or more fourth MLMs comprise one or more analytical functions having at least one parameter, the one or more analytical functions being trained to generate the fourth prediction of the one or more control sequences of the machine based at least on the at least one trajectory.

[0272] 7. The method according to any one of 1-6, wherein determining the third prediction of the at least one trajectory of the machine comprises: determining a plurality of candidate trajectories of the machine based at least on the one or more correspondence scores; determining cost values ​​of the plurality of candidate trajectories using one or more analytical functions of the one or more third MLMs, the one or more analytical functions being used to generate at least one prediction corresponding to the cost values; and selecting the at least one trajectory from the plurality of candidate trajectories based at least on the cost values.

[0273] 8. A system comprising: one or more processors configured to perform operations including: determining one or more control sequences for a machine using a machine learning model (MLM) sequence that predicts differentiable output sequences, the differentiable output sequences comprising: correspondence data between one or more object detection results and one or more object tracks, one or more object motion predictions corresponding to the correspondence data, one or more motion plans corresponding to the one or more object motion predictions, and the one or more control sequences corresponding to the one or more motion plans; and performing one or more control operations on the machine based at least on the one or more control sequences.

[0274] 9. The system of claim 8, wherein the one or more object traces are applied to at least one MLM to generate the one or more object motion predictions, and the one or more object traces are generated based at least on: matching the one or more object detection results with the one or more object traces using one or more differentiable combinatorial solvers and the correspondence data; and updating one or more previous versions of the one or more object traces based at least on the matching.

[0275] 10. The system according to any one of 8-9, wherein the one or more object detection results and the correspondence data are applied to at least one MLM to generate the object motion prediction.

[0276] 11. The system according to any one of 8-10, wherein the MLM sequence is trained at least based on backpropagating the loss corresponding to the one or more control sequences through the MLM sequence.

[0277] 12. The system according to any one of 8-11, wherein the MLM sequence comprises one or more analytic functions having at least one parameter, the one or more analytic functions being trained to predict the one or more motion plans based at least on the one or more object motion predictions.

[0278] 13. The system according to any one of 8-12, wherein the MLM sequence comprises one or more analytical functions having at least one parameter, the one or more analytical functions being used to predict the one or more control sequences of the machine based at least on the one or more motion plans.

[0279] 14. The system according to any one of 8-13, wherein the system comprises at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more analog operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more visual language models (VLMs); a system for performing operations using one or more multimodal language models; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0280] 15. At least one processor, comprising: one or more circuits, said one or more circuits being configured to: perform one or more control operations on said virtual machine within a simulated environment, based at least on one or more control sequences determined in response to a machine learning model (MLM) sequence that at least processes sensor data generated using one or more virtual sensors of said virtual machine, said machine learning model sequence being trained in an end-to-end process, and said simulated environment being generated using one or more optical transmission simulation algorithms, said one or more optical transmission simulation algorithms being configured to simulate at least one of lighting, shading, or shadow within said simulated environment.

[0281] 16. The at least one processor according to claim 15, wherein the MLM sequences predict differentiable output sequences, the differentiable output sequences comprising: correspondence data between one or more object detection results and one or more object traces, one or more object motion predictions corresponding to the one or more object detection results, one or more motion plans corresponding to the one or more object motion predictions, and the one or more control sequences corresponding to the one or more motion plans.

[0282] 17. At least one processor according to any one of 15-16, wherein the MLM sequences predict differentiable output sequences, the differentiable output sequences comprising correspondence data between one or more object detection results and one or more object traces and one or more object motion predictions, and the one or more object detection results and the correspondence data are applied to at least one MLM to generate the one or more object motion predictions.

[0283] 18. At least one processor according to any one of 15-17, wherein the MLM sequence is trained at least based on backpropagation of loss corresponding to the one or more control sequences for the virtual machine through the MLM sequence.

[0284] 19. At least one processor according to any one of 15-18, wherein the MLM sequence comprises one or more analytic functions having at least one parameter, the one or more analytic functions being trained to predict one or more motion plans based at least on one or more object motion predictions.

[0285] 20. The at least one processor according to any one of 15-19, wherein the at least one processor comprises at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more analog operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more visual language models (VLMs); a system for performing operations using one or more multimodal language models; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

Claims

1. A method comprising: Using one or more first machine learning models (MLMs) and sensor data obtained using one or more sensors associated with the machine, determine a first prediction of one or more correspondence scores between one or more object detection results and one or more object traces, wherein the one or more correspondence scores are differentiable with respect to the one or more object detection results and the one or more object traces; Using one or more second MLMs, determine a second prediction of one or more future movements associated with the one or more object detection results, wherein the one or more future movements are differentiable with respect to the one or more correspondence scores; Using one or more third MLMs, a third prediction of at least one trajectory of the machine is determined, wherein the at least one trajectory is differentiable with respect to the one or more future movements; Using one or more fourth MLMs, determine fourth predictions for one or more control sequences of the machine, wherein the one or more control sequences are differentiable with respect to the at least one trajectory; and One or more control operations are performed on the machine based on at least one or more control sequences.

2. The method of claim 1, wherein the one or more object traces are applied to the one or more second MLMs to generate the second prediction, and the one or more object traces are generated based on at least the following: The one or more object detection results are matched with the one or more object traces using one or more differentiable combinatorial solvers and the one or more correspondence scores; and At least based on the matching, update one or more previous versions of the one or more object traces.

3. The method of claim 1, wherein the one or more object detection results and the one or more correspondence scores are applied to the one or more second MLMs to generate the second prediction.

4. The method of claim 1, wherein the one or more first MLMs, the one or more second MLMs, the one or more third MLMs, and the one or more fourth MLMs are trained at least based on backpropagating the loss corresponding to the one or more control sequences through the one or more fourth MLMs, the one or more third MLMs, the one or more second MLMs, and the one or more first MLMs.

5. The method of claim 1, wherein the one or more third MLMs comprise one or more analytical functions having at least one parameter, the one or more analytical functions being trained to generate the third prediction of the at least one trajectory of the machine based at least on the one or more future moves.

6. The method of claim 1, wherein the one or more fourth MLMs comprise one or more analytical functions having at least one parameter, the one or more analytical functions being trained to generate the fourth prediction for the one or more control sequences of the machine based at least on the at least one trajectory.

7. The method of claim 1, wherein determining the third prediction of the at least one trajectory of the machine comprises: Multiple candidate trajectories of the machine are determined based on at least one or more correspondence scores; The cost values ​​of the plurality of candidate trajectories are determined using one or more analytical functions of the one or more third MLMs, wherein the one or more analytical functions are used to generate at least one prediction corresponding to the cost values; as well as At least one trajectory is selected from the plurality of candidate trajectories based on at least the cost value.

8. A system comprising: One or more processors, said one or more processors being configured to perform operations, said operations including: A machine learning model (MLM) sequence predicting a differentiable output sequence is used to determine one or more control sequences for the machine, the differentiable output sequence comprising: Correspondence data between the detection results of one or more objects and the traces of one or more objects. Motion prediction of one or more objects corresponding to the aforementioned correspondence data. One or more motion plans corresponding to the motion predictions of the one or more objects, and The one or more control sequences corresponding to the one or more motion plans; and One or more control operations are performed on the machine based on at least one or more control sequences.

9. The system of claim 8, wherein the one or more object traces are applied to at least one MLM to generate the one or more object motion predictions, and the one or more object traces are generated based on at least the following: The one or more object detection results are matched with the one or more object traces using one or more differentiable combinatorial solvers and the corresponding data; and At least based on the matching, update one or more previous versions of the one or more object traces.

10. The system of claim 8, wherein the one or more object detection results and the correspondence data are applied to at least one MLM to generate the object motion prediction.

11. The system of claim 8, wherein the MLM sequence is trained at least based on backpropagating the loss corresponding to the one or more control sequences through the MLM sequence.

12. The system of claim 8, wherein the MLM sequence comprises one or more analytic functions having at least one parameter, the one or more analytic functions being trained to predict the one or more motion plans based at least on the one or more object motion predictions.

13. The system of claim 8, wherein the MLM sequence comprises one or more analytical functions having at least one parameter, the one or more analytical functions being used to predict the one or more control sequences for the machine based at least on the one or more motion plans.

14. The system of claim 8, wherein the system comprises at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for collaborative content creation of 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing operations using one or more large language model LLMs; A system for performing operations using one or more visual language models (VLMs); A system for performing operations using one or more multimodal language models; A system for performing one or more conversational AI operations; A system for generating synthetic data; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

15. At least one processor, comprising: One or more circuits, said one or more circuits being configured to: perform one or more control operations on said virtual machine within a simulated environment, based at least on one or more control sequences determined in response to a machine learning model (MLM) sequence that has been trained in an end-to-end process, and said simulated environment being generated using one or more optical transmission simulation algorithms, said one or more optical transmission simulation algorithms being configured to simulate at least one of lighting, shading, or shadow within said simulated environment.

16. The at least one processor of claim 15, wherein the MLM sequences respectively predict differentiable output sequences, the differentiable output sequences comprising: The correspondence data between one or more object detection results and one or more object traces, one or more object motion predictions corresponding to the one or more object detection results, one or more motion plans corresponding to the one or more object motion predictions, and one or more control sequences corresponding to the one or more motion plans.

17. The at least one processor of claim 15, wherein the MLM sequences predict differentiable output sequences, the differentiable output sequences comprising correspondence data between one or more object detection results and one or more object traces and one or more object motion predictions, and the one or more object detection results and the correspondence data are applied to at least one MLM to generate the one or more object motion predictions.

18. The at least one processor of claim 15, wherein the MLM sequence is trained at least based on backpropagation of loss corresponding to the one or more control sequences for the virtual machine through the MLM sequence.

19. The at least one processor of claim 15, wherein the MLM sequence comprises one or more analytic functions having at least one parameter, the one or more analytic functions being trained to predict one or more motion plans based at least on one or more object motion predictions.

20. The at least one processor according to claim 15, wherein the at least one processor is comprised of at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for collaborative content creation of 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing operations using one or more large language model LLMs; A system for performing operations using one or more visual language models (VLMs); A system for performing operations using one or more multimodal language models; A system for performing one or more conversational AI operations; A system for generating synthetic data; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2