Guided uncertainty-aware policy optimization: Combining model-free and model-based strategies for sample-efficient learning
Patent Information
- Application Number
- DE102020129425
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-03
- Filing Date
- 2020-11-09
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2040-11-09
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] At least one embodiment relates to training robots to perform tasks. For example, at least one embodiment relates to training robots using models and artificial intelligence according to various novel techniques described herein. BACKGROUND
[0002] Training robots to perform tasks accurately can use considerable memory, time, or computational resources. Sometimes, this training may require extreme amounts of training data, which may be unavailable for some tasks or prohibitively expensive to obtain. In some examples, training may result in an excessively brittle or unstable system that does not reliably converge on a solution to a task. EP 3 549 725 A1 describes an apparatus for controlling a robot arm. DE11 2017 002 114 T5 discloses an object recognizer that recognizes a position and pose of an object from data measured by sensors. Finding ways to train more effectively and efficiently is a significant problem, and it is an object of the invention to improve the control of a robot.
[0003] This problem is solved by the features of the independent claims. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 illustrates a process that instructs a robot to complete a task using a combination of model-based and model-free methods, according to at least one embodiment; Fig. 2 illustrates an overview of a perception module for a model-based robot control system according to at least one embodiment; Fig. 3 illustrates an overview of a perception module for a model-free robot control system according to at least one embodiment; Fig. 4 illustrates examples of training images for a robot control system according to at least one embodiment; Fig. 5 illustrates an example of test results for various robot control methods according to at least one embodiment; Fig.6 illustrates a table of test results for the performance of a task according to at least one embodiment; Fig. 7 illustrates a process performed as a result of causing a computer system to instruct a robot to perform a task using a combination of model-based and model-free methods, according to at least one embodiment; Fig. 8A illustrates inference and / or training logic according to at least one embodiment; Fig. 8B illustrates inference and / or training logic according to at least one embodiment; Fig. 9 illustrates training and deployment of a neural network according to at least one embodiment; Fig. 10 illustrates an example data center system, according to at least one embodiment; Fig.11A illustrates an example of an autonomous vehicle according to at least one embodiment; Fig. 11B illustrates an example of camera locations and fields of view for the autonomous vehicle of Fig. 11A according to at least one embodiment; Fig. Figure 11C is a block diagram illustrating an example system architecture for the autonomous vehicle of Fig. 11A illustrates, according to at least one embodiment; Fig. 11D is a diagram illustrating a system for communication between a cloud-based server(s) and the autonomous vehicle of Fig. 11A illustrates, according to at least one embodiment; Fig. 12 is a block diagram illustrating a computer system according to at least one embodiment; Fig. 13 is a block diagram illustrating a computer system according to at least one embodiment; Fig.14 illustrates a computer system according to at least one embodiment; Fig. 15 illustrates a computer system according to at least one embodiment; Fig. 16A illustrates a computer system according to at least one embodiment; Fig. 16B illustrates a computer system according to at least one embodiment; Fig. 16C illustrates a computer system according to at least one embodiment; Fig. 16D illustrates a computer system according to at least one embodiment; Fig. 16E and Fig. 16F illustrate a shared programming model according to at least one embodiment; Fig. 17 illustrates example integrated circuits and associated graphics processors according to at least one embodiment; Fig. 18A and Fig.18B illustrate example integrated circuits and associated graphics processors according to at least one embodiment; Fig. 19A and Fig. 19B illustrate additional exemplary graphics processor logic according to at least one embodiment; Fig. 20 illustrates a computer system according to at least one embodiment; Fig. 21A illustrates a parallel processor according to at least one embodiment; Fig. 21B illustrates a partition unit according to at least one embodiment; Fig. 21C illustrates a processing cluster according to at least one embodiment; Fig. 21D illustrates a graphics multiprocessor according to at least one embodiment; Fig. 22 illustrates a multi-graphics processing unit (GPU) system according to at least one embodiment; Fig.23 illustrates a graphics processor according to at least one embodiment; Fig. 24 is a block diagram illustrating a processor microarchitecture for a processor, according to at least one embodiment; Fig. 25 illustrates a deep learning application processor according to at least one embodiment; Fig. 26 is a block diagram illustrating an exemplary neuromorphic processor, according to at least one embodiment; Fig. 27 illustrates at least portions of a graphics processor according to one or more embodiments; Fig. 28 illustrates at least portions of a graphics processor according to one or more embodiments; Fig. 29 illustrates at least portions of a graphics processor according to one or more embodiments; Fig.30 is a block diagram of a graphics processing engine of a graphics processor according to at least one embodiment; Fig. 31 is a block diagram of at least portions of a graphics processor core according to at least one embodiment; Fig. 32A and Fig. 32B illustrate thread execution logic with an arrangement of processing elements of a graphics processor core, according to at least one embodiment; Fig. 33 illustrates a parallel processing unit ("PPU") according to at least one embodiment; Fig. 34 illustrates a general processing cluster (“GPC”) according to at least one embodiment; Fig. 35 illustrates a memory partition unit of a parallel processing unit ("PPU") according to at least one embodiment; Fig.36 illustrates a streaming multiprocessor according to at least one embodiment. Fig. 37 is an example data flow diagram for an enhanced compute pipeline according to at least one embodiment; Fig. 38 is a system diagram for an example system for training, adapting, instantiating, and deploying machine learning models into an advanced compute pipeline, according to at least one embodiment; Fig. 39 includes an exemplary illustration of an enhanced compute pipeline for processing imaging data in accordance with at least one embodiment; Fig. 40A includes an example data flow diagram of a virtual device supporting an ultrasound device, according to at least one embodiment; Fig.40B includes an example data flow diagram of a virtual device supporting a CT scanner, according to at least one embodiment; Fig. 41A illustrates a data flow diagram for a process to train a machine learning model, according to at least one embodiment; and Fig. 41B is an example illustration of a client-server architecture for enhancing annotation tools with pre-trained annotation models, according to at least one embodiment. DETAILED DESCRIPTION
[0004] In at least one embodiment, techniques described herein demonstrate a robot control algorithm that combines strengths of a model-based method (MBM) with strengths of a model-free method (MFM). In at least one embodiment, an MBM is leveraged to provide efficient movement in a free-space environment. For example, in at least one embodiment, an MBM is used to operate a robot in a space where collisions with the environment or humans are easily avoided. In at least one embodiment, an MBM is used in combination with an MFM, which adds the capacity to learn an environment using a loosely defined target. In at least one embodiment, a perception system capable of predicting pose uncertainty is used to help the system fuse the MBM and MFM.
[0005] Techniques such as deep reinforcement learning (“RL”) enable robots to learn and act from raw sensory data. For example, RL methods can be successfully applied to contact-rich manipulation tasks, such as inserting, pushing, and grasping objects. However, in some situations, the RL method can be sample-inefficient, requiring many training interactions with the environment to be practical for real-world applications. To mitigate this limitation, various embodiments can carefully tune densely formed rewards, which often requires explicit knowledge of the state of the world, such as a goal location. On the other hand, given an accurate model and current state of the world, there are many algorithms to plan, develop, or search for policies to accomplish the task.This type of model-based ("MB") strategy can be used for many manipulation tasks, such as pin insertion, grasping, and reaching. Nevertheless, in some examples, MB strategies can be hampered by model bias and state estimation errors and often achieve lower asymptotic performance.
[0006] At least one embodiment described herein combines the rehearsal efficiency of a model-based policy and overcomes the errors in the dynamics and perceptual models with a reinforcement learning policy that closes the loop on raw sensory data. At least one embodiment leverages the efficiency of a model-based method to move into free space, where collisions and contact with the environment and humans are impossible. At least one embodiment uses the capacity of an RL method to learn from the agent's interactions with the environment and a loosely defined target. To switch between the MBM and RL, various examples introduce a perception system that can predict pose uncertainty to help the system fuse the two policies.In at least one embodiment, the task is initialized with an MB policy, moving the robot within the uncertainty region of the object of interest, e.g., the box where the pen is to be inserted. In at least one embodiment, when the system reaches the uncertainty region, it switches to an RL policy to complete the task. At training time, we leverage information from the RL to reduce the uncertainties of the perceptual system. At least one embodiment leverages uncertainty to blend MB and RL policies to accomplish complicated tasks that would be quite challenging using either method alone.
[0007] In various embodiments, robotic approaches rely on an accurate model of an environment, a detailed description of how to perform a task, and a robust perceptual system to keep track of a current state. In some other embodiments, reinforcement learning ("RL") approaches operate directly from raw sensory inputs, using a reward signal to describe the task. In one embodiment, a system is developed to obtain a general method capable of overcoming inaccuracies of elements in a conventional pipeline while requiring minimal interaction with an environment.In at least one embodiment, this is achieved by leveraging uncertainty estimates to partition the space into regions where the given model-based policy is reliable and regions where it is defective or not well defined. In at least one embodiment, in these hard regions, a local model-free policy is learned directly from raw sensory input. These "hard regions" may also be referred to as "uncertain regions" because, in one embodiment, the model used in the model-based method may be uncertain. In at least one embodiment, the system enables faster construction of robotic systems from simple and inexpensive components and only a high-level description of the task.In at least one embodiment, an algorithm called Guided Uncertainty-Aware Policy Optimization (“GUAPO”) is used in a real-world robot to perform a tight pin insertion task.
[0008] In at least one embodiment, the task is initialized with the MBM, which method moves the robot within the range of uncertainties of the object of interest, e.g., the box into which a pen is to be inserted. In at least one embodiment, after reaching a defined region near the goal, the system switches the control method to an MFM to complete the task. In at least one embodiment, information from the MFM task completion is leveraged at learn time to reduce perceptual system uncertainties. In at least one embodiment, the system mixes MBM and MFM to perform a complicated task requiring environmental interactions. In various examples, the switching between MBM and MFM is learned through RL or optimization.
[0009] In at least one embodiment, a simple yet efficient way is used to express pose uncertainties for a keypoint-based pose estimator. In at least one embodiment, this extension augments the peak estimation algorithm by fitting a 2D Gaussian around each detected peak. In at least one embodiment, the PnP algorithm is executed on n sets of keypoints, where each set of keypoints is constructed by sampling all of the 2D Gaussians. In at least one embodiment, this provides n possible poses of the object consistent with the detection algorithm, which are treated as equally likely in some examples described herein.
[0010] In at least one embodiment, a perception module is used to sense and track the state of the world. In some examples, a simple perception failure in this context can be catastrophic for a robot, as its motion generator may rely on it. Furthermore, in some situations, classical motion generators are rigid and inflexible in how they perform a task. This can lead to an inability to adapt to changing conditions. For example, a robot controlled by a classical motion generator may only be able to select an object in a specific way and be unable to recover if grasping fails. In at least one embodiment, these problems can make robotic systems based on such control algorithms unstable in new domains and difficult to scale.To expand the robotics reach, several robust, adaptive and flexible systems can be used, as described here.
[0011] In at least one embodiment, a robot system leverages the capacity of robotics operators to model its environment. An MBM can be used to navigate within robot-free space, such as path following. In various embodiments, using an MBM when the robot needs to interact with its environment—e.g., grasping an object, placing an object, object insertion, etc.—can be difficult due to limitations or inaccuracies in the model. Real-world physical systems can be complex and stochastic. Within a robotics simulator, modeling the correct parameters that vividly represent real robot behavior can be a non-trivial task. In various embodiments, using the MBM in common tasks can be difficult and complex. Furthermore, the MBM often relies on perceptual systems that are imperfect.Managing these potential errors can be difficult when mixed with various complex physical systems. Furthermore, the MBM can be time-consuming and costly to use because it can be tied to the roboticist's expertise and ingenuity.
[0012] In various embodiments, an MFM has the capacity to adapt and immediately deal with raw sensory inputs that cannot be subject to estimation errors. In at least one embodiment, the strength of the MFM derives from its capacity to define a task at a higher level by a reward function that specifies what to do, rather than by an explicit set of control actions that describe how the task should be performed. In at least one embodiment, an MFM does not require specific physical modeling because it implicitly learns such from interaction with an environment, allowing the method to be deployed in different settings.In at least one embodiment, MFM methods have various limitations, including that random interaction with an environment for human users as well as for different materials can be complex, and that MFM is sometimes not sample-efficient. In some examples, introducing MFM into a new environment is difficult and complex.
[0013] In at least one embodiment, a system implements an algorithm that combines the strengths of MBM and MFM. In at least one embodiment, the efficiency of MBM to navigate in a free-space environment—e.g., a space where collisions with the environment or people are easily avoided—is leveraged with the capacity of MFM to learn from its environment of a loosely defined target. In at least one embodiment, a perception system predicts pose uncertainties to help the system fuse MBM and MFM.
[0014] Model-based methods are control methods that rely on a physical model of the environment to control the robot. A physical model (sometimes simply referred to as a model) provides locations and poses for various objects in the robot's environment and, in some examples, a pose of the robot itself. For example, a physical model can be a model of a fixed object of the robot and objects in the robot's vicinity. Such methods plan movements based on the physical model, which in various examples includes object avoidance and object manipulation. Such systems often use a perception system, such as a camera or depth camera, to locate objects and determine an orientation and position, or pose, of the model.When under the control of a model-based method, the robot's performance is generally limited by the accuracy of the model and the accuracy of the perception system. An alternative to the model-based method is a model-free method, such as reinforcement learning. Model-free methods are control methods that do not rely on a physical model of the environment to operate. Such systems can operate directly from sensor data, without the use of a model, in various examples. For example, a model-free method can manipulate an object using image data from a handheld camera and based on future movements on the images rather than an explicit physical model. However, model-free systems can be very time-consuming and difficult to train, especially for complex tasks.
[0015] Fig.1 illustrates a process that instructs a robot to complete a task using a combination of model-based and model-free methods, according to at least one embodiment. Fig.1 illustrates an example 100 of a real-world setting for pin insertion. In at least one embodiment, a robot 102 attempts to complete a task involving inserting a pin into a hole 104. In at least one embodiment, a first perception system 106 (such as a camera) provides the approximate position of the relevant objects, and a model-based method 110 drives the system within the uncertainty region 112. Once within this uncertainty region 112, the model cannot be trusted, and a model-free policy 114 is immediately learned from the raw sensory inputs of a second perception system 108, which provides enough information to complete the task 116. In various examples, a perception system may be a camera, infrared camera, radar, lidar, depth camera, or 2D or 3D imaging system.
[0016] Fig.1 illustrates an overview of a system according to various embodiments in which a task 116 may be initialized with the MFM, which may move the robot within the range of uncertainties of the object of interest, e.g., a box, where the pin is to be inserted. In at least one embodiment, the MFM may then be used to complete the task. In at least one embodiment, information from the completion of the MFM task is used at training time to reduce perceptual system uncertainties. In at least one embodiment, the system uses MBM and MFM to perform complicated tasks that require environmental interactions. In various embodiments, the system outperforms MFM or MBM classical approaches for various complicated tasks, such as pin insertion.In at least one embodiment, the system is used to express pose uncertainties for a keypoint-based pose estimator. In at least one embodiment, the system is sample-efficient for learning methods on real-world robots.
[0017] At least one embodiment described herein solves the problem of learning to perform an a priori unknown operation in a domain for which there can only be an estimated location and no exact model. In at least one embodiment, the problem can be formalized as solving a Markov Decision Process (“MDP”) in which a particular reward r = S → ℝ is obtained by learning a policy π : S × A → ℝ +which is a probability distribution over actions a ∈ A at any given state s ∈ S. In at least one embodiment, a first assumption used within the system is partial knowledge of the transition function P : S × A × S → ℝ + which dictates the probability of next states when a particular action is applied to the current state. In at least one embodiment, the transition is assumed to occur only within a subspace of the state space S open⊂ S. is available. This is useful in various robot systems where it can be determined how the robot moves while in open space, but where there are no reliable and accurate models of general contacts and interactions with the robot's environment. In at least one embodiment, this submodel is combined with various methods capable of planning and executing trajectories that S open through, but completing tasks that require action in S hard = S\S open require, can still be difficult and complex.
[0018] In various embodiments, tasks that can be solved include reaching a certain state or configuration through interaction with the environment, such as pin insertion, toggling switches, or grasping. In various embodiments, these tasks are represented by a binary reward function r(s) = 1[s ∈ S g ] which determines the successful achievement of a set goal S g ⊂ S hard This reward can be extremely sparse, and detecting random actions may therefore require a large number of samples. Furthermore, in many examples, the exact location of S hard but only a noise estimate of it. In at least one embodiment, the system learns to efficiently solve the full task by interacting with the environment using various imperfect perceptual systems and dynamics.
[0019] In at least one embodiment, a model-free method ("MFM") is used to perform a task. In some embodiments, performing an MFM is inefficient and brittle because it may be necessary to learn how to control the robot everywhere, and if the position of the target changes, the task may look different. In at least one embodiment, a partial model is leveraged with a model-based method ("MBM"), which determines the uncertainty in the perception and actuation systems that guide the agent to the region of interest, thus reducing the area where the MFM policy may need to be optimized and making it more invariant to the absolute target location. In this document, various embodiments of the algorithm are referred to as guided uncertainty-aware policy optimization ("GUAPO").
[0020] In at least one embodiment, the GUAPO comprises a method for generating a superset S^ hard by S hard based on the perception system uncertainty estimate. In at least one embodiment, the set is used to partition the space into the regions where the MBM is used and regions where the MFM is trained. In at least one embodiment, an MBM is determined that is outside of S^ hard is used to bring the robot into the crowd. In at least one embodiment, the MFM is defined and learning is made more efficient by making its inputs local.
[0021] In many examples, it may be faster to set up coarse perception systems because they may require simpler hardware, such as RGB cameras, and can be used out of the box without excessive tuning and calibration efforts. When such a system is used tohard to locate immediately, the perceptual errors can misleadingly indicate that a particular area belongs to S open which may lead to attempting to apply the MBM and potentially not being able to learn how to recover from it. In various embodiments, a perception system is used that also provides an uncertainty estimate. In at least one embodiment, various methods are used to estimate the uncertainty by a nonparametric distribution with n possible poses of the region {Shardi}i=1n and their assigned weights p(Shardi) If these weights are interpreted as the probabilities, the probability of a particular state s leading to S hard heard, can be expressed as: p(s∈Shard)=∑i=1n1[s∈Shardi]p(Shardi)
[0022] In at least one embodiment, the perception system provides a parametric distribution and the above probability can be calculated by integration or approximated in such a way that the set S^ hard = {s : p(s ∈ S hard ) > ε} a superset of S hard for a corresponding set ε by a user. In at least one embodiment, a more accurate perception system S^ hard a narrower superset of S hard , thus further reducing the area where the model-free method is required. α(s) can be written as α(s) = 1[s ∈ S^ hard ] and the entire policy used can be: π(a|s)=(1−α(s))∗πMB(a|s)+α(s)∗πMF(a|s) where π MB (a|s) and π MF(a|s) may be the model-based and model-free policies, respectively. In at least one embodiment, a switch between these two policies is therefore used based at least in part on the uncertainty estimate. For example, a system may transition from MBM to MFM if it is determined that the robot is within S^ hard with a margin based on the uncertainty of the perception module. This would, in various examples, cause the robot to be determined to be within a subregion of S^ hard A subregion is a region that is completely contained in S^ hard which has a total volume smaller than S^ hard is.
[0023] In at least one embodiment, S^ hardthe region where it is known that there is a certain reward for achieving the task, but not how the task is to be achieved. In at least one embodiment, it is assumed that outside of this region, the environment model is well known and it is therefore amenable to use a model-based approach. In at least one embodiment, therefore, outside of S^ hard , the model-based approach returns the robot to it. In at least one embodiment, the system selects a specific point within S hard , such as its center of gravity, which can be found as the most likely location based on the probability distribution obtained from the perception system, and sets this location as a target for the model-based method. In at least one embodiment, this can ensure that the robot moves in the direction S^ hardgoes whenever he is outside of it.
[0024] In at least one embodiment, the formulation is extended to consider multiple rewards. For example, if there is an obstacle that must be avoided and there is uncertainty about its location, S^ hard as Shardgoal∪Shardobst be described, where Sg⊂Shardgoal and r(s)=−1∀s∈Shardobst. Then, in at least one embodiment, an obstacle-avoiding MBM may be used to get to the area where the target is while avoiding the regions where the obstacle might be.
[0025] In at least one embodiment, once π MB the system in S^ hard brought the control to π MF passed as expressed in the policy. In at least one embodiment, the switch definition can go both ways and therefore, if π MFundertakes exploration actions outside of S^ hard , move, the MBM will act again to push the state into the region of interest. In at least one embodiment, this provides a framework for safe learning in the case where there may be obstacles to be avoided, as explained above. In at least one embodiment, there may be several advantages to having a more restricted region where the MFM must learn how to act: first, exploration may become easier, and second, the policy may be local. For example, in at least one embodiment, only images π MF The images are from a wrist-mounted camera and their current speeds, as shown in Fig. 3 is clearly illustrated.
[0026] In at least one embodiment that does not include global information such as that provided by the perception system in Fig.2, the MFM policy generalizes better over locations of S^ hard . In at least one embodiment, a policyless MFM algorithm is used such that all of the observed transitions in the replay buffer can be added, regardless of whether they are from π MB or π MF come.
[0027] In at least one embodiment, this framework uses each newly captured experience to generate S^ hard such that subsequent implementations can use the model-based procedure in larger regions of the state space. For example, in the pen insertion task, once the reward of fully inserting the pen is received, the location of the opening can be immediately known and S^ hard = S hard can be updated, whereby the model-free procedure now only has to do the actual insertion and does not have to search for the opening any further.
[0028] In at least one embodiment, a GUAPO algorithm for a close-fitting pin insertion task is developed using Franka Panda, a 7-DoF torque-controlled robot, although in various embodiments, any robot may be used. In at least one embodiment, a state estimation module is used and an uncertainty estimate is obtained to locate S^. In at least one embodiment, a model-based policy is used to locate S open to navigate while avoiding obstacles, and an RL algorithm and a model-free policy architecture are also used.
[0029] In at least one embodiment, a deep object pose estimator (DOPE) is used as a perception system. In at least one embodiment, the DOPE uses a simple neural network architecture that can be rapidly trained with synthetic data and domain randomization.
[0030] Fig. 4 illustrates examples of training images for a robot control system according to at least one embodiment. Fig.Figure 4 illustrates generated images with heavy use of domain randomization used to train a perception system. In some examples, the model of the object that DOPE needs to capture may not be very detailed, consisting primarily of the shape. In at least one embodiment, no depth sensing is used to supplement the RGB information. In at least one embodiment, the algorithm first finds the keypoints of the cuboid-shaped object using a local vertex on the map. Given the real dimensions of the cuboid, camera intrinsics, and the keypoint locations, DOPE may execute a "Perspective-n-Point" (PnP) algorithm to, in one embodiment, find the final object pose in the camera frame. In the Fig. In the example shown in Figure 4, the training images illustrate a hole box with different occlusions or backgrounds.
[0031] Fig.2 illustrates an overview of a perception module for a model-based robot control system according to at least one embodiment. In at least one embodiment, Fig. 2 DOPE perception and uncertainty to S hard In at least one embodiment, an image 202 obtained from a camera is provided to a DOPE perception system 204, which provides an estimate of locations for objects in the environment of a robot.
[0032] In at least one embodiment, the DOPE perception system is extended to obtain uncertainty estimates of the object pose. In at least one embodiment, the extension augments the peak estimation algorithm by fitting a 2D Gaussian function around each detected peak, as represented by the dark contour maps 206 in Fig.2. In at least one embodiment, a PnP algorithm is executed on a set n of keypoints, where each set of keypoints is constructed by sampling the 2D Gaussian functions. In at least one embodiment, this provides n possible poses 208, 210, and 212 of the object consistent with the detection algorithm, as shown in Fig. 2. In at least one embodiment, they may be considered equally likely.
[0033] In some examples, access to a rough description of the area of interest S hard214 around the object where an operation must be performed. In one embodiment of the pin insertion task, this is a rectangle centered at the opening of the hole. In at least one embodiment, for each of the n pose samples given by the extended DOPE perception algorithm, its associated hole opening position {pi}i=1n be calculated, which are in Fig. 2. In at least one embodiment, these points are then fitted by a 3D Gaussian function with diagonal covariance, shown in blue in the same figure. In at least one embodiment, the mean μhole=1n∑i=1npi as the center of S^ hard used and the probability of a certain state leading to S hard belongs, is moved by S hard along the axis by one standard deviation.
[0034] In at least one embodiment, the full setting in Fig. 1 is clearly shown, where the camera for DOPE (640x480x3 RGB images) is mounted on the workspace and an example image is provided.
[0035] In at least one embodiment, a model-based controller uses target attractors defined by Riemannian motion policies (“RMPs”) to move the robot toward a desired end-effector location. In at least one embodiment, the RMPs assume a desired end-effector position x x ∈ ℝ 3 in Cartesian space. In at least one embodiment, the target is set to be the center of gravity of S^ hard which corresponds to the opening of the hole µ^ holeIn at least one embodiment, a coarse model of the object may be used to train a perception module capable of providing this localization estimate and its uncertainty. In at least one embodiment, the RMPs also use a model of the robot. In at least one embodiment, these two features may contribute to the "model-based" component of the system. In at least one embodiment, the GUAPO algorithm does not require these models to be extremely accurate. In at least one embodiment, if obstacles must be avoided to achieve S^, barrier-type RMPs may be defined.
[0036] In various embodiments, the policies send end effector position commands at 20 Hz. In at least one embodiment, the RMPs calculate desired joint positions q dat 1000 Hz. In at least one embodiment, given that the impedance end-effector controller is an action space that improves the sampling efficiency for policy learning for RL, the interface of the RMPs can be used as a model-free action space.
[0037] Fig. Figure 3 illustrates an overview of a perception module for a model-free robot control system, according to at least one embodiment. In at least one embodiment, a variational autoencoder for a model-free system is used to complete a task. In at least one embodiment, an RGB image 302 and velocities are provided to an autoencoder 306 and 308.
[0038] In at least one embodiment, a model-free, off-policy RL algorithm, such as Soft Actor Critic, is used. In at least one embodiment, the model-free policy acts directly on raw sensory inputs. In at least one embodiment, this consists of joint velocities and images from a camera mounted on a robot's wrist (e.g., 64x64x3 RGB images from a Logitech Carl Zeiss Tessar) (see Fig. 1). In at least one embodiment, as in Fig.3, inputs are fed to a β-variational autoencoder ("VAE") that gives a low-dimensional latent space representation of the state. In at least one embodiment, the parameters of this VAE are trained in advance on an offline collected dataset. In at least one embodiment, the portion that can be learned by the RL algorithm is a two-layer multi-level perception ("MLP") 314 and 316 that takes as input the 64-dimensional latent representation 310 given by the VAE and generates a 3D position displacement Δx of the robot end effector. In at least one embodiment, the autoencoder is provided with a robot action 312 that estimates a next velocity 318 and a reconstructed image 320.
[0039] In at least one embodiment, the VAE is trained with at least 160,000 data points for 12 epochs on a graphics processing unit, such as the Titan XP GPU. In at least one embodiment, the DOPE is trained for 8 hours on four P100 GPUs, although any processing unit may be used in various embodiments. In at least one embodiment, GUAPO, SAC, and Residual are trained for 60 training episodes of 1,000 steps each, which in some examples takes approximately 90 minutes.
[0040] In at least one embodiment, a sparse reward is used for the GUAPO when the policy completes the task (inserting the pen). In at least one embodiment, the policy receives -1 everywhere and 0 when it completes the task. In at least one embodiment, a negative L2 norm is applied to the perceptual estimate of the target location µ^ for SAC and residual. hole, a 0 reward if they S^ hard reached and uses 1 when she finishes the task.
[0041] The performance of the GUAPO algorithm can be measured in various examples. In one example of performance metrics, three overall baselines are used. Initially, there are baselines that do not involve learning. In at least one embodiment, these are what can be described as static "model-based" methods that do not leverage real-world interactions to update their policies. Thus, such embodiments cannot recover from failures and can be represented in various graphs as horizontal dotted lines. A second type of baseline is a similar model-free algorithm used in one or more embodiments of the GUAPO method, but without leveraging a model-based policy to reduce its working space. Here, one example uses the Soft Actor Critic ("SAC") to compare with GUAPO.Finally, residual learning techniques can be compared, which may also attempt to combine model-based and model-free methods.
[0042] Within the category of static model-based baselines, the performance of a script-guided method may be considered. In at least one embodiment, the most direct method may use the same attractor-based control that may be used in the GUAPO method to determine the center of gravity of S^ hardimmediately, and then use a scripted policy that terminates to perform the task. In at least one embodiment, for the pen insertion task, this may be hard-coding a direct lowering of the pen. In at least one embodiment, this method is very sensitive to errors in the perception system because even a few millimeters of error can cause the pen to be inserted improperly. In at least one embodiment, to provide improved chances of inserting the pen and consequently completing the task, random actions may be added to this downward movement. In at least one embodiment, this is shown by the curve of a model-based policy with random actions using DOPE goal estimates (MB-RA-DOPE-EST), as shown in Fig.5 is clearly illustrated.
[0043] Fig. Figure 5 illustrates an example of test results for various robot control methods according to at least one embodiment. A first graph 502 illustrates task completion versus a number of training iterations for task success, and a second graph 504 illustrates the number of steps required for task completion versus training iterations. Fig.Figure 5 illustrates task results comparing GUAPO with five other baselines: (1) Model-Based Policy with Perfect Goal Estimate (MBPERF-EST), (2) Model-Based Policy with Random Actions with Perfect Goal Estimate (MB-RA-PERF-EST), (3) Model-Based Policy with Random Actions using DOPE Goal Estimates (MB-RA-DOPE-EST), (4) Model-Free Soft Actor Critic (SAC), and (5) Residual policy. In this example, GUAPO, SAC, and Residual are trained for 60 training episodes. In at least one embodiment, after 60 episodes (approximately 90 minutes of training time), GUAPO is able to insert the stylus 100% of the time and reduce the number of steps required for stylus insertion.In at least one embodiment, when DOPE is used, a model-based algorithm ("MB") fails to solve the task due to errors in the perceptual system, and SAC and Residual are unable to complete the task with only 60 episodes.
[0044] To provide an oracle and to demonstrate that this scripted approach can operate under perfect state estimation, the performance of this policy under perfect state estimation can be examined using Model-Based Policy with Perfect Goal Estimate (MB-PERF-EST) and Model-Based Policy with Random Actions with Perfect Goal Estimate (MB-RA-PERF-EST). MB-PERF-EST can walk straight to the hole and push down without performing any random actions. In some examples, random actions do not unduly degrade the full performance achieved by the same scripted approach, but without any added random action. When there are perceptual errors, MB-RA-DOPE-EST may be able to perform better than MB-DOPE-EST in some examples, as in Fig. 6 seen. Fig. 6 illustrates a table of test results for performing a task according to at least one embodiment.
[0045] In at least one embodiment, a model-based method, such as SAC, is compared without various model-based components. In some examples, SAC may not be able to achieve arbitrary deployment. In some examples, this is due to an extremely low data regime. In some examples, RL may require several orders of magnitude of additional data.
[0046] In at least one embodiment, residual learning may also be considered. In at least one embodiment, this method applies random actions to a given policy. In some examples, this is a scripted policy. In various embodiments, so-called "residual learning" resulted in various errors. In at least one embodiment, a known perturbation case is applying large perturbations too far from the hole opening and therefore landing on the side of the box where the hole is, and then pushing against the side of the box instead. In at least one embodiment, the system only turns on the model-free section once it is already close to the region of interest, avoiding this perturbation case.
[0047] Results are shown in Fig. 5 and in Fig.6 in one embodiment. In at least one embodiment, the model-based policies with perfect perceptual estimates ("MB-PERF-EST") can outperform GUAPO because the policy knows exactly where the box's hole is and can use a hand-scripted movement to complete the task. However, in at least one embodiment, when DOPE is used as the perception system, which can have approximately 2.5 to 3.5 cm of noise and error values, the performance of MB-DOPE-EST and MB-RA-DOPE-EST drops drastically. In at least one embodiment, MB-RA-DOPE-EST performs 26.6% better than MB-DOPE-EST because the random actions offset the perceptual error. In at least one embodiment, both Residual and SAC are unable to perform the task within the training timeframe. However, in at least one embodiment, Residual is able to hard100% of the time after 60 training episodes, whereas SAC is still unable to reach this region.
[0048] In at least one embodiment of robotic manipulation, there are different paradigms for performing a task, such as model-based and model-free paradigms. In at least one embodiment, the first category of methods may rely on an accurate description of the task, such as the exact CAD model of all objects, as well as various perception systems. In some examples, it may be used with various search algorithms, such as motion planning. In at least one embodiment, this approach is limited by implementation, and both may be associated with irrecoverable interference if the perception system contains some noise.
[0049] In at least one embodiment, a model-free approach does not require detailed description, but instead requires access to interaction with the environment, as well as a reward that can indicate success. In at least one embodiment, such binary rewards may be easy to describe, but they may make RL methods extremely rehearsal-inefficient, and extremely sophisticated rewards may be used, requiring considerable tuning and precise perceptual systems. In at least one embodiment, automatic curriculum generation or the use of demonstrations may be used and may require large amounts of interaction with the environment. Furthermore, if the position of objects in the scene changes or there are new distractors in the background, these methods may need to be retrained in various embodiments.In at least one embodiment, the developed system is even sample-efficient when using only a sparse reward for success and is robust to these variations due to the model-based component.
[0050] In at least one embodiment, object pose estimation is used in various robotics and computer vision applications. In at least one embodiment, inference is used to keypoints on the object or on a cuboid surrounding the object. In at least one embodiment, keypoints are first detected by a neural network, then PnP is used to predict the object's pose. In at least one embodiment, uncertainty is exploited by leveraging a random sample matching (Ransac) voting algorithm to find regions where a keypoint might be detected. In at least one embodiment, PnP allows for the use of a probabilistic PnP, where different keypoints are weighted according to their spatial distribution.In at least one embodiment, regression may be performed on a vector tuning map, and the line intersection may then be used to find keypoints. In some embodiments, pose uncertainty is not considered in the final prediction.
[0051] In one embodiment, a GUAPO algorithm is developed that combines the generalization capabilities of model-based methods and the adaptability of model-free methods. In at least one embodiment, it may allow the task to be loosely defined by providing only a rough model of the objects and a rough description of the region where an operation needs to be performed. In at least one embodiment, the model-based system may leverage high-level information and low-cost state estimation systems to generate a funnel around the region of interest.In at least one embodiment, an uncertainty estimate provided by perception systems can be used to automatically switch between a model-based policy and a model-free policy that may be capable of learning from a sparse reward, which can overcome the model and estimation errors of the model-based part. In at least one embodiment, learning is achieved in the real world of a close-matching pin insertion task.
[0052] Fig. Figure 7 illustrates a process which may be performed as a result of being performed by a computer system such as the one described below in Fig.8-41 and associated description, the computer system causes a robot to perform a task using a combination of model-based and model-free methods as described above, according to at least one embodiment. In at least one embodiment, the process begins at block 702, with the computer system receiving information from a first perception system. A perception system may be a camera, depth camera, LIDAR, RADAR, or other sensory system that provides position or localization information, for example, as listed below. In at least one embodiment, the perception system is a stationary camera that observes the robot and the surrounding environment.In at least one embodiment, at block 704, the information is processed with a pose estimator such as DOPE, as described above, to generate a physical model of the environment surrounding the robot. In at least one embodiment, the physical model is a 2D or 3D model of objects and surfaces around the robot. In at least one embodiment, the physical model includes pose and localization information of objects and, in some embodiments, of the robot itself. The robot may be an articulated robot, an autonomous vehicle, a pick and place machine, or other machine under the control of the computer system. In at least one embodiment, an uncertainty of the information is determined 706 by the first perception system, as described above. In at least one embodiment, the uncertainty is a distribution of possible object and / or robot poses.In at least one embodiment, the uncertainty is based on a camera resolution and an error of the first perception system.
[0053] In at least one embodiment, the computer system uses the physical model at block 708 to implement a model-based method, such as the MBM described above, to move the robot into a region of space, which in some examples is determined based on the uncertainty. In at least one embodiment, the region of space is an area where the computer system determines that the model is no longer sufficiently accurate to enable the robot to complete a task. In at least one embodiment, the computer system determines whether the robot is in the region based on the uncertainty at decision block 710. For example, using the uncertainty, the system may determine that the probability that the robot is within the region is greater than a threshold or percentage.If the robot is not in the region, execution returns to block 702, but if it is determined that the robot is within the region with the required security, execution proceeds to block 712.
[0054] In at least one embodiment, at block 712, the computer system receives information from a second perception system. In at least one embodiment, the second perception system is an in-hand camera mounted on a robot manipulator or a camera mounted on the wrist of the robot. In at least one embodiment, at block 714, the computer system provides the information from the second perception system to a model-free method that generates control instructions to the robot from the information without relying on a physical model. In at least one embodiment, at block 716, the robot is moved using the control instructions generated using the model-free method. In at least one embodiment, at decision block 718, the computer system determines whether the task performed by the model-free method has been completed.If not, execution returns to block 712. In at least one embodiment, when the task is complete, execution proceeds to block 720 and the process ends. In various examples, the task may be inserting a pin into a hole, as described above, or operating an autonomous vehicle, as described below. INFERENCE AND TRAINING LOGIC
[0055] Fig. 8A illustrates the inference and / or training logic 815 used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described below in connection with Fig. 8A and / or 8B provided.
[0056] In at least one embodiment, the inference and / or training logic 815 may include, but is not limited to, code and / or data storage 801 for storing feedforward and / or output weights and / or input / output data and / or other parameters for configuring neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 815 may include or be coupled to code and / or data storage 801 for storing graphics code or other software for controlling the timing and / or order in which weight and / or other parameter information is to be loaded to configure logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)).In at least one embodiment, code, such as graphics code, loads weight or other parameter information into processor ALUs based on a neural network architecture to which the code corresponds. In at least one embodiment, code and / or data storage 801 stores weight parameters and / or input / output data of each layer of a neural network trained or used in connection with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 801 may be included in other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.
[0057] In at least one embodiment, any portion of the code and / or data storage 801 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or the code and / or data storage 801 may be cache memory, dynamic RAM ("DRAM"), static RAM ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or the code and / or data storage 801 is, for example, internal or external to a processor, or is DRAM, SRAM, flash memory, or another type of storage, may be on-chip versus off-chip.off-chip available storage, latency requirements of the training and / or inference functions performed, the batch size of the data used in inferencing and / or training a neural network, or a combination of these factors.
[0058] In at least one embodiment, the inference and / or training logic 815 may include, but is not limited to, code and / or data storage 805 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the code and / or data storage 805 stores weight parameters and / or input / output data of each layer of a neural network trained or used in connection with one or more embodiments during backpropagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments.In at least one embodiment, training logic 815 may include or be coupled to code and / or data storage 805 to store graphics code or other software for controlling the timing and / or order in which weight and / or other parameter information is to be loaded for configuring logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)).
[0059] In at least one embodiment, code, such as graphics code, loads weight or other parameter information into processor ALUs based on a neural network architecture to which the code conforms. In at least one embodiment, any portion of code and / or data storage 805 may be coupled to other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 805 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 805 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether the code and / or data storage 805 is, for example, internal or external to a processor, or consists of DRAM, SRAM, flash memory, or another storage type, may depend on on-chip versus off-chip available memory, latency requirements of training and / or inference functions performed, the batch size of the data used in inferencing and / or training a neural network, or a combination of these factors.
[0060] In at least one embodiment, code and / or data storage 801 and code and / or data storage 805 may be separate storage structures. In at least one embodiment, code and / or data storage 801 and code and / or data storage 805 may be a same storage structure. In at least one embodiment, code and / or data storage 801 and code and / or data storage 805 may be partially a same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 801 and code and / or data storage 805 may be combined with other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.
[0061] In at least one embodiment, the inference and / or training logic 815 may include, but is not limited to, one or more arithmetic logic units ("ALU(s)") 810, including integer and / or floating point units, to perform logical and / or mathematical operations based at least in part on or indicated by training and / or inference code (e.g., graphics code), the result of which may produce activations (e.g., output values of layers or neurons within a neural network) stored in an activation storage 820 that are functions of input / output and / or weight parameter data stored in the code and / or data storage 801 and / or the code and / or data storage 805.In at least one embodiment, activations stored in activation storage 820 are generated in accordance with linear algebraic and / or matrix-based mathematics performed by ALU(s) 810 in response to executing instructions or other code, using weight values located in code and / or data storage 805 and / or data 801 as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 805 or code and / or code and / or data storage 801 or other storage on or off-chip.
[0062] In at least one embodiment, the ALU(s) 810 is / are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, the ALU(s) 810 may be external to a processor or other hardware logic device or circuit that uses it (e.g., a co-processor). In at least one embodiment, ALUs 810 may be included in the execution units of a processor or otherwise in a bank of ALUs that can be accessed by the execution units of a processor, either within the same processor or distributed among various processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.).In at least one embodiment, code and / or data storage 801, code and / or data storage 805, and enable storage 820 may share a processor or other hardware logic device or circuit, whereas in another embodiment, they may be in different processors or other hardware logic devices or circuits, or a combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of enable storage 820 may be included with other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.Furthermore, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry, and may be retrieved and / or processed using the fetch, decode, scheduling, execution, retirement, and / or other logic circuitry of a processor.
[0063] In at least one embodiment, the activation storage 820 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation storage 820 may be wholly or partially internal or external to one or more processors or other logic circuitry. In at least one embodiment, a choice of whether the activation storage 820 is, for example, internal or external to a processor or consists of DRAM, SRAM, flash memory, or another type of storage may depend on on-chip versus off-chip available storage, latency requirements of the training and / or inference functions to be performed, the batch size of the data used in inferencing and / or training a neural network, or a combination of these factors.
[0064] In at least one embodiment, the Fig.8A may be used in conjunction with an application-specific integrated circuit ("ASIC"), such as Google's Tensorflow® processing unit, a Graphcore™ inference processing unit (IPU), or a Nervana® (e.g., "Lake Crest") processor from Intel Corp. In at least one embodiment, the inference and / or training logic 815 illustrated in FIG. 8A may be used in conjunction with central processing unit ("CPU") hardware, graphics processing unit ("GPU") hardware, or other hardware, such as field-programmable gate arrays ("FPGAs").
[0065] Fig.8B illustrates the inference and / or training logic 815 according to at least one different embodiment. In at least one embodiment, the inference and / or training logic 815 may include, but is not limited to, hardware logic in which computational resources are dedicated or otherwise used exclusively in connection with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, the inference and / or training logic 815 may Fig. 8B may be used in conjunction with an application-specific integrated circuit (ASIC), such as Google's Tensorflow® processing unit, a Graphcore™ inference processing unit (IPU), or a Nervana® (e.g., "Lake Crest") processor from Intel Corp. In at least one embodiment, the inference and / or training logic 815 illustrated in Fig.8B may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware, such as field programmable gate arrays (FPGAs). In at least one embodiment, the inference and / or training logic 815 includes, but is not limited to, code and / or data storage 801 and code and / or data storage 805, which may be used to store code (e.g., graphics code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment, Fig.8B, each of the code and / or data storage 801 and the code and / or data storage 805 is each associated with a dedicated computing resource, such as the computing hardware 802 and the computing hardware 806. In at least one embodiment, the computing hardware 802 and the computing hardware 806 each include one or more ALUs that perform mathematical functions, such as linear algebraic functions, only on information stored in the code and / or data storage 801 and the code and / or data storage 805, respectively, the result of which is stored in the activation storage 820.
[0066] In at least one embodiment, each of the code and / or data storage 801 and 805 and corresponding computational hardware 802 and 806 corresponds to different layers of a neural network, such that an activation resulting from one "storage / compute pair 801 / 802" of the code and / or data storage 801 and the computational hardware 802 is provided as an input to the next "storage / compute pair 805 / 806" of the code and / or data storage 805 and the computational hardware 806 to mirror the conceptual organization of a neural network. In at least one embodiment, each of the storage / compute pairs 801 / 802 and 805 / 806 may correspond to more than one neural network layer. In at least one embodiment, additional memory / compute pairs (not shown) may be included in the inference and / or training logic 815 after or in parallel with the memory / compute pairs 801 / 802 and 805 / 806. TRAINING AND DEPLOYMENT OF A NEURAL NETWORK
[0067] Fig.9 illustrates training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, an untrained neural network 906 is trained using a training dataset 902. In at least one embodiment, the training framework 904 is a PyTorch framework, whereas in other embodiments, the training framework 904 is a Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, the training framework 904 trains an untrained neural network 906 and facilitates its training using processing resources described herein to produce a trained neural network 908. In at least one embodiment, weights may be chosen randomly or by pre-training using a deep belief network.In at least one embodiment, the training may be conducted in either a supervised, semi-supervised, or unsupervised manner.
[0068] In at least one embodiment, an untrained neural network 906 is trained using supervised learning, where the training data set 902 includes an input paired with a desired output for an input, or where the training data set 902 includes an input with a known output and an output of the neural network 906 is manually ranked. In at least one embodiment, an untrained neural network 906 is trained in a supervised manner and processes inputs from the training data set 902 and compares resulting outputs against a set of expected or desired outputs. In at least one embodiment, errors are then backpropagated through the untrained neural network 906. In at least one embodiment, the training framework 904 adjusts weights that control the untrained neural network 906.In at least one embodiment, the training framework 904 includes tools to monitor how well the untrained neural network 906 converges to a model, such as the trained neural network 908, that is capable of generating correct answers, such as in the output 914, based on known input data, such as a new dataset 912. In at least one embodiment, the training framework 904 repeatedly trains the untrained neural network 906 while adjusting weights to refine an output of the untrained neural network 906 using a loss function and a tuning algorithm, such as stochastic gradient descent. In at least one embodiment, the training framework 904 trains the untrained neural network 906 until the untrained neural network 906 achieves a desired accuracy.In at least one embodiment, the trained neural network 908 may then be used to implement any number of machine learning operations.
[0069] In at least one embodiment, the untrained neural network 906 is trained using unsupervised learning, where the untrained neural network 906 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training dataset 902 will include input data without any associated output or ground truth data. In at least one embodiment, the untrained neural network 906 may learn groupings within the training dataset 902 and may determine how individual inputs relate to the untrained dataset 902.In at least one embodiment, unsupervised training may be used to generate a self-organizing map in the trained neural network 908 capable of performing operations useful in reducing the dimensionality of a new dataset 912. In at least one embodiment, unsupervised training may also be used to perform anomaly detection, which enables identification of data points in a new dataset 912 that deviate from normal patterns of the new dataset 912.
[0070] In at least one embodiment, semi-supervised learning may be used, which is a technique in which a training dataset 902 includes a mixture of labeled and unlabeled data. In at least one embodiment, the training framework 904 may be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning allows the trained neural network 908 to adapt to a new dataset 912 without forgetting the knowledge introduced into the trained neural network 908 during initial training. DATA CENTER
[0071] Fig.10 illustrates an example data center 1000 in which at least one embodiment may be used. In at least one embodiment, the data center 1000 includes a data center infrastructure layer 1010, a framework layer 1020, a software layer 1030, and an application layer 1040.
[0072] In at least one embodiment, as in Fig.10, the data center infrastructure layer 1010 may include a resource orchestrator 1012, clustered compute resources 1014, and node compute resources (“Node CR”) 1016(1)-1016(N), where “N” represents a positive integer (which may be a different integer “N” than that used in other figures). In at least one embodiment, the node CRs 1016(1)-1016(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory storage devices 1018(1)-1018(N) (e.g., dynamic read-only memory), storage devices (e.g., solid-state storage or hard disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules and cooling modules, etc.In at least one embodiment, one or more of the node CRs 1016(1)-1016(N) may be a server having one or more of the computing resources mentioned above.
[0073] In at least one embodiment, grouped computing resources 1014 may include separate groupings of node CRs housed in one or more racks (not shown) or in many racks housed in data centers in different geographic locations (also not shown). In at least one embodiment, separate groupings of node CRs within grouped computing resources 1014 may include grouped computing, networking, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads.In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and in any combination.
[0074] In at least one embodiment, resource orchestrator 1012 may configure or otherwise control one or more node CRs 1016(1)-1016(N) and / or clustered computing resources 1014. In at least one embodiment, resource orchestrator 1012 may include a software design infrastructure ("SDI") management entity for data center 1000. In at least one embodiment, the resource orchestrator may include hardware, software, or a combination thereof.
[0075] In at least one embodiment, as in Fig.10, the framework layer 1020 includes a job scheduler 1022, a configuration manager 1024, a resource manager 1026, and a distributed file system 1028. In at least one embodiment, the framework layer 1020 may include a framework for supporting the software 1032 of the software layer 1030 and / or one or more applications 1042 of the application layer 1040. In at least one embodiment, the software 1032 or the application(s) 1042 may each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure.In at least one embodiment, but not limited to, the framework layer 1020 may be a type of framework for a free and open source software web application framework such as Apache Spark™ (hereinafter referred to as "Spark"), which may utilize the distributed file system 1028 for large-scale computing (e.g., "big data"). In at least one embodiment, the job scheduler 1032 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 1000. In at least one embodiment, the configuration manager 1024 may be capable of configuring various layers, such as the software layer 1030 and the framework layer 1020, including Spark and the distributed file system 1028, to support large-scale computing.In at least one embodiment, resource manager 1026 may be capable of managing clustered or grouped computing resources mapped or allocated to support distributed file system 1028 and job scheduler 1022. In at least one embodiment, clustered or grouped computing resources may include clustered computing resource 1014 at data center infrastructure layer 1010. In at least one embodiment, resource manager 1026 may coordinate with resource orchestrator 1012 to manage these mapped or allocated computing resources.
[0076] In at least one embodiment, the software 1032 included in software layer 1030 may include software used by at least portions of node CRs 1016(1)-1016(N), clustered computing resources 1014, and / or the distributed file system 1028 of framework layer 1020. In at least one embodiment, one or more types of software may include, but are not limited to, Internet web page search software, email virus scanner software, database software, and streaming video content software.
[0077] In at least one embodiment, the application(s) 1042 included in the application layer 1040 may include one or more types of applications used by at least portions of the node CRs 1016(1)-1016(N), clustered computing resources 1014, and / or the distributed file system 1028 of the framework layer 1020. In at least one embodiment, one or more types of applications may include, but are not limited to, any number of genomic applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in connection with one or more embodiments.
[0078] In at least one embodiment, each of the configuration manager 1034, the resource manager 1036, and the resource orchestrator 1012 may implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible manner. In at least one embodiment, self-modifying actions may free an operator of the data center 1000 from making potentially poor configuration decisions and potentially avoid underutilized and / or malfunctioning portions of a data center.
[0079] In at least one embodiment, data center 1000 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models, according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 1000.In at least one embodiment, trained machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1000 using weight parameters calculated by one or more of the training techniques described herein.
[0080] In at least one embodiment, the data center may use CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to allow users to train or infer information, such as image recognition, speech recognition, or other artificial intelligence services.
[0081] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described herein in connection with Fig.8A and / or 8B. In at least one embodiment, in the system of Fig. 10 the inference and / or training logic 915 may be used to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0082] In at least one embodiment, a computer system as described herein may be used to create a robot control system implementing model-based and model-free control. For example, a computer system as described above may include memory storing executable instructions that, as a result of being executed by a processor of the computer system, cause the computer system to implement a model-based and model-free control system as described herein. AUTONOMOUS VEHICLE
[0083] Fig.11A illustrates an example of an autonomous vehicle 1100 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1100 (alternatively referred to herein as "vehicle 1100") may be, but is not limited to, a passenger vehicle, such as a car, a truck, a bus, and / or another type of vehicle capable of accommodating one or more passengers. In at least one embodiment, a vehicle 1100 may be a semi-trailer truck used to pull cargo. In at least one embodiment, a vehicle 1100 may be an aircraft, a robotic vehicle, or another type of vehicle.
[0084] Autonomous vehicles may generally be described in terms of levels of automation defined by the National Highway Traffic Safety Administration ("NHTSA"), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers ("SAE") standard "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and prior and future versions of that standard). In one or more embodiments, a vehicle 1100 may be capable of operating in accordance with one or more of autonomous driving levels 1 through 5. For example, in at least one embodiment, a vehicle 1100 may be capable of performing conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5).
[0085] In at least one embodiment, a vehicle 1100 may include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. In at least one embodiment, a vehicle 1100 may include a propulsion system 1150, such as an internal combustion engine, a hybrid electric system, an all-electric motor, and / or another type of propulsion system. In at least one embodiment, the propulsion system 1150 may be connected to a drivetrain of the vehicle 1100, which may include a transmission to enable propulsion of the vehicle 1100. In at least one embodiment, the propulsion system 1150 may be controlled in response to receiving signals from a throttle / accelerator(s) 1152.
[0086] In at least one embodiment, a steering system 1154, which may include, but is not limited to, a steering wheel, is used to steer a vehicle 1100 (e.g., along a desired path or route) when the propulsion system 1150 is operating (e.g., when a vehicle 1100 is in motion). In at least one embodiment, the steering system 1154 may receive signals from a steering actuator(s) 1156. A steering wheel may be optional for full automation (Level 5) functionality. In at least one embodiment, the brake sensor system 1146 may be used to apply vehicle brakes in response to receiving signals from the brake actuator(s) 1148 and / or brake sensors.
[0087] In at least one embodiment, the controller(s) 1136 comprising one or more Systems on Chips (“SoCs”) (in Fig.11A not shown) and / or graphics processing units ("GPU(s)"), provide signals (e.g., representing commands) to one or more components and / or systems of the vehicle 1100. For example, the controller(s) 1136 may send signals to apply vehicle brakes via one or more brake actuators 1148, to actuate the steering system 1154 via one or more steering actuators 1156, and to actuate the propulsion system 1150 via one or more throttles / accelerators 1152. In at least one embodiment, the controller(s) 1136 may include one or more built-in (e.g., integrated) computing devices that process sensor signals and issue operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 1100.In at least one embodiment, controller(s) 1136 may include a first controller for autonomous driving functions, a second controller for functional safety functions, a third controller for artificial intelligence functions (e.g., computer vision), a fourth controller for infotainment functionality, a fifth controller for emergency redundancy, and / or other controllers. In at least one embodiment, a single controller may handle two or more of the above functionalities, two or more controllers may handle a single functionality, and / or any combination thereof.
[0088] In at least one embodiment, the controller(s) 1136 may provide signals to control one or more components and / or systems of the vehicle 1100 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, the sensor data may be obtained from, for example, and without limitation, global navigation satellite system sensor(s) 1158 (e.g., Global Positioning System Sensor(s); “GNSS”), RADAR sensor(s) 1160, ultrasonic sensor(s) 1162, LIDAR sensor(s) 1164, Inertial Measurement Unit (IMU) sensor(s) 1166 (e.g., accelerometer(s), gyroscope(s), magnetic compass(es), magnetometer(s), etc.), microphone(s) 1196, stereo camera(s) 1168, wide-angle camera(s) 1170 (e.g., fisheye cameras), infrared camera(s) 1172, surround camera(s) 1174 (e.g., 360-degree cameras), long-range cameras (in Fig. 11A not shown), mid-range camera(s) (in Fig.11A not shown), speed sensor(s) 1144 (e.g., for measuring the speed of the vehicle 1100), vibration sensor(s) 1142, steering sensor(s) 1140, brake sensor(s) (e.g., as part of the brake sensor system 1146), and / or other types of sensors.
[0089] In at least one embodiment, one or more controllers 1136 may receive inputs (e.g., represented by input data) from an instrument cluster 1132 of the vehicle 1100 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (“HMI”) display 1134, an audible annunciator, a speaker, and / or via other components of the vehicle 1100. In at least one embodiment, the outputs may include information such as vehicle vector speed, speed, time, map data (e.g., a high-definition map (in Fig.11A not shown), location data (e.g., the location of the vehicle, such as on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and status of objects as perceived by the controller(s) 1136, etc. For example, in at least one embodiment, the HMI display 1134 may display information about the presence of one or more objects (e.g., a road sign, a warning sign, a traffic light change, etc.) and / or information about maneuvers a vehicle has performed, is currently performing, or will perform (e.g., currently changing lanes, taking exit 34B in two miles, etc.).
[0090] In at least one embodiment, a vehicle 1100 further includes a network interface 1124 that may utilize one or more wireless antenna(s) 1126 and / or modem(s) to communicate over one or more networks. For example, in at least one embodiment, a network interface 1124 may be capable of communicating over Long-Term Evolution ("LTE"), Wide Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communication ("GSM"), ("CDMA2000"), IMT-CDMA Multi-Carrier ("CDMA2000") networks, etc. In at least one embodiment, the wireless antenna(s) 1126 may also facilitate communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area network(s), such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, etc., and / or low-power wide area network(s) (“LPWANs”) such as LoRaWAN, SigFox, etc. protocols.
[0091] Inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 815 are described herein in connection with Fig. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 in the system of Fig. 11A may be used to infer or predict operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0092] In at least one embodiment, the techniques described herein may be used to control an autonomous vehicle. For example, a model-based control algorithm may be used to position a vehicle for parking, or a model-free RL network may be used to park a vehicle.
[0093] Fig. 11B illustrates an example of camera locations and fields of view for the example autonomous vehicle 1100 of Fig. 11A according to at least one embodiment. In at least one embodiment, the cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different locations on a vehicle 1100.
[0094] In at least one embodiment, the camera types may include, but are not limited to, digital cameras that may be adapted for use with the components and / or systems of the vehicle 1100. In at least one embodiment, the camera(s) may operate at Automotive Safety Integrity Level (ASIL) B and / or at another ASIL. The camera types may be capable of any image capture rate, e.g., 60 frames per second (fps), 1120 fps, 240 fps, etc., depending on the environment. In at least one embodiment, the cameras may be capable of using rolling shutters, global shutters, another type of shutter, or a combination thereof.In at least one embodiment, a color filter array may include a Red Clear ("RCCC") color filter array, a Red Clear Blue ("RCCB") color filter array, a Red Blue Green Clear ("RBGC") color filter array, a Foveon X3 color filter array, a Bayer Sensor ("RGGB") color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In at least one embodiment, clear pixel cameras, such as cameras with an RCCC, an RCCB, and / or an RBGC color filter array, may be used in an effort to increase light sensitivity.
[0095] In at least one embodiment, one or more of the cameras may be used to perform Advanced Driver Assistance Systems ("ADAS") functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multifunction mono camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more of the cameras (e.g., all cameras) may record and provide image data (e.g., video) simultaneously.
[0096] In at least one embodiment, one or more cameras may be mounted in a mounting assembly, such as a custom (three-dimensional ("3D") printed) assembly, to reduce stray light and reflections from a vehicle interior (e.g., reflections from the dashboard reflected in windshield mirrors) that may impair the camera's image data detection capabilities. With respect to the exterior mirror mounting assemblies, in at least one embodiment, the exterior mirror assemblies may be custom 3D printed such that a camera mounting plate conforms to the shape of an exterior mirror. In at least one embodiment, the camera(s) may be integrated into exterior mirrors. In at least one embodiment, for side-view cameras, the camera(s) may also be integrated within four pillars at each corner of a cab.
[0097] In at least one embodiment, cameras with a field of view that includes portions of an environment in front of a vehicle 1100 (e.g., forward-facing cameras) may be used for surround vision to help identify forward paths and obstacles, as well as to help provide, with the assistance of one or more controllers 1136 and / or control SoCs, important information for generating an occupancy grid and / or determining preferred vehicle paths. In at least one embodiment, forward-facing cameras may be used to perform many of the same ADAS functions as LIDAR, including, but not limited to, emergency braking, pedestrian detection, and collision avoidance.In at least one embodiment, forward-facing cameras may also be used for ADAS features and systems, including Lane Departure Warnings (“LDW”), Autonomous Cruise Control (“ACC”), and / or other features such as traffic sign recognition.
[0098] In at least one embodiment, a variety of cameras may be used in a forward-facing configuration, including, for example, a monocular camera platform including a CMOS (complementary metal oxide semiconductor) color imager. In at least one embodiment, a wide-angle camera 1170 may be used to perceive objects coming into view from a periphery (e.g., pedestrians, cross traffic, or bicycles). Although only one wide-angle camera 180 may be used in Fig.11B, in other embodiments, there may be any number (including zero) of wide-angle cameras on a vehicle 1100. In at least one embodiment, any number of long-range cameras 1198 (e.g., a long-view stereo camera pair) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. In at least one embodiment, a long-range camera(s) 1198 may also be used for object detection and classification, as well as basic object tracking.
[0099] In at least one embodiment, one or more stereo cameras 1168 may also be included in a forward-facing configuration. In at least one embodiment, one or more stereo cameras 1168 may include an integrated controller unit comprising a scalable processing unit that may provide a field-programmable logic (“FPGA”) and a multi-core microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the environment of the vehicle 1100 that includes a distance estimate for all points in an image.In at least one embodiment, an alternative stereo camera(s) 1168 may include a compact stereo vision sensor(s) that may include, but is not limited to, two camera lenses (one each on the left and right) and an image processing chip that can measure the distance from a vehicle 1100 to the target object and use the generated information (e.g., metadata) to activate the autonomous features of emergency braking and lane departure warning. In at least one embodiment, other types of stereo camera(s) 1168 may be used in addition to, or alternatively to, those described herein.
[0100] In at least one embodiment, cameras with a field of view that includes portions of the environment to the side of the vehicle 1100 (e.g., side view cameras) may be used for the surround view, which provides information used to generate and update the occupancy grid, as well as to generate side impact warnings. For example, in at least one embodiment, the surround camera(s) 1174 (e.g., four surround cameras 1174, as in Fig.11B) may be positioned on a vehicle 1100. In at least one embodiment, the surround camera(s) 184 may include, but are not limited to, any number and combination of wide-angle camera(s) 180, fisheye camera(s), 360-degree camera(s), and / or similar cameras. For example, in at least one embodiment, four fisheye cameras may be positioned on a front, a rear, and the sides of the vehicle 1100. In at least one embodiment, a vehicle 1100 may utilize three surround cameras 1174 (e.g., left, right, and rear) and may leverage one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.
[0101] In at least one embodiment, cameras with a field of view encompassing portions of an environment behind a vehicle 1100 (e.g., rearview cameras) may be used for parking assistance, surround view, rear collision warnings, and for generating and updating an occupancy grid. In at least one embodiment, a wide variety of cameras may be used, including, but not limited to, cameras that are also suitable as a forward-facing camera(s) (e.g., long-range and / or mid-range camera(s) 1176, stereo camera(s) 1168, infrared cameras 1172, etc.), as described herein.
[0102] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 815 are described below in connection with Fig.8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 in the system of Fig. 11B may be used to infer or predict operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0103] In at least one embodiment, robotic control systems as described herein may be used to navigate an autonomous vehicle. For example, a model-based control system may navigate the automobile using a map and GPS signals, while a model-based controller is used to provide lane maintenance during driving.
[0104] Fig.11C is a block diagram illustrating an example system architecture for the autonomous vehicle 1100 of Fig. 11A, according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 1100 is Fig.11C as connected via a bus 1102. In at least one embodiment, bus 1102 may include a CAN (Controller Area Network) data interface (alternatively referred to herein as a "CAN bus"). In at least one embodiment, a CAN may be a network within vehicle 1100 used to assist in controlling various features and functionality of vehicle 1100, such as the application of brakes, acceleration, braking, steering, windshield wipers, etc. In at least one embodiment, bus 1102 may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). In at least one embodiment, bus 1102 may be read to determine steering wheel angle, ground speed, engine speeds per minute (RPM), switch positions, and / or other vehicle status indicators.In at least one embodiment, bus 1102 may be a CAN bus that is ASIL B compliant.
[0105] In at least one embodiment, in addition to or alternatively to CAN, FlexRay, and / or Ethernet protocols may be used. In at least one embodiment, there may be any number of buses forming bus 1102, which may include, but is not limited to, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero and / or other types of buses with a different protocol. In at least one embodiment, two or more buses 1102 may be used to perform different functions and / or may be used for redundancy.
[0106] For example, a first bus may be used for collision avoidance functionality and a second bus for actuation control. In at least one embodiment, each bus 1102 may communicate with any of the components of the vehicle 1100, and two or more buses 1102 may communicate with corresponding components. In at least one embodiment, each of any number of system-on-chip(s) (“SoC(s)”) (such as SoC 1104(A) and SoC 1104(B), each of the controllers 1136, and / or each computer in the vehicle may have access to the same input data (e.g., inputs from sensors of the vehicle 1100) and be connected to a common bus, such as a CAN bus.
[0107] In at least one embodiment, a vehicle 1100 may include one or more controllers 1136, such as those described herein with respect to Fig.11A. In at least one embodiment, the controller(s) 1136 may be used for a variety of functions. In at least one embodiment, the controller(s) 1136 may be coupled to any of various other components and systems of the vehicle 1100 and may be used to control the vehicle 1100, the artificial intelligence of the vehicle 1100, the infotainment for a vehicle 1100, and / or the like.
[0108] In at least one embodiment, a vehicle 1100 may include any number of SoCs 1104. In at least one embodiment, each of the SoCs 1104 may include, but is not limited to, central processing units ("CPU(s)") 1106, graphics processing units ("GPU(s)") 1108, processor(s) 1110, cache memory 1112, accelerators 1114, data storage 1116, and / or other unillustrated components and features. In at least one embodiment, the SoC(s) 1104 may be used to control the vehicle 1100 in a variety of platforms and systems. For example, in at least one embodiment, the SoC(s) 1104 may be combined in a system (e.g., the system of the vehicle 1100) with a high-definition (“HD”) map 1122 that may receive map refreshes and / or updates via a network interface 1124 from one or more servers (in Fig.11C not shown).
[0109] In at least one embodiment, the CPU(s) 1106 may comprise a CPU cluster or CPU complex (alternatively referred to herein as a "CCPLEX"). In at least one embodiment, the CPU(s) may comprise multiple cores and / or Level 2 ("L2") caches. For example, in at least one embodiment, the CPU(s) 1106 may comprise eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU(s) 1106 may comprise four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 megabyte (MB) L2 cache). In at least one embodiment, the CPU(s) 1106 (e.g., the CCPLEX) may be configured to support simultaneous cluster operations that allow any combination of clusters of the CPU(s) 1106 to be active at any given time.
[0110] In at least one embodiment, one or more of the CPU(s) 1106 may implement power management capabilities, including, but not limited to, one or more of the following features: individual hardware blocks may be clock-gated automatically when idle to conserve dynamic power; each core clock may be gated when the core is not actively executing instructions due to the execution of Wait for Interrupt ("WFI") / Wait for Event ("WFE") instructions; each core may be independently power-gated; each core cluster may be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster may be independently power-gated when all cores are power-gated.In at least one embodiment, the CPU(s) 1106 may further implement an advanced power state management algorithm where allowable power states and expected wake-up times are specified, and the hardware / microcode determines the best power state to enter for the core, cluster, and CCPLEX. In at least one embodiment, the processing cores may support simplified power state entry sequences in software, offloading the work to microcode.
[0111] In at least one embodiment, the GPU(s) 1108 may comprise an integrated GPU(s) (alternatively referred to herein as an "iGPU"). In at least one embodiment, the GPU(s) 1108 may be programmable and efficient for parallel workloads. In at least one embodiment, the GPU(s) 1108 may utilize an enhanced tensor instruction set. In at least one embodiment, the GPU(s) 1108 may comprise one or more streaming microprocessors, where each streaming microprocessor may comprise a Level 1 ("L1") cache (e.g., an L1 cache with at least 116 KB of memory capacity), and two or more of the streaming microprocessors may share a Level 2 ("L2") cache (e.g., an L2 cache with a memory capacity of 512 KB). In at least one embodiment, the GPU(s) 1108 may include at least eight streaming microprocessors.In at least one embodiment, the GPU(s) 1108 may utilize an application programming interface(s) (“API(s)”). In at least one embodiment, the GPU(s) 1108 may utilize one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA model).
[0112] In at least one embodiment, the GPU(s) 1108 may be power-optimized for best performance in automotive and embedded use cases. For example, in one embodiment, the GPU(s) 1108 may be fabricated on a fin field-effect transistor ("FinFET"). In at least one embodiment, each streaming microprocessor may include a number of mixed-precision processing cores divided into multiple blocks. For example, and not limited to, 64 PF32 cores and 32 PF64 cores could be divided into four processing blocks. In at least one embodiment, each processing block could be allocated to 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two NVIDIA mixed-precision Tensor Cores for deep learning matrix arithmetic, a Level 0 ("L") instruction cache, a warp scheduler, a scheduling unit, and / or a 64KB register file.In at least one embodiment, streaming microprocessors may include independent parallel integer and floating-point datapaths to provide efficient execution of workloads with a mix of arithmetic and addressing calculations. In at least one embodiment, streaming microprocessors may include independent thread scheduling functionality to enable finer synchronization and collaboration between parallel threads. In at least one embodiment, streaming microprocessors may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0113] In at least one embodiment, one or more of the GPU(s) 1108 may include high bandwidth memory (“HBM”) and / or a 16 GB HBM2 memory subsystem to provide, in some examples, a peak memory bandwidth of approximately 900 GB / second. In at least one embodiment, in addition to or as an alternative to the HBM memory, a synchronous graphics random-access memory (“SGRAM”) may be used, such as a graphics double data rate type five synchronous random-access memory (“GDDR5”).
[0114] In at least one embodiment, the GPU(s) 1108 may include unified memory technology. In at least one embodiment, support for Address Translation Services (“ATS”) may be used to enable the GPU(s) 1108 to directly access page tables of the CPU(s) 1106. In at least one embodiment, when the Memory Management Unit (“MMU”) of the GPU(s) 1108 experiences a miss, an address translation request may be sent to the CPU(s) 1106. In response, in at least one embodiment, the CPU(s) 1106 may look up the virtual-to-physical address mapping for the address in their page tables and transmit the translation back to the GPU(s) 1108.In at least one embodiment, the unified memory technology may enable a single unified virtual address space for memory of both the CPU(s) 1106 and the GPU(s) 1108, thereby simplifying programming of the GPU(s) 1108 and porting applications to the GPU(s) 1108.
[0115] In at least one embodiment, the GPU(s) 1108 may include any number of access counters that may track the frequency of access by the GPU(s) 1108 to the memory of other processors. In at least one embodiment, the access counter(s) may contribute to moving memory pages into the physical memory of the processor that accesses pages most frequently, thereby improving efficiency for memory shared between processors.
[0116] In at least one embodiment, one or more of the SoCs 1104 may include any number of caches 1112, including those described herein. For example, in at least one embodiment, the cache(s) 1112 may include a Level 3 ("L3") cache available to both the CPU(s) 1106 and the GPU(s) 1108 (e.g., connected to both the CPU(s) 1106 and the GPU(s) 1108). In at least one embodiment, the cache(s) 1112 may include a write-back cache that may track line states, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, an L3 cache may include 4 MB or more, depending on the embodiment, although smaller cache sizes may be used.
[0117] In at least one embodiment, one or more of the SoCs 1104 may include one or more accelerators 1114 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoC(s) 1104 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4 MB SRAM) may enable a hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to supplement the GPU(s) 1108 and offload some of the tasks of the GPU(s) 1108 (e.g., to free up more cycles of the GPU(s) 1108 to perform other tasks). In at least one embodiment, the accelerator(s) 1114 could be optimized for targeted workloads (e.g.,Perception, convolutional neural networks (CNNs), recurrent neural networks (RNNs), etc.) that are robust enough to be amenable to acceleration may be used. In at least one embodiment, a CNN may include region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., as used for object detection).
[0118] In at least one embodiment, the accelerator(s) 1114 (e.g., the hardware acceleration cluster) may include a deep learning accelerator (“DLA(s)”). DLA(s) may include, but are not limited to, one or more tensor processing units (“TPUs”) that may be configured to provide an additional tens of trillion operations per second for deep learning applications and inference. The TPUs may be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, the DLA(s) may be further optimized for a particular set of neural network types and floating-point operations, as well as for inference.In at least one embodiment, the design of the DLA(s) may provide more performance per millimeter than a typical general-purpose graphics processor, typically far exceeding the performance of a CPU. In at least one embodiment, the TPU(s) may perform multiple functions, including a single-instance convolution function that supports, for example, but is not limited to, features and weights on INT8, INT16, and FP16 data types, as well as post-processing functions.In at least one embodiment, DLA(s) may quickly and efficiently execute neural networks, particularly CNNs, on processed or unprocessed data for any of a variety of functions, including, for example, and not limited to: a CNN for object identification and recognition using data from camera sensors; a CNN for distance estimation using data from camera sensors; a CNN for vehicle detection and identification using data from microphones 1196; a CNN for facial recognition and vehicle owner identification using data from camera sensors; and / or a CNN for security and / or security-related events.
[0119] In at least one embodiment, DLA(s) may perform any function of GPU(s) 1108, and by using an inference accelerator, for example, a designer may target either the DLA(s) or the GPU(s) 1108 for each function. For example, in at least one embodiment, the designer may focus on processing CNNs and floating-point operations on the DLA(s) and leave other functions to the GPU(s) 1108 and / or other accelerator(s) 1114.
[0120] In at least one embodiment, the accelerator(s) 1114 may comprise a programmable vision accelerator (“PVA”), which may alternatively be referred to herein as a computer vision accelerator. In at least one embodiment, the PVA may be designed and configured to accelerate image processing algorithms for advanced driver assistance systems (“ADAS”), autonomous driving, augmented reality (“AR”), and / or virtual reality (“VR”) applications. In at least one embodiment, a PVA may provide a balance between performance and flexibility.For example, in at least one embodiment, each PVA may include, for example, and not limited to, any number of Reduced Instruction Set Computer (“RISC”) cores, Direct Memory Access (“DMA”) cores, and / or any number of vector processors.
[0121] In at least one embodiment, RISC cores may interact with image sensors (e.g., image sensors from any of the cameras described herein), image signal processor(s), etc. In at least one embodiment, RISC cores may include any amount of memory. In at least one embodiment, RISC cores may use any number of protocols depending on the embodiment. In at least one embodiment, RISC cores may execute a real-time operating system (“RTOS”). In at least one embodiment, RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, the RISC cores may include an instruction cache and / or tightly coupled RAM.
[0122] In at least one embodiment, the DMA may enable components of the PVA to access system memory independently of CPU(s) 1106. In at least one embodiment, the DMA may support any number of features used to provide optimization to a PVA, including, but not limited to, support for multidimensional addressing and / or circular addressing. In at least one embodiment, the DMA may support up to six or more dimensions of addressing, which may include, but are not limited to, block width, block height, block depth, horizontal block gradation, vertical block gradation, and / or depth gradation.
[0123] In at least one embodiment, the vector processors may be programmable processors that may be configured to efficiently and flexibly perform programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA may include a processor subsystem, one(s) (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, the vector processing subsystem may operate as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., vector memory; "VMEM").In at least one embodiment, the VPU may include a digital signal processor, such as a single instruction, multiple data (SIMD) digital signal processor and a very long instruction word (VLIW) digital signal processor. In at least one embodiment, the combination of SIMD and VLIW may increase throughput and speed.
[0124] In at least one embodiment, each of the vector processors may include an instruction cache and be coupled to dedicated memory. As a result, in at least one embodiment, each of the vector processors may be configured to execute independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA may be configured to utilize data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm but on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may concurrently execute different computer vision algorithms on the same image or even execute different algorithms on sequential images or portions of an image.In at least one embodiment, among other things, any number of PVAs may be included in the hardware acceleration cluster and any number of vector processors may be included in each PVA. In at least one embodiment, the PVA may additionally include memory for error correcting code (ECC) to increase overall system security.
[0125] In at least one embodiment, the accelerator(s) 1114 may include an on-chip computer vision network and static random-access memory (“SRAM”) to provide high-bandwidth, low-latency SRAM for the accelerator(s) 1114. In at least one embodiment, the on-chip memory may include at least 4 MB of SRAM, consisting of, for example, and without limitation, eight field-configurable memory blocks accessible by both a PVA and a DLA. In at least one embodiment, each memory block pair may include an Advanced Peripheral Bus interface (“APB”), configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may be connected via a backbone.A backbone that provides high-speed access to the memory to a PVA and a DLA. In at least one embodiment, the backbone may include an on-chip computer vision network that connects the PVA and the DLA to the memory (e.g., using the APB).
[0126] In at least one embodiment, the on-chip computer vision network may include an interface that determines that both the PVA and the DLA are providing ready and valid signals prior to transmitting any control signals / addresses / data. In at least one embodiment, an interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. In at least one embodiment, an interface may conform to International Organization for Standardization ("ISO") 26262 or International Electrotechnical Commission ("IEC") 61508 standards, although other standards and protocols may be used.
[0127] In at least one embodiment, one or more of the SoC(s) 1104 may include a real-time ray tracing hardware accelerator. In at least one embodiment, the real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the positions and dimensions of objects (e.g., within a world model) to generate real-time visualization simulations for radar signal interpretation, sound propagation synthesis and / or analysis, simulation of sonar systems, simulation of general wave propagation, comparison with lidar data for localization and / or other functions, and / or for other applications.
[0128] In at least one embodiment, accelerator(s) 1114 may have a wide range of applications for autonomous driving. In at least one embodiment, a PVA may be used for processing stages in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of a PVA are a good match for algorithmic domains that require predictable processing with low power and low latency. In other words, the PVA can perform well in semi-dense or dense regular computing, even on small datasets, requiring predictable runtimes with low latency and low power. In at least one embodiment, in autonomous vehicles, such as vehicle 1100, PVAs are configured to execute classical computer vision algorithms because they are efficient at object detection and operate on integer mathematics.
[0129] For example, according to at least one embodiment of the technology, the PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching-based algorithm may be used, although this is not intended to be limiting. In at least one embodiment, Level 3-5 autonomous driving applications require motion estimation / on-the-fly stereo matching (e.g., structure from motion, pedestrian detection, lane detection, etc.). In at least one embodiment, the PVA may perform a computer stereo vision function on inputs from two monocular cameras.
[0130] In at least one embodiment, the PVA may be used to perform dense optical flow. For example, in at least one embodiment, the PVA may process raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar data. In at least one embodiment, a PVA is used for deep time-of-flight processing by processing raw time-of-flight data to provide, for example, processed time-of-flight data.
[0131] In at least one embodiment, a DLA may be used to power any type of network to enhance control and driving safety, including, for example, and not limited to, a neural network that outputs a confidence level for each object detection. In at least one embodiment, such a confidence level may be interpreted as a probability or as providing a relative "weight" of each detection compared to other detections. In at least one embodiment, a confidence level allows a system to make further decisions regarding which detections should be considered true positives rather than false positives.For example, in at least one embodiment, a system may set a confidence threshold and consider only detections exceeding the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections would cause a vehicle to automatically perform emergency braking, which is obviously undesirable. In at least one embodiment, highly confident detections may be considered as triggers for AEB. In at least one embodiment, the DLA may run a neural network to regress the confidence value. In at least one embodiment, the neural network may use at least a subset of parameters as its input, such as dimensions of a bounding box, a ground plane estimate (e.g.,from another subsystem), an output from sensors of the inertial measurement unit (IMU) 1166 that correlates with the orientation of the vehicle 1100, a distance, 3D location estimates of the object that come from, among other things, the neural network and / or other sensors (e.g., LIDAR sensor(s) 1164 or RADAR sensor(s) 1160).
[0132] In at least one embodiment, one or more of the SoC(s) 1104 may include a data store 1116 (e.g., memory). In at least one embodiment, the data store(s) 1116 may be on-chip memory of the SoC(s) 1104 that may store neural networks to be executed on the GPU(s) 1108 and / or a DLA. In at least one embodiment, the data store(s) 1116 may be large enough in capacity to store multiple instances of neural networks for redundancy and security. In at least one embodiment, the data store(s) 1116 may include one or more L2 or L3 cache(s) 1112.
[0133] In at least one embodiment, one or more SoC(s) 1104 may include any number of processor(s) 1110 (e.g., embedded processors). In at least one embodiment, processor(s) 1110 may include a boot and power management processor, which may be a dedicated processor and subsystem to handle boot power and management functions and associated security enforcement. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC(s) 1104 and provide runtime power management services. In at least one embodiment, the boot and power management processor may provide clock and voltage programming, low-power system transition support, management of thermal and temperature sensors of SoC(s) 1104, and / or management of the power states of SoC(s) 1104.In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the SoC(s) 1104 may use ring oscillators to sense temperatures of the CPU(s) 1106, the GPU(s) 1108, and / or the accelerator(s) 1114. In at least one embodiment, if it is determined that temperatures exceed a threshold, then the boot and power management processor may enter a temperature fault routine and place the SoC(s) 1104 into a lower power state and / or place a vehicle 1100 into a chauffeur-to-safe stop mode (e.g., bring a vehicle 1100 to a safe stop).
[0134] In at least one embodiment, the processor(s) 1110 may further comprise a set of embedded processors that may serve as an audio processing engine, which may be an audio subsystem enabling full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0135] In at least one embodiment, the processor(s) 1110 may further include an always-on processor engine that may provide necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, the always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0136] In at least one embodiment, processor(s) 1110 may further include a security cluster engine, including, but not limited to, a dedicated processor subsystem to handle security management for automotive applications. In at least one embodiment, a security cluster engine may include, but is not limited to, two or more processor cores, tightly coupled RAM, peripheral support (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a security mode, two or more cores may, in at least one embodiment, operate in a lockstep mode, acting as a single core with comparison logic to detect any differences between their operations.In at least one embodiment, processor(s) 1110 may further include, but are not limited to, a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, processor(s) 1110 may further include, but are not limited to, a high dynamic range signal processor, which may include, but are not limited to, an image signal processor that is a hardware engine that is part of the camera processing pipeline.
[0137] In at least one embodiment, processor(s) 1110 may include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate the final image for a player window. In at least one embodiment, a video image compositor may perform lens distortion correction on a wide-angle camera(s) 1170, a surround-view camera(s) 1174, and / or on an in-cabin camera sensor(s). In at least one embodiment, the in-cabin surveillance camera sensor(s) is / are preferably monitored by a neural network running on another instance of SoC 1104 and configured to identify and respond to events in the cabin.In at least one embodiment, an in-cabin system may perform, but is not limited to, lip reading to activate a cellular network and place a call, dictate emails, change a vehicle's destination, activate or change a vehicle's infotainment system settings, or provide voice-activated internet browsing. In at least one embodiment, certain features are available to the driver when a vehicle is operating in an autonomous mode and are disabled otherwise.
[0138] In at least one embodiment, a video image compositor may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, when motion occurs in a video, the noise reduction weights spatial information accordingly and reduces the weight of information provided by neighboring frames. In at least one embodiment, when an image or portion of an image does not include motion, the temporal noise reduction performed by the video image compositor may use information from the previous frame to reduce noise in the current frame.
[0139] In at least one embodiment, the video image compositor may also be configured to perform stereo rectification on input stereo lens frames. In at least one embodiment, the video image compositor may further be used for user interface composition when a desktop operating system is used and the GPU(s) 1108 are not required to continuously render new surfaces. In at least one embodiment, when the GPU(s) 1108 are powered on and actively performing 3D rendering, the video image compositor may be used to offload the GPU(s) 1108 to improve performance and responsiveness.
[0140] In at least one embodiment, one or more SoCs 1104 may further include a serial MIPI (Mobile Industry Processor Interface; "MIPI") camera interface for receiving video and inputs from cameras, a high-speed interface, and / or a video input block that may be used for camera and associated pixel input functions. In at least one embodiment, one or more SoCs 1104 may further include an input / output controller that may be controlled by software and may be used to receive I / O signals that are not tied to a specific role.
[0141] In at least one embodiment, the SoC(s) 1104 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio encoders / decoders ("codecs"), power management, and / or other devices. The SoC(s) 1104 may be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., LIDAR sensor(s) 1164, RADAR sensor(s) 1160, etc., which may be connected via Ethernet), data from the bus 1102 (e.g., vehicle speed 1100, steering wheel position, etc.), data from GNSS sensor(s) 1158 (e.g., connected via Ethernet or CAN bus).In at least one embodiment, one or more SoC(s) 1104 may further include dedicated high-performance mass storage controllers, which may include their own DMA engines and which may be used to free the CPU(s) 1106 from routine data management tasks.
[0142] In at least one embodiment, one or more SoC(s) 1104 may be an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture that leverages and efficiently deploys computer vision and ADAS techniques for diversity and redundancy, as well as providing a platform for a flexible, reliable driver software stack along with deep learning tools. In at least one embodiment, the SoC(s) 1104 may be faster, more reliable, and even more power and space efficient than conventional systems. For example, in at least one embodiment, the accelerator(s) 1114 in combination with the CPU(s) 1106, the GPU(s) 1108, and the data memory(s) 1116 may provide a fast, efficient platform for Level 3-5 autonomous vehicles.
[0143] In at least one embodiment, computer vision algorithms may be executed on CPUs configured with a high-level programming language, such as the C programming language, to perform a wide variety of processing algorithms on a wide variety of visual data. However, in at least one embodiment, CPUs are often unable to meet the performance requirements of many image processing applications, such as those related to execution time and power consumption. In at least one embodiment, many CPUs are unable to execute complex object detection algorithms in real time, which are used for in-vehicle ADAS applications and practical Level 3-5 autonomous vehicles.
[0144] Embodiments described herein enable multiple neural networks to be used concurrently and / or sequentially and the results combined together to enable Level 3-5 autonomous driving functionality. For example, in at least one embodiment, a CNN executing on the DLA or a discrete GPU (e.g., GPU(s) 1120) may include text and word recognition enabling reading and understanding of traffic signs, including signs for which the neural network has not been specifically trained. In at least one embodiment, the DLA may further include a neural network capable of identifying, interpreting, and semantically understanding a sign and passing this semantic understanding to path planning modules running on a CPU complex.
[0145] In at least one embodiment, multiple neural networks may be executed simultaneously, as required for Level 3, 4, or 5 driving. For example, in at least one embodiment, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with an electric light may be interpreted independently or jointly by multiple neural networks. In at least one embodiment, such a warning sign may itself be identified as a traffic sign by a first deployed neural network (e.g., a neural network that has been trained), the text "Flashing lights indicate icy conditions" may be interpreted by a second deployed neural network, which informs the vehicle's path planning software (preferably on the CPU complex) that icy conditions exist upon detection of flashing lights.In at least one embodiment, a flashing light may be identified by running a third deployed neural network over multiple frames, which informs the vehicle's path planning software of the presence (or absence) of flashing lights. In at least one embodiment, all three neural networks may run concurrently, for example, within the DLA and / or on the GPU(s) 1108.
[0146] In at least one embodiment, a CNN for facial recognition and vehicle owner identification may use data from camera sensors to identify the presence of an authorized driver and / or an owner of the vehicle 1100. In at least one embodiment, the always-on sensor processing engine may be used to unlock a vehicle and turn on lights when an owner approaches a driver-side door, and to disable a vehicle in security mode when an owner exits a vehicle. In this way, the SoC(s) 1104 provide protection against theft and / or vehicle theft.
[0147] In at least one embodiment, a CNN for detecting and identifying emergency vehicles may use data from microphones 1196 to detect and identify emergency vehicle sirens. In at least one embodiment, the SoC(s) 1104 uses a CNN to classify environmental and urban sounds, as well as to classify visual data. In at least one embodiment, a CNN running on a DLA is trained to characterize the relative approach speed of the emergency vehicle (e.g., by using the Doppler effect). In at least one embodiment, a CNN may also be trained to identify emergency vehicles specific to the local area in which a vehicle is operating, as identified by GNSS sensor(s) 1158.For example, in at least one embodiment, when operating in Europe, the CNN will attempt to detect European sirens, and when operating in the United States, the CNN will attempt to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program may be used to execute an emergency vehicle safety routine, slow a vehicle, pull over to the side of the road, park a vehicle, and / or idle a vehicle using ultrasonic sensor(s) 1162 until emergency vehicles pass by.
[0148] In at least one embodiment, a vehicle 1100 may include a CPU(s) 1118 (e.g., discrete CPU(s) or dCPU(s)) that may be coupled to the SoC(s) 1104 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, the CPU(s) 1118 may comprise, for example, an x86 processor. The CPU(s) 1118 may be used, for example, to perform a variety of functions, including arbitrating potentially inconsistent results between ADAS sensors and the SoC(s) 1104 and / or exemplary monitoring of the status and health of the controller(s) 1136 and / or an infotainment system-on-chip ("infotainment SoC") 1130.
[0149] In at least one embodiment, a vehicle 1100 may include one or more GPU(s) 1120 (e.g., discrete GPU(s) or dGPU(s)) that may be coupled to the SoC(s) 1104 via a high-speed interconnect (e.g., NVIDIA's NVLINK channel). The GPU(s) 1120 may provide additional artificial intelligence functionality, such as by executing redundant and / or distinct neural networks, and may be used to train and / or update neural networks based in part on inputs (e.g., sensor data) from sensors of a vehicle 1100.
[0150] In at least one embodiment, a vehicle 1100 may further include a network interface 1124, which may include one or more wireless antennas 1126 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, a network interface 1124 may be used to enable wireless connectivity to Internet cloud services (e.g., to one or more servers and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). In at least one embodiment, to communicate with other vehicles, a direct connection between a vehicle 1100 and another vehicle and / or an indirect connection (e.g., via networks and via the Internet) may be established.In at least one embodiment, direct connections may be provided via a vehicle-to-vehicle communication link. A vehicle-to-vehicle communication link may provide a vehicle 1100 with information about vehicles in the vicinity of the vehicle 1100 (e.g., vehicles in front of, to the side of, and / or behind a vehicle 1100). In at least one embodiment, the aforementioned functionality may be part of a cooperative adaptive cruise control feature of a vehicle 1100.
[0151] In at least one embodiment, a network interface 1124 may include an SoC that provides modulation and demodulation functionality and enables a controller(s) 1136 to communicate over wireless networks. In at least one embodiment, a network interface 1124 may include a radio frequency front end for upconverting from baseband to radio frequency and for downconverting from radio frequency to baseband. In at least one embodiment, the frequency conversions may be performed by any technically feasible method. For example, frequency conversions could be performed by well-known methods and / or by using superheterodyne methods. In at least one embodiment, the radio frequency front end functionality may be provided by a separate chip.In at least one embodiment, a network interface may include wireless capabilities for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0152] In at least one embodiment, a vehicle 1100 may further include, but is not limited to, data storage 1128, which may also include off-chip memory (e.g., outside of the SoC(s) 1104). In at least one embodiment, the data storage 1128 may include, but is not limited to, one or more memory elements including RAM, SRAM, dynamic random access memory ("DRAM"), video random access memory ("VRAM"), flash, hard drives, and / or other components and / or devices capable of storing at least one bit of data.
[0153] In at least one embodiment, a vehicle 1100 may further include one or more GNSS sensors 1158 (e.g., GPS and / or assisted GPS sensors) to assist with mapping, sensing, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1158 may be used, including, for example, and not limited to, a GPS with a USB connector and an Ethernet-to-serial (e.g., RS-232) bridge.
[0154] In at least one embodiment, a vehicle 1100 may further include a RADAR sensor(s) 1160. In at least one embodiment, a RADAR sensor(s) 1160 may be used by a vehicle 1100 for long-range vehicle detection even in darkness and / or extreme weather conditions. In at least one embodiment, the functional safety levels of the RADAR may be ASIL B. In at least one embodiment, a RADAR sensor(s) 1160 may use a CAN bus and / or a bus 1102 (e.g., to transmit data generated by RADAR sensors 1160) for control and access to object tracking data, and with access to Ethernet for accessing raw data. In at least one embodiment, a wide variety of RADAR sensor types may be used. For example, and without limitation, a RADAR sensor(s) 1160 may be suitable for front, rear, and side RADAR deployment.In at least one embodiment, one or more sensors are pulse Doppler radar sensors.
[0155] In at least one embodiment, the RADAR sensor(s) 1160 may include different configurations, such as long range with a narrow field of view, short range with a wide field of view, short range side coverage, etc. In at least one embodiment, the long range RADAR may be used for an adaptive cruise control function. In at least one embodiment, long range RADAR systems may provide a wide field of view realized by two or more independent scans, such as within a range of 250 m. In at least one embodiment, the RADAR sensor(s) 1160 may help distinguish between static and moving objects and may be used by an ADAS system 1138 for emergency braking assistance and forward collision warning.In at least one embodiment, sensor(s) 1160 included in a long-range RADAR system may include, but are not limited to, monostatic multimodal RADAR sensors with multiple (e.g., six or more) fixed RADAR antennas and a high-speed CAN and FlexRay interface. In at least one embodiment with six antennas, the central four antennas may create a focused beam pattern designed to capture surroundings of the vehicle 1100 at higher speeds with minimal interference from traffic in adjacent lanes. In at least one embodiment, the other two antennas may expand the field of view, allowing vehicles entering or exiting the lane of the vehicle 1100 to be quickly detected.
[0156] For example, in at least one embodiment, medium-range radar systems may include a range of up to 1160 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 150 degrees (rear). Short-range radar systems may include, but are not limited to, radar sensors configured for installation at either end of the rear bumper. When installed at either end of the rear bumper, in at least one embodiment, such a radar sensor system may generate two beams that continuously monitor the blind spot area to the rear and to the side of a vehicle. In at least one embodiment, short-range radar systems may be used in an ADAS system for blind spot detection and / or lane change assistance.
[0157] In at least one embodiment, a vehicle 1100 may further include one or more ultrasonic sensors 1162. In at least one embodiment, an ultrasonic sensor(s) 1162, which may be positioned at the front, rear, and / or sides of the vehicle 1100, may be used for parking assistance and / or for generating and updating an occupancy grid. In at least one embodiment, a wide variety of ultrasonic sensors 1162 may be used, and different ultrasonic sensors 1162 may be used for different detection ranges (e.g., 2.5 m; 4 m). In at least one embodiment, an ultrasonic sensor(s) 1162 may operate at functional safety levels of ASIL B.
[0158] In at least one embodiment, a vehicle 1100 may include one or more LIDAR sensors 1164. A LIDAR sensor(s) 1164 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, a LIDAR sensor(s) may be of functional safety level ASIL B. In at least one embodiment, a vehicle 1100 may include multiple LIDAR sensors 1164 (e.g., two, four, six, etc.) that may use an Ethernet channel (e.g., to provide data to a Gigabit Ethernet switch).
[0159] In at least one embodiment, a LIDAR sensor(s) 1164 may be capable of providing a list of objects and their distances for a 360-degree field of view. For example, commercially available LIDAR sensors 1164 may have an advertised range of approximately 100 m with an accuracy of 2 cm to 3 cm and support for a 100 Mbps Ethernet connection. In at least one embodiment, one or more non-protruding LIDAR sensors 1164 may be used. In such an embodiment, the LIDAR sensor(s) 1164 may comprise a small device that may be embedded in a front, rear, side, and / or corner of the vehicle 1100.In at least one embodiment, a LIDAR sensor(s) 1164, in such an embodiment, may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees with a range of 200 m even for low-reflectance objects. In at least one embodiment, a front-mounted LIDAR sensor(s) 1164 may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0160] In at least one embodiment, LIDAR technologies, such as 3D flash LIDAR, may also be used. In at least one embodiment, 3D flash LIDAR uses a laser flash as a transmission source to illuminate surroundings of a vehicle up to approximately 200 m. In at least one embodiment, a flash LIDAR unit includes, but is not limited to, a receptor that detects the laser pulse time of flight and reflected light on each pixel, corresponding to a range from a vehicle 1100 to objects. In at least one embodiment, flash LIDAR may enable highly accurate and distortion-free images of surroundings to be generated with each laser flash. In at least one embodiment, four flash LIDAR sensors may be deployed, one on each side of the vehicle 1100.In at least one embodiment, 3D flash lidar systems include, but are not limited to, a 3D solid-state lidar camera with a fixed array and no moving parts other than a fan (e.g., a non-scanning lidar device). In at least one embodiment, a flash lidar device(s) may use a Class I (eye-safe) laser with 5-nanosecond pulses per frame and collect the reflected laser light as 3D range point clouds and co-registered intensity data.
[0161] In at least one embodiment, a vehicle may further include one or more IMU sensors 1166. In at least one embodiment, an IMU sensor(s) 1166 may be located at a center of the rear axle of the vehicle 1100. In at least one embodiment, an IMU sensor(s) 1166 may include, for example, and not limited to, accelerometer(s), magnetometer(s), gyroscope(s), magnetic compass(es), and / or other sensor types. In at least one embodiment, such as in nine-axis applications, an IMU sensor(s) 1166 may include accelerometer(s) and gyroscope(s), while in nine-axis applications, an IMU sensor(s) 1166 may include accelerometer(s), gyroscope(s), and magnetometer(s).
[0162] In at least one embodiment, an IMU sensor(s) 1166 may be implemented as a miniaturized, high-performance GPS-aided inertial navigation system (GPS / INS) that combines micro-electro-mechanical systems (MEMS) of inertial sensors, a high-sensitivity GPS receiver, and extended Kalman filter algorithms to provide estimates of position, velocity vector, and altitude. In at least one embodiment, an IMU sensor(s) 1166 may enable a vehicle 1100 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in the velocity vector from a GPS to an IMU sensor(s) 1166. In at least one embodiment, an IMU sensor(s) 1166 and a GNSS sensor(s) 1158 be combined into a single integrated unit.
[0163] In at least one embodiment, a vehicle 1100 may include one or more microphones 1196 disposed in and / or around a vehicle 1100. In at least one embodiment, a microphone(s) 1196 may be used, among other things, for detecting and identifying emergency vehicles.
[0164] In at least one embodiment, a vehicle 1100 may further include any number of camera types, including one or more stereo cameras 1168, one or more wide-angle cameras 1170, one or more infrared cameras 1172, one or more surround-view cameras 1174, one or more long-range and / or medium-range cameras 1198, and / or other camera types. In at least one embodiment, cameras may be used to capture image data about an entire perimeter of the vehicle 1100. In at least one embodiment, the camera types used may depend on the embodiments and requirements for a vehicle 1100. In at least one embodiment, any combination of camera types may be used to provide the required coverage around a vehicle 1100. In at least one embodiment, the number of cameras may vary depending on the embodiment.For example, in at least one embodiment, a vehicle could include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. In at least one embodiment, cameras can support, for example, and not limited to, Gigabit Multimedia Serial Link ("GMSL") and / or Gigabit Ethernet. In at least one embodiment, each camera could be as described hereinabove with reference to FIG. Fig. 11A and Fig. 11B is described in more detail.
[0165] In at least one embodiment, a vehicle 1100 may further include one or more vibration sensors 1142. In at least one embodiment, vibration sensor(s) 1142 may measure the vibrations of components of the vehicle, such as an axle(s). For example, in at least one embodiment, changes in vibrations may indicate a change in the road surface. In at least one embodiment, when two or more vibration sensors 1142 are used, differences between vibrations may be used to determine friction or slippage of the road surface (e.g., when there is a vibration difference between a driven axle and a free-spinning axle).
[0166] In at least one embodiment, a vehicle 1100 may include an ADAS system 1138. In at least one embodiment, an ADAS system 1138 may include, in some examples but not limited to, an SoC. In at least one embodiment, an ADAS system 1138 may include, but is not limited to, any number and combination of an autonomous / adaptive / automatic cruise control (“ACC”) system, a cooperative adaptive cruise control (“CACC”) system, a forward collision warning (“FCW”) system, an automatic emergency braking (“AEB”) system, a lane departure warning (“LDW”) system, a lane keep assist (“LKA”) system, a blind spot warning (“BSW”) system, a rear cross traffic alert (“RCTW”) system, a collision warning (“CW”) system, a lane centering (“LC”) system, and / or other systems, features, and / or functionality.
[0167] In at least one embodiment, an ACC system may use one or more RADAR sensors 1160, one or more LIDAR sensors 1164, and / or any number of cameras. In at least one embodiment, an ACC system may include a longitudinal ACC and / or a lateral ACC system. In at least one embodiment, a longitudinal ACC system monitors and controls a distance to the vehicle immediately ahead of a vehicle 1100 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicles ahead. In at least one embodiment, a lateral ACC system performs follow-through and recommends a vehicle 1100 to change lanes if necessary. In at least one embodiment, a lateral ACC system is associated with other ADAS applications such as LC and CW.
[0168] In at least one embodiment, a CACC system utilizes information from other vehicles that may be received via a network interface 1124 and / or one or more wireless antennas 1126 from other vehicles over a wireless connection or indirectly via a network connection (e.g., over the Internet). In at least one embodiment, direct connections may be provided by a vehicle-to-vehicle ("V2V") communication link, while indirect connections may be provided by an infrastructure-to-vehicle ("I2V") communication link. In general, the V2V communication concept provides information about the immediately preceding vehicles (e.g., vehicles immediately in front of and in the same lane as a vehicle 1100), while the I2V communication concept may provide information about more distant traffic.In at least one embodiment, a CACC system may include one or both of the I2V and V2V information sources. In at least one embodiment, given information from the vehicles ahead of a vehicle 1100, a CACC system may be more reliable and has the potential to improve traffic flow smoothness and reduce congestion on the road.
[0169] In at least one embodiment, an FCW system is configured to warn a driver of a hazard so that the driver can take corrective action. In at least one embodiment, an FCW system uses a forward-facing camera and / or one or more RADAR sensors 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to provide driver feedback, such as a display, speaker, and / or vibration component. In at least one embodiment, an FCW system can provide a warning, such as in the form of a sound, a visual warning, a vibration, and / or a rapid braking pulse.
[0170] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object and may automatically apply the brakes if a driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, an AEB system may utilize one or more forward-facing cameras and / or one or more radar sensors 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when an AEB system detects a hazard, it will first alert a driver to take corrective action to avoid a collision, and if a driver does not take corrective action, an AEB system may automatically apply brakes in an effort to prevent or at least mitigate an impact of a predicted collision.In at least one embodiment, an AEB system may include techniques such as dynamic brake assist and / or imminent collision braking.
[0171] In at least one embodiment, an LDW system provides visual, audible, and / or tactile alerts, such as steering wheel or seat vibrations, to warn a driver when a vehicle 1100 crosses lane markings. In at least one embodiment, an LDW system is not activated when a driver indicates an intentional lane departure, such as by activating a turn signal. In at least one embodiment, an LDW system may utilize forward / side-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibration component. In at least one embodiment, an LKA system is a variant of an LDW system.In at least one embodiment, an LKA system provides steering input or braking to correct a vehicle 1100 when a vehicle 1100 begins to depart from its lane.
[0172] In at least one embodiment, a BSW system detects vehicles in a vehicle's blind spot and warns a driver of them. In at least one embodiment, a BSW system may provide a visual, audible, and / or tactile alert to indicate that merging or changing lanes is unsafe. In at least one embodiment, a BSW system may provide an additional warning when a driver uses a turn signal. In at least one embodiment, a BSW system may utilize one or more rear-facing cameras and / or one or more RADAR sensors 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0173] In at least one embodiment, an RCTW system may provide visual, audible, and / or tactile notification when an object is detected outside of the rearview camera range when a vehicle 1100 is reversing. In at least one embodiment, an RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a crash. In at least one embodiment, an RCTW system may utilize one or more rear-facing RADAR sensors 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0174] In at least one embodiment, conventional ADAS systems may be prone to false positives, which may be annoying and disruptive to a driver, but are typically not catastrophic because the ADAS systems warn a driver and allow the driver to decide whether a safety condition actually exists and act accordingly. In at least one embodiment, even in the case of conflicting results, a vehicle 1100 decides whether to heed a result from a primary computer or a secondary computer (e.g., a first controller or a second controller 1136). For example, in at least one embodiment, an ADAS system 1138 may be a backup and / or secondary computer to provide perception information to a rationality module of a backup computer.In at least one embodiment, a rationality monitor of the backup computer may execute redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. In at least one embodiment, outputs from an ADAS system 1138 may be provided to a supervisor MCU. In at least one embodiment, if outputs from a primary computer and outputs from a secondary computer conflict, the supervisor MCU determines how to resolve the conflict to ensure safe operation.
[0175] In at least one embodiment, a primary computer may be configured to provide a supervisor MCU with a confidence score indicating a primary computer's confidence in the selected outcome. In at least one embodiment, if the confidence score exceeds a threshold, a supervisor MCU may follow the direction of a primary computer regardless of whether the secondary computer provides a conflicting or inconsistent outcome. In at least one embodiment, if a confidence score does not meet a threshold and where primary and secondary computers indicate different outcomes (e.g., a conflict), a supervisor MCU may arbitrate between computers to determine an appropriate outcome.
[0176] In at least one embodiment, a supervisor MCU may be configured to operate one or more neural networks trained and configured to determine, based in part on outputs from a primary computer and a secondary computer, conditions under which the secondary computer will provide false alarms. In at least one embodiment, neural network(s) in a supervisor MCU may learn when output from the secondary computer can be trusted and when it cannot. For example, if the secondary computer is a radar-based FCW system in at least one embodiment, neural network(s) in the supervisor MCU may learn when an FCW system identifies metallic objects that are not actually hazards, such as a drainage grate or manhole cover, that triggers an alarm.In at least one embodiment, when a secondary computer is a camera-based LDW system, a neural network in the supervisor MCU may learn to override the LDW when cyclists or pedestrians are present and lane departure is actually the safest maneuver. In at least one embodiment, a supervisor MCU may include at least one of a DLA or a GPU capable of executing neural network(s) with associated memory. In at least one embodiment, a supervisor MCU may comprise a component and / or be included as a component of the SoC(s) 1104.
[0177] In at least one embodiment, an ADAS system 1138 may include a secondary computer that executes ADAS functionality using conventional computer vision rules. In at least one embodiment, the secondary computer may use classic computer vision (if-then) rules, and the presence of one or more neural networks in the supervisor MCU may improve reliability, safety, and performance. For example, in at least one embodiment, diverse implementation and intentional non-identity makes an overall system more fault-tolerant, particularly against errors caused by a software (or software-hardware interface) functionality.For example, in at least one embodiment, if there is a software bug or error in software running on the primary computer and non-identical software code running on a secondary computer produces a consistent overall result, then a supervisor MCU may have more confidence that an overall result is correct and a bug in the software or hardware on the primary computer does not cause a significant error.
[0178] In at least one embodiment, an output of an ADAS system 1138 may be fed into a perception block of a primary computer and / or a dynamic driving task block of a primary computer. For example, in at least one embodiment, if an ADAS system 1138 indicates a forward collision warning due to an immediately preceding object, a perception block may use this information in identifying objects. In at least one embodiment, a secondary computer may have its own neural network trained, thus reducing the risk of false positives, as described herein.
[0179] In at least one embodiment, a vehicle 1100 may further include an infotainment SoC 1130 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC in at least one embodiment, the infotainment system, in at least one embodiment, may not be an SoC and may include, but is not limited to, two or more discrete components. In at least one embodiment, an infotainment SoC 1130 may include, but is not limited to, a combination of hardware and software that may be used to provide audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g.,Navigation systems, rear parking assistance, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, oil level, door open / closed, air filter information, etc.) to a vehicle 1100. For example, an infotainment SoC 1130 could include radios, disk players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, WiFi, steering wheel audio controls, hands-free calling, a head-up display ("HUD"), an HMI display 1134, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, an infotainment SoC 1130 may be further used to provide information (e.g.,visual and / or audible) to a user(s) of a vehicle 1100, such as information from an ADAS system 1138, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0180] In at least one embodiment, an infotainment SoC 1130 may include any amount and type of GPU functionality. In at least one embodiment, an infotainment SoC 1130 may communicate with other devices, systems, and / or components of a vehicle 1100 via a bus 1102 (e.g., CAN bus, Ethernet, etc.). In at least one embodiment, an infotainment SoC 1130 may be coupled to a supervisor MCU so that a GPU of the infotainment system may assume some self-driving functions in the event that the primary controller(s) 1136 (e.g., the primary and / or backup computers of the vehicle 1100) fail(s). In at least one embodiment, an infotainment SoC 1130 may place a vehicle 1100 into a chauffeur-to-safe-stop mode, as described herein.
[0181] In at least one embodiment, a vehicle 1100 may further include an instrument cluster 1132 (e.g., a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). In at least one embodiment, the instrument cluster 1132 may include, but is not limited to, a controller and / or a supercomputer (e.g., a discrete controller or a supercomputer). In at least one embodiment, the instrument cluster 1132 may include any number and combination of a set of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn signal, shift position indicator, one or more seat belt warning lights, one or more parking brake warning lights, one or more engine trouble lights, supplemental restraint system (e.g.,Airbag information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between an infotainment SoC 1130 and an instrument cluster 1132. In at least one embodiment, an instrument cluster 1132 may be integrated as part of an infotainment SoC 1130, or vice versa.
[0182] An inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 815 are described below in connection with Fig. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 in the system of Fig.11C may be used to infer or predict operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0183] Fig. 11D is a system diagram 1176 for communication between one or more cloud-based servers and an autonomous vehicle 1100 of Fig.11A according to at least one embodiment. In at least one embodiment, a system 1176 may include, but is not limited to, one or more servers 1178, one or more networks 1190, and any number and type of vehicles, including a vehicle 1100. In at least one embodiment, a server 1178 may include multiple GPUs 1184(A)-1184(H) (collectively referred to herein as GPUs 1184), PCIe switches 1182(A)-1182(H) (collectively referred to herein as PCIe switches 1182), and / or CPUs 1180(A)-1180(B) (collectively referred to herein as CPUs 1180). In at least one embodiment, GPUs 1184, CPUs 1180, and PCIe switches may be connected to high-speed interconnects, such as, but not limited to, NVLink interfaces 1188 developed by NVIDIA and / or PCIe connectors 1186.In at least one embodiment, GPUs 1184 are connected via NVLink and / or NVSwitch SoC, and the GPUs 1184 and the PCIe switches 1182 are connected via PCIe interconnects. Although eight GPUs 1184, two CPUs 1180, and two PCIe switches are illustrated in at least one embodiment, this is not intended to be limiting. In at least one embodiment, each of the servers 1178 may include, but is not limited to, any number of GPUs 1184, CPUs 1180, and / or PCIe switches. For example, in at least one embodiment, one or more servers 1178 could each include eight, sixteen, thirty-two, and / or more GPUs 1184.
[0184] In at least one embodiment, a server 1178 may receive, via a network(s) 1190 and from vehicles, image data representing images depicting unexpected or changed road conditions, such as recently commenced roadwork. In at least one embodiment, a server 1178 may transmit, via a network(s) 1190 and to the vehicles, neural networks 1192, updated or otherwise, neural networks 1192 and / or map information 1194, including, but not limited to, information regarding traffic and road conditions. In at least one embodiment, updates to the map information 1194 may include updates to the HD map 1122, such as information about construction, potholes, detours, flooding, and / or other obstacles.In at least one embodiment, neural networks 1192 and / or map information 1194 may have resulted from new training and / or from experience represented by data from any number of vehicles in the environment and / or based on training performed in a data center (e.g., using server(s) 1178 and / or other server(s).
[0185] In at least one embodiment, a server 1178 may be used to train machine learning models (e.g., neural networks) based in part on training data. In at least one embodiment, training data may be generated by vehicles and / or in a simulation (e.g., with a gaming machine). In at least one embodiment, any number of training data items are labeled (e.g., if the neural network benefits from supervised learning) and / or undergo other preprocessing. In at least one embodiment, any number of training data items are unlabeled and / or preprocessed (e.g., if the neural network does not require supervised learning). In at least one embodiment, once machine learning models are trained, vehicle machine learning models may be used (e.g.,transmitted to vehicles over a network(s) 1190), and / or machine learning models may be used by a server(s) 1178 for remote vehicle monitoring.
[0186] In at least one embodiment, a server 1178 may receive data from vehicles and apply data to real-time neural networks for intelligent real-time inferencing. In at least one embodiment, a server 1178 may include deep learning supercomputers and / or dedicated AI computers powered by one or more GPUs 1184, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, a server 1178 may include a deep learning infrastructure using CPU-powered data centers.
[0187] In at least one embodiment, a deep learning infrastructure from a server(s) 1178 may be capable of rapid, real-time inference and may utilize this capability to assess and verify the health of the processors, software, and / or associated hardware in the vehicle 1100. For example, in at least one embodiment, a deep learning infrastructure may receive periodic updates from a vehicle 1100, such as a sequence of images and / or objects that a vehicle 1100 has located in that sequence of images (e.g., through computer vision and / or other machine learning techniques for classifying learning objects).In at least one embodiment, the deep learning infrastructure may run its own neural network to label objects and compare them to the objects identified by a vehicle 1100, and if results do not match and a deep learning infrastructure concludes that AI in the vehicle 1100 is not working, then a server 1178 may send a signal to a vehicle 1100 instructing a fail-safe computer of a vehicle 1100 to take over control, notify passengers, and perform a safe parking maneuver.
[0188] In at least one embodiment, a server 1178 may include one or more GPU(s) 1184 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3 devices). In at least one embodiment, a combination of GPU-powered servers and inference acceleration may enable real-time responsiveness. In at least one embodiment, such as where performance is less critical, servers powered by CPUs, FPGAs, and other processors may be used for inference. In at least one embodiment, hardware structures are used to perform one or more embodiments. Details regarding hardware structure(s) 815 are described herein in connection with Fig. 8A and / or 8B provided. COMPUTER SYSTEMS
[0189] Fig.12 is a block diagram illustrating an example computer system, which may be a system with interconnected devices and components, a system on a chip (SOC), or a combination thereof, formed with a processor that may include execution units for executing an instruction, according to at least one embodiment. In at least one embodiment, a computer system 1200 may include, but is not limited to, a component, such as a processor 1202, for utilizing execution units with logic for executing algorithms on process data in accordance with the present disclosure, such as the embodiments described herein. In at least one embodiment, the computer system 1200 may include processors, such as the
[0190] PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™ or Intel® Nervana™ microprocessors, available from Intel Corporation of Santa Clara, California, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, and the like) may be used. In at least one embodiment, computer system 1200 may run a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0191] Embodiments may be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions, according to at least one embodiment.
[0192] In at least one embodiment, computer system 1200 may include, but is not limited to, processor 1202, which may include, but is not limited to, one or more execution units 1208 to perform machine learning model training and / or inference in accordance with techniques described herein. In at least one embodiment, computer system 1200 is a single-processor desktop or server system; however, in another embodiment, computer system 1200 may be a multiprocessor system.In at least one embodiment, processor 1202 may include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 1202 may be coupled to a processor bus 1210 that can communicate data signals between processor 1202 and other components in computer system 1200.
[0193] In at least one embodiment, processor 1202 may include, but is not limited to, an internal Level 1 ("L1") cache ("cache") 1204. In at least one embodiment, processor 1202 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache may be external to processor 1202. Other embodiments may also include a combination of internal and external caches, depending on the implementation and needs. In at least one embodiment, a register file 1206 may store various types of data in various registers, including, but not limited to, integer registers, floating-point registers, state registers, and an instruction pointer register.
[0194] In at least one embodiment, execution unit 1208, including, but not limited to, logic for performing integer and floating-point operations, is also located in processor 1202. In at least one embodiment, processor 1202 may also include microcode ("ucode") read-only memory ("ROM") that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 1208 may include logic for handling a packed instruction set 1209. In at least one embodiment, by including packed instruction set 1209 in the instruction set of a general-purpose processor, along with associated instruction execution circuitry, operations used by many multimedia applications may be performed using packed data in a general-purpose processor 1202.In one or more embodiments, many multimedia applications may be accelerated and executed more efficiently by using the full width of a processor data bus to perform operations on packed data, which may eliminate the need to transfer smaller units of data across the processor data bus to perform one or more operations one data element at a time.
[0195] In at least one embodiment, execution unit 1208 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 1200 may include, but is not limited to, memory 1220. In at least one embodiment, memory 1220 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or another storage device. In at least one embodiment, memory 1220 may store one or more instructions 1219 and / or data 1221 represented by data signals that may be executed by processor 1202.
[0196] In at least one embodiment, a system logic chip may be coupled to processor bus 1210 and memory 1220. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub ("MCH") 1216, and processor 1202 may communicate with MCH 1216 via processor bus 1210. In at least one embodiment, MCH 1216 may provide a high-bandwidth memory path 1218 to memory 1220 for instruction and data storage, as well as for storing graphics instructions, data, and textures. In at least one embodiment, MCH 1216 may route data signals between processor 1202, memory 1220, and other components in computer system 1200, and may bridge data signals between processor bus 1210, memory 1220, and a system I / O interface 1222.In at least one embodiment, a system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 1216 may be coupled to memory 1220 via a high-bandwidth memory path 1218, and a graphics / video card 1218 may be coupled to MCH 1216 via an Accelerated Graphics Port ("AGP") interconnect 1214.
[0197] In at least one embodiment, computer system 1200 may use system I / O interface 1222 as a proprietary hub interface bus to connect MCH 1216 to an I / O controller hub ("ICH") 1230. In at least one embodiment, ICH 1230 may provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, a local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripherals to memory 1220, a chipset, and processor 1202.Examples may include, but are not limited to, an audio controller 1229, a firmware hub ("flash BIOS") 1228, a wireless transceiver 1226, data storage 1224, a legacy I / O controller 1223 with user input and keyboard interfaces, a serial expansion port 1227, such as a Universal Serial Bus ("USB") port, and a network controller 1234. In at least one embodiment, data storage 1224 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0198] In at least one embodiment, Fig. 12 a system comprising interconnected hardware devices or “chips”, while in other embodiments Fig. 12 may represent an exemplary system-on-chip ("SoC"). In at least one embodiment, Fig.12 may be connected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of computer system 1200 are connected using Compute Express Link (CXL) interconnects.
[0199] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described herein in connection with Fig. 8A and / or 8B. In at least one embodiment, in the system of Fig.12 the inference and / or training logic 815 may be used to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0200] In at least one embodiment, a computer system as described above may be used to create a robot control system that implements model-based and model-free control. For example, a computer system as described above may include memory storing executable instructions that, as a result of being executed by a processor of the computer system, cause the computer system to implement model-based and model-free control as described herein.
[0201] Fig.13 is a block diagram illustrating an electronic device 1300 for using a processor 1310, according to at least one embodiment. In at least one embodiment, the electronic device 1300 may be, for example, and not limited to, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop computer, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0202] In at least one embodiment, electronic device 1300 may include, but is not limited to, processor 1310 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1310 is connected to the processor 1310 via a bus or interface, such as an I 2 C-Bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advance Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3, etc.), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, Fig. 13 a system comprising interconnected hardware devices or “chips”, while in other embodiments Fig.13 may represent an exemplary system on a chip (“SoC”). In at least one embodiment, the Fig. 13 may be connected to proprietary connections, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of Fig. 13 interconnected using Compute Express Link (CXL) interconnects.
[0203] In at least one embodiment, Fig.13 a display 1324, a touchscreen 1325, a touchpad 1330, a near-field communications unit (“NFC”) 1345, a sensor hub 1340, a thermal sensor 1346, an express chipset (“EC”) 1335, a trusted platform module (“TPM”) 1338, BIOS / firmware / flash memory (“BIOS, FW-Flash”) 1322, a DSP 1360, a drive (“SSD or HDD”) 1312, such as a solid-state disk (“SSD”) or a hard disk (“HDD”), a wireless local area network unit (“WLAN”) 1350, a Bluetooth unit 1352, a wireless wide area network unit (“WWAN”) 1356, a global positioning system (GPS) unit 1355, a camera (“USB 3.0 Camera”) 1354, such as a USB 3.0 camera, or a Low Power Double Data Rate (“LPDDR”) memory unit (“LPDDR3”) 1315, implemented, for example, in an LPDDR3 standard.These components can each be implemented in any suitable manner.
[0204] In at least one embodiment, other components may be communicatively coupled to processor 1310 through the components described herein. In at least one embodiment, an accelerometer 1341, an ambient light sensor (ALS) 1342, a compass 1343, and a gyroscope 1344 may be communicatively coupled to sensor hub 1340. In at least one embodiment, a thermal sensor 1339, a fan 1337, a keyboard 1336, and a touchpad 1330 may be communicatively coupled to EC 1335. In at least one embodiment, speakers 1363, headphones 1364, and a microphone (“mic”) 1365 may be communicatively coupled to an audio unit (“audio codec and Class D amp”) 1362, which in turn may be communicatively coupled to DSP 1360.In at least one embodiment, an audio unit 1362 may include, for example, and not limited to, an audio encoder / decoder ("codec") and a Class D amplifier. In at least one embodiment, the SIM card ("SIM") 1357 may be communicatively coupled to the WWAN unit 1356. In at least one embodiment, components such as the WLAN unit 1350 and the Bluetooth unit 1352, as well as the WWAN unit 1356, may be implemented in a next-generation form factor (NGFF).
[0205] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described herein in connection with Fig.8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 in the system of Fig. 13 be used to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0206] In at least one embodiment, a computer system as described above may be used to create a robot control system that implements model-based and model-free control. For example, a computer system as described above may include memory storing executable instructions that, as a result of being executed by a processor of the computer system, cause the computer system to implement model-based and model-free control as described herein.
[0207] Fig. 14 illustrates a computer system 1400 according to at least one embodiment. In at least one embodiment, the computer system 1400 is configured to implement various processes and methods described throughout this disclosure.
[0208] In at least one embodiment, computer system 1400 includes, but is not limited to, at least one central processing unit ("CPU") 1402 connected to a communications bus 1410 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or other bus or one or more point-to-point communications protocols. In at least one embodiment, computer system 1400 includes, but is not limited to, main memory 1404 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data is stored in main memory 1404, which may take the form of random access memory ("RAM").In at least one embodiment, a network interface subsystem (“network interface”) 1422 provides an interface to other computing devices and networks for receiving data from and transmitting data to other systems with the computer system 1400.
[0209] In at least one embodiment, computer system 1400 includes, but is not limited to, input devices 1408, a parallel processing system 1412, and display devices 1406, which may be implemented using a conventional cathode ray tube ("CRT"), liquid crystal display ("LCD"), light-emitting diode ("LED"), plasma display, or other suitable display technology. In at least one embodiment, user input is received from input devices 1428 such as a keyboard, mouse, touchpad, microphone, etc. In at least one embodiment, each module described herein may be arranged on a single semiconductor platform to form a processing system.
[0210] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described herein in connection with Fig. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 in the system of Fig. 14 be used to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0211] In at least one embodiment, a computer system as described above may be used to create a robot control system that implements model-based and model-free control. For example, a computer system as described above may include memory storing executable instructions that, as a result of being executed by a processor of the computer system, cause the computer system to implement model-based and model-free control as described herein.
[0212] Fig.15 illustrates a computer system 1500 according to at least one embodiment. In at least one embodiment, the computer system 1500 includes, but is not limited to, a computer 1510 and a USB flash drive 1520. In at least one embodiment, the computer 1510 may include, but is not limited to, any number and type of processor(s) (not shown) and memory (not shown). In at least one embodiment, the computer 1510 includes, but is not limited to, a server, a cloud instance, a laptop, and a desktop computer.
[0213] In at least one embodiment, USB flash drive 1520 includes, but is not limited to, a processing unit 1530, a USB interface 1540, and USB interface logic 1550. In at least one embodiment, processing unit 1530 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, processing unit 1530 may include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, processing core 1530 includes an application-specific integrated circuit ("ASIC") optimized to perform any set and type of machine learning-related operations.For example, in at least one embodiment, processing unit 1530 is a tensor processing unit ("TPC") optimized for performing machine learning inference operations. In at least one embodiment, processing unit 1530 is a vision processing unit ("VPU") optimized for performing machine vision and machine learning inference operations.
[0214] In at least one embodiment, USB interface 1540 may be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 1540 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 1540 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 1550 may include any amount and type of logic that enables processing unit 1530 to communicate with devices (e.g., computer 1510) via USB connector 1540.
[0215] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described herein in connection with Fig.8A and / or 8B. In at least one embodiment, in the system of Fig. 15 the inference and / or training logic 815 may be used to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0216] In at least one embodiment, a computer system as described above may be used to create a robot control system that implements model-based and model-free control. For example, a computer system as described above may include memory storing executable instructions that, as a result of being executed by a processor of the computer system, cause the computer system to implement model-based and model-free control as described herein.
[0217] Fig.16A illustrates an example architecture in which multiple GPUs 1610(1)-1610(N) are communicatively coupled to multiple multi-core processors 1605(1)-1605(M) via high-speed interconnects 1640(1)-1640(N) (e.g., buses, point-to-point interconnects, etc.). In one embodiment, high-speed interconnects 1640(1)-1640(N) support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or higher. In at least one embodiment, various interconnect protocols may be used, including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0. In various figures, ("N") and ("M") represent positive integers, the values of which may vary from figure to figure.
[0218] Additionally, in one embodiment, two or more GPUs 1610 are interconnected via high-speed interconnects 1629(1)-1629(2), which may be implemented using similar or different protocols / connections than those used for high-speed interconnects 1640(1)-1640(N). Similarly, two or more multi-core processors 1605 may be interconnected via high-speed interconnect 1628, which may be symmetric multi-core processor (SMP) buses operating at 12 GB / s, 30 GB / s, 120 GB / s, or more. Alternatively, all communication between Fig. 16A, the various system components are executed using the same protocols / connections (e.g., via a common interconnect architecture).
[0219] In one embodiment, each multi-core processor 1605 is coupled to processor memory 1601(1)-1601(M) via memory interconnects 1626(1)-1626(M), respectively, and each GPU 1610(1)-1610(N) is communicatively coupled to GPU memory 1620(1)-1620(N) via GPU memory interconnects 1650(1)-1650(N), respectively. In at least one embodiment, memory interconnects 1626 and 1650 may utilize the same or different memory access technologies. By way of example and not limitation, processor memories 1601(1)-1601(M) and GPU memory 1620 may be volatile memories such as dynamic random access memories (DRAMs) (including stacked DRAMs), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (“HBM”) and / or non-volatile memories such as 3D XPoint or Nano-Ram.In one embodiment, a portion of processor memory 1601 may be volatile memory and another portion may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0220] As described herein, although various multi-core processors 1605 and GPUs 1610 may be physically coupled to a particular memory 1601, 1620, and / or a unified memory architecture may be implemented in which a virtual system address space (also referred to as "effective address space") is distributed among various physical memories. For example, processor memories 1601(1)-1601(M) may each comprise 64 GB of system memory address space, and GPU memories 1620(1)-1620(N) may each comprise 32 GB of system memory address space, resulting in a total of 256 GB of addressable memory when M=2 and N=4.
[0221] Fig.16B illustrates additional details for an interconnect between a multi-core processor 1607 and a graphics acceleration module 1646 according to an example embodiment. In at least one embodiment, the graphics acceleration module 1646 may include one or more GPU chips integrated on a wiring card coupled to the processor 1607 via a high-speed interconnect 1640 (e.g., a PCI bus, NVLink, etc.). Alternatively, in at least one embodiment, the graphics acceleration module 1646 may be integrated on a package or die with the processor 1607.
[0222] In at least one embodiment, processor 1607 includes a plurality of cores 1660A-1660D, each having a translation lookaside buffer ("TLB") 1661A-1661D and one or more caches 1662A-1662D. In at least one embodiment, cores 1660A-1660D may include various other components for executing instructions and processing data, which are not illustrated. In at least one embodiment, caches 1662A-1662D may include level 1 (L1) and level 2 (L2) caches. Additionally, one or more shared caches 1656 may be included in caches 1662A-1662D and shared by sets of cores 1660A-1660D. For example, one embodiment of processor 1607 includes 24 cores, each with its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, one or more L2 and L3 caches are shared between two adjacent cores.In at least one embodiment, processor 1607 and graphics acceleration module 1646 couple to system memory 1614, which includes processor memories 1601(1)-1601(M) of FIG. Fig. 16A may include.
[0223] In at least one embodiment, coherency for data and instructions stored in various caches 1662A-1662D, 1656, and system memory 1614 is maintained via inter-core communication over a coherency bus 1664. For example, cache coherency logic / circuitry may be associated with each cache to communicate with the coherency bus 1664 in response to detected read or write operations to specific cache lines. In at least one embodiment, a cache observation protocol is implemented over the coherency bus 1664 to observe cache accesses.
[0224] In at least one embodiment, a proxy circuit (PROXY) 1625 communicatively couples the graphics acceleration module 1646 to the coherence bus 1664 so that the graphics acceleration module 1646 can participate in a cache coherence protocol as a peer of the cores 1660A-1660D. Specifically, an interface (INTF) 1635 provides connectivity to the proxy circuit 1625 via the high-speed interconnect 1640, and an interface (INTF) 1637 connects the graphics acceleration module 1646 to the high-speed interconnect 1640.
[0225] In at least one embodiment, an accelerator integration circuit 1636 provides cache management, memory access, context management, and interrupt management services on behalf of multiple graphics processing engines 1631(1)-1631(N) of the graphics acceleration module 1646. In at least one embodiment, the graphics processing engines 1631(1)-1631(N) may each comprise a separate graphics processing unit (GPU). Alternatively, in at least one embodiment, the graphics processing engines 1631(1)-1631(N) may comprise various types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines.In at least one embodiment, the graphics acceleration module 1646 may be a GPU with multiple graphics processing engines 1631(1)-1631(N), or the graphics processing engines 1631(1)-1631(N) may be individual GPUs integrated in or on a common package, wiring board, or chip.
[0226] In at least one embodiment, accelerator integration circuit 1636 includes a memory management unit (MMU) 1639 for performing various memory management functions, such as virtual-to-physical memory translations (also referred to as effective-to-real memory translations) and memory access protocols for accessing system memory 1614. In at least one embodiment, MMU 1639 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations. In at least one embodiment, a cache 1638 may cache instructions and data for efficient access by graphics processing engines 1631(1)-1631(N).In one embodiment, the data stored in cache 1638 and graphics memories (GFX MEM) 1633(1)-1633(M) is kept coherent with core caches 1662A-1662D, 1656 and system memory 1614, possibly using a fetch unit 1644. As mentioned, this may be accomplished via proxy circuitry 1625 on behalf of cache 1638 and memories 1633(1)-1633(M) (e.g., sending updates to cache 1638 regarding modifications / accesses to cache lines on processor caches 1662A-1662D, 1656 and receiving updates from cache 1638).
[0227] In at least one embodiment, a set of registers 1645 stores context data for threads executed by graphics processing engines 1631(1)-1631(N), and a context management circuit 1648 manages thread contexts. For example, context management circuit 1648 may perform save and restore operations to save and restore contexts of different threads during context switches (e.g., when a first thread is saved and a second thread is saved so that a second thread can be executed by a graphics processing engine). For example, upon a context switch, context management circuit 1648 may save current register values to a specific location in memory (identified, for example, by a context pointer). Upon returning to a context, it may then restore the register values.In one embodiment, an interrupt management circuit (INTRPT MGMT) 1647 receives and processes interrupts received from system devices.
[0228] In one implementation, virtual / effective addresses from a graphics processing engine 1631 are translated into real / physical addresses in system memory 1614 by the MMU 1639. One embodiment of the accelerator integration circuit 1636 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 1646 and / or other acceleration devices. In at least one embodiment, the graphics accelerator module 1646 may be associated with a single application executing on the processor 1607 or may be shared between multiple applications. In at least one embodiment, a virtualized graphics execution environment is presented in which resources of the graphics processing engines 1631(1)-1631(N) are shared among multiple applications or virtual machines (VMs). In at least one embodiment, resources may be sliced.be divided into “slices” that are allocated to different VMs and / or applications based on processing requirements and priorities associated with VMs and / or applications.
[0229] In at least one embodiment, accelerator integration circuitry 1636 acts as a bridge to a system for graphics acceleration module 1646 and provides address translation and system memory caching services. Additionally, accelerator integration circuitry 1636 may provide virtualization facilities for a host processor to manage the virtualization of graphics processing engines 1631-1632, interrupts, and memory management.
[0230] In at least one embodiment, because hardware resources of graphics processing engines 1631(1)-1631(N) are explicitly mapped to a real address space seen by host processor 1607, each host processor can directly address these resources using an effective address value. In at least one embodiment, a function of accelerator integration circuit 1636 is to physically separate graphics processing engines 1631(1)-1631(N) so that they appear to a system as independent entities.
[0231] In at least one embodiment, one or more graphics memories 1633(1)-1633(M) are coupled to each of the graphics processing engines 1631(1)-1631(N). In at least one embodiment, the graphics memories 1633(1)-1633(M) store instructions and data processed by each of the graphics processing engines 1631(1)-1631(N). In at least one embodiment, the graphics memories 1633(1)-1633(M) may be volatile memories, such as DRAMs (including stacked DRAMs), GDDR memories (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memories such as 3D XPoint or Nano-Ram.
[0232] In one embodiment, to reduce data traffic over high-speed interconnect 1640, biasing techniques are used to ensure that data stored in graphics memories 1633(1)-1633(M) is data most frequently used by graphics processing engines 1631(1)-1631(N) and preferably not used (at least not frequently) by cores 1660A-1660D. Similarly, in at least one embodiment, a biasing mechanism attempts to keep data needed by the cores (and preferably not by graphics processing engines 1631(1)-1631(N)) in caches 1662A-1662D, 1656, and system memory 1614.
[0233] Fig.16C illustrates another exemplary embodiment in which accelerator integration circuitry 1636 is integrated with processor 1607. In this embodiment, graphics processing engines 1631(1)-1631(N) communicate directly over high-speed interconnect 1640 with accelerator integration circuitry 1636 via interface 1637 and interface 1635 (which, in turn, may use any form of bus or interface protocol). In at least one embodiment, accelerator integration circuitry 1636 may perform similar operations to those described with respect to Fig.16B, but potentially with higher throughput due to its close proximity to the coherence bus 1664 and caches 1662A-1662D, 1656. One embodiment supports different programming models, including a dedicated process programming model (no graphics acceleration module virtualization) and shared programming models (with virtualization), which may include programming models controlled by the accelerator integration circuit 1636 and programming models controlled by the graphics acceleration module 1646.
[0234] In at least one embodiment, the graphics processing engines 1631(1)-1631(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application may direct other application requests to the graphics processing engines 1631(1)-1631(N) to provide virtualization within a VM / partition.
[0235] In at least one embodiment, the graphics processing engines 1631(1)-1631(N) may be shared between multiple VM / application partitions. In at least one embodiment, shared models may use a system hypervisor to virtualize the graphics processing engines 1631(1)-1631(N) to enable access by any operating system. In at least one embodiment, for single-partition systems without a hypervisor, the graphics processing engines 1631(1)-1631(N) are owned by an operating system. In at least one embodiment, an operating system may virtualize the graphics processing engines 1631(1)-1631(N) to enable access by any process or application.
[0236] In at least one embodiment, the graphics acceleration module 1646 or an individual graphics processing engine 1631(1)-1631(N) selects a process element using a process handle. In one embodiment, process elements are stored in system memory 1614 and are addressable using the effective address to real address translation techniques described herein. In at least one embodiment, a process handle may be an implementation-specific value provided to a host process upon registering its context with the graphics processing engine 1631(1)-1631(N) (i.e., invoking system software to add a process element to a linked list of process elements). In at least one embodiment, the lower 16 bits of a process handle may be an offset of the process element within a linked list of process elements.
[0237] Fig.16D illustrates an exemplary accelerator integration slice 1690. In at least one embodiment, a "slice" comprises a particular portion of the processing resources of accelerator integration circuit 1636. In at least one embodiment, an application-effective address space 1682 in system memory 1614 stores process elements 1683. In at least one embodiment, process elements 1683 are stored in response to GPU calls 1681 from applications 1680 executing on processor 1607. In at least one embodiment, a process element 1683 contains the process state for the corresponding application 1680. A work descriptor (WD) 1684 contained in process element 1683 may be a single job requested by an application or may contain a pointer to a queue of jobs.In at least one embodiment, WD 1684 is a pointer to a job request queue in the effective address space 1682 of an application.
[0238] In at least one embodiment, a graphics acceleration module 1646 and / or individual graphics processing engines 1631(1)-1631(N) may be shared by all or a subset of processes in a system. In at least one embodiment, an infrastructure for establishing process state and sending a WD 1684 to a graphics acceleration module 1646 to start a job in a virtualized environment may be included.
[0239] In at least one embodiment, a dedicated process programming model is implementation-specific. In at least one embodiment, in this model, a single process owns the graphics acceleration module 1646 or a single graphics processing engine 1631. In at least one embodiment, when the graphics acceleration module 1646 is owned by a single process, a hypervisor initializes the accelerator integration circuit 1636 for an owning partition, and an operating system initializes the accelerator integration circuit 1636 for an owning process when the graphics acceleration module 1646 is allocated.
[0240] In operation, in at least one embodiment, a WD fetch unit 1691 in the accelerator integration slice 1690 fetches the next WD 1684, which includes an indication of the work to be performed by one or more graphics processing engines of the graphics acceleration module 1646. In at least one embodiment, data from the WD 1684 may be stored in registers 1645 and used by the MMU 1639, the interrupt management circuitry 1647, and / or the context management circuitry (CONTEXT MGMT) 1648, as illustrated. For example, one embodiment of the MMU 1639 includes segment / page walkthrough circuitry for accessing segment / page tables 1686 within an operating system (OS) virtual address space 1685. In at least one embodiment, interrupt management circuitry 1647 may process interrupt events 1692 received from graphics acceleration module 1646.In at least one embodiment, when performing graphics operations, an effective address 1693 generated by a graphics processing engine 1631(1)-1631(N) is translated into a real address by the MMU 1639.
[0241] In one embodiment, the same set of registers 1645 is duplicated for each graphics processing engine 1631(1)-1631(N) and / or the graphics acceleration module 1646 and may be initialized by a hypervisor or operating system. Each of these duplicated registers may be included in an accelerator integration slice 1690. Example registers that may be initialized by a hypervisor are shown in Table 1. Table 1 - Hypervisor-initialized registers register Description 1 Slice control register 2 Pointer to real address (RA) of the area of scheduled processes 3 Register for overriding authorization masks 4 Offset interrupt vector table entry 5 Limit interrupt vector table entry 6 Condition register 7 ID of the logical partition 8 Pointer to real address (RA) of the hypervisor accelerator utilization entry 9 Memory description register
[0242] Example registers that can be initialized by an operating system are shown in Table 2. Table 2 - Operating system initialized registers register Description 1 Process and thread identification 2 Pointer to effective address (EA) of the context save / restore 3 Pointer to virtual address (VA) of the accelerator utilization entry 4 Pointer to virtual address (VA) of the memory segment table 5 Authorization mask 6 Work descriptor
[0243] In at least one embodiment, each WD 1684 is specific to a particular graphics acceleration module 1646 and / or graphics processing engines 1631(1)-1631(N). In at least one embodiment, it contains all the information required by a graphics processing engine 1631(1)-1631(N) to perform work, or it may be a pointer to a memory location where an application has established a command queue for work to be completed.
[0244] Fig. 16E illustrates additional details for an exemplary embodiment of a shared model. This embodiment includes a hypervisor real address space 1698 in which a process element list 1699 is stored. The hypervisor real address space 1698 is accessible via a hypervisor 1696 that virtualizes graphics acceleration engine engines for the operating system 1695.
[0245] In at least one embodiment, shared programming models enable the use of a graphics acceleration module 1646 for all or a subset of processes from all or a subset of partitions in a system. There are two programming models in which the graphics acceleration module 1646 is shared among multiple processes and partitions: time-sliced shared and graphics-oriented shared.
[0246] In this model, the system hypervisor 1696 has the graphics acceleration module 1646 and makes its functionality available to all operating systems 1695. In order for a graphics acceleration module 1646 to support virtualization through the system hypervisor 1696, the graphics acceleration module 1646 can be used as follows: 1) An application's job request must be autonomous (ie, state does not need to be maintained between jobs), or the graphics acceleration module 1646 must provide a mechanism for saving and restoring context. 2) The graphics acceleration module 1646 guarantees that an application's job request will be completed within a specified amount of time, including any translation errors, or the graphics acceleration module 1646 provides an opportunity to preempt the processing of a job. 3) The graphics acceleration module 1646 must be guaranteed fairness between processes when operating in a targeted shared programming model.
[0247] In at least one embodiment, the application 1680 must make a system call to the operating system 1695 with a graphics acceleration module type 1646, a work descriptor (WD), an authorization mask register (AMR) value, and a context save / restore area (CSRP) pointer. In at least one embodiment, the graphics acceleration module type 1646 describes a target acceleration function for a system call. In at least one embodiment, the graphics acceleration module type 1646 may be a system-specific value.In at least one embodiment, the WD is specifically formatted for the graphics acceleration module 1646 and may be in the form of an instruction of the graphics acceleration module 1646, an effective address pointer to a user-defined structure, an effective address pointer to an instruction queue, or any other data structure for describing the work to be performed by the graphics acceleration module 1646.
[0248] In one embodiment, an AMR value is an AMR state to be used for a current process. In at least one embodiment, a value passed to an operating system is comparable to an application setting an AMR. In at least one embodiment, if the implementations of accelerator integration circuit 1636 and graphics acceleration module 1646 do not support a User Authority Mask Override Register (“UAMOR”), an operating system may apply a current UAMOR value to an AMR value before passing an AMR in a hypervisor call. Hypervisor 1696 may optionally apply a current Authority Mask Override Register (AMOR) value before placing an AMR into process element 1683.In at least one embodiment, CSRP is one of registers 1645 that contain an effective address of a region in an application's address space 1682 for the graphics acceleration module 1646 to save and restore context state. In at least one embodiment, this pointer is optional if no state needs to be saved between jobs or if a job is preempted. In at least one embodiment, the context save / restore region may serve as fixed system memory.
[0249] Upon receiving a system call, the operating system 1695 may verify that the application 1680 is registered and has been granted permission to use the graphics acceleration module 1646. In at least one embodiment, the operating system 1695 then invokes the hypervisor 1696 with the information shown in Table 3. Table 3 - Parameters for hypervisor call by operating system parameter Description 1 A work descriptor (WD) 2 An Authorization Mask Register (AMR) value (potentially masked). 3 A pointer to an effective address (EA) of the context save / restore area (CSRP) 4 A process ID (PID) and an optional thread ID (TID). 5 A pointer to a virtual address (VA) of the accelerator utilization entry (AURP) 6 Pointer to virtual address of the memory segment table (SSTP) 7 A logical interrupt service number (LISN)
[0250] In at least one embodiment, upon receiving a hypervisor call, the hypervisor 1696 verifies that the operating system 1695 has registered and is authorized to use the graphics acceleration module 1646. In at least one embodiment, the hypervisor 1696 then places the process element 1683 in a linked list of process elements for a type of corresponding graphics acceleration module 1646. In at least one embodiment, a process element may include information shown in Table 4. Table 4 - Information on process elements element Description 1 A work descriptor (WD) 2 An Authorization Mask Register (AMR) value (potentially masked) 3 A pointer to an effective address (EA) of the context save / restore area (CSRP) 4 A process ID (PID) and an optional thread ID (TID) 5 A pointer to a virtual address (VA) of the accelerator utilization entry (AURP) 6 Pointer to virtual address of the memory segment table (SSTP) 7 A logical interrupt service number (LISN) 8 Interrupt vector table derived from hypervisor call parameters 9 A status register (SR) value 10 A logical partition ID (LPID) 11 A pointer to a real address (RA) of the hypervisor accelerator utilization entry 12 Memory Descriptor Register (SDR)
[0251] In at least one embodiment, the hypervisor initializes a plurality of registers 1645 of the accelerator integration slice 1690.
[0252] As in Fig.16F, in at least one embodiment, a unified memory addressable via a common virtual memory address space is used to access physical processor memories 1601(1)-1601(N) and GPU memories 1620(1)-1620(N). In this implementation, operations performed on GPUs 1620(1)-1620(N) use the same virtual / effective memory address space to access processor memories 1601(1)-1601(N) and vice versa, thereby simplifying programmability. In at least one embodiment, a first portion of a virtual / effective address space is allocated to processor memory 1601, a second portion is allocated to second processor memory 1602, a third portion is allocated to GPU memory 1612, and so on.In at least one embodiment, this distributes an entire virtual / effective memory space (sometimes referred to as an effective address space) across each of the processor memories 1601 and GPU memories 1620 such that each processor or GPU can access each physical memory, with a virtual address mapped to that memory.
[0253] In one embodiment, bias / coherence management circuits 1694A-1694E within one or more MMUs 1639A-1639E ensure cache coherence between caches of one or more host processors (e.g., 1605) and the GPUs 1610 and implement biasing techniques that indicate physical memories in which certain types of data should be stored. In at least one embodiment, while in Fig.16F illustrates multiple instances of bias / coherence management circuits 1694A-1694E, the bias / coherence circuits may be implemented within an MMU of one or more host processors 1605 and / or within accelerator integration circuit 1636.
[0254] One embodiment enables GPU memory 1620 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology, but without incurring performance penalties associated with full system cache coherence. In at least one embodiment, the ability to access GPU-bound memory 1620 as system memory without burdensome cache coherence overhead provides a beneficial operating environment for GPU offloading. This arrangement enables host processor 1605 software to set up operands and access computation results without the overhead of traditional I / O DMA data copies. In at least one embodiment, such traditional copies involve driver calls, interrupts, and memory mapped I / O (MMIO) accesses, all of which are inefficient compared to simple memory accesses.In at least one embodiment, the ability to access GPU-bound memory 1620 without cache coherence overheads may be critical to the execution time of an offloaded computation. For example, in at least one embodiment, in cases with significant streaming memory write traffic, cache coherence overhead may significantly reduce effective write bandwidth seen by a GPU 1610. In at least one embodiment, operand facility efficiency, result access efficiency, and GPU computation efficiency may all play a role in determining the effectiveness of GPU offloading.
[0255] In at least one embodiment, the selection of GPU bias and host processor bias is controlled by a bias tracker data structure. For example, in at least one embodiment, a bias table may be used, which may be a page-granular structure (i.e., controlled to a granularity of a memory page) containing 1 or 2 bits per GPU-bound memory page. In at least one embodiment, a bias table may be implemented in a stolen memory region of one or more GPU-bound memories 1620 with or without a bias cache in a GPU 1610 (e.g., to cache frequently / recently used bias table entries). Alternatively, in at least one embodiment, an entire bias table may be maintained within a GPU.
[0256] In at least one embodiment, a bias table entry associated with each access to GPU-bound memory 1620 is accessed before actually accessing any GPU memory, which initiates the following operations. First, local requests from a GPU 1610 that find their page in GPU bias are forwarded directly to a corresponding GPU memory 1620. In at least one embodiment, local requests from a GPU that find its page in host bias are forwarded to processor 1605 (e.g., over a high-speed interconnect as described herein). In at least one embodiment, requests from processor 1605 that find a requested page in host processor bias complete a request like a normal memory read. Alternatively, requests directed to a GPU-biased page may be forwarded to GPU 1610.In at least one embodiment, a GPU may then transition a page to host processor bias if it is not currently using a page. In at least one embodiment, a page's bias state may be changed either by a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited number of cases, a purely hardware-based mechanism.
[0257] In at least one embodiment, a mechanism for changing bias state uses an API call (e.g., OpenCL), which in turn calls a GPU's device driver, which in turn sends a message to a GPU (or queues a command descriptor) instructing it to change bias state and, for some transitions, perform a cache flush operation in a host. In at least one embodiment, a cache flush operation is used for a transition from the host processor 1605 bias to the GPU bias, but not for an opposite transition.
[0258] In one embodiment, cache coherence is maintained by temporarily rendering GPU-biased pages that cannot be cached by host processor 1605. To access these pages, processor 1605 may request access from GPU 1610, which may or may not grant access immediately. Thus, to reduce communication between processor 1605 and GPU 1610, it is advantageous to ensure that GPU-biased pages are those required by a GPU but not by host processor 1605, and vice versa.
[0259] Hardware structure(s) 815 are used to carry out one or more embodiments. Details regarding hardware structure(s) 815 are described herein in connection with Fig. 8A and / or 8B provided.
[0260] Fig.Figure 17 illustrates exemplary integrated circuits and associated graphics processors that may be fabricated using one or more IP cores, according to various embodiments described herein. In addition to what is illustrated, other logic and circuitry may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0261] Fig.17 is a block diagram illustrating an exemplary system on a chip integrated circuit 1700, which may be fabricated from one or more IP cores, according to at least one embodiment. In at least one embodiment, the integrated circuit 1700 includes one or more application processors 1705 (e.g., CPUs), at least one graphics processor 1710, and may additionally include an image processor 1715 and / or a video processor 1720, each of which may be a modular IP core. In at least one embodiment, the integrated circuit 1700 includes peripheral or bus logic including a USB controller 1725, a UART controller 1730, an SPI / SDIO controller 1735, and an I 2 2S / I 22C controller 1740. In at least one embodiment, the integrated circuit 1700 may include a display device 1745 coupled to one or more of a High Definition Multimedia Interface (HDMI) controller 1750 and a Mobile Industry Processor Interface (MIPI) display interface 1755. In at least one embodiment, memory may be provided by a flash memory subsystem 1760 including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1765 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits additionally include an embedded security engine 1770.
[0262] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 815 are described herein in connection with Fig. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 in the integrated circuit 1700 may be used to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0263] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0264] Fig. 18A-18B illustrate example integrated circuits and associated graphics processors that may be fabricated using one or more IP cores, according to various embodiments described herein. In addition to what is illustrated, other logic and circuitry may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0265] Fig. 18A-18B are block diagrams illustrating example graphics processors for use within an SoC, according to embodiments described herein. Fig. 18A illustrates an exemplary graphics processor 1810 of a system on an integrated circuit chip that may be fabricated using one or more IP cores, according to at least one embodiment. Fig. 18B illustrates an additional exemplary graphics processor 1840 of a system on an integrated circuit chip, which may be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, the graphics processor 1810 is Fig. 18A is a low-power graphics processor core. In at least one embodiment, the graphics processor 1840 is Fig.18B, a more powerful graphics processor core. In at least one embodiment, each of the graphics processors 1810, 1840 may be a variant of the graphics processor 1710 of Fig. be 17.
[0266] In at least one embodiment, graphics processor 1810 includes a vertex processor 1805 and one or more fragment processors 1815A-1815N (e.g., 1815A, 1815B, 1815C, 1815D, through 1815N-1, and 1815N). In at least one embodiment, graphics processor 1810 may execute different shader programs via separate logic, such that vertex processor 1805 is optimized to perform operations for vertex shader programs, while one or more fragment processors 1815A-1815N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, vertex processor 1805 performs a vertex processing phase of a 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, fragment processor(s) 1815A-1815N use primitives and vertex data generated by vertex processor 1805 to generate a frame or fragment.To generate a frame buffer that is displayed on a display device. In at least one embodiment, the fragment processor(s) 1815A-1815N are optimized to execute fragment shader programs as provided in an OpenGL API, which can be used to perform similar operations as a pixel shader program as provided in a Direct 3D API.
[0267] In at least one embodiment, graphics processor 1810 additionally includes one or more memory management units (MMUs) 1820A-1820B, cache(s) 1825A-1825B, and circuit interconnect(s) 1830A-1830B. In at least one embodiment, one or more MMU(s) 1820A-1820B provide virtual-to-physical address mapping for graphics processor 1810, including vertex processor 1805 and / or fragment processor(s) 1815A-1815N, which may reference vertex or image / texture data stored in memory in addition to vertex or image / texture data stored in one or more cache(s) 1825A-1825B. In at least one embodiment, one or more MMU(s) 1820A-1820B may be synchronized with other MMU(s) within the system, including one or more MMU(s) associated with one or more application processor(s) 1705, image processor(s) 1715, and / or video processor(s) 1712 of Fig.17, so that each processor 1705-1712 can participate in a shared or pooled virtual memory system. In at least one embodiment, one or more circuit interconnects 1830A-1830B enable the graphics processor 1810 to interface with other IP cores within the SoC, either via an internal bus of the SoC or via a direct connection.
[0268] In at least one embodiment, the graphics processor 1840 includes one or more shader cores 1855A-1855N (e.g., 1855A, 1855B, 1855C, 1855D, 1855E, 1855F, through 1855N-1 and 1855N), as shown in Fig.18B, which provide a unified shader core architecture in which a single core or type of core can execute all types of programmable shader code, including shader code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, a number of shader cores may vary. In at least one embodiment, the graphics processor 1840 includes an inter-core task manager 1845 acting as a thread dispatcher to dispatch execution threads to one or more shader cores 1855A-1855N, and a tiling unit 1858 to accelerate tiling operations for tile-based rendering, in which rendering operations for a scene are partitioned in image space, for example, to exploit local spatial coherence within a scene or to optimize the use of internal caches.
[0269] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 815 are described herein in connection with Fig. 8A and / or 8B. In at least one embodiment, inference and / or training logic 815 in integrated circuit 18A and / or 18B may be used to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0270] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0271] Fig. 19A-19B illustrate additional exemplary graphics processor logic according to embodiments described herein. Fig. 19A illustrates a graphics core 1900 included in at least one embodiment in the graphics processor 1710 of Fig. 17 and in at least one embodiment, a unified shader core 1855A-1855N as in Fig. 18B in at least one embodiment. Fig.19B illustrates a highly parallel general purpose graphics processing unit (“GPGPU”) 1930 suitable for use on a multi-chip module in at least one embodiment.
[0272] In at least one embodiment, the graphics core 1900 includes a shared instruction cache 1902, a texture unit 1918, and a cache / shared memory 1920 that are common to execution resources within the graphics core 1900. In at least one embodiment, the graphics core 1900 may include multiple slices 1901A-1901N or partitions for each core, and a graphics processor may include multiple instances of the graphics core 1900. In at least one embodiment, the slices 1901A-1901N may include support logic including a local instruction cache 1904A-1904N, a thread scheduler 1906A-1906N, a thread dispatcher 1908A-1908N, and a set of registers 1910A-1910N.In at least one embodiment, slices 1901A-1901N may include a set of additional functional units (AFUs 1912A-1912N), floating point units (FPU 1914A-1914N), integer arithmetic logic units (ALUs 1916A-1916N), address calculation units (ACU 1913A-1913N), double precision floating point units (DPFPU 1915A-1915N), and matrix processing units (MPU 1917A-1917N).
[0273] In at least one embodiment, FPUs 1914A-1914N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while DPFPUs 1915A-1915N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, ALUs 1916A-1916N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. In at least one embodiment, MPUs 1917A-1917N can also be configured for mixed-precision matrix operations, including half-precision floating-point and 8-bit integer operations. In at least one embodiment, the MPUs 1917-1917N may perform a variety of matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general purpose orGeneral matrix-to-matrix multiplication (GEMM). In at least one embodiment, AFUs 1912A-1912N may perform additional logical operations not supported by floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).
[0274] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 815 are described herein in connection with Fig.8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 in the graphics core 1900 may be used to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0275] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0276] Fig.19B illustrates a general-purpose processing unit (GPGPU) 1930 that may be configured to enable highly parallel computational operations to be performed by an array of graphics processing units, in at least one embodiment. In at least one embodiment, the GPGPU 1930 may be directly linked to other instances of the GPGPU 1930 to create a multi-GPU cluster to improve training speed for deep neural networks. In at least one embodiment, the GPGPU 1930 includes a host interface 1932 to enable connection to a host processor. In at least one embodiment, the host interface 1932 is a PCI Express interface. In at least one embodiment, the host interface 1932 may be a vendor-specific communication interface or communication fabric.In at least one embodiment, GPGPU 1930 receives instructions from a host processor and uses a global scheduler 1934 to distribute the execution threads associated with those instructions to a set of compute clusters 1936A-1936H. In at least one embodiment, compute clusters 1936A-1936H share a cache 1938. In at least one embodiment, cache 1938 may serve as a master cache for caches within compute clusters 1936A-1936H.
[0277] In at least one embodiment, GPGPU 1930 includes memory 1944A-1944B coupled to compute clusters 1936A-1936H via a set of memory controllers 1942A-1942B. In at least one embodiment, memory 1944A-1944B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including double data rate (GDDR) graphics memory.
[0278] In at least one embodiment, the compute clusters 1936A-1936H each include a set of graphics cores, such as the graphics core 1900 of Fig.19A, which may include multiple types of integer and floating-point logic units capable of performing computational operations at a range of precision levels, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of floating-point units in each of compute clusters 1936A-1936H may be configured to perform 16-bit or 32-bit floating-point operations, while another subset of floating-point units may be configured to perform 64-bit floating-point operations.
[0279] In at least one embodiment, multiple instances of the GPGPU 1930 may be configured to operate as a compute cluster. In at least one embodiment, the communication used by the compute clusters 1936A-1936H for synchronization and data exchange varies depending on the embodiment. In at least one embodiment, multiple instances of the GPGPU 1930 communicate via the host interface 1932. In at least one embodiment, the GPGPU 1930 includes an I / O hub 1939 that couples the GPGPU 1930 to a GPU interconnect 1940 that enables direct connection to other instances of the GPGPU 1930. In at least one embodiment, the GPU interconnect 1940 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple GPGPU 1930 instances.In at least one embodiment, GPU interconnect 1940 couples to a high-speed interconnect to send and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 1930 reside in separate computing systems and communicate via a network device accessible via host interface 1932. In at least one embodiment, GPU interconnect 1940 may be configured to enable connection to a host processor in addition to, or alternatively to, host interface 1932.
[0280] In at least one embodiment, the GPGPU 1930 may be configured to train neural networks. In at least one embodiment, the GPGPU 1930 may be used within an inference platform. In at least one embodiment where the GPGPU 1930 is used for inference, the GPGPU may include fewer compute clusters 1936A-1936H than when the GPGPU 1930 is used to train a neural network. In at least one embodiment, the memory technology associated with the memory 1944A-1944B may differ between inference and training configurations, with higher bandwidth memory technologies provided for training configurations. In at least one embodiment, the inference configuration of the GPGPU 1930 may support inference-specific instructions.For example, in at least one embodiment, an inference configuration may support one or more 8-bit integer dot product instructions that may be used during inference operations for deployed neural networks.
[0281] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 815 are described herein in connection with Fig.8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 in the GPGPU 1930 may be used to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0282] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0283] Fig.20 is a block diagram illustrating a computer system 2000 according to at least one embodiment. In at least one embodiment, the computer system 2000 includes a processing subsystem 2001 having one or more processors 2002 and a system memory 2004 communicating via an interconnect path that may include a memory hub 2005. In at least one embodiment, the memory hub 2005 may be a separate component within a chipset component or integrated with one or more processors 2002. In at least one embodiment, the memory hub 2005 couples to an I / O subsystem 2011 via a communications link 2006. In at least one embodiment, the I / O subsystem 2011 includes an I / O hub 2007 that may enable the computer system 2000 to receive input from one or more input devices 2008.In at least one embodiment, the I / O hub 2007 may enable a display controller, which may be included in one or more processors 2002, to provide outputs to one or more display devices 2010A. In at least one embodiment, one or more display devices 2010A coupled to the I / O hub 2007 may comprise a local, internal, or embedded display device.
[0284] In at least one embodiment, the processing subsystem 2001 includes one or more parallel processors 2012 coupled to the storage hub 2005 via a bus or other communication link 2013. In at least one embodiment, the communication link 2013 may be one of any number of standards-based communication interconnect technologies or protocols, such as, but not limited to, PCI Express, or may be a vendor-specific communication interface or communication fabric. In at least one embodiment, one or more parallel processors 2012 form a computationally focused parallel or vector processing system, which may include a large number of processing cores and / or processing clusters, such as a Many Integrated Core (MIC) processor.In at least one embodiment, one or more parallel processors 2012 form a graphics processing subsystem that can output pixels to one or more display devices 2010A coupled via the I / O hub 2007. In at least one embodiment, one or more parallel processors 2012 may also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 2010B.
[0285] In at least one embodiment, a system storage unit 2014 may connect to the I / O hub 2007 to provide a storage mechanism for the computer system 2000. In at least one embodiment, an I / O switch 2016 may be used to provide an interface mechanism to enable connections between the I / O hub 2007 and other components, such as a network adapter 2018 and / or a wireless network adapter 2019 that may be integrated into the platform, and various other devices that may be added via one or more add-in devices 2012. In at least one embodiment, the network adapter 2018 may be an Ethernet adapter or other wired network adapter.In at least one embodiment, the wireless network adapter 2019 may include one or more Wi-Fi, Bluetooth, near field communication (NFC), or other network devices that include one or more wireless radios.
[0286] In at least one embodiment, computer system 2000 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, and the like, which may also be connected to I / O hub 2007. In at least one embodiment, communication paths connecting various components in Fig.20 interconnect using any suitable protocols, such as PCI (Peripheral Component Interconnect)-based protocols (e.g., PCI Express), or other bus or point-to-point communication interfaces and / or protocols, such as NV-Link high-speed links or interconnect protocols.
[0287] In at least one embodiment, one or more parallel processors 2012 include circuitry optimized for graphics and video processing, including, for example, video output circuitry and forming a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 2012 include circuitry optimized for general processing. In at least one embodiment, components of computer system 2000 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 2012, memory hub 2005, processor(s) 2002, and I / O hub 2007 may be integrated into a system-on-chip (SoC) integrated circuit.In at least one embodiment, components of computer system 2000 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of components of computer system 2000 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules to form a modular computer system.
[0288] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 815 are described herein in connection with Fig.8A and / or 8B. In at least one embodiment, inference and / or training logic 815 in system 2000 may be used to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0289] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein. PROCESSORS
[0290] Fig. 21A illustrates a parallel processor 2100 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 2100 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 2100 is a variant of one or more in Fig. 20 shown parallel processors 2012 according to an exemplary embodiment.
[0291] In at least one embodiment, parallel processor 2100 includes a parallel processing unit 2102. In at least one embodiment, parallel processing unit 2102 includes an I / O unit 2104 that enables communication with other devices, including other instances of parallel processing unit 2102. In at least one embodiment, I / O unit 2104 may be directly connected to other devices. In at least one embodiment, I / O unit 2104 connects to other devices using a hub or switch interface, such as storage hub 2105. In at least one embodiment, connections between storage hub 2105 and I / O unit 2104 form a communication link 2113.In at least one embodiment, the I / O unit 2104 connects to a host interface 2106 and a memory cross-bus 2116, where the host interface 2106 receives commands intended to perform processing operations and the memory cross-bus 2116 receives commands intended to perform memory operations.
[0292] In at least one embodiment, when the host interface 2106 receives a command buffer via the I / O unit 2104, the host interface 2106 may instruct work operations to execute those commands at a frontend 2108. In at least one embodiment, the frontend 2108 couples to a scheduler 2110 configured to schedule commands or other work items to a
[0293] Processing cluster array 2112. In at least one embodiment, scheduler 2110 ensures that cluster array 2112 is properly configured and in a valid state before dispatching tasks to processing cluster array 2112 of processing cluster array 2112. In at least one embodiment, scheduler 2110 is implemented via firmware logic executing on a microcontroller. In at least one embodiment, microcontroller-implemented scheduler 2110 is configurable to perform complex scheduling and work dispatch operations at coarse and fine granularity, enabling rapid anticipation and context switching of threads executing on processing array 2112.In at least one embodiment, the host software may allocate workloads for scheduling on the processing array 2112 across one of multiple graphics processing doorbells. In at least one embodiment, workloads may then be automatically distributed across the processing array 2112 by the logic of the scheduler 2110 within a microcontroller including the scheduler 2110.
[0294] In at least one embodiment, the processing cluster arrangement 2112 may include up to "N" processing clusters (e.g., cluster 2114A, cluster 2114B, through cluster 2114N).
[0295] In at least one embodiment, each cluster 2114A-2114N of processing cluster array 2112 may execute a large number of concurrent threads. In at least one embodiment, scheduler 2110 may allocate work to clusters 2114A-2114N of processing cluster array 2112 using various scheduling and / or work distribution algorithms, which may vary depending on the workload associated with each type of program or computation. In at least one embodiment, scheduling may be handled dynamically by scheduler 2110 or may be partially assisted by compiler logic during compilation of the program logic configured for execution by processing cluster array 2112.In at least one embodiment, different clusters 2114A-2114N of the processing cluster arrangement 2112 may be allocated to process different types of programs or to perform different types of computations.
[0296] In at least one embodiment, processing cluster assembly 2112 may be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster assembly 2112 is configured to perform general parallel computing operations. For example, in at least one embodiment, processing cluster assembly 2112 may include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations, including physical operations, and performing data transformations.
[0297] In at least one embodiment, the processing cluster assembly 2112 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster assembly 2112 may include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, the processing cluster assembly 2112 may be configured to execute graphics processing-related shader programs, such as vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2102 may transfer data from system memory via the I / O unit 2104 for processing.In at least one embodiment, data transferred during processing may be stored in on-chip memory (e.g., memory of parallel processor 2122) during processing and subsequently written back to system memory.
[0298] In at least one embodiment, when the parallel processing unit 2102 is used to perform graphics processing, the scheduler 2110 may be configured to divide a processing workload into approximately equal-sized tasks to better facilitate the distribution of graphics processing operations across multiple clusters 2114A-2114N of the processing cluster assembly 2112. In at least one embodiment, portions of the processing cluster assembly 2112 may be configured to perform different types of processing. For example, in at least one embodiment, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen-space operations to generate a rendered image for display.In at least one embodiment, intermediate data generated by one or more of clusters 2114A-2114N may be stored in buffers so that intermediate data may be transferred between clusters 2114A-2114N for further processing.
[0299] In at least one embodiment, the processing cluster arrangement 2112 may receive processing tasks to be executed via the scheduler 2110, which receives commands from the front end 2108 that define processing tasks. In at least one embodiment, processing tasks may include indices of data to be processed, e.g., surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands that define how data is to be processed (e.g., which program is to be executed). In at least one embodiment, the scheduler 2110 may be configured to fetch indices corresponding to tasks or may receive indices from the front end 2108. In at least one embodiment, the front end 2108 may be configured to ensure that the processing cluster arrangement 2112 is configured to a valid state before a command buffer specified by incoming command buffers (e.g.,stack buffer, push buffer, etc.) specified workload is initiated.
[0300] In at least one embodiment, each of one or more instances of parallel processing unit 2102 may be coupled to parallel processor memory 2122. In at least one embodiment, parallel processor memory 2122 may be accessed via memory cross-bus 2116, which may receive memory requests from processing cluster arrangement 2112 as well as I / O unit 2104. In at least one embodiment, memory cross-bus 2116 may access parallel processor memory 2122 via a memory interface 2118. In at least one embodiment, memory interface 2118 may include multiple partitioning units (e.g., partitioning unit 2120A, partitioning unit 2120B through partitioning unit 2120N), each of which may couple to a portion (e.g., the memory unit) of parallel processor memory 2122.In at least one embodiment, a number of partitioning units 2120A-2122N is configured to be equal to a number of storage units, such that a first partitioning unit 2120A has a corresponding first storage unit 2124A, a second partitioning unit 2120B has a corresponding storage unit 2124B, and an Nth partitioning unit 2120N has a corresponding Nth storage unit 2124N. In at least one embodiment, a number of partitioning units 2120A-2120N may not be equal to a number of storage devices.
[0301] In at least one embodiment, memory units 2124A-2124N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including double data rate graphics memory (GDDR). In at least one embodiment, memory units 2124A-2124N may also include 3D stack memories, including, but not limited to, high bandwidth memories (HBM). In at least one embodiment, render targets, such as frame buffers or texture maps, may be stored across memory units 2124A-2124N so that partition units 2120A-2120N can write portions of each render target in parallel to efficiently utilize available bandwidth of parallel processor memory 2122.In at least one embodiment, a local instance of parallel processor memory 2122 may be eliminated in favor of a unified memory design that utilizes system memory in conjunction with local cache memory.
[0302] In at least one embodiment, any of the clusters 2114A-2114N of the processing cluster arrangement 2112 may process data written to any of the storage units 2124A-2124N in the parallel processor memory 2122. In at least one embodiment, the memory cross-bus 2116 may be configured to transfer an output of each cluster 2114A-2114N to any partition unit 2112A-2112N or to another cluster 2114A-2114N that may perform additional processing operations on an output. In at least one embodiment, each cluster 2114A-2114N may communicate with the memory interface 2118 via the memory cross-bus 2116 to read from or write to various external storage devices.In at least one embodiment, the memory cross-bus 2116 includes a connection to the memory interface 2118 for communication with the I / O unit 2104, as well as a connection to a local instance of the parallel processor memory 2122, so that processing units within different processing clusters 2114A-2114N can communicate with system memory or other memory that is not local to the parallel processing unit 2102. In at least one embodiment, the memory cross-bus 2116 can use virtual channels to separate streams of data traffic between the clusters 2114A-2114N and the partitioning units 2120A-2120N.
[0303] In at least one embodiment, multiple instances of the parallel processing unit 2102 may be provided on a single expansion card, or multiple expansion cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 2102 may be configured to interoperate with each other even if different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of the parallel processing unit 2102 may include higher-precision floating-point units relative to other instances.In at least one embodiment, systems including one or more instances of the parallel processing unit 2102 or the parallel processor 2100 may be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.
[0304] Fig. 21B is a block diagram of a partitioning unit 2120 according to at least one embodiment. In at least one embodiment, the partitioning unit 2120 is an instance of one of the partitioning units 2120A-2120N of Fig.21A. In at least one embodiment, partitioning unit 2120 includes an L2 cache 2121, a frame buffer interface 2125, and a raster operation unit (ROP) 2126. L2 cache 2121 is a read / write cache configured to perform load and store operations received from memory cross-bus 2116 and ROP 2126. In at least one embodiment, read misses and urgent write-back requests are issued from L2 cache 2121 to frame buffer interface 2125 for processing. In at least one embodiment, updates may also be sent to a frame buffer via frame buffer interface 2125 for processing. In at least one embodiment, frame buffer interface 2125 is coupled to one of the memory units in the parallel processor memory, such as memory units 2124A-2124N of Fig.21 (e.g. within the parallel processor memory 2122).
[0305] In at least one embodiment, ROP 2126 is a processing unit that performs raster operations such as stenciling, Z-testing, blending, and the like. In at least one embodiment, ROP 2126 then outputs processed graphics data stored in graphics memory. In at least one embodiment, ROP 2126 includes compression logic for compressing depth or color data written to memory and for decompressing depth or color data read from memory. In at least one embodiment, the compression logic may be lossless compression logic using one or more of multiple compression algorithms. The type of compression performed by ROP 2126 may vary based on statistical characteristics of the data to be compressed.For example, in at least one embodiment, delta color compression is performed on depth and color data on a tile-by-tile basis.
[0306] In at least one embodiment, the ROP 2126 is in each processing cluster (e.g., clusters 2114A-2114N of Fig. 21A) instead of in the partitioning unit 2120. In at least one embodiment, read and write requests for pixel data are transmitted over the memory cross-bus 2116 instead of pixel fragment data. In at least one embodiment, processed graphics data may be displayed on a display device, such as one or more display devices 2110 of Fig. 20, forwarded for further processing by the processor(s) 2002, or for further processing by one of the processing entities within the parallel processor 2100 from Fig. 21A.
[0307] Fig. 21C is a block diagram of a processing cluster 2114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is an instance of one of the processing clusters 2114A-2114N of Fig.21A. In at least one embodiment, the processing cluster 2114 may be configured to execute many threads in parallel, where "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single instruction, multiple data (SIMD) instruction issue techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction, multiple thread (SIMT) techniques are used to support the parallel execution of a large number of generally synchronized threads using a common instruction unit configured to issue instructions to a number of processing engines within each of the processing clusters.
[0308] In at least one embodiment, the operation of the processing cluster 2114 may be controlled by a pipeline manager 2132 that dispatches processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 2132 receives instructions from the scheduler 2110 of Fig.21A and manages the execution of these instructions via a graphics multiprocessor 2134 and / or a texture unit 2136. In at least one embodiment, the graphics multiprocessor 2134 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors of different architectures may be included within the processing cluster 2114. In at least one embodiment, one or more instances of the graphics multiprocessor 2134 may be included in a processing cluster 2114. In at least one embodiment, the graphics multiprocessor 2134 may process data and may use a data cross-rail 2140 to distribute processed data to one of several possible destinations, including other shader units.In at least one embodiment, the pipeline manager 2132 may facilitate the distribution of the processed data by specifying destinations for processed data to be distributed across the data cross-rail 2140.
[0309] In at least one embodiment, each graphics multiprocessor 2134 within the processing cluster 2114 may include an identical set of functional execution logic (e.g., arithmetic logic units, load-store units, etc.). In at least one embodiment, functional execution logic may be configured in a pipelined manner, in which new instructions may be issued before previous instructions are completed. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifting, and the computation of various algebraic functions. In at least one embodiment, the same functional unit hardware may be leveraged to perform various operations, and any combination of functional units may be present.
[0310] In at least one embodiment, instructions transferred to processing cluster 2114 constitute a thread. In at least one embodiment, a set of threads executing across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program on different input data. In at least one embodiment, each thread within a thread group may be associated with a different processing engine within a graphics multiprocessor 2134. In at least one embodiment, a thread group may include fewer threads than a number of processing engines within the graphics multiprocessor 2134.In at least one embodiment, when a thread group includes fewer threads than a number of processing engines, one or more of the processing engines may be idle during the cycles in which that thread group is processing. In at least one embodiment, a thread group may include more threads than a number of processing engines within the graphics multiprocessor 2134. In at least one embodiment, when a thread group includes more threads than a number of processing engines within the graphics multiprocessor 2134, processing may be performed over consecutive clock cycles. In at least one embodiment, multiple thread groups may execute concurrently on a graphics multiprocessor 2134.
[0311] In at least one embodiment, the graphics multiprocessor 2134 includes an internal cache to perform load and store operations. In at least one embodiment, the graphics multiprocessor 2134 may forgo an internal cache and utilize a cache (e.g., the L1 cache 2148) within the processing cluster 2114. In at least one embodiment, each graphics multiprocessor 2134 also has access to L2 caches within partition units (e.g., the partition units 2120A-2120N of Fig.21A), which are shared among all processing clusters 2114 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 2134 can also access off-chip global memory, which can include one or more parallel processor local memories and / or system memories. In at least one embodiment, any memory external to the parallel processing unit 2102 can be used as global memory. In at least one embodiment, the processing cluster 2114 includes multiple instances of the graphics multiprocessor 2134, which can share common instructions and data, which can be stored in the L1 cache 2148.
[0312] In at least one embodiment, each processing cluster 2114 may include a memory management unit (MMU) 2145 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2145 may reside within the memory interface 2118 of Fig.21A. In at least one embodiment, the MMU 2145 includes a set of page table entries (PTEs) used to map a virtual address to a physical address of a tile and optionally a cache line index. In at least one embodiment, the MMU 2145 may include address translation lookaside buffers (TLBs) or caches that may be located in the graphics multiprocessor 2134 or in the L1 cache 2148 or in the processing cluster 2114. In at least one embodiment, a physical address is processed to distribute surface data access locally to enable efficient interleaving of requests between partitioning units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or a miss.
[0313] In at least one embodiment, a processing cluster 2114 may be configured such that each graphics multiprocessor 2134 is coupled to a texture unit 2136 for performing texture mapping operations, e.g., determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 2134 and fetched from an L2 cache, local parallel processor memory, or system memory as needed.In at least one embodiment, each graphics multiprocessor 2134 issues processed tasks to the data cross-bus 2140 to provide the processed task to another processing cluster 2114 for further processing or to store the processed task in an L2 cache, local parallel processor memory, or system memory via the memory cross-bus 2116. In at least one embodiment, a pre-raster operations unit (preROP) 2142 is configured to receive data from the graphics multiprocessor 2134 and direct data to ROP units, which may be arranged with partitioning units as described herein (e.g., partitioning units 2120A-2120N of FIG. Fig. 21A). In at least one embodiment, the PreROP unit 2142 may perform color mixing optimizations to organize pixel color data and perform address translations.
[0314] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described herein in connection with Fig. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 in the graphics processing cluster 2114 may be used to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0315] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0316] Fig.21D illustrates a graphics multiprocessor 2134 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 2134 couples to the pipeline manager 2132 of the processing cluster 2114. In at least one embodiment, the graphics multiprocessor 2134 has an execution pipeline that includes, but is not limited to, an instruction cache 2152, an instruction unit 2154, an address mapping unit 2156, a register file 2158, one or more general-purpose graphics processing unit GPGPU cores 2162, and one or more load / store units 2166. The GPGPU cores 2162 and the load / store units 2166 are coupled to the cache memory 2172 and the shared memory 2170 via a memory and cache interconnect 2168.
[0317] In at least one embodiment, instruction cache 2152 receives a stream of instructions to be executed by pipeline manager 2132. In at least one embodiment, instructions are cached in instruction cache 2152 and provided for execution by instruction unit 2154. In at least one embodiment, instruction unit 2154 may dispatch instructions as thread groups (e.g., warps), with each thread of the thread group associated with a different execution unit within GPGPU core 2162. In at least one embodiment, an instruction may access any of a local, shared, or global address space by specifying an address within a unified address space.In at least one embodiment, the address mapping unit 2156 may be used to translate addresses in a unified address space into a unique memory address accessible by the load / store units 2166.
[0318] In at least one embodiment, register file 2158 provides a set of registers for functional units of graphics multiprocessor 2134. In at least one embodiment, register file 2158 provides temporary storage for operands associated with data paths of functional units (e.g., GPGPU cores 2162, load / store units 2166) of graphics multiprocessor 2134. In at least one embodiment, register file 2158 is partitioned among each of the functional units such that each functional unit is assigned a dedicated portion of register file 2158. In at least one embodiment, register file 2158 is partitioned between different chains or warps executed by graphics multiprocessor 2134.
[0319] In at least one embodiment, GPGPU cores 2162 may each include floating-point units (FPUs) and / or integer arithmetic logic units (ALUs) used to execute instructions of graphics multiprocessor 2134. GPGPU cores 2162 may be similar in architecture or different in architecture. In at least one embodiment, a first portion of GPGPU cores 2162 includes a single-precision FPU and an integer ALU, while a second portion of GPGPU cores includes a double-precision FPU. In at least one embodiment, FPUs may implement the IEEE 754-1208 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic.In at least one embodiment, graphics multiprocessor 2134 may additionally include one or more fixed-function or special-function units for performing specific functions, such as copy rectangle or pixel blending operations. In at least one embodiment, one or more GPGPU cores 2162 may also include logic for a fixed or special function.
[0320] In at least one embodiment, GPGPU cores 2162 include SIMD logic capable of executing a single instruction on multiple data sets. In at least one embodiment, GPGPU cores 2162 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for GPGPU cores can be generated at compile time by a shader compiler or automatically generated upon execution of programs written and compiled for Single Program, Multiple Data (SPMD) or SIMT architectures. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can execute via a single SIMD instruction.For example, in at least one embodiment, eight SIMT threads performing the same or similar operations may be executed in parallel via a single SIMD8 logic unit.
[0321] In at least one embodiment, the memory and cache interconnect 2168 is an interconnect network that connects each functional unit of the graphics multiprocessor 2134 to the register file 2158 and the shared memory 2170. In at least one embodiment, the memory and cache interconnect 2168 is a cross-rail interconnect that enables the load / store unit 2166 to implement load and store operations between the shared memory 2170 and the register file 2158. In at least one embodiment, the register file 2158 may operate at the same frequency as the GPGPU cores 2162, such that data transfer between the GPGPU cores 2162 and the register file 2158 has very low latency.In at least one embodiment, shared memory 2170 may be used to enable communication between threads executing on functional units within graphics multiprocessor 2134. In at least one embodiment, cache 2172 may be used, for example, as a data cache to cache texture data exchanged between functional units and texture unit 2136. In at least one embodiment, shared memory 2170 may also be used as a program-managed cache. In at least one embodiment, threads executing on GPGPU cores 2162 may programmatically store data within shared memory in addition to automatically cached data stored within cache 2172.
[0322] In at least one embodiment, a parallel processor or GPGPU, as described herein, is communicatively coupled to host / processor cores to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host / processor cores via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated into or on the same package or die as the cores and communicatively coupled to cores via an internal processor bus / interconnect (e.g., internal to the package or die).In at least one embodiment, processor cores, regardless of how the GPU is connected, may allocate work to the GPU in the form of sequences of commands / instructions encompassed in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0323] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described herein in connection with Fig.8A and / or 8B. In at least one embodiment, inference and / or training logic 815 in graphics multiprocessor 2134 may be used to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0324] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0325] Fig.22 illustrates a multi-GPU computer system 2200 according to at least one embodiment. In at least one embodiment, the multi-GPU computer system 2200 may include a processor 2202 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 2206A-D via a host interface switch 2204. In at least one embodiment, the host interface switch 2204 is a PCI Express switch device that couples the processor 2202 to a PCI Express bus over which the processor 2202 can communicate with the GPGPUs 2206A-D. In at least one embodiment, the GPGPUs 2206A-D can interconnect via a number of high-speed point-to-point GPU-to-GPU links 2216. In at least one embodiment, the P2P connections 2216 connect to each of the GPGPUs 2206A-D via a dedicated GPU connection.In at least one embodiment, the P2P GPU connections 2216 enable direct communication between each of the GPGPUs 2206A-D without requiring communication over the host interface bus 2204 to which the processor 2202 is connected. In at least one embodiment, with GPU-to-GPU traffic directed to the P2P GPU connections 2216, the host interface bus 2204 remains available for system memory access or for communication with other instances of the multi-GPU computer system 2200, for example, over one or more network devices. While in at least one embodiment, the GPGPUs 2206A-D connect to the processor 2202 via the host interface switch 2204, in at least one embodiment, the processor 2202 includes direct support for the P2P GPU connections 2216 and can be connected directly to the GPGPUs 2206A-D.
[0326] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described herein in connection with Fig. 8A and / or 8B. In at least one embodiment, inference and / or training logic 815 may be used in multi-GPU computing system 2200 to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0327] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0328] Fig.Figure 31 is a block diagram of a graphics processor 2300 according to at least one embodiment. In at least one embodiment, the graphics processor 2300 includes a ring interconnect 2302, a pipelined front end 2304, a media engine 2337, and graphics cores 2380A-2380N. In at least one embodiment, the ring interconnect 2302 couples the graphics processor 2300 to other processing units, including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, the graphics processor 2300 is one of many processors integrated within a multicore processing system.
[0329] In at least one embodiment, graphics processor 2300 receives batches of instructions via ring interconnect 2302. In at least one embodiment, incoming instructions are interpreted by an instruction streamer 2303 in pipeline front end 2304. In at least one embodiment, graphics processor 2300 includes scalable execution logic for performing 3D geometry processing and media processing via graphics core(s) 2380A-2380N. In at least one embodiment, instruction streamer 2303 provides instructions to geometry pipeline 2336 for 3D geometry processing instructions. In at least one embodiment, instruction streamer 2303 provides instructions to a video front end 2334 coupled to a media engine 2337 for at least some media processing instructions.In at least one embodiment, the media engine 2337 includes a video quality engine (VQE) 2330 for video and image post-processing and a multi-format encoding / decoding engine (MFX) 2333 for hardware-accelerated media data encoding and decoding. In at least one embodiment, the geometry pipeline 2336 and the media engine 2337 each generate execution threads for thread execution resources provided by at least one graphics core 2380A.
[0330] In at least one embodiment, graphics processor 2300 includes scalable threaded execution resources with modular cores 2380A-2380N (sometimes referred to as core slices), each having a plurality of sub-cores 2350A-2350N, 2360A-2360N (sometimes referred to as core sub-slices). In at least one embodiment, graphics processor 2300 may include any number of graphics cores 2380A-2380N. In at least one embodiment, graphics processor 2300 includes a graphics core 2380A having at least a first sub-core 2350A and a second sub-core 2360A. In at least one embodiment, graphics processor 2300 is a low-power processor with a single sub-core (e.g., 2350A). In at least one embodiment, the graphics processor 2300 includes a plurality of graphics cores 2380A-2380N, each including a set of first sub-cores 2350A-2350N and a set of second sub-cores 2360A-2360N.In at least one embodiment, each subcore in the first subcores 2350A-2350N includes at least a first set of execution units 2352A-2352N and media / texture samplers 2354A-2354N. In at least one embodiment, each subcore in the second subcores 2360A-2360N includes at least a second set of execution units 2362A-2362N and samplers 2364A-2364N. In at least one embodiment, each subcore 2350A-2350N, 2360A-2360N shares a set of common resources 2370A-2370N. In at least one embodiment, shared resources include shared cache memory and pixel operation logic.
[0331] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described herein in connection with Fig.8A and / or 8B. In at least one embodiment, inference and / or training logic 815 in graphics processor 2300 may be used to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0332] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0333] Fig.24 shows a processor 2400 that may include logic circuitry for executing instructions, according to at least one embodiment. In at least one embodiment, the processor 2400 may execute instructions, including x86 instructions, ARM instructions, special instructions for application-specific integrated circuits (ASICs), etc. In at least one embodiment, the processor 2400 may include registers for storing packed data, such as 64-bit wide MMX™ registers in microprocessors equipped with MMX technology from Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers, available in both integer and floating-point forms, may operate on packed data elements accompanying single-instruction multiple data ("SIMD") and streaming SIMD extensions ("SSE").In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or beyond (commonly referred to as "SSEx") may include such packed data operands. In at least one embodiment, the processor 2400 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inferencing.
[0334] In at least one embodiment, the processor 2400 includes a front end ("front end") 2401 to fetch instructions to be executed and to prepare instructions to be used later in the processor pipeline. In at least one embodiment, the front end 2401 may include multiple units. In at least one embodiment, an instruction prefetcher 2426 fetches instructions from memory and passes instructions to an instruction decoder 2428, which in turn decodes or interprets instructions. For example, in at least one embodiment, the instruction decoder 2428 decodes a received instruction into one or more operations, referred to as "micro-instructions" or "micro-operations" (also referred to as "micro-ops" or "uops"), that a machine may perform. In at least one embodiment, the front end parses orInstruction decoder 2428 parses an instruction into an opcode and corresponding data and control fields that can be used by the microarchitecture to perform operations according to at least one embodiment. In at least one embodiment, a trace cache 2430 may assemble decoded uops into program-ordered sequences or traces in a uop queue 2434 for execution. In at least one embodiment, when trace cache 2430 encounters a complex instruction, a microcode ROM 2432 provides the uops required to complete the operation.
[0335] In at least one embodiment, some instructions may be converted into a single micro-op, while others may require multiple micro-operations to complete the full operation. In at least one embodiment, if more than four micro-ops are required to complete an instruction, instruction decoder 2421 may access microcode ROM 2432 to execute the instruction. In at least one embodiment, an instruction may be decoded into a small number of micro-ops for processing at instruction decoder 2421. In at least one embodiment, an instruction may be stored in microcode ROM 2432 if a number of micro-operations are required to perform the operation.In at least one embodiment, trace cache 2430 refers to a programmable entry point logic array ("PLA") for determining a correct microinstruction pointer for reading microcode sequences to complete one or more instructions from microcode ROM 2432, according to at least one embodiment. In at least one embodiment, microcode ROM 2432 completes sequencing micro-ops for an instruction, and machine front-end 2401 may resume fetching micro-ops from trace cache 2430.
[0336] In at least one embodiment, an out-of-order execution engine ("out-of-order engine") 2403 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic includes a number of buffers to smooth and reorder the flow of instructions to optimize performance as they proceed down a pipeline and are scheduled for execution. The out-of-order execution engine 2403 includes, but is not limited to, an allocator / register renamer 2440, a memory uop queue 2442, an integer / floating point uop queue 2444, a memory scheduler 2446, a fast scheduler 2402, a slow / general purpose floating point scheduler (“slow / general purpose FP scheduler”) 2404, and a simple floating point scheduler (“simple FP scheduler”) 2406.In at least one embodiment, the fast scheduler 2402, the slow / universal floating-point scheduler 2404, and the simple floating-point scheduler 3126 are also collectively referred to herein as "uop schedulers 2402, 2404, 2406." In at least one embodiment, the allocator / register renamer 2440 allocates engine buffers and resources required by each uop to execute. In at least one embodiment, the allocator / register renamer 2440 renames logic registers to entries in a register file. In at least one embodiment, the allocator / register renamer 2440 also allocates an entry for each uop in one of two uop queues, the memory uop queue 2442 for memory operations and the integer / floating point uop queue 2444 for non-memory operations, prior to the memory scheduler 2446 and the uop schedulers 2402, 2404, 2406.In at least one embodiment, the uop schedulers 2402, 2404, 2406 determine when a uop is ready to execute based on the readiness of their dependent input register operand sources and the availability of execution resources that uops require to complete their operation. In at least one embodiment, the fast scheduler 2402 may schedule on each half of a main clock cycle, while the slow / universal floating-point scheduler 2404 and the simple floating-point scheduler 2406 may schedule once per main processor clock cycle. In at least one embodiment, the uop schedulers 2402, 2404, 2406 arbitrate for transmit ports to schedule uops for execution.
[0337] In at least one embodiment, an execution block 2411 includes, but is not limited to, an integer register file / bypass network 2408, a floating-point register file / bypass network ("FP register file / bypass network") 2410, address generation units ("AGUs") 2412 and 2414, fast arithmetic logic units (ALUs) ("fast ALUs") 2416 and 2418, a slow arithmetic logic unit ("slow ALU") 2412, a floating-point ALU ("FP") 2422, and a floating-point move unit ("FP move") 2424. In at least one embodiment, the integer register file / bypass network 2408 and the floating-point register file / bypass network 2410 are also referred to herein as "register files 2408, 2410." designated.In at least one embodiment, the AGUSs 2412 and 2414, the fast ALUs 2416 and 2418, the slow ALU 2412, the floating-point ALU 2422, and the floating-point move unit 2424 are also referred to as "execution units 2412, 2414, 2416, 2418, 2412, 2422, and 2424." In at least one embodiment, the execution block 2411 may include, but is not limited to, any number (including zero) and type of register files, bypass networks, address generation units, and execution units in any combination.
[0338] In at least one embodiment, register files 2408, 2410 may be located between uop schedulers 2402, 2404, 2406 and execution units 2412, 2414, 2416, 2418, 2412, 2422, and 2424. In at least one embodiment, integer register file / bypass network 2408 performs integer operations. In at least one embodiment, floating-point register file / bypass network 2410 performs floating-point operations. In at least one embodiment, each of register networks 2408, 2410 may include, but is not limited to, a bypass network that may bypass just-completed results that have not yet been written to the register file or forward them to new dependent uops. In at least one embodiment, register files 2408, 2410 may communicate data with each other.In at least one embodiment, the integer register file / bypass network 2408 may include, but is not limited to, two separate register files, one register file for 32 low-order data bits and a second register file for 32 high-order data bits. In at least one embodiment, the floating-point register file / bypass network 2410 may include, but is not limited to, 128-bit wide entries because floating-point instructions typically have operands 64 to 128 bits wide.
[0339] In at least one embodiment, execution units 2412, 2414, 2416, 2418, 2412, 2422, 2424 may execute instructions. In at least one embodiment, register files 2408, 2410 store integer and floating-point data operand values that microinstructions must execute. In at least one embodiment, processor 2400 may include, but is not limited to, any number and combination of execution units 2412, 2414, 2416, 2418, 2412, 2422, 2424. In at least one embodiment, floating-point ALU 2422 and floating-point move unit 2424 may execute floating-point, MMX, SIMD, AVX, and SSE operations, or other operations, including special machine learning instructions. In at least one embodiment, the floating point ALU 2422 may include, but is not limited to, a 64-bit by 64-bit floating point divider to perform division, square root, and remainder micro-operations.In at least one embodiment, instructions involving a floating-point value may be handled by floating-point hardware. In at least one embodiment, ALU operations may be passed to fast ALUs 2416, 2418. In at least one embodiment, fast ALUs 2416, 2418 may perform fast operations with an effective latency of half a clock cycle. In at least one embodiment, the most complex integer operations are offloaded to slow ALU 2412, as slow ALU 2412 may include, but is not limited to, integer execution hardware for long-latency operations, such as a multiplier, a shifter, flag logic, and branch processing. In at least one embodiment, memory load / store operations may be performed by ALUs 2412, 2414.In at least one embodiment, the fast ALU 2416, the fast ALU 2418, and the slow ALU 2412 may perform integer operations on 64-bit data operands. In at least one embodiment, the fast ALU 2416, the fast ALU 2418, and the slow ALU 2420 may be implemented to support a plurality of data bit sizes, including sixteen, thirty-two, 128, 326, etc. In at least one embodiment, the floating-point ALU 2422 and the floating-point move unit 2424 may be implemented to support a number of operands with different bit widths, such as 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.
[0340] In at least one embodiment, the uop schedulers 2402, 2404, 2406 dispatch dependent operations before a parent load completes execution. In at least one embodiment, because uops may be speculatively scheduled and executed in the processor 2400, the processor 2400 may also include logic to handle memory misses. In at least one embodiment, if a data load is missing from a data cache, there may be dependent operations in the pipeline that have left a scheduler with temporarily incorrect data. In at least one embodiment, a replay mechanism tracks instructions that use incorrect data and reexecutes them. In at least one embodiment, dependent operations may need to be replayed, and independent operations may complete.In at least one embodiment, schedulers and a replay mechanism of at least one embodiment of a processor may also be configured to intercept instruction sequences for text string comparison operations.
[0341] In at least one embodiment, "registers" may refer to on-board processor memory locations that may be used as part of instructions to identify operands. In at least one embodiment, registers may be those usable (from a programmer's perspective) from outside the processor. In at least one embodiment, registers may not be limited to a particular circuit type. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein.In at least one embodiment, registers described herein may be implemented by circuitry within a processor using any number of different techniques, such as dedicated physical registers, dynamically allocated physical registers using register renaming, combinations of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, integer registers store 32-bit integer data. A register file of at least one embodiment further includes eight multimedia SIMD packed data registers.
[0342] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details of the inference and / or training logic 815 are described below in connection with Fig.8A and / or 8B. In at least one embodiment, portions or all of the inference and / or training logic 815 may be incorporated into an execution block 2411 and other memory or registers, which may or may not be shown. For example, in at least one embodiment, training and / or inference techniques described herein may utilize one or more of the ALUs illustrated in execution block 2411.
[0343] Additionally, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure ALUs of execution block 2411 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0344] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0345] Fig.25 illustrates a deep learning application processor 2500 according to at least one embodiment. In at least one embodiment, the deep learning application processor 2500 uses instructions that, when executed by the deep learning application processor 2500, cause the deep learning application processor 2500 to perform some or all of the processes and techniques described in this disclosure. In at least one embodiment, the deep learning application processor 2500 is an application-specific integrated circuit (ASIC). In at least one embodiment, the application processor 2500 performs matrix multiplication operations either "hard-wired" in hardware or as a result of the execution of one or more instructions, or both.In at least one embodiment, the deep learning application processor 2500 includes, but is not limited to, processing clusters 2510(1)-2510(12), inter-chip interconnects (“ICLs”) 2512(1)-2512(12), inter-chip controllers (“ICCs”) 2530(1)-2530(2), second-generation high-bandwidth memory (“HBM2”) 2540(1)-2540(4), memory controllers (“Mem Ctrlrs”) 2542(1)-2542(4), a high-bandwidth memory physical layer (“HBM PHY”) 2544(1)-2544(4), a management controller central processing unit (“Management Controller CPU”) 2550, a serial peripheral interface, an inter-chip integrated circuit, and a general-purpose input / output block (“SPI, I. 2 C, GPIO”) 2560, a Peripheral Interconnect Express Controller and Direct Memory Access Block (“PCIe Controller and DMA”) 2570, and a sixteen-channel Peripheral Interconnect Express Port (“PCI Express × 16”) 2580.
[0346] In at least one embodiment, processing clusters 2510 may perform deep learning operations, including inference or prediction operations based on weight parameters calculated using one or more training techniques, including those described herein. In at least one embodiment, each processing cluster 2510 may include, but is not limited to, any number and type of processors. In at least one embodiment, deep learning application processor 2500 may include any number and type of processing clusters 2500. In at least one embodiment, inter-chip interconnects 2512 are bidirectional.In at least one embodiment, inter-chip interconnects 2512 and inter-chip controllers 2530 enable multiple deep learning application processors 2500 to exchange information, including activation information resulting from the execution of one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, deep learning application processor 2500 may include any number (including zero) and type of ICLs 2512 and ICCs 2530.
[0347] In at least one embodiment, the HBM2s 2540 provide a total of 32 gigabytes (GB) of memory. The HBM2 2540(i) is associated with both the memory controller 2542(i) and the HBM PHY 2544(i). In at least one embodiment, any number of HBM2s 2540 may provide any type and total amount of high-bandwidth memory and may be associated with any number (including zero) and type of memory controllers 2542 and HBM PHYs 2544. In at least one embodiment, SPI, I 2 C, GPIO 2560, PCIe controller and DMA 2570 and / or PCIe 2580 may be replaced by any number and type of blocks enabling any number and type of communication standards in any technically feasible manner.
[0348] The inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 815 are described herein in connection with Fig.8A and / or 8B. In at least one embodiment, the deep learning application processor is used to train a machine learning model, such as a neural network, to predict, or infer information provided to the deep learning application processor 2500. In at least one embodiment, the deep learning application processor 2500 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) trained by another processor or system or by the deep learning application processor 2500. In at least one embodiment, the processor 3300 may be used to perform one or more of the neural network use cases described herein.
[0349] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0350] Fig.26 is a block diagram of a neuromorphic processor 2600 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2600 may receive one or more inputs from sources external to the neuromorphic processor 2600. In at least one embodiment, these inputs may be communicated to one or more neurons 2602 within the neuromorphic processor 2600. In at least one embodiment, the neurons 2602 and their components may be implemented using circuitry or logic, including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2600 may include, but is not limited to, thousands or millions of instances of neurons 2602, although any number of neurons 2602 may be used.In at least one embodiment, each instance of neuron 2602 may include a neuron input 2604 and a neuron output 2606. In at least one embodiment, neurons 2602 may generate outputs that may be transferred to inputs of other instances of neurons 2602. For example, in at least one embodiment, neuron inputs 2604 and neuron outputs 2606 may be connected to each other via synapses 2608.
[0351] In at least one embodiment, neurons 2602 and synapses 2608 may be interconnected such that neuromorphic processor 2600 is used to process or analyze information received by neuromorphic processor 2600. In at least one embodiment, neurons 2602 may send an output pulse (or "fire" or "spike") when inputs received via neuron input 2604 exceed a threshold. In at least one embodiment, neurons 2602 may sum or integrate signals received at neuron inputs 2604.For example, in at least one embodiment, neurons 2602 may be implemented as leaky integrate-and-fire neurons, where when a sum (referred to as a "membrane potential") exceeds a threshold, a neuron 2602 may generate an output (or "firing") using a transfer function, such as a sigmoid or threshold function. In at least one embodiment, a leaky integrate-and-fire neuron may sum signals received at neuron inputs 2604 to a membrane potential and may further apply a decay factor (or leak) to decrease a membrane potential. In at least one embodiment, a leaky integrate-and-fire neuron may fire if multiple input signals are received at neuron inputs 2604 quickly enough to exceed a threshold (i.e., before a membrane potential becomes too low to fire).In at least one embodiment, neurons 2602 may be implemented using circuitry or logic that receives inputs, integrates inputs to a membrane potential, and decays a membrane potential. In at least one embodiment, inputs may be averaged, or any other suitable transfer function may be used. Further, in at least one embodiment, but not limited to, neurons 2602 may include comparator circuitry or logic that generates an output spike at neuron output 2606 when a result of applying a transfer function to neuron input 2604 exceeds a threshold. In at least one embodiment, after neuron 2602 fires, it may ignore previously received input information, for example, by resetting a membrane potential to 0 or another suitable default value.In at least one embodiment, after the membrane potential is reset to 0, the neuron 2602 may resume normal operation after a suitable period of time (or refractory period).
[0352] In at least one embodiment, neurons 2602 may be interconnected by synapses 2608. In at least one embodiment, synapses 2608 may be operable to transmit signals from an output of a first neuron 2602 to an input of a second neuron 2602. In at least one embodiment, neurons 2602 may transmit information via more than one instance of synapse 2608. In at least one embodiment, one or more instances of neuron output 2606 may be connected via an instance of synapse 2608 to an instance of neuron input 2604 in the same neuron 2602. In at least one embodiment, an instance of neuron 2602 that generates an output to be transmitted via an instance of synapse 2608 may be referred to as a "presynaptic neuron" with respect to that instance of synapse 2608.In at least one embodiment, an instance of neuron 2602 that receives input transmitted across an instance of synapse 2608 may be referred to as a "postsynaptic neuron" with respect to that instance of synapse 2608. Because an instance of neuron 2602 may receive inputs from one or more instances of synapse 2608 and may also transmit outputs across one or more instances of synapse 2608, a single instance of neuron 2602 may therefore be both a "presynaptic neuron" and a "postsynaptic neuron" with respect to different instances of synapses 2608 in at least one embodiment.
[0353] In at least one embodiment, neurons 2602 may be organized into one or more layers. Each instance of neuron 2602 may have a neuron output 2606 that may propagate through one or more synapses 2608 to one or more neuron inputs 2604. In at least one embodiment, neuron outputs 2606 of neurons 2702 in a first layer 2710 may be connected to neuron inputs 2704 of neurons 2702 in a second layer 2712. In at least one embodiment, layer 2610 may be referred to as a "feed-forward layer." In at least one embodiment, each instance of neuron 2602 in an instance of first layer 2610 may propagate to each instance of neuron 2602 in second layer 2612. In at least one embodiment, the first layer 2610 may be referred to as a “fully connected feed-forward layer.”In at least one embodiment, each instance of neuron 2602 in an instance of second layer 2612 may spread to fewer than all instances of neuron 2602 in a third layer 2614. In at least one embodiment, second layer 2612 may be referred to as a "sparsely connected feed-forward layer." In at least one embodiment, neurons 2602 in second layer 2612 may spread to neurons 2602 in multiple other layers, including neurons 2602 in the (same) second layer 2612. In at least one embodiment, second layer 2612 may be referred to as a "recurrent layer." Neuromorphic processor 2600 may include, but is not limited to, any suitable combination of recurrent layers and feed-forward layers, including, but not limited to, both sparsely connected feed-forward layers and fully connected feed-forward layers.
[0354] In at least one embodiment, neuromorphic processor 2600 may include, but is not limited to, a reconfigurable interconnect architecture or dedicated hard-wired interconnects to connect synapse 2608 to neurons 2602. In at least one embodiment, neuromorphic processor 2600 may include, but is not limited to, circuitry or logic that allows synapses to be assigned to different neurons 2602 as needed based on neural network topology and neuron fan-in / out. For example, in at least one embodiment, synapses 2608 may be connected to neurons 2602 using an interconnect structure, such as a network on chip, or with dedicated interconnects. In at least one embodiment, synapse interconnects and components thereof may be implemented using circuitry or logic.
[0355] In at least one embodiment, a processor or GPU, as described herein, may be used to create a robot control system that implements model-based and model-free control. For example, a processor or GPU, as described above, may be used to execute executable instructions that cause the processor to implement model-based and model-free control, as described herein.
[0356] Fig.27 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 2700 includes one or more processors 2702 and one or more graphics processors 2708, and may be a desktop system with a single processor, a multiprocessor workstation system, or a server system with a large number of processors 2702 or processor cores 2707. In at least one embodiment, system 2700 is a processing platform integrated into a system-on-a-chip (SoC) integrated circuit for use in mobile, wearable, or embedded devices.
[0357] In at least one embodiment, system 2700 may include, couple with, or be integrated with a gaming console, including a gaming and media console, a mobile gaming console, a wearable gaming console, or an online gaming console within a server-based gaming platform. In at least one embodiment, system 2700 is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, processing system 2700 may also include, couple with, or be integrated with a wearable device, such as a wearable smart watch device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 2700 is a television or set-top box device having one or more processors 2702 and a graphics interface generated by one or more graphics processors 2708.
[0358] In at least one embodiment, one or more processors 2702 each include one or more processor cores 2707 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each one or more processor cores 2707 is configured to process a particular instruction sequence 2709. In at least one embodiment, the instruction sequence 2709 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Comp...
Claims
[1] Computer-implemented method comprising: moving a robot (102) to be within a region under the control of a first method using a physical model based at least in part on information from a first perception system; Determining an uncertainty of the information generated by the first perception system (106); Determining that the robot (102) is in the region based at least in part on the uncertainty; and as a result of determining that the robot is in the region, moving the robot to perform a task under the control of a second method using information generated by a second perception system (108), The second method does not rely on the physical model and is a model-free method. [2] A method according to any one of the preceding claims, wherein the first method is a model-based method. [3] A method according to any one of the preceding claims, wherein: the first perception system (106) is a stationary camera; and the second perception system (108) is a camera attached to the robot. [4] A method according to any one of the preceding claims, further comprising determining the region based at least in part on the uncertainty of the information. [5] Method according to one of the preceding claims, further comprising: Determining that the robot (102) is outside the region; and as a result of determining that the robot is outside the region, moving the robot to be within the region using the first method. [6] A method according to any one of the preceding claims, wherein the uncertainty is a non-parametric distribution of a plurality of poses of the region and associated weights for each pose. [7] A method according to any one of the preceding claims, wherein the uncertainty is a parametric distribution. [8] A method according to any one of the preceding claims, wherein the region is a sub-region of a region in which the second method is useful to complete a task. [9] A method according to any one of the preceding claims, wherein the second method is performed using an autoencoder (306, 308) trained to complete a task to which an input is input from the second perceptual system. [10] Computer system comprising: one or more processors; and a computer-readable memory storing executable instructions that, as a result of being executed by the one or more processors, cause the computer system to: moving a robot (102) into a region using a model of the robot's environment, the model being oriented with image data from a first camera; determine an uncertainty of the model using uncertainty information associated with the first camera; Determining that the robot is in the region based at least in part on the uncertainty of the model; and as a result of determining that the robot is in the region, to perform a task, wherein the robot is under the control of a machine learning system trained with image data from a second camera, wherein the machine learning system does not rely on a physical model but on a model-free algorithm. [11] The computer system of claim 10, wherein the second camera is a wrist-mounted camera on the robot. [12] The computer system of claim 10 or 11, wherein, as a result of completing the task, the computer system updates the uncertainty of the model using a result of the task. [13] The computer system of claim 12, wherein the result of the task indicates a pose for the model. [14] A computer system according to any one of claims 10 to 13, wherein the model is oriented at least by processing the image data from the first camera using a deep object pose estimator. [15] The computer system of any one of claims 10 to 14, wherein the first camera and the second camera are the different cameras. [16] A computer system according to any one of claims 10 to 15, wherein the image data from the first camera is used to generate a plurality of possible poses consistent with the image data from the first camera. [17] A computer system according to any one of claims 10 to 16, wherein the robot (102) is moved into the region using a model-based controller (1136) that uses target attractors defined by motion policies of the robot. [18] Computer-readable media storing executable instructions which, as a result of being executed on one or more processors of a computer system, cause the computer system to at least: moving a robot (102) within a region under the control of a model-based method using a physical model oriented using information from a first perception system (106); to determine an uncertainty of the information generated by the first perception system (106); determine that the robot is within the region by a margin based at least in part on the uncertainty; and as a result of determining that the robot is in the region, instructing the robot to perform a task under the control of a second method based on information generated by a second perception system (108), wherein the second method does not rely on the physical model and is a model-free method. [19] Computer-readable media according to claim 18, wherein: the first perception system (106) is a stationary camera; and the second perception system (108) is a camera that moves with the robot. [20] Computer-readable media according to claim 18 or 19, wherein a dimension of the region is determined based at least in part on the uncertainty of the information generated by the first perception system (106). [21] Computer-readable media according to one of claims 18 to 20, wherein the uncertainty of the information generated by the first perception system (106) is determined as a distribution of multiple poses of the region. [22] Computer-readable media according to any one of claims 18 to 20, wherein the executable instructions further cause the computer system to: to determine that the robot is outside the region; and as a result of determining that the robot is outside the region, to move the robot to within the region using the model-based method. [23] Computer-readable media according to any one of claims 18 to 20, wherein the second method is a model-free method implemented using a machine-learned model trained with input from the second perception system. [24] Computer-readable media according to claim 23, wherein the input comprises simulated images generated by a simulation of the task. [25] Computer-readable media according to any one of claims 18 to 20, wherein the task controls an autonomous vehicle.
Citation Information
Patent Citations
Object handling equipment and object handling procedures
DE112017002114T5
Autonomous mobile robot and method for operating the same
EP2713232A1
Upper limb motion assisting device and upper limb motion assisting system
EP3549725A1
Computer program for recommending jewelry product
KR1020210022194A
System and method for scanning a region using a low discrepancy sequence
US20020146152A1