Machine learning control of object handoff
By generating robot grasping postures that do not interfere with human hands using depth cameras and deep networks, and combining them with a robust logic-dynamic system, the reliability problem of object exchange in robot-human interaction is solved, achieving efficient and reliable object handover, suitable for collaborative manufacturing and home assistance tasks.
Patent Information
- Application Number
- CN202110852773.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-28
- Filing Date
- 2021-07-27
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2041-07-27
AI Technical Summary
When robots interact with humans, the process of exchanging objects can easily interfere with the human hand, causing objects to fall or injure the human operator. Existing technologies make it difficult to achieve reliable object handover.
3D images of a human hand and an object are acquired using a depth camera, point clouds are generated, and the point cloud portions of the human hand and the object are separated. A trained deep network is used for classification to generate a robot grasping posture that does not interfere with the human hand. A robust logic-dynamic system is combined for reactive planning to generate a robot grasping posture to receive the object.
It enables reliable receiving of objects from the human hand without interfering with the human hand, improving the success rate and time efficiency of object handover, adapting to various human grasping postures, and is suitable for collaborative manufacturing and home assistance tasks.
Smart Images

Figure CN114004329B_ABST
Abstract
Description
BACKGROUND
[0001] Robotic automation of tasks is an important area of development. However, some tasks involve collaboration with a human operator. Tasks such as personal care tasks, handover of objects, or joining operations between a robot and a manual task often involve interaction between a robot and a human. One problem in the collaborative task space is the exchange of objects between the human operator and the robot. As part of performing a task, an exchange can be made from the robot to the human and / or from the human to the robot. When picking up an object from the human hand, the exchange can encounter difficulties if the grasp chosen by the robot interferes with the human hand while being able to reliably grasp the object. Failure can result in the object falling, in some cases also injuring the human operator. Therefore, developing a reliable handover technique that allows the robot to receive an object from a human operator is an important problem within the collaborative task space. BRIEF DESCRIPTION OF DRAWINGS
[0002] Various techniques will be described with reference to the drawings, in which:
[0003] Figure 1 An example of a human-robot interaction pose is shown in which the human performs a palm-down pinch grasp of an object, in accordance with an embodiment;
[0004] Figure 2 An example of a human-robot interaction pose is shown in which the human performs a downward grasp of an object, in accordance with an embodiment;
[0005] Figure 3 An example of a human-robot interaction pose is shown in which the human performs a palm-up pinch grasp of an object, in accordance with an embodiment;
[0006] Figure 4 An example of a human-robot interaction pose is shown in which the human performs a horizontal pinch grasp of an object, in accordance with an embodiment;
[0007] Figure 5 An example of a framework for performing a handover between a robot and a human hand is shown in accordance with an embodiment;
[0008] Figure 6 An example of a human hand pose that can be used to prompt a transfer of an object from a robot is shown in accordance with an embodiment;
[0009] Figure 7 An example of a hand pose that can be used to grasp an object is shown in accordance with an embodiment;
[0010] Figure 8A An example of a robot gripper is shown in accordance with an embodiment;
[0011] Figure 8BAn example of a robotic gripper with four fingers is shown, in accordance with an embodiment;
[0012] Figure 9 An example of a human-robot interaction is shown, in accordance with an embodiment;
[0013] Figure 10 Data describing performance of a human-robot interaction system is shown, in accordance with an embodiment;
[0014] Figure 11 An example of a hand pose that can be used to present an object to a robot is shown, in accordance with an embodiment;
[0015] Figure 12 An example of a process to transfer an object between a robot and a human hand as a result of execution by a computer system is shown, in accordance with an embodiment;
[0016] Figure 13A Inference and / or training logic is shown, in accordance with at least one embodiment;
[0017] Figure 13B Inference and / or training logic is shown, in accordance with at least one embodiment;
[0018] Figure 14 Training and deployment of a neural network is shown, in accordance with at least one embodiment;
[0019] Figure 15 An example data center system is shown, in accordance with at least one embodiment;
[0020] Figure 16A An example of an autonomous vehicle is shown, in accordance with at least one embodiment;
[0021] Figure 16B An example of a camera position and field of view of an autonomous vehicle is shown, in accordance with at least one embodiment; Figure 16A
[0022] A block diagram showing an example system architecture of an autonomous vehicle is shown, in accordance with at least one embodiment; Figure 16C Figure 16A A diagram showing a system for communication between one or more cloud-based servers and an autonomous vehicle is shown, in accordance with at least one embodiment;
[0023] Figure 16D Figure 16A A block diagram showing a computer system is shown, in accordance with at least one embodiment;
[0024] Figure 17 A block diagram showing a computer system is shown, in accordance with at least one embodiment;
[0025] Figure 18 A block diagram showing a computer system is shown, in accordance with at least one embodiment;
[0026] Figure 19 A computer system is shown, in accordance with at least one embodiment;
[0027] Figure 20 A computer system is shown, in accordance with at least one embodiment;
[0028] Figure 21A A computer system is shown, in accordance with at least one embodiment;
[0029] Figure 21B A computer system is shown, in accordance with at least one embodiment;
[0030] Figure 21C A computer system is shown, in accordance with at least one embodiment;
[0031] Figure 21D A computer system is shown, in accordance with at least one embodiment;
[0032] Figure 21E and Figure 21F A shared programming model is shown, in accordance with at least one embodiment;
[0033] Figure 22 An exemplary integrated circuit and associated graphics processor are shown, in accordance with at least one embodiment.
[0034] Figure 23A and Figure 23B An exemplary integrated circuit and associated graphics processor are shown, in accordance with at least one embodiment.
[0035] Figure 24A and Figure 24B Additional exemplary graphics processor logic is shown, in accordance with at least one embodiment;
[0036] Figure 25 A computer system is shown, in accordance with at least one embodiment;
[0037] Figure 26A A parallel processor is shown, in accordance with at least one embodiment;
[0038] Figure 26B A partition unit is shown, in accordance with at least one embodiment;
[0039] Figure 26C A processing cluster is shown, in accordance with at least one embodiment;
[0040] Figure 26D A graphics multiprocessor is shown, in accordance with at least one embodiment;
[0041] Figure 27A multi-Graphics Processing Unit (GPU) system is shown in accordance with at least one embodiment;
[0042] Figure 28 A graphics processor is shown in accordance with at least one embodiment;
[0043] Figure 29 is a block diagram illustrating a processor micro-architecture for a processor in accordance with at least one embodiment;
[0044] Figure 30 A deep learning application processor is shown in accordance with at least one embodiment;
[0045] Figure 31 is a block diagram illustrating an example neuromorphic processor in accordance with at least one embodiment;
[0046] Figure 32 At least a portion of a graphics processor is shown in accordance with one or more embodiments;
[0047] Figure 33 At least a portion of a graphics processor is shown in accordance with one or more embodiments;
[0048] Figure 34 At least a portion of a graphics processor is shown in accordance with one or more embodiments;
[0049] Figure 35 is a block diagram of a graphics processing engine of a graphics processor in accordance with at least one embodiment;
[0050] Figure 36 is a block diagram of at least a portion of a graphics processor core in accordance with at least one embodiment;
[0051] Figure 37A and Figure 37B Thread execution logic including an array of processing elements of a graphics processor core is shown in accordance with at least one embodiment.
[0052] Figure 38 A parallel processing unit (“PPU”) is shown in accordance with at least one embodiment;
[0053] Figure 39 A general processing cluster (“GPC”) is shown in accordance with at least one embodiment;
[0054] Figure 40 A memory partition unit of a parallel processing unit (“PPU”) is shown in accordance with at least one embodiment;
[0055] Figure 41 A streaming multiprocessor is shown in accordance with at least one embodiment. DETAILED DESCRIPTION
[0056] This document describes a vision-based system that allows a robot to receive an object presented in a human hand. In one example, a human grasps an object and presents it in the field of view of a depth camera monitor. The depth camera takes a 3D image of the hand grasping the object and provides it to the system. The system generates a point cloud from the image and separates the portion of the point cloud associated with the human hand and the portion of the point cloud associated with the object. Using this information, the system is able to determine the pose of the human hand and the pose of the object. The system generates a set of grasps that the robot can perform to grasp the object, and then selects a grasp from the set of grasps that does not interfere with the human hand. In one implementation, the human’s grasp is classified into various types of hand poses to help the system select an appropriate robot grasp to obtain the object. In another implementation, given a human-object handoff dataset annotated with ground truth hand and object poses, a deep network is trained that takes as input a color point cloud observed by a depth camera, segments the human hand from the object, and proposes good grasps and controls to the robot so that it can receive the object from the human hand without pinching the person’s fingers. In various embodiments, the techniques described herein can be applied to systems that hand off objects from one robot to another robot or from an animal to a robot, or in some examples when taking objects from appendages other than a human hand (e.g., taking off a hat).
[0057] Object transfer between a human and a robot is an important capability for robots to collaborate with humans. Both robot-to-human handoff and human-to-robot handoff can be difficult. Some techniques described herein describe a method of human-to-robot handoff in which a robot encounters a person mid-way, classifies the person’s grasp of an object, and quickly plans a trajectory accordingly to take the object from the person’s hand according to their intent. At least one embodiment collects a human grasp dataset that covers typical ways of holding an object with various hand shapes and poses, and learns a depth model on this dataset to classify hand grasps into one of these categories. At least one embodiment provides a planning and execution method that takes an object from a person’s hand according to the detected grasp and hand position, and replans as needed if the handoff is interrupted. Through system evaluations, this document demonstrates that various embodiments produce improved handoff relative to two baselines.
[0058] Providing and retrieving objects to and from humans is a fundamental capability of collaborative robots across applications from manufacturing to physical assistance in the home. Some techniques focus on transferring objects from the robot to the human, assuming the human can place the object in the robot's gripper for reverse operation. This approach is sometimes infeasible in situations where the human needs to focus on a task at hand (e.g., performing surgery) or the human is impaired and has limited arm motion due to injury. Accordingly, at least one embodiment provides more reactive handoff, which can accommodate ways in which a human presents an object to a robot and meets them halfway to take the object.
[0059] In at least one embodiment, one of the challenges in making human-robot handoff reactive is reliable and continuous perception of the object and the human. At least one embodiment estimates human hand pose and 6D object pose by methods that leverage computer vision. Various techniques can be used, including techniques that estimate the pose of the hand and the pose of the object separately. At least some embodiments estimate hand and object pose as the hand interacts with the object.
[0060] In at least one embodiment, the techniques described herein solve the perception problem for human-to-robot handoff by formulating it as a hand grasp classification problem. In one example, these techniques discretize the ways in which a human can hold a small object into several categories and collect a dataset to learn a deep model that classifies a given hand holding an object into one of these grasp categories. The handoff task is modeled as a robust, logic-dynamic system that generates a motion plan that avoids contact between the gripper and the human hand given the human grasp classification. At least one embodiment compares to two baseline methods, one that does not infer the human hand pose and another that relies on independent hand and object pose estimation. At least one embodiment demonstrates that our approach has higher success rate and temporal efficiency over both baselines. A user study (N=9) presents effectiveness of naive users when focusing on the robot and focusing on secondary tasks.
[0061] At least one embodiment provides a vision-based reliable algorithm that can generate a grasp to avoid collision for human-to-robot handoff or robot-to-robot handoff tasks, e.g., how a robot takes an object from a human or another robot. At least one embodiment adjusts the way a robot grasps an object based on how a human hand holds the object. As a result, at least one embodiment is more reliable for not touching the human hand.
[0062] Human-robot handover is an important topic in human-robot collaboration, involving numerous application domains from collaborative manufacturing to home assistance. At least one embodiment focuses on robot-to-human handover, where the robot starts with an object in hand and transfers it to the human. One challenge is to choose parameters of the robot motion to optimize for smooth handover. This includes the choice of object pose and the robot’s grasp of the object, taking into account the user’s comfort, preferences based on subjective feedback, functional affordances and intended use of the object after handover, motion constraints of the human, social roles of the human, and the configuration of the object when it is grasped before handover. At least one embodiment emphasizes trajectory parameters to reach the handover pose, exploration of approach angles, starting pose of the trajectory in contrast to the handover pose, motion smoothness, object release time, estimated human wrist pose, relative timing of the handover phase, and human ergonomic preferences. While some embodiments focus on offline computation of handover parameters, at least one embodiment involves human perception to enable reactive handover.
[0063] The technology described herein focuses on human perception to enable reactive handover and leverage human hand pose estimation technology. Pose estimation of the human hand can be done through two-dimensional RGB monocular images or 3D depth camera information. In general, 3D depth cameras provide additional information that allows for more accurate estimation of hand pose compared to technology that relies only on two-dimensional RGB images.
[0064] The technology described herein enables a system for performing tasks based on a robust logic-dynamic system, a method of automatically creating a reactive task plan for a robot. In at least one embodiment, the idea is to constantly identify the current logic state and reactively re-plan to handle uncertainty and changes in the logic state, which is a useful approach for handling partially observable environments. In some examples, the task model can be considered in a manner similar to a behavior tree, which is a method for representing complex tasks useful for human-robot collaboration.
[0065] At least one embodiment partially addresses this issue by classifying human grasp poses. In at least one embodiment, human grasp poses are discretized into seven categories, which can not cover all ways in which a human hand can grasp an object.
[0066] In at least one embodiment, a dataset of human-object handovers is generated, annotated with ground truth hand poses and object poses. The dataset is used to train a deep network that takes as input a color point cloud obtained by an RGB depth camera to segment the human hand from the object and suggest a grasp and control scheme for the robot so that the robot can receive the object from the human hand without pinching or touching the human fingers.
[0067] One advantage of the technology described herein is that the model can generate a robotic grasp that can receive an object from a human hand / another robot while learning a hand-object manipulation dataset without pinching the human hand (or the other robot's gripper). Some examples provide a low-cost solution based on an RGB-D camera, which is both inexpensive and lightweight compared to many wearable sensors. Various embodiments enable reliable handoff, which is important when building more complex robotic systems. For example, the technology described herein can be suitable for an elder care robot or a cooking robot that can work closely with humans.
[0068] In various embodiments, the technology described herein provides: (1) hand-object interaction reasoning for handoff posed as a classification problem via a dataset that covers a wide range of hand shapes and poses; (2) a system that adaptively plans a robotic grasp to acquire an object from a human, enabling the robot to respond to humans fluidly and naturally; and (3) experimental results demonstrating improvement over baseline methods, as well as a user study that validates our method with naive users.
[0069] Humans hand off objects in different ways. They can present an object on the palm, or they can pinch grasp and present the object in different orientations. The technology described herein can determine which grasp the human is using and adjust accordingly, enabling reactive human-robot handoff.
[0070] Figure 1 An example of a human-robot interaction pose is shown in which a human performs a palm-down pinch grasp of an object, in accordance with an embodiment. A robot 100 is positioned to take an object 102 from a human hand 104. In at least one embodiment, a depth camera takes an image of the human hand 104 grasping the object 102, generates a point cloud from the image, and provides the point cloud to a trained neural network. The neural network generates an appropriate grasp for a robotic gripper 106 connected to the robot 100, such that the robotic gripper 106 is able to grasp the object 102 without interfering with the human hand 104.
[0071] Figure 2 An example of a human-robot interaction pose is shown in which a human performs a down grasp of an object, in accordance with an embodiment. A robot 200 is positioned to take an object 202 from a human hand 204. In at least one embodiment, a depth camera takes an image of the human hand 204 grasping the object 202, generates a point cloud from the image, and provides the point cloud to a trained neural network. The neural network generates an appropriate grasp for a robotic gripper 206 connected to the robot 200, such that the robotic gripper 206 is able to grasp the object 202 without interfering with the human hand 204.
[0072] Figure 3An example of a human-robot interaction pose is shown in which a human performs a palm-up pinch grasp on an object. A robot 300 is positioned to take an object 302 from a human hand 304. In at least one embodiment, a depth camera takes an image of the human hand 304 grasping the object 302, generates a point cloud from the image, and provides the point cloud to a trained neural network. The neural network generates an appropriate grasp for a robotic gripper 306 connected to the robot 300 such that the robotic gripper 306 is able to grasp the object 302 without interfering with the human hand 304.
[0073] Figure 4 An example of a human-robot interaction pose is shown in which a human performs a horizontal pinch grasp on an object. A robot 400 is positioned to take an object 402 from a human hand 404. In at least one embodiment, a depth camera takes an image of the human hand 404 grasping the object 402, generates a point cloud from the image, and provides the point cloud to a trained neural network. The neural network generates an appropriate grasp for a robotic gripper 406 connected to the robot 400 such that the robotic gripper 406 is able to grasp the object 402 without interfering with the human hand 404.
[0074] In at least one embodiment, when a robot takes an object from a human hand, the robot’s motion is adjusted according to the way the human hand is grasping the object. Generally, this can prevent the robot from behaving in a nonintuitive way, or in a way that interferes with or even harms human fingers. In at least one embodiment, the handoff framework solves this problem by taking a point cloud centered on a human hand detected with an Azure Body Tracking Software Development Kit (“SDK”), then estimating the hand’s grasp class based on how the object is being grasped by the human hand. In at least one embodiment, the robot grasp is then adaptively planned.
[0075] Figure 5 An example of a framework for performing a handoff between a robot and a human hand is shown according to an embodiment. In one example, the framework obtains an RGBD image 502 showing a hand holding an object. Using this RGBD image, a point cloud 504 of the hand and object is generated. In at least one embodiment, the framework takes the hand-detection centered point cloud and then classifies it into one of seven grasp types using a model 506, which encompass the various ways that human users tend to grasp objects. The task model then adaptively plans 508 a robot grasp.
[0076] Figure 6Examples of hand poses that can be used to signal a human hand pose to transfer an object from a robot are shown. For example, a first hand pose 602 can be designated as a pose that indicates a human is ready to receive an object, and a second hand pose 604 can be designated as a pose that indicates a human is not ready to receive an object from a robot.
[0077] At least one embodiment defines a set of discrete human grasps that describe the way a human hand grasps an object to complete a handoff task. At least one embodiment discretizes common human grasps for human-robot handoff tasks into seven categories, such as Figure 7 shown. For example, if a hand is grasping a block, the pose of the hand can be categorized as open palm, pinch bottom, pinch top, pinch side, or lifting. In another example, if a hand is not grasping anything, it can be waiting for a robot to hand off an object, or just not doing anything (other).
[0078] Figure 7 Examples of hand poses that can be used to grasp an object are shown, according to embodiments. A first hand pose 704 shows an object being held in an open palm. A second hand pose 704 shows an object being held in a pinch bottom position. A third hand pose 706 shows an object being held in a lifting pose. A fourth hand pose 708 shows an object being held in a pinch top position. A fifth hand pose 710 shows an object being held as it is pinched from the side. These and other hand poses can be used by a human when presenting an object to a robot. In some implementations, a categorization of hand poses can be used, such as Figure 7 the hand pose categories shown in FIG. 7. In other implementations, the hand pose can be freeform, which can include Figure 7 various types of hand poses not shown in FIG. 7. In such examples, the system can use a point cloud or a skeletal pose of the hand to construct an appropriate pose for the robot to take the object from the human hand.
[0079] Figure 8A Examples of a robot gripper are shown, according to embodiments. In one example, a robot gripper 802 includes a set of jaws 804 that can be closed and separated. Some examples can include tactile sensors on the surface of the jaws 804, such that the system can obtain a measure of force. In some examples, the gripper 802 includes a wrist joint that can articulate and position the set of jaws 804 under the control of the system.
[0080] Figure 8BAn example of a robotic gripper 806 with four fingers that can be used in embodiments is shown. The robotic gripper 806 includes a first finger 808, a second finger 810, a third finger 812, and an opposing finger 814 that approximate a human hand. The robotic gripper 806 can be used in accordance with the various embodiments described and illustrated herein. For example, the robotic gripper 806 can be used to take an object held by a human hand in a manner that does not disturb the human hand. In addition, other types of robotic grippers with more or fewer fingers can also be used. In one example, a robotic gripper can have 2, 3, 4, 5, or more fingers, and each finger can have a frictional surface to help grip an object. In one example, the fingers of the robotic gripper can have tactile sensors that assist the system by indicating contact with an object. In one example, the gripper can include magnetic elements to help take ferrous objects from a human hand. In another example, the gripper can include vacuum pick-up elements to capture objects.
[0081] In one example, to learn a model that classifies human grasps, an embodiment created a dataset by using an Azure Kinect RGBD camera that covered eight subjects with various hand shapes and hand poses. For example, in one implementation, a subject was shown example images of hand grasps, and the subject performed similar poses that were recorded for twenty to sixty seconds. The image sequences were labeled with the corresponding human grasp category. During the recording process, the subject can move his or her body and hands to different positions to diversify the camera’s perspective. Both the left and right hands of each subject were recorded. In one example, the dataset consisted of 151,551 images.
[0082] Rather than using ConvNets to learn depth features on depth images, at least one embodiment employs PointNet++ on point clouds for human grasp classification. In at least one embodiment, the backbone network consists of four set abstraction layers for learning point features and three layers of perceptrons with batch normalization, ReLu, and Dropout for global feature learning and human grasp classification. Given a point cloud cropped around a hand, the network classifies it into one of the defined grasp categories, which can be used for further robot grasp planning.
[0083] At least one embodiment associates each type of human grasp with a canonical robot grasp direction in order to minimize the human’s effort during the human-to-robot handoff. As shown, the coordinate represents a canonical robot grasp frame in the camera frame. The motivation is to reduce the chances of the robot grabbing the human hand while keeping its motion and trajectory as natural and smooth as possible. Figure 3
[0084] At least one embodiment of the task model is based on robust logic-dynamic systems. This represents a task as a list of reactive operators o with certain properties. Each operator is a tuple o = {LP, LR, LE, π}, where LP is a set of logical preconditions to enter o, LR is a set of run conditions that must be held while execution of o proceeds, and LE is a set of logical effects that will be true. The operator is also associated with a policy π that generates the necessary controls to achieve the effects LE. In at least one embodiment, the policies and predicates are learned from data, but in other embodiments they are manually specified. Given a plan, at least one embodiment selects the highest priority operator that satisfies the preconditions, checking the conditions at 10hz so that it can quickly respond to changes.
[0085] Figure 9 An example of a robot-human interaction is shown in accordance with an embodiment. Figure 9 An overview of the different steps in the final task plan is given. In at least one embodiment, the system must adapt to different possible grasps, reactively choosing the right way to approach the human user and take the object from them. In one example, it will hold in the "home" position and wait until it gets a stable estimate of how the human wants the object to be presented.
[0086] In various embodiments, the handoff from human to robot is done in four stages, where the robot waits for the human to present the object in the appropriate pose 902, then makes a plan to position the robot gripper so that it can grasp the object 904, grasps the object 906, and then in some examples, drops the object 908.
[0087] Some implementations do not just use reactive local planning, but plan and make intelligent decisions based on a large number of possible grasps in order to find a grasp that is natural for the human user. The table below shows one embodiment of a task plan, ordered by priority in descending order. In the table below, the operators and corresponding preconditions LP for task execution and reactive execution are shown. The operators are listed in descending order of priority; if all preconditions are true, the associated operator will be executed by the techniques described herein regardless of what operator was executed previously.
[0088]
[0089] Wait for human. At least one embodiment computes several predicates that determine how the robot should interact with the hand: stable, hand_over_table, hand_has_obj, and too_close_to_hand. The hand_over_table predicate corresponds to whether the observations are within the specified body of the table described above, stable() is true if the hand has not moved and the hand has been observed for at least 5 time steps (0.5 seconds). For a threshold λ, this is defined based at least in part on the velocity with position x and time t:
[0090] stable() = ||xt-1 - xt||2 < λ
[0091] If these conditions are not true, the robot will wait in place.
[0092] Avoid human. In at least one embodiment, if too_close_to_hand() is true for either hand and the robot is not in the near zone corresponding to the particular grasp, the robot will attempt to avoid the hand and will move back to the home position. At least one embodiment defines too_close_to_hand() to be true if the Euclidean distance between the end effector and the hand is less than 20 cm.
[0093] Find a feasible goal. To ensure that the robot's motion is safe, rather than a purely reactive policy, the techniques described herein plan an entire trajectory for execution. If it is stable(), hand_over_table(), and hand_has_obj(), then the robot will attempt to use the grasp pose described above to take the object from the hand.
[0094] In at least one embodiment, to find a valid trajectory ξ, the robot first finds a valid grasp pose, so the system adds has_goal and is_goal_valid. If either of these is false, the system searches for a reasonable goal pose.
[0095] In at least one embodiment, the planner creates a list of goal pose candidates and associated standoff positions. In one example, there are ten options, around Figure 3the y-axis in the world frame by θ_y ∈ {-π / 4, -π / 8, 0, π / 8, π / 4} and a rotation around the z-axis by θ_z ∈ {0, π}. In at least one embodiment, both the grasp position and the opposition position must be collision-free and have a valid IK solution in order to be considered a viable goal option. Various examples also add a constraint that the robot should not occlude its view of the object when determining whether a state is valid.
[0096] Find a plan. In at least one embodiment, if the planner has a list of goal options, it will order them by their distance from the current joint configuration and attempt to find a motion plan to the opposition position using RRT-Connect
[58] . If the system can find both a grasp pose and a motion plan, the robot executes a sub-policy to follow this motion plan.
[0097] However, the human can move their hand or change the way they are holding the object. If the goal has an associated motion plan and the object has not moved more than some threshold from the time it was first observed, the goal is considered valid (according to the is goal valid predicate). If the object moves too much, the robot stops and the task model transitions back to finding a new grasp.
[0098] Grasp the object. In at least one embodiment, once the motion plan is complete, the robot should be in the opposition pose and have an associated goal pose - the expected position of the object in the human’s hand. These two poses define a near zone - a cone in which the robot can move to approach the object. Once the gripper is closed, the has obj predicate is set to true if the robot is in its goal pose. The grasping operator can occlude the object, so this is only performed when blocking, open-loop actions.
[0099] Open the gripper. In at least one embodiment, if has obj is true, indicating that the robot believes it is holding an object, this perception can be incorrect in the case that the object moves or the pose estimate is incorrect. Some examples add a gripper fully closed predicate, indicating that the gripper has been closed the entire time. If both conditions are true, has obj is set to false and the robot reverts to a different state.
[0100] Move to drop-off and drop the object. In at least one embodiment, the drop-off position is a single joint space position; the robot finds a safe, collision-free motion plan to it. If it is at the drop-off position, the system instructs the robot to open the gripper and place the object on the table.
[0101] One example of our entire system is in Figure 7A series of different hand positions and grasps were systematically tested, including the above classification model and task model. One embodiment used two different Franka Panda robots mounted on the same table in different locations. A human user handed the robots four colored blocks, one at a time. During the system evaluation, each of three methods for determining which grasp pose to use to take the object from the human was tested:
[0102] Simple baseline: Wait until it sees the block in the human’s hand and take it from the hand using a fixed grasp orientation. The human hand was detected via a Microsoft Azure body tracker.
[0103] Hand pose estimation: A state-estimation-based version of the system, where the human hand pose from the Azure body tracker was used to infer the grasp orientation.
[0104] Embodiment of the proposed system: The proposed system, which classifies human grasps based on depth information as described above.
[0105] Variants performed the same task model, as described above. The order in which the three test cases were provided was random. The user presented the blocks to the robot with the right hand.
[0106] System performance was evaluated through a set of metrics computed during the trial. These were computed and logged automatically as the user performed the task.
[0107] Planning success rate: The number of times the follow_plan operator was able to successfully execute a trajectory that brought the robot to its standoff pose, and measures the certainty of both the human and the system.
[0108] Grasp success rate: The frequency with which the robot successfully took the object from the human, as a ratio of the total number of times it attempted to grasp.
[0109] Action execution time: The time taken to execute a single planned trajectory, grasp the block, and place it on the table. This value is higher if the robot had to take a longer path to grab the block from the human hand.
[0110] Total execution time: The amount of time taken to execute all planned trajectories, including the time spent replanning due to human movement or due to changes in grasp approach.
[0111] Trial duration: The time from the first detection of the human hand to the completion of the trial.
[0112] Figure 10Results for various metrics during system evaluation are shown. The techniques described herein consistently improve success rates and reduce total execution time and trial duration compared to the other two baselines, demonstrating the effectiveness and reliability of the approach. In a first graph 1002, accuracy of human grasp classification is shown. In a second graph 1004, a comparison of object miss rate between our hand state classification and PoseCNN is illustrated. In many cases, the hand will occlude the object, which means it is difficult to obtain an accurate pose estimate.
[0113] One exception is the action execution time, where the simple baseline is sometimes faster because the simple baseline does not adaptively plan like the other baselines; it will not attempt unusual grasps. This means that the time from a successful approach to discarding the object can sometimes be significantly shortened on average.
[0114] Evaluation of human grasp classification: One embodiment evaluates the hand grasp classification model on a validation set that was collected with unseen subjects during the training process. Classification accuracy is reported in the first graph 1002, which indicates that our model has good generalization capabilities to unseen subjects.
[0115] In addition, an experiment was conducted to evaluate the detection rate, i.e., whether there is an object in the hand, to understand the robustness of the handoff system to occlusions. The detection rate of our hand grasp classification model (object present / absent) was compared to the detection rate of another object detection method. The results are reported in the second graph 1004. One embodiment of the human grasp model described herein achieves a higher detection rate and is more robust compared to the alternative, especially when severe occlusions occur (e.g., 87.5% vs. 6.8% for pinch side, and 94.4% vs. 11.9% for pick up).
[0116] A user study was conducted to verify whether the system allows for a smooth human-robot collaboration. Nine users were recruited, aged between 20 and 36 years. Two females and seven males. The average age was 30.44 ± 4.74 years. The study included three rounds:
[0117] Freeform: The users were given four blocks and instructed to stand in front of the table and hand over the blocks to the robot one at a time. They were told that the robot would only take the blocks when their hands were still, but they could hold the blocks in any way they liked.
[0118] Note: The five human grasps shown in FIG. 1 1002 were presented Figure 3 The participants were asked to hand over the four blocks again. They were encouraged to try the predefined hand grasps, but they could also use any other way.
[0119] Distraction: User performance was tested in the presence of distraction.
[0120] Our quantitative metrics of handoff performance results are shown in the table below. The plan success rate indicates how often the system needed to re-plan its approach, while the grasp success rate is the number of times the system successfully took the object.
[0121]
[0122] In addition to the above metrics, the following statistics were calculated during the user study: (a) the number of times the robotic hand touched a human finger, (b) the number of times the user changed the grasp they were using, and (c) the number of times they changed the position of their hand. After each trial, participants were asked to describe any problems they encountered in handing over the object to the robot. After completing all three trials, participants were asked to fill out a Likert scale questionnaire and explain their answers.
[0123] There was a range of responses, but users indicated that they were able to fluidly collaborate with the robot and trust that it would do the right thing, despite pointing out several common issues when asked to provide feedback. They also trusted that the robot knew their actions.
[0124] Quantitative metrics of user data are shown in the table above. Approach and grasp were less successful when the user was distracted, but the time was similar. Users counted an average of 12:88±3:48 faces in the music video, while the correct number was 13. This means that many of them had some degree of confidence in the handoff system and were very focused on the video.
[0125] In one embodiment, the classification of human grasps described herein covers 77% of typical user grasps. At least one embodiment can handle most unseen human grasps; they tend to result in higher uncertainty, sometimes causing the robot to back off and re-plan. Some of these unseen grasps are shown in Figure 11 Figure 11 An example of an unusual grasp that did not appear in our training dataset is shown, and is an example of the type of grasp for which our system exhibited higher uncertainty, which resulted in slightly worse handoff performance.
[0126] During the final distraction test, users had to reposition or change their grasp more frequently compared to the first two rounds. Some complained that their fingers were pinched or saw the robot fail to grasp the object. One in particular "chose to use a palm-up hand pose to minimize the risk of failure"; another "had to look at the robot every 10 seconds or so."
[0127] The table below provides quantitative results from the user study. Even when users were distracted and had to focus on a different scenario, they were able to quickly complete the tasks.
[0128]
[0129]
[0130] At least in one example, the system is more clear about which blocks the robot wants to move to and how it wants to get there.
[0131] In general, during the testing, users quickly noticed that the robot was trying to grasp the blocks in a way that was unobtrusive. They also sometimes noticed slight inaccuracies in the robot’s grasp and approach, but generally adjusted on their own. After the second round of experiments, when shown how to grasp the objects, the system was more reliable and easier to use.
[0132] This document describes embodiments of a system that enables human-robot handovers by classifying different types of grasps. Other embodiments produce grasps from point clouds of human hands holding objects by training neural networks, making the planning system more flexible and supporting general grasp types. Variations of these techniques can be applied to many other types of human-robot collaboration, such as medical procedures, manufacturing, and personal care work.
[0133] Figure 12 An example of a process for transferring an object between a robot and a human hand as a result of execution by a computer system is shown in accordance with an embodiment. At least one embodiment can be implemented using a computer system, processor, GPU, or machine learning network, such as those shown in FIGS. 13-41, and described in related descriptions. At least one embodiment is implemented using a processor that reads executable instructions from a computer-readable storage that stores the instructions, as a result of execution of the instructions by one or more processors of the computer system, cause the computer system to perform the following operations.
[0134] In at least one embodiment, at block 1202, an image of a hand holding an object is obtained using a depth camera. In various examples, the depth camera can be a binocular RGB camera, such as a medical imaging device such as a three-dimensional x-ray, ultrasound, CAT scan, or magnetic resonance imaging (“MRI”). In some examples, such as those involving an autonomous vehicle, the image can be generated using a radar imaging device or a laser imaging device (“LIDAR”). In at least one embodiment, at block 1204, the system generates a point cloud from the image. The point cloud provides a set of three-dimensional points representing the object and the hand. In one example, the point cloud is a color point cloud from an RGB depth camera.
[0135] In at least one example, at block 1206, the system processes the point cloud and identifies a first portion of the point cloud that represents a hand holding the object. At block 1208, the system identifies a second portion of the point cloud that represents the object. Using the appropriate portions of the point cloud, at block 1210, in at least one example, the pose of the object is determined from the second portion of the point cloud and the pose of the hand is determined 1212 from the first portion of the point cloud. In one example, the pose of the hand includes identifying the skeletal structure of the hand, which includes the joint and segment lengths of the fingers of the hand.
[0136] In at least one embodiment, at block 1214, the system generates a set of grasping poses for the robotic gripper that will allow the gripper to grasp the object. This can be accomplished using the portions of the point cloud that represent the object and / or the object pose information determined at block 1210. In at least one embodiment, the object pose information includes the size, shape, and orientation of the object in space. The orientation of the object can include vertical and horizontal rotations and tilts. The set of grasping poses can include many that would interfere with the hand of a person, but are still preferable if the hand were not present.
[0137] Accordingly, at block 1216, the system identifies a particular grasp from the set of grasping poses that does not interfere with the hand. Interfering with the hand means contacting the hand or pinching the hand or a portion of the hand with the robotic gripper. In at least one embodiment, the system selects a grasp that satisfies adequate safety criteria for the object while maximizing the distance from the fingers of the hand. After identifying an appropriate grasp that does not contact the hand, at block 1218, the system instructs the robot to perform the identified grasp to remove the object from the hand of the person.
[0138] Inference and training logic
[0139] Figure 13A Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Inference and / or training logic 1315 are used to process electrical signals received from one or more sensors (e.g., microphones, cameras, infrared sensors, etc.). Inference and / or training logic 1315 can process the electrical signals to generate information about the environment of the system. Figure 13A And / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B.
[0140] In at least one embodiment, inference and / or training logic 1315 can include, without limitation, code and / or data storage 1301 for storing forward and / or output weights and / or input / output data, and / or other parameters configuring neurons or layers of a neural network being trained and / or used for inferencing, in aspects of one or more embodiments. In at least one embodiment, training logic 1315 can include or be coupled to code and / or data storage 1301 for storing graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, code and / or data storage 1301 stores weight parameters and / or input / output data for each layer of a neural network being trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 1301 can be included with other on-chip or off-chip data storage, including a processor’s Ll, L2, or L3 cache or system memory.
[0141] In at least one embodiment, any portion of code and / or data storage 1301 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 1301 can be cache memory, dynamic random addressable memory (“DRAM”), static random addressable memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 1301 is internal or external to a processor, e.g., or comprised of DRAM, SRAM, flash or some other storage type, can depend on available storage space to store on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.
[0142] In at least one embodiment, inference and / or training logic 1315 can include, without limitation, code and / or data storage 1305 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 1305 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, training logic 1315 can include or be coupled to code and / or data storage 1305 for storing graph code or other software to control timing and / or order, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, any portion of code and / or data storage 1305 can be included with other on-chip or off-chip data storage, including a processor’s L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 1305 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 1305 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., Flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 1305 is internal or external to a processor, e.g., or made up of DRAM, SRAM, Flash memory, or some other storage type, can depend on whether available storage is on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data used in inference and / or training of a neural network, or some combination of these factors.
[0143] In at least one embodiment, code and / or data storage 1301 and code and / or data storage 1305 can be separate storage structures. In at least one embodiment, code and / or data storage 1301 and code and / or data storage 1305 can be the same storage structure. In at least one embodiment, code and / or data storage 1301 and code and / or data storage 1305 can be partly the same storage structure and partly separate storage structures. In at least one embodiment, any portion of code and / or data storage 1301 and code and / or data storage 1305 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.
[0144] In at least one embodiment, inference and / or training logic 1315 can include, without limitation, one or more arithmetic logic units (“ALUs”) 1310 (including integer and / or floating point units) for performing logical and / or mathematical operations based, at least in part, on training and / or inference code (e.g., graph code) or instructions therefrom. Results of such operations can result in activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 1320, which are functions of input / output and / or weight parameter data stored in code and / or data storage 1301 and / or code and / or data storage 1305. In at least one embodiment, activations stored in activation storage 1320 are generated by ALUs 1310 in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALUs 1310, where weight values stored in code and / or data storage 1305 and / or code and / or data storage 1301 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which can be stored in code and / or data storage 1305 or code and / or data storage 1301 or other on-chip or off-chip storage.
[0145] In at least one embodiment, one or more ALUs 1310 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 1310 may be located outside the processor or other hardware logic device or the circuitry that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 1310 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, code and / or data storage 1301, code and / or data storage 1305, and activation storage 1320 may be on the same processor or other hardware logic device or circuitry, while in another embodiment, they may be in different processors or other hardware logic devices or circuitries, or in some combination of the same and different processors or other hardware logic devices or circuitries. In at least one embodiment, any portion of activation storage 1320 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.
[0146] In at least one embodiment, the active memory 1320 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 1320 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 1320 is internal to or external to the processor may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of the data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or other memory types. In at least one embodiment, Figure 13A The inference and / or training logic 1315 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 13AThe illustrated inference and / or training logic 1315 can be used in combination with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware (e.g., field programmable gate arrays (“FPGAs”)).
[0147] Figure 13B Inference and / or training logic 1315 are illustrated as a portion of the system 1301 in the example of FIG. 13. In other examples, the inference and / or training logic 1315 can be a separate and distinct component from the system 1301. Figure 13B The inference and / or training logic 1315 illustrated in FIG. 13 can be used in combination with application-specific integrated circuit (ASIC) hardware such as Google’s Tensor Processing Unit (TPU), Graphcore’s AI processing units from Graphcore’s AI TM Inference Processing Unit (IPU) from Intel Corp., or Intel’s “Lake Crest” processor. In at least one embodiment, the inference and / or training logic 1315 illustrated in FIG. 13 can be used in combination with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as field programmable gate arrays (FPGAs). Figure 13B In at least one embodiment, the inference and / or training logic 1315 includes, without limitation, code and / or data storage 1301 and code and / or data storage 1305, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment, each of the code and / or data storage 1301 and the code and / or data storage 1305 are respectively associated with dedicated computing resources (e.g., computing hardware 1302 and computing hardware 1306). In at least one embodiment, each of the computing hardware 1302 and the computing hardware 1306 includes one or more ALUs that only perform mathematical functions (e.g., linear algebraic functions) on information stored in the code and / or data storage 1301 and the code and / or data storage 1305, respectively, and results of performing the functions are stored in activation storage 1320. Figure 13B
[0148] In at least one embodiment, each of code and / or data stores 1301 and 1305 and corresponding compute hardware 1302 and 1306, respectively, correspond to different layers of a neural network, such that activations resulting from one “storage / compute pair 1301 / 1302” of code and / or data store 1301 and compute hardware 1302 provide input to next “storage / compute pair 1305 / 1306” of code and / or data store 1305 and compute hardware 1306, in order to reflect a conceptual organization of a neural network. In at least one embodiment, each storage / compute pair 1301 / 1302 and 1305 / 1306 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in inference and / or training logic 1315 after or in parallel with storage compute pairs 1301 / 1302 and 1305 / 1306.
[0149] Neural network training and deployment
[0150] Figure 14 Training and deployment of a deep neural network is shown, in accordance with at least one embodiment. In at least one embodiment, an untrained neural network 1406 is trained using a training dataset 1402. In at least one embodiment, training framework 1404 is a PyTorch framework, while in other embodiments, training framework 1404 is a TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, training framework 1404 trains untrained neural network 1406 and enables it to train using processing resources described herein to generate a trained neural network 1408. In at least one embodiment, weights can be chosen at random or by pre-training using a deep belief network. In at least one embodiment, training can be performed in a supervised, partially supervised, or unsupervised manner.
[0151] In at least one embodiment, an untrained neural network 1406 is trained using supervised learning, where a training dataset 1402 includes inputs paired with desired outputs for inputs, or where a training dataset 1402 includes inputs with known outputs and the output of the neural network 1406 is manually graded. In at least one embodiment, an untrained neural network 1406 is trained in a supervised manner and processes an input from a training dataset 1402 and compares a resulting output to a set of expected or desired outputs. In at least one embodiment, an error is then propagated back through the untrained neural network 1406. In at least one embodiment, a training framework 1404 adjusts weights that control the untrained neural network 1406. In at least one embodiment, a training framework 1404 includes tools to monitor how well an untrained neural network 1406 is converging towards a model (e.g., a trained neural network 1408) that is suitable for generating correct answers (e.g., results 1414) based on input data (e.g., new datasets 1412). In at least one embodiment, a training framework 1404 repeatedly trains an untrained neural network 1406 while adjusting weights to improve outputs of the untrained neural network 1406 using a loss function and an adjustment algorithm (e.g., stochastic gradient descent). In at least one embodiment, a training framework 1404 trains an untrained neural network 1406 until the untrained neural network 1406 reaches a desired accuracy. In at least one embodiment, a trained neural network 1408 can then be deployed to implement any number of machine learning operations.
[0152] In at least one embodiment, an untrained neural network 1406 is trained using unsupervised learning, where an untrained neural network 1406 attempts to train itself using unlabeled data. In at least one embodiment, an unsupervised learning training dataset 1402 will include input data without any associated output data or “ground truth” data. In at least one embodiment, an untrained neural network 1406 can learn groupings within a training dataset 1402 and can determine how individual inputs relate to the untrained dataset 1402. In at least one embodiment, unsupervised training can be used to generate a self-organizing map, which is a type of trained neural network 1408 that can perform operations useful for reducing a dimensionality of new data 1412. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for identification of data points in a new dataset 1412 that deviate from a normal pattern of the new dataset 1412.
[0153] In at least one embodiment, semi-supervised learning can be used, which is a technique in which a mix of labeled and unlabeled data is included in training dataset 1402. In at least one embodiment, training framework 1404 can be used to perform incremental learning, for example, through a transferred learning technique. In at least one embodiment, incremental learning enables trained neural network 1408 to adapt to new data 1412 without forgetting knowledge that was imprinted into network during initial training.
[0154] data center
[0155] Figure 15 An example data center 1500 that can use at least one embodiment is shown. In at least one embodiment, data center 1500 includes a data center infrastructure layer 1510, a framework layer 1520, a software layer 1530, and an application layer 1540.
[0156] In at least one embodiment, as Figure 15 shown, data center infrastructure layer 1510 can include a resource orchestrator 1512, grouped computing resources 1514, and node computing resources (“node C.R.s”) 1516(1)-1516(N), where “N” represents a positive integer. In at least one embodiment, node C.R.s 1516(1)-1516(N) can include, but are not limited to, any number of central processing units (“CPUs” or “processors”), including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc., memory devices (e.g., dynamic random access memory), storage devices (e.g., solid state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more of node C.R.s 1516(1)-1516(N) can be a server having one or more of the above computing resources.
[0157] In at least one embodiment, grouped computing resources 1514 can include individual groups of node C.R.s housed within one or more racks (not shown), or housed within a number of racks (also not shown) within data centers at various geographic locations. Individual groups of node C.R.s within grouped computing resources 1514 can include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R.s including CPUs or processors can be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks can also include any quantity of power modules, cooling modules, and network switches, in any combination.
[0158] In at least one embodiment, resource coordinator 1512 may be configured or otherwise control one or more nodes CR1516(1)-1516(N) and / or grouped computing resources 1514. In at least one embodiment, resource coordinator 1512 may include a Software Design Infrastructure (“SDI”) management entity for data center 1500. In at least one embodiment, resource coordinator may include hardware, software, or some combination thereof.
[0159] In at least one embodiment, such as Figure 15 As shown, framework layer 1520 includes a job scheduler 1532, a configuration manager 1534, a resource manager 1536, and a distributed file system 1538. In at least one embodiment, framework layer 1520 may include a framework of software 1532 supporting software layer 1530 and / or one or more applications 1542 supporting application layer 1540. In at least one embodiment, software 1532 or application 1542 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 1520 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 1538 for large-scale data processing (e.g., "big data"). TM (Hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 1532 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of data center 1500. In at least one embodiment, the configuration manager 1534 may be able to configure different layers, such as software layer 1530 and framework layer 1520 including Spark and a distributed file system 1538 for supporting large-scale data processing. In at least one embodiment, the resource manager 1536 is able to manage cluster or group computing resources mapped to or allocated to support distributed file system 1538 and job scheduler 1532. In at least one embodiment, cluster or group computing resources may include group computing resources 1514 on data center infrastructure layer 1510. In at least one embodiment, the resource manager 1536 may coordinate with resource coordinator 1512 to manage these mapped or allocated computing resources.
[0160] In at least one embodiment, software 1532 included in software layer 1530 can include software used by at least portions of node C.R.s 1516(l)-1516(N), grouped computing resources 1514, and / or distributed file system 1538 of framework layer 1520. One or more types of software can include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0161] In at least one embodiment, one or more applications 1542 included in application layer 1540 can include one or more types of applications used by at least portions of node C.R.s 1516(l)-1516(N), grouped computing resources 1514, and / or distributed file system 1538 of framework layer 1520. One or more types of applications can include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0162] In at least one embodiment, any of configuration manager 1534, resource manager 1536, and resource orchestrator 1512 can implement any number and type of self-modifying actions based on any amount and type of data obtained in any technologically feasible manner. In at least one embodiment, self-modifying actions can relieve data center operators of data center 1500 from making possibly poor configuration decisions and can avoid underutilized and / or poorly performing portions of a data center.
[0163] In at least one embodiment, data center 1500 can include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information in accordance with one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by computing weight parameters according to a neural network architecture using software and computing resources described above with respect to data center 1500. In at least one embodiment, using weight parameters computed by one or more training techniques described herein, a trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using resources described above with respect to data center 1500.
[0164] In at least one embodiment, data center can use CPUs, application specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using resources described above. Moreover, one or more software and / or hardware resources described above can be configured as a service to allow users to train or perform inferencing of information, such as image recognition, speech recognition, or other artificial intelligence services.
[0165] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Inferences and / or determinations can be made based, at least in part, on the training of neural networks, neural network functions, and / or architecture, or weight parameters computed using neural network use cases described herein. Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided. In at least one embodiment, inference and / or training logic 1315 can be used in a system Figure 15 to infer or predict operations based, at least in part, on weight parameters computed using neural network training operations, neural network functions, and / or architecture, or neural network use cases described herein.
[0166] The techniques described above can be used, for example, to implement a system for performing human-robot object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0167] Autonomous vehicle
[0168] Figure 16A An example of an autonomous vehicle 1600 is shown, in accordance with at least one embodiment. In at least one embodiment, autonomous vehicle 1600 (alternatively referred to herein as “vehicle 1600”) can be, but is not limited to, a passenger vehicle such as a car, truck, bus, and / or another type of vehicle that can accommodate one or more passengers. In at least one embodiment, vehicle 1600 can be a semi-truck tractor-trailer used to haul cargo. In at least one embodiment, vehicle 1600 can be an airplane, a robotic vehicle, or another type of vehicle.
[0169] Autonomous vehicles can be described in terms of automation levels defined by the National Highway Traffic Safety Administration (“NHTSA”), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (“SAE”) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (e.g., Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and previous and future versions of this standard). In one or more embodiments, vehicle 1600 can be capable of functioning according to one or more of Levels 1 through 5 of the automated driving levels. For example, in at least one embodiment, vehicle 1600 can be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.
[0170] Various embodiments of human-robot interaction systems can be integrated into vehicles to assist with tasks such as package delivery and warehouse automation. For example, one embodiment can be used to receive a package from a person for delivery, or to receive a physical cash payment from a customer.
[0171] In at least one embodiment, vehicle 1600 can include, without limitation, components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. In at least one embodiment, vehicle 1600 can include, without limitation, a propulsion system 1650, such as a combustion engine, a hybrid electric device, a fully electric motor, and / or another type of propulsion system. In at least one embodiment, propulsion system 1650 can be connected to a drivetrain of vehicle 1600, which can include, without limitation, a transmission to enable propulsion of vehicle 1600. In at least one embodiment, propulsion system 1650 can be controlled in response to receiving signals from throttle / accelerator 1652.
[0172] In at least one embodiment, steering system 1654 (which can include, without limitation, a steering wheel) is used to steer vehicle 1600 (e.g., along a desired path or route) when propulsion system 1650 is operating (e.g., when the vehicle is in motion). In at least one embodiment, steering system 1654 can receive signals from steering actuator 1656. A steering wheel can be optional for full automation (Level 5) functionality. In at least one embodiment, brake sensor system 1646 can be used to operate vehicle brakes in response to signals received from brake actuator 1648 and / or brake sensors.
[0173] In at least one embodiment, controller 1636 may include, but is not limited to, one or more system-on-chips (“SoCs”). Figure 16A A controller 1636 (not shown) and / or a graphics processing unit (“GPU”) provides signals (e.g., signals representing commands) to one or more components and / or systems of vehicle 1600. For example, in at least one embodiment, controller 1636 may send signals to operate vehicle braking via brake actuator 1648, steering system 1654 via steering actuator 1656, and propulsion system 1650 via one or more throttles / accelerators 1652. One or more controllers 1636 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a driver in driving vehicle 1600. In at least one embodiment, one or more controllers 1636 may include a first controller 1636 for autonomous driving functions, a second controller 1636 for functional safety functions, a third controller 1636 for artificial intelligence functions (e.g., computer vision), a fourth controller 1636 for infotainment functions, a fifth controller 1636 for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller 1636 may handle two or more of the functions described above, and two or more controllers 1636 may handle a single function and / or any combination thereof.
[0174] In at least one embodiment, one or more controllers 1636 provide signals for controlling one or more components and / or systems of vehicle 1600 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data can be received from sensors, including but not limited to one or more Global Navigation Satellite System (“GNSS”) sensors 1658 (e.g., one or more Global Positioning System sensors), one or more RADAR sensors 1660, one or more ultrasonic sensors 1662, one or more LIDAR sensors 1664, one or more inertial measurement unit (IMU) sensors 1666 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetic compasses, one or more magnetometers, etc.), one or more microphones 1696, one or more stereo cameras 1668, one or more wide-angle cameras 1670 (e.g., fisheye cameras), one or more infrared cameras 1672, one or more surround cameras 1674 (e.g., 360-degree cameras), and remote cameras (…). Figure 16A (not shown in the image), medium-range camera ( Figure 16Aone or more speed sensors 1644 (e.g., to measure a speed of the vehicle 1600), one or more vibration sensors 1642, one or more steering sensors 1640, one or more brake sensors (e.g., as part of a brake sensor system 1646), and / or other sensor types.
[0175] In at least one embodiment, the one or more controllers 1636 can receive input (e.g., represented by input data) from an instrument cluster 1632 of the vehicle 1600 and provide output (e.g., represented by output data, display data, etc.) through a human-machine interface (“HMI”) display 1634, audible annunciators, speakers, and / or other components of the vehicle 1600. In at least one embodiment, output can include information such as vehicle speed, velocity, time, map data (e.g., high-definition map Figure 16A In at least one embodiment, the HMI display 1634 can display information about the presence of one or more objects (e.g., a street sign, a warning sign, a traffic signal change, etc.) and / or information about a driving maneuver that the vehicle has made, is making, or will make (e.g., change lanes now, take exit 34B in two miles, etc.). In at least one embodiment, the HMI display 1634 can also display information about the vehicle’s surroundings (e.g., as shown in FIG. 16B), such as a map of the vehicle’s surroundings, a location of the vehicle 1600 (e.g., on a map), a direction of the vehicle 1600, a location of other vehicles (e.g., an occupancy grid), information about objects, a state of objects perceived by the one or more controllers 1636, etc.
[0176] In at least one embodiment, the vehicle 1600 further includes a network interface 1624 that can communicate over one or more networks using one or more wireless antennas 1626 and / or one or more modems. For example, in at least one embodiment, the network interface 1624 can be capable of communicating over Long-Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”), etc. In at least one embodiment, the one or more wireless antennas 1626 can also use one or more local area networks (e.g., Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or one or more low power wide area networks (hereinafter “LPWANs”) (e.g., LoRaWAN, SigFox, etc.) to communicate between objects (e.g., vehicles, mobile devices) in an environment.
[0177] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Inference and / or training logic 1315 are used to process electrical signals received from one or more sensors (e.g., image sensors, microphones, etc.), and provide appropriate signals to one or more output devices (e.g., display devices, speakers, etc.). Figure 13A and / orFigure 13B Details regarding the inference and / or training logic 1315 are provided. In at least one embodiment, the inference and / or training logic 1315 can be implemented in the system. Figure 16A The operation is used to infer or predict the operation based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0178] The techniques described above can be used, for example, to implement systems for performing human-machine object handover. Some examples use inference and / or training logic to create neural networks that are trained to generate grasping of objects held by a human hand as described above.
[0179] Figure 16B The illustration shows an embodiment according to at least one of the embodiments. Figure 16A Examples of camera positions and fields of view for an autonomous vehicle 1600. In at least one embodiment, the camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 1600.
[0180] In at least one embodiment, the camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 1600. One or more cameras may operate at Automotive Safety Integrity Level (“ASIL”) B and / or other ASILs. In at least one embodiment, the camera type may have any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc. In at least one embodiment, the camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-to-clear (“RCCC”) color filter array, a red-to-clear-blue (“RCCB”) color filter array, a red-blue-green (“RBGC”) color filter array, a Foveon X3 color filter array, a Bayer sensor (“RGGB”) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a transparent pixel camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used to improve photosensitivity.
[0181] In at least one embodiment, one or more cameras can be used to perform advanced driver assistance system (“ADAS”) functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-functional mono camera can be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. In at least one embodiment, one or more cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0182] In at least one embodiment, one or more cameras can be mounted in mounting assemblies, such as custom designed (three-dimensional (“3D”) printed) assemblies, in order to cut out stray light and reflections from within a car (e.g., dashboard reflections in rearview mirror reflections), which can interfere with a camera’s image data capture capabilities. With regard to rearview mirror mounting assemblies, in at least one embodiment, a rearview mirror assembly can be 3D printed custom designed such that a camera mounting plate matches a shape of a rearview mirror. In at least one embodiment, one or more cameras can be integrated into a rearview mirror. In at least one embodiment, for side view cameras, one or more cameras can also be integrated within four pillars at each corner of a cabin.
[0183] In at least one embodiment, a camera with a field of view that includes a portion of an environment in front of vehicle 1600 (e.g., a front-facing camera) can be used for surround view, as well as to help identify a path and obstacles ahead with the help of one or more controllers 1636 and / or a control SoC, thereby providing information that is critical to generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, a front-facing camera can be used to perform many of the same ADAS functions as a LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, a front-facing camera can also be used for ADAS functions and systems including, but not limited to, lane departure warning (“LDW”), automatic cruise control (“ACC”), and / or other functions (e.g., traffic sign recognition).
[0184] In at least one embodiment, various cameras can be used in a front-facing configuration, including, for example, a monocular camera platform including a CMOS (“complementary metal-oxide semiconductor”) color imager. In at least one embodiment, a wide-view camera 1670 can be used to perceive objects (e.g., pedestrians, crossing or bicycles) entering from the periphery. Although in Figure 16BOnly one wide-angle camera 1670 is shown; however, in other embodiments, the vehicle 1600 may have any number (including zero) of wide-angle cameras 1670. In at least one embodiment, any number of remote cameras 1698 (e.g., a pair of remote stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not yet been trained. In at least one embodiment, the remote camera 1698 can also be used for object detection and classification, as well as basic object tracking.
[0185] In at least one embodiment, any number of stereo cameras 1668 may also be included in a forward configuration. In at least one embodiment, one or more stereo cameras 1668 may include an integrated control unit comprising a scalable processing unit that may provide programmable logic (“FPGA”) and a multi-core microprocessor with a controller area network (“CAN”) or Ethernet interface integrated on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the environment of the vehicle 1600, including distance estimates for all points in the image. In at least one embodiment, one or more stereo cameras 1668 may include, but are not limited to, a compact stereo vision sensor, which may include, but is not limited to, two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle 1600 to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 1668 may also be used in addition to those described herein.
[0186] In at least one embodiment, a camera (e.g., a side-view camera) having a field of view including a portion of the environment on the side of the vehicle 1600 can be used for surround viewing, thereby providing information for creating and updating the occupied grid, and generating a side collision warning. For example, in at least one embodiment, a surround camera 1674 (e.g., as...) Figure 16B The four surround cameras 1674 shown can be positioned on the vehicle 1600. One or more surround cameras 1674 can include, but are not limited to, any number and combination of wide-angle cameras 1670, one or more fisheye lenses, one or more 360-degree cameras, and / or similar cameras. For example, in at least one embodiment, the four fisheye lens cameras can be located at the front, rear, and sides of the vehicle 1600. In at least one embodiment, the vehicle 1600 can use three surround cameras 1674 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.
[0187] In at least one embodiment, a camera with a field of view that includes portions of an environment behind vehicle 1600 (e.g., a rearview camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating an occupancy grid. In at least one embodiment, a wide variety of cameras can be used, including but not limited to cameras that are also suitable as one or more forward facing cameras (e.g., long range cameras 1698 and / or one or more mid-range cameras 1676, one or more stereo cameras 1668, one or more infrared cameras 1672, etc.), as described herein.
[0188] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, inference and / or training logic 1315 can be used in Figure 13A and / or Figure 13B Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, inference and / or training logic 1315 can be used in Figure 16B a system to infer or predict operations based, at least in part, on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0189] The above-described techniques can be used, for example, to implement a system for performing human-robot handover. Some examples use inference and / or training logic to create a neural network trained to generate a grasp on an object held by a human hand as described above.
[0190] Figure 16C FIG. 16 illustrates an example system architecture of an autonomous vehicle 1600, in accordance with at least one embodiment. Figure 16A In at least one embodiment, each of one or more components, one or more features, and one or more systems of vehicle 1600 in Figure 16C In at least one embodiment, bus 1602 can include, without limitation, a CAN data interface (alternatively referred to herein as a “CAN bus”). In at least one embodiment, CAN can be a network within vehicle 1600 that helps control various features and functions of vehicle 1600, such as actuation of brakes, acceleration, braking, steering, wipers, etc. In one embodiment, bus 1602 can be configured with tens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). In at least one embodiment, bus 1602 can be read to find steering wheel angle, ground speed, engine revolutions per minute (“RPM”), button positions, and / or other vehicle status indicators. In at least one embodiment, bus 1602 can be an ASIL B compliant CAN bus.
[0191] In at least one embodiment, FlexRay and / or Ethernet can be used in addition to or instead of CAN. In at least one embodiment, there can be any number of form bus 1602, which can include, but are not limited to, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses 1602 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 1602 can be used for collision avoidance functions, and a second bus 1602 can be used for actuation control. In at least one embodiment, each bus 1602 can be in communication with any component of vehicle 1600, and two or more buses 1602 can be in communication with the same components. In at least one embodiment, each of any number of system on a chip (“SoC”) 1604, each of one or more controllers 1636, and / or each computer within a vehicle can have access to the same input data (e.g., input from sensors of vehicle 1600), and can be connected to a common bus, such as a CAN bus.
[0192] In at least one embodiment, vehicle 1600 can include one or more controllers 1636, such as those described herein with respect to Figure 16A In at least one embodiment, controllers 1636 can be used for a variety of functions. In at least one embodiment, controllers 1636 can be coupled to any of various other components and systems of vehicle 1600, and can be used to control vehicle 1600, artificial intelligence of vehicle 1600, infotainment of vehicle 1600, etc.
[0193] In at least one embodiment, vehicle 1600 can include any number of SoCs 1604. Each of SoCs 1604 can include, without limitation, central processing units (“one or more CPUs”) 1606, graphics processing units (“one or more GPUs”) 1608, one or more processors 1610, one or more caches 1612, one or more accelerators 1614, one or more data stores 1616, and / or other non- shown components and features. In at least one embodiment, one or more SoCs 1604 can be used to control vehicle 1600 in a variety of platforms and systems. For example, in at least one embodiment, one or more SoCs 1604 can be combined with a high definition (“HD”) map 1622 in a system (e.g., a system of vehicle 1600), which can obtain map refreshes and / or updates from one or more servers (not shown in FIG. 16) via a network interface 1624. Figure 16C
[0194] In at least one embodiment, CPU(s) 1606 can include a CPU cluster or CPU complex (alternatively referred to herein as a “CCPLEX”). In at least one embodiment, CPU(s) 1606 can include multiple cores and / or level two (“L2”) caches. For example, in at least one embodiment, CPU(s) 1606 can include eight cores in a multi-processor configuration coupled to one another. In at least one embodiment, CPU(s) 1606 can include four dual-core clusters with each cluster having a dedicated L2 cache (e.g., 2 MB L2 cache). In at least one embodiment, CPU(s) 1606 (e.g., CCPLEX) can be configured to support simultaneous cluster operation such that any combination of clusters of CPU(s) 1606 can be active at any given time.
[0195] In at least one embodiment, CPU(s) 1606 can implement power management functionality including, without limitation, one or more of the following features: individual hardware modules can be automatically clock-gated at idle to conserve dynamic power; each core clock can be gated when that core is not actively executing instructions due to execution of a wait for interrupt (“WFI”) / wait for event (“WFE”) instruction; each core can be independently powered; each core cluster can be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster can be independently power-gated when all cores are power-gated. In at least one embodiment, CPU(s) 1606 can further implement enhanced algorithms for managing power states with allowed power states and expected wake-up times specified and hardware / microcode determining optimal power states for core, cluster, and CCPLEX inputs. In at least one embodiment, processing cores can support a simplified power state input sequence in software with work offloaded to microcode.
[0196] In at least one embodiment, GPU(s) 1608 can include an integrated GPU (also referred to herein as an “iGPU”). In at least one embodiment, GPU(s) 1608 can be programmable and efficient for parallel workloads. In at least one embodiment, GPU(s) 1608 can use an enhanced tensor instruction set. In one embodiment, GPU(s) 1608 can include one or more streaming microprocessors, where each streaming microprocessor can include a level one (“LI”) cache (e.g., with at least 96 KB of storage capacity), and two or more streaming microprocessors can share an L2 cache (e.g., with 512 KB storage capacity). In at least one embodiment, GPU(s) 1608 can include at least eight streaming microprocessors. In at least one embodiment, GPU(s) 1608 can use a compute application programming interface (“API”). In at least one embodiment, GPU(s) 1608 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA).
[0197] In at least one embodiment, GPU(s) 1608 can be power-optimized to achieve best performance in automotive and embedded use cases. For example, in one embodiment, GPU(s) 1608 can be fabricated on a finned field effect transistor (“FinFET”). In at least one embodiment, each streaming microprocessor can contain a plurality of mixed-precision processing cores divided into a plurality of blocks. For example and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor Cores for deep learning matrix arithmetic, a level zero (“L0”) instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. In at least one embodiment, a streaming microprocessor can include independent parallel integer and floating point data paths to provide efficient execution of workloads that mix compute and address operations. In at least one embodiment, a streaming microprocessor can include independent thread scheduling capabilities to enable finer-grain synchronization and cooperation between parallel threads. In at least one embodiment, a streaming microprocessor can include a combined LI data cache and shared memory unit to improve performance while simplifying programming.
[0198] In at least one embodiment, one or more GPU(s) 1608 can include high bandwidth memory (“HBM”) and / or 16 GB HBM2 memory subsystems to provide, in some examples, a peak memory bandwidth of about 900 GB / sec. In at least one embodiment, in addition to, or instead of, HBM memory, synchronous graphics random access memory (“SGRAM”) can be used, for example, graphics double data rate type five synchronous random access memory (“GDDR5”).
[0199] In at least one embodiment, one or more GPU(s) 1608 can include unified memory technology. In at least one embodiment, address translation services (“ATS”) support can be used to allow one or more GPU(s) 1608 to directly access one or more CPU(s) 1606 page tables. In at least one embodiment, when a memory management unit (“MMU”) of a GPU of one or more GPU(s) 1608 experiences a miss, an address translation request can be sent to one or more CPU(s) 1606. In response, one or more CPU(s) 1606 can look up a virtual-to-physical mapping for an address in their page tables and transmit the translation back to one or more GPU(s) 1608, in at least one embodiment. In at least one embodiment, unified memory technology can allow a single unified virtual address space to be used for memory of both one or more CPU(s) 1606 and one or more GPU(s) 1608, simplifying programming of one or more GPU(s) 1608 and porting of applications to one or more GPU(s) 1608.
[0200] In at least one embodiment, one or more GPU(s) 1608 can include any number of access counters that can track how frequently one or more GPU(s) 1608 is accessing memory of other processors. In at least one embodiment, one or more access counters can help ensure that memory pages are moved into physical memory of a processor that most frequently accesses the page, improving efficiency of memory ranges shared between processors.
[0201] In at least one embodiment, one or more SoC(s) 1604 can include any number of caches 1612, including those described herein. For example, in at least one embodiment, one or more cache(s) 1612 can include a level three (“L3”) cache that can be available to one or more CPU(s) 1606 and one or more GPU(s) 1608 (e.g., connected to both CPU(s) 1606 and GPU(s) 1608). In at least one embodiment, one or more cache(s) 1612 can include a write-back cache that can track state of lines, for example, by using a cache coherence protocol (e.g., MESI, MSI, etc.). In at least one embodiment, although smaller cache sizes can be used, L3 cache can include 4MB or more, depending on embodiment.
[0202] In at least one embodiment, one or more SoC(s) 1604 can include one or more accelerator(s) 1614 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, one or more SoC(s) 1604 can include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4MB of SRAM) can enable hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, hardware acceleration cluster can be used to supplement and offload some tasks of one or more GPU(s) 1608 (e.g., freeing up more cycles of one or more GPU(s) 1608 to perform other tasks). In at least one embodiment, one or more accelerator(s) 1614 can be used for target workloads that are stable enough for the acceleration to be worth the silicon real estate (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.). In at least one embodiment, CNNs can include region-based or region with convolutional neural networks (“RCNNs”) and fast RCNNs (e.g., as used for object detection) or other types of CNNs.
[0203] In at least one embodiment, one or more accelerators 1614 (e.g., hardware acceleration clusters) can include a deep learning accelerator (“DLA”). One or more DLAs can include, without limitation, one or more Tensor Processing Units (“TPUs”) that can be configured to provide an additional 100 trillion operations per second for deep learning applications and inferencing. In at least one embodiment, a TPU can be an accelerator configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, one or more DLAs can be further optimized for a particular set of neural network types and floating point operations and inferencing. Design of one or more DLAs can provide higher performance per mm than a typical general purpose GPU, and often greatly exceeds performance of a CPU. In at least one embodiment, one or more TPUs can perform several functions including support for INT8, INT16, and FP16 data types for features and weights, single instance convolution functions, and post-processor functions, for example. In at least one embodiment, one or more DLAs can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data for any of a variety of functions including, for example and without limitation: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and identification and detection using data from microphones 1696; CNNs for face recognition and vehicle owner identification using data from camera sensors; and / or CNNs for safety and / or safety related events.
[0204] In at least one embodiment, a DLA can perform any of functions of GPU(s) 1608, and by using an inferencing accelerator, for example, a designer can target one or more DLAs or GPU(s) 1608 for any function. For example, in at least one embodiment, a designer can concentrate processing and floating point operations of a CNN on one or more DLAs, and leave other functions to GPU(s) 1608 and / or one or more other accelerators 1614.
[0205] In at least one embodiment, one or more accelerators 1614 (e.g., hardware acceleration clusters) can include programmable vision accelerators (“PVAs”), which can alternatively be referred to herein as computer vision accelerators. In at least one embodiment, one or more PVAs can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (“ADAS”) 1638, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. In at least one embodiment, one or more PVAs can strike a balance between performance and flexibility. For example, in at least one embodiment, each of one or more PVAs can include, without limitation, any number of reduced instruction set computer (“RISC”) cores, direct memory access (“DMA”), and / or any number of vector processors.
[0206] In at least one embodiment, RISC cores can interact with image sensors (e.g., image sensors of any camera described herein), image signal processors, etc. In at least one embodiment, each RISC core can include any number of memories. In at least one embodiment, RISC cores can use any of a number of protocols, depending on embodiment. In at least one embodiment, RISC cores can execute a real-time operating system (“RTOS”). In at least one embodiment, RISC cores can be implemented using one or more integrated circuit devices, application specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, RISC cores can include instruction caches and / or tightly coupled RAM.
[0207] In at least one embodiment, DMA can enable components of a PVA to access system memory independently of one or more CPUs 1606. In at least one embodiment, DMA can support any number of features for providing optimizations to a PVA, including but not limited to, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more dimensions of addressing, which can include, but are not limited to, block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.
[0208] In at least one embodiment, vector processors can be programmable processors that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, a PVA can include a PVA core and two vector processing subsystem partitions. In at least one embodiment, a PVA core can include a processor subsystem, DMA engines (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, a vector processing subsystem can function as a primary processing engine for a PVA and can include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, a VPU core can include a digital signal processor, such as a single instruction multiple data (“SIMD”), very long instruction word (“VLIW”) digital signal processor. In at least one embodiment, a combination of SIMD and VLIW can improve throughput and speed.
[0209] In at least one embodiment, each vector processor can include an instruction cache and can be coupled to a dedicated memory. As a result, in at least one embodiment, each vector processor can be configured to execute independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA can be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA can execute the same computer vision algorithm, except on different regions of an image. In at least one embodiment, vector processors included in a particular PVA can execute different computer vision algorithms on one image at a time, or even different algorithms on a sequence of images or portions of images. In at least one embodiment, any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each PVA, among other things. In at least one embodiment, a PVA can include additional error-correcting code (“ECC”) memory to enhance overall system security.
[0210] In at least one embodiment, one or more accelerators 1614 (e.g., hardware acceleration clusters) can include an on-chip computer vision network and static random access memory (“SRAM”) for providing high bandwidth, low latency SRAM for one or more accelerators 1614. In at least one embodiment, on-chip memory can include at least 4 MB of SRAM that includes, for example and without limitation, eight field-programmable memory blocks that are accessible by both PVA and DLA. In at least one embodiment, each pair of memory blocks can include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory can be used. In at least one embodiment, PVA and DLA can access memory via a backbone that provides PVA and DLA with high-speed access to memory. In at least one embodiment, a backbone can include an on-chip computer vision network that interconnects PVA and DLA to memory (e.g., using APB).
[0211] In at least one embodiment, an on-chip computer vision network can include an interface that determines that both PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transmission. In at least one embodiment, although other standards and protocols can be used, an interface can comply with International Organization for Standardization (“ISO”) 26262 or International Electrotechnical Commission (“IEC”) 61508 standards.
[0212] In at least one embodiment, one or more SoC 1604 can include a real-time line-of-sight tracking hardware accelerator. In at least one embodiment, a real-time line-of-sight tracking hardware accelerator can be used to quickly and efficiently determine locations and ranges of objects (e.g., within a world model) to generate real-time visualizations simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulations of SONAR systems, for general wave propagation simulations, for comparison with LIDAR data for positioning and / or other functions, and / or for other uses.
[0213] In at least one embodiment, one or more accelerators 1614 (e.g., hardware acceleration clusters) have broad utility for autonomous driving. In at least one embodiment, a PVA can be a programmable vision accelerator that can be used for key processing stages in ADAS and autonomous vehicles. In at least one embodiment, capabilities of a PVA at low power and low latency are well matched to algorithmic domains that require predictable processing. In other words, PVAs excel at semi-dense or dense regular computations, even on small data sets that require predictable run-time with low latency and low power. In at least one embodiment, autonomous vehicles such as vehicle 1600, PVAs are designed to run classic computer vision algorithms because they are efficient at object detection and integer math operations.
[0214] For example, in accordance with at least one embodiment of technology, a PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching based algorithm can be used in some examples, although this is not meant to be limiting. In at least one embodiment, applications for level 3-5 autonomous driving use dynamic estimation / stereo matching in run (e.g., structure from motion, pedestrian recognition, lane detection, etc.). In at least one embodiment, a PVA can perform computer stereo vision functions on inputs from two monocular cameras.
[0215] In at least one embodiment, a PVA can be used to perform dense optical flow. For example, in at least one embodiment, a PVA can process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, a PVA is used for time-of-flight depth processing, e.g., by processing raw time-of-flight data to provide processed time-of-flight data.
[0216] In at least one embodiment, DLA can be used to run any type of network to enhance control and driving safety, including, for example and without limitation, a neural network that outputs a confidence level for each object detection. In at least one embodiment, confidence level can be represented or interpreted as a probability, or as providing a relative “weight” of each detection relative to other detections. In at least one embodiment, confidence levels enable a system to make further decisions as to which detections should be considered as true positive detections and not false positive detections. For example, in at least one embodiment, a system can set a threshold for confidence levels, and only consider detections that exceed a threshold as true positive detections. In embodiments using automatic emergency braking (“AEB”) systems, false positive detections would result in a vehicle automatically performing an emergency brake, which is obviously undesirable. In at least one embodiment, highly confident detections can be considered as triggers for AEB. In at least one embodiment, DLA can run a neural network for regression of confidence values. In at least one embodiment, a neural network can take as its input at least some subset of parameters, such as bounding box size, ground plane estimates obtained (e.g., from another subsystem), outputs of one or more IMU sensors 1666 related to object’s vehicle 1600 direction, distance, 3D position estimates obtained from a neural network and / or other sensors (e.g., one or more LIDAR sensors 1664 or one or more RADAR sensors 1660), etc.
[0217] In at least one embodiment, one or more SoC(s) 1604 can include one or more data stores 1616 (e.g., memory). In at least one embodiment, one or more data stores 1616 can be on-chip memory of one or more SoC(s) 1604 that can store neural networks to be executed on one or more GPU(s) 1608 and / or DLA. In at least one embodiment, one or more data stores 1616 can have a capacity large enough to store multiple instances of a neural network for redundancy and safety. In at least one embodiment, one or more data stores 1616 can include L2 or L3 cache.
[0218] In at least one embodiment, one or more SoC(s) 1604 can include any number of processor(s) 1610 (e.g., embedded processors). One or more processor(s) 1610 can include a boot and power management processor that can be a dedicated processor and subsystem to handle boot power and management functions and related security enforcement. In at least one embodiment, boot and power management processor can be part of a one or more SoC(s) 1604 boot sequence and can provide run-time power management services. In at least one embodiment, boot power and management processor can provide clock and voltage programming, assist system low power state transitions, one or more SoC(s) 1604 thermal and temperature sensor management, and / or one or more SoC(s) 1604 power state management. In at least one embodiment, each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to temperature, and one or more SoC(s) 1604 can use ring oscillators to detect temperature of one or more CPU(s) 1606, one or more GPU(s) 1608, and / or one or more accelerator(s) 1614. In at least one embodiment, if a temperature is determined to exceed a threshold, boot and power management processor can enter a temperature fault routine and place one or more SoC(s) 1604 into a lower power state and / or place vehicle 1600 into a safe park pattern for the driver (e.g., cause vehicle 1600 to safely park).
[0219] In at least one embodiment, one or more processor(s) 1610 can further include a set of embedded processors that can function as an audio processing engine. In at least one embodiment, an audio processing engine can be an audio subsystem that is capable of providing full hardware support for multi-channel audio to hardware through a number of interfaces as well as a broad and flexible range of audio I / O interfaces. In at least one embodiment, an audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0220] In at least one embodiment, one or more processor(s) 1610 can further include an always-on processor engine that can provide necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, a processor on an always-on processor engine can include, but is not limited to, a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0221] In at least one embodiment, one or more processors 1610 can further include a safety cluster engine including, without limitation, a dedicated processor subsystem for handling safety management for automotive applications. In at least one embodiment, safety cluster engine can include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a safety mode, in at least one embodiment, two or more cores can operate in a lockstep mode and can function as a single core with comparison logic to detect any differences between their operations. In at least one embodiment, one or more processors 1610 can further include a real-time camera engine that can include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, one or more processors 1610 can further include a high dynamic range signal processor that can include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.
[0222] In at least one embodiment, one or more processors 1610 can include a video image compositor that can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions needed by a video playback application to produce a final video to produce a final image for a player window. In at least one embodiment, video image compositor can perform lens distortion correction on one or more wide-view cameras 1670, one or more surround cameras 1674, and / or one or more in-cabin monitoring camera sensors. In at least one embodiment, preferably, in-cabin monitoring camera sensors are monitored by a neural network running on another instance of SoC 1604 that is configured to identify cabin events and respond accordingly. In at least one embodiment, in-cabin systems can perform, without limitation, lip reading to activate cellular service and place a phone call, dictate an email, change a destination of a vehicle, activate or change an infotainment system and settings of a vehicle, or provide voice-activated web surfing. In at least one embodiment, certain functionality is available to a driver when a vehicle is operating in an autonomous mode, otherwise it is disabled.
[0223] In at least one embodiment, video image compositor can include enhanced temporal noise reduction for simultaneous spatial and temporal noise reduction. For example, in at least one embodiment, where motion occurs in a video, noise reduction appropriately weights spatial information, reducing a weight of information provided by adjacent frames. In at least one embodiment, where an image or portion of an image does not include motion, temporal noise reduction performed by video image compositor can use information from a previous image to reduce noise in a current image.
[0224] In at least one embodiment, video image compositor can also be configured to perform stereo correction on input stereoscopic lens frames. In at least one embodiment, when using an operating system desktop, video image compositor can also be used for user interface composition and one or more GPUs 1608 are not required to continuously render new surfaces. In at least one embodiment, when one or more GPUs 1608 are powered and active for 3D rendering, video image compositor can be used to offload one or more GPUs 1608 to improve performance and responsiveness.
[0225] In at least one embodiment, one or more SoCs in SoC 1604 can further include mobile industry processor interface (“MIPI”) camera serial interfaces for receiving video and input from cameras, high-speed interfaces, and / or video input blocks that can be used for camera and related pixel input functionality. In at least one embodiment, one or more SoCs 1604 can further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.
[0226] In at least one embodiment, one or more SoCs in SoC 1604 can further include a wide range of peripheral interfaces to enable communication with peripherals, audio encoders / decoders (“codecs”), power management, and / or other devices. One or more SoCs 1604 can be used to process data from cameras (e.g., over gigabit multimedia serial link and Ethernet connections), sensors (e.g., one or more LIDAR sensors 1664, one or more RADAR sensors 1660, etc., which can be connected over Ethernet), data from bus 1602 (e.g., speed of vehicle 1600, steering wheel position, etc.), data from one or more GNSS sensors 1658 (e.g., over Ethernet or CAN bus connections), etc. In at least one embodiment, one or more SoCs in SoC 1604 can further include dedicated high-performance mass storage controllers that can include their own DMA engines and can be used to free one or more CPUs 1606 from regular data management tasks.
[0227] In at least one embodiment, SoC(s) 1604 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS technology for diversity and redundancy, which provides a platform that can provide a flexible, reliable driving software stack, as well as deep learning tools. In at least one embodiment, SoC(s) 1604 can be faster, more reliable, and even more energy and spatial efficient than conventional systems. For example, in at least one embodiment, accelerator(s) 1614, when combined with CPU(s) 1606, GPU(s) 1608, and data storage device(s) 1616, can provide a fast, efficient platform for level 3-5 autonomous vehicles.
[0228] In at least one embodiment, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages (e.g., C programming language) to perform a variety of processing algorithms on a variety of visual data. However, in at least one embodiment, CPUs typically cannot meet performance requirements of many computer vision applications, such as performance requirements related to execution time and power consumption. In at least one embodiment, many CPUs cannot execute complex object detection algorithms in real-time, which are used in on-board ADAS applications and actual level 3-5 autonomous vehicles.
[0229] Embodiments described herein allow for simultaneous and / or sequential execution of multiple neural networks, and allow for combining results together to enable level 3-5 autonomous driving functionality. For example, in at least one embodiment, CNNs executed on DLAs or discrete GPUs (e.g., GPU(s) 1620) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs that a neural network has not been specifically trained for. In at least one embodiment, DLAs can also include neural networks capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing that semantic understanding to a path planning module running on a CPU Complex.
[0230] In at least one embodiment, multiple neural networks can be run simultaneously for a level 3, 4, or 5 drive. For example, in at least one embodiment, a warning sign consisting of a “Caution: flashing lights indicate icy conditions” sign with flashing lights can be interpreted independently or collectively by multiple neural networks. In at least one embodiment, the sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a neural network that has already been trained), the text “flashing lights indicate icy conditions” can be interpreted by a second deployed neural network that informs a vehicle’s path planning software (preferably executing on a CPU Complex) that icy conditions exist when flashing lights are detected. In at least one embodiment, flashing lights can be recognized by a third deployed neural network operating over multiple frames, informing a vehicle’s path planning software of the existence (or non-existence) of flashing lights. In at least one embodiment, all three neural networks can be run simultaneously, e.g., within a DLA and / or on one or more GPU(s) 1608.
[0231] In at least one embodiment, a CNN for facial recognition and vehicle owner identification can use data from a camera sensor to identify presence of an authorized driver and / or owner of vehicle 1600. In at least one embodiment, when an owner approaches a driver door and opens a light, a normally open sensor processor engine can be used to unlock a vehicle, and, in a safe mode, when an owner leaves the vehicle, can be used to disable the vehicle. In this way, one or more SoC(s) 1604 provide a safeguard against theft and / or carjacking.
[0232] In at least one embodiment, a CNN for emergency vehicle detection and identification can use data from microphones 1696 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 1604 use a CNN to classify ambient and urban sounds, as well as to classify visual data. In at least one embodiment, a CNN running on a DLA is trained to identify relative proximity of emergency vehicles (e.g., by using Doppler effect). In at least one embodiment, a CNN can also be trained to identify emergency vehicles for regions in which a vehicle is operating, as identified by one or more GNSS sensors 1658. In at least one embodiment, when operating in Europe, a CNN will seek to detect European sirens, while in the United States, a CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used to execute emergency vehicle safety routines, slow vehicle down, pull vehicle to side of road, stop, and / or idle vehicle until emergency vehicle passes, with assistance of one or more ultrasonic sensors 1662.
[0233] In at least one embodiment, vehicle 1600 can include one or more CPUs 1618 (e.g., one or more discrete CPUs or one or more dCPUs) that can be coupled to one or more SoCs 1604 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, one or more CPUs 1618 can include an X86 processor, such as one or more CPUs 1618 can be used to perform any of a variety of functions, such as including arbitrating inconsistent results between ADAS sensors and one or more SoCs 1604, and / or one or more supervisory controllers 1636 state and health and / or an information on a chip system (“Info SoC”) 1630.
[0234] In at least one embodiment, vehicle 1600 can include one or more GPUs 1620 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to one or more SoCs 1604 via a high-speed interconnect (e.g., NVIDIA’s NVLINK). In at least one embodiment, one or more GPUs 1620 can provide additional artificial intelligence functionality, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based at least in part on input (e.g., sensor data) from sensors of vehicle 1600.
[0235] In at least one embodiment, vehicle 1600 can further include network interface 1624, which can include, without limitation, one or more wireless antennas 1626 (e.g., one or more wireless antennas for different communication protocols such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, network interface 1624 can be used to enable wireless connectivity over the Internet with a cloud (e.g., with servers and / or other network devices), with other vehicles, and / or with computing devices (e.g., client devices of passengers). In at least one embodiment, to communicate with other vehicles, a direct link can be established between vehicle 1600 and other vehicles and / or an indirect link can be established (e.g., through a network and the Internet). In at least one embodiment, a direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 1600 with information about vehicles in a vicinity of vehicle 1600 (e.g., vehicles in front of, to the side of, and / or behind vehicle 1600). In at least one embodiment, this aforementioned functionality can be part of a cooperative adaptive cruise control functionality of vehicle 1600.
[0236] In at least one embodiment, network interface 1624 can include a SoC that provides modulation and demodulation functionality and enables one or more controllers 1636 to communicate over wireless networks. In at least one embodiment, network interface 1624 can include radio frequency front ends for up-conversion from baseband to radio frequency and for down-conversion from radio frequency to baseband. In at least one embodiment, frequency conversion can be performed in any technically feasible way. For example, frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In at least one embodiment, radio frequency front end functionality can be provided by a separate chip. In at least one embodiment, a network interface can include wireless functionality to communicate over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0237] In at least one embodiment, vehicle 1600 can further include one or more data stores 1628, which can include, without limitation, off-chip (e.g., of SoC(s) 1604) storage. In at least one embodiment, one or more data stores 1628 can include, without limitation, one or more storage elements including RAM, SRAM, dynamic random access memory (“DRAM”), video random access memory (“VRAM”), flash memory, hard disks, and / or other components and / or devices that can store at least one bit of data.
[0238] In at least one embodiment, vehicle 1600 can further include one or more GNSS sensors 1658 (e.g., GPS and / or assisted GPS sensors) to assist in mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1658 can be used including, for example and without limitation, a GPS using a USB connector with Ethernet connected to a serial interface (e.g., RS-232) bridge.
[0239] In at least one embodiment, vehicle 1600 can further include one or more RADAR sensors 1660. One or more RADAR sensors 1660 can be used by vehicle 1600 for long range vehicle detection, even in darkness and / or adverse weather conditions. In at least one embodiment, a RADAR functional safety level can be ASIL B. One or more RADAR sensors 1660 can use CAN and / or bus 1602 (e.g., to transmit data generated by one or more RADAR sensors 1660) for control and access to object tracking data, and in certain examples can access Ethernet for access to raw data. In at least one embodiment, a wide variety of RADAR sensor types can be used. For example and without limitation, one or more of RADAR sensors 1660 can be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more RADAR sensors 1660 are pulse Doppler RADAR sensors.
[0240] In at least one embodiment, RADAR sensor(s) 1660 can include different configurations, such as long-range with narrow field of view, short-range with wide field of view, short-range side coverage, etc. In at least one embodiment, long-range RADAR can be used for adaptive cruise control functionality. In at least one embodiment, long-range RADAR systems can provide a wide field of view enabled by two or more independent scans (e.g., up to 250 m range). In at least one embodiment, RADAR sensor(s) 1660 can help distinguish between static and moving objects, and can be used by ADAS system 1638 for emergency brake assist and forward collision warning. One or more sensors 1660 included in a long-range RADAR system can include, without limitation, a monostatic multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas, as well as high-speed CAN and FlexRay interfaces. In at least one embodiment, with six antennas, a central four antennas can create a focused beam pattern designed to record the environment around vehicle 1600 at higher speeds with minimal interference from traffic in adjacent lanes. In at least one embodiment, other two antennas can expand the field of view, such that vehicles entering or leaving a lane of vehicle 1600 can be quickly detected.
[0241] In at least one embodiment, as an example, a mid-range RADAR system can include, for example, a range of up to 160 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, a short-range RADAR system can include, without limitation, any number of RADAR sensors 1660 designed to be mounted at either end of a rear bumper. When mounted at either end of a rear bumper, in at least one embodiment, a RADAR sensor system can produce two beams that constantly monitor the blind spot behind and near vehicle 1600. In at least one embodiment, a short-range RADAR system can be used in ADAS system 1638 for blind spot detection and / or lane change assist.
[0242] In at least one embodiment, vehicle 1600 can further include ultrasonic sensor(s) 1662. Ultrasonic sensor(s) 1662, which can be positioned in front, rear, and / or side locations of vehicle 1600, can be used for parking assist and / or to create and update occupancy grids. In at least one embodiment, a wide variety of ultrasonic sensors 1662 can be used, and different ultrasonic sensors 1662 can be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, ultrasonic sensors 1662 can operate at a functional safety level of ASIL B.
[0243] In at least one embodiment, vehicle 1600 can include one or more LIDAR sensors 1664. One or more LIDAR sensors 1664 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, one or more LIDAR sensors 1664 can be at a functional safety level of ASIL B. In at least one embodiment, vehicle 1600 can include multiple (e.g., two, four, six, etc.) LIDAR sensors 1664 that can use Ethernet (e.g., provide data to a Gigabit Ethernet switch).
[0244] In at least one embodiment, one or more LIDAR sensors 1664 can be capable of providing a list of objects and their distances for a 360 degree field of view. In at least one embodiment, one or more LIDAR sensors 1664 that are commercially available can have an advertised range of approximately 100 m, have an accuracy of 2 cm - 3 cm, and support a 100 Mbps Ethernet connection, for example. In at least one embodiment, one or more non-protruding LIDAR sensors can be used. In such embodiments, one or more LIDAR sensors 1664 can be implemented as small devices that can be embedded into front, rear, side, and / or corner locations of vehicle 1600. In at least one embodiment, one or more LIDAR sensors 1664, in such embodiments, can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, with a range of 200 m, even for low reflectivity objects. In at least one embodiment, a forward-facing one or more LIDAR sensors 1664 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0245] In at least one embodiment, LIDAR technology such as 3D Flash LIDAR can also be used. 3D Flash LIDAR uses a laser flash as a transmission source to illuminate approximately 200 m around vehicle 1600. In at least one embodiment, a flash LIDAR unit includes, without limitation, a receiver that records laser pulse travel time and reflected light on each pixel, which in turn corresponds to a range from vehicle 1600 to an object. In at least one embodiment, flash LIDAR can allow for highly accurate and distortion-free images of a surrounding environment to be generated with each laser flash. In at least one embodiment, four flash LIDAR sensors can be deployed, one on each side of vehicle 1600. In at least one embodiment, a 3D flash LIDAR system includes, without limitation, a solid-state 3D line-of-sight array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, a flash LIDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame, and can capture reflected laser light in the form of 3D ranging point clouds and co-registered intensity data.
[0246] In at least one embodiment, vehicle 1600 can also include one or more IMU sensors 1666. In at least one embodiment, one or more IMU sensors 1666 can be located at a center of a rear axle of vehicle 1600. In at least one embodiment, one or more IMU sensors 1666 can include, without limitation, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In at least one embodiment, such as in a six-axis application, one or more IMU sensors 1666 can include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, such as in a nine-axis application, one or more IMU sensors 1666 can include, without limitation, an accelerometer, a gyroscope, and a magnetometer.
[0247] In at least one embodiment, one or more IMU sensors 1666 can be implemented as a miniature, high-performance GPS-aided inertial navigation system (“GPS / INS”) that combines micro-electro-mechanical systems (“MEMS”) inertial sensors, high-sensitivity GPS receivers, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude; in at least one embodiment, one or more IMU sensors 1666 can enable vehicle 1600 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from GPS to one or more IMU sensors 1666. In at least one embodiment, one or more IMU sensors 1666 and one or more GNSS sensors 1658 can be combined in a single integrated unit.
[0248] In at least one embodiment, vehicle 1600 can include one or more microphones 1696 placed within and / or around vehicle 1600. In at least one embodiment, additionally, one or more microphones 1696 can be used for emergency vehicle detection and identification.
[0249] In at least one embodiment, vehicle 1600 can further include any number of camera types, including one or more stereo cameras 1668, one or more wide-view cameras 1670, one or more infrared cameras 1672, one or more surround-view cameras 1674, one or more long-range cameras 1698, one or more mid-range cameras 1676, and / or other camera types. In at least one embodiment, cameras can be used to capture image data around entire periphery of vehicle 1600. In at least one embodiment, type of cameras used depends on vehicle 1600. In at least one embodiment, any combination of camera types can be used to provide necessary coverage around vehicle 1600. In at least one embodiment, number of cameras can vary from embodiment to embodiment. For example, in at least one embodiment, vehicle 1600 can include six cameras, seven cameras, ten cameras, twelve cameras, or other number of cameras. In at least one embodiment, cameras can support Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet, by way of example and without limitation. In at least one embodiment, cameras can be described in greater detail herein previously with reference to FIG. 6. Figure 16A and Figure 16B Each camera can be described in greater detail.
[0250] In at least one embodiment, vehicle 1600 can further include one or more vibration sensors 1642. One or more vibration sensors 1642 can measure vibrations of components of vehicle 1600 (e.g., axles). For example, in at least one embodiment, changes in vibration can be indicative of changes in road surface. In at least one embodiment, when two or more vibration sensors 1642 are used, differences between vibrations can be used to determine friction or slippage of a road surface (e.g., when there is a difference in vibration between a power driven axle and a freely rotating axle).
[0251] In at least one embodiment, vehicle 1600 can include ADAS system 1638. ADAS system 1638 can include, without limitation, an SoC. In at least one embodiment, ADAS system 1638 can include, without limitation, any number of adaptive / autonomous / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward collision warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keep assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functionality, and combinations thereof.
[0252] In at least one embodiment, an ACC system can use one or more RADAR sensors 1660, one or more LIDAR sensors 1664, and / or any number of cameras. In at least one embodiment, an ACC system can include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, a longitudinal ACC system monitors and controls the distance to a vehicle immediately in front of vehicle 1600 and automatically adjusts speed of vehicle 1600 to maintain a safe distance from the vehicle in front. In at least one embodiment, a lateral ACC system performs distance keeping and advises vehicle 1600 to change lanes when necessary. In at least one embodiment, lateral ACC is relevant to other ADAS applications such as LC and CW.
[0253] In at least one embodiment, a CACC system uses information from other vehicles that can be received from other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet) via network interface 1624 and / or one or more wireless antennas 1626. In at least one embodiment, a direct link can be provided by a vehicle-to-vehicle (“V2V”) communication link, while an indirect link can be provided by an infrastructure-to-vehicle (“I2V”) communication link. Generally, a V2V communication concept provides information about the immediately preceding vehicle (e.g., a vehicle immediately ahead of and in the same lane as vehicle 1600), while an I2V communication concept provides information about traffic further ahead. In at least one embodiment, a CACC system can include one or both of I2V and V2V information sources. In at least one embodiment, a CACC system can be more reliable given information about vehicles ahead of vehicle 1600, and has potential to improve smoothness of traffic flow and reduce road congestion.
[0254] In at least one embodiment, an FCW system is designed to alert a driver of a hazard so that the driver can take corrective action. In at least one embodiment, an FCW system uses a forward facing camera and / or one or more RADAR sensors 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to a driver feedback such as a display, speaker, and / or vibrating component. In at least one embodiment, an FCW system can provide a warning, for example, in the form of a sound, visual warning, vibration, and / or quick brake pulse.
[0255] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object and can automatically apply brakes if a driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, an AEB system can use one or more forward facing cameras and / or one or more RADAR sensors 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when an AEB system detects a hazard, an AEB system typically first alerts a driver to take corrective action to avoid a collision, and if that driver does not take corrective action, the AEB system can automatically apply brakes in an attempt to prevent or at least mitigate the effects of a predicted collision. In at least one embodiment, an AEB system can include technologies such as dynamic brake support and / or brake to avoid a collision.
[0256] In at least one embodiment, an LDW system provides visual, audible, and / or tactile warnings, e.g., steering wheel or seat vibration, to warn the driver when vehicle 1600 crosses lane markers. In at least one embodiment, an LDW system does not activate when the driver indicates an intentional lane departure, such as by activating turn signals. In at least one embodiment, an LDW system can use a forward facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. In at least one embodiment, an LKA system is a variation of an LDW system. If vehicle 1600 begins to deviate from a lane, an LKA system provides a steering input or brake to correct vehicle 1600.
[0257] In at least one embodiment, a BSW system detects and warns vehicle drivers of vehicles in a car’s blind spot. In at least one embodiment, a BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, a BSW system can provide additional warnings when a driver uses turn signals. In at least one embodiment, a BSW system can use one or more rear-facing side-facing cameras and / or one or more RADAR sensors 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.
[0258] In at least one embodiment, a RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside of a rear camera range while vehicle 1600 is backing up. In at least one embodiment, a RCTW system includes an AEB system to ensure application of vehicle brakes to avoid a collision. In at least one embodiment, a RCTW system can use one or more rear-facing RADAR sensors 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to provide driver feedback such as displays, speakers, and / or vibrating components.
[0259] In at least one embodiment, conventional ADAS systems can be prone to false positives, which can annoy and distract drivers, but are typically not catastrophic because conventional ADAS systems warn the driver and allow that driver to decide whether a safety situation is truly present and take appropriate action. In at least one embodiment, in the event of a result conflict, vehicle 1600 itself decides whether to heed the results of a primary computer or a secondary computer (e.g., first controller 1636 or second controller 1636). For example, in at least one embodiment, ADAS system 1638 can be a backup and / or secondary computer for providing perception information to a backup computer plausibility module. In at least one embodiment, a backup computer plausibility monitor can run redundant varieties of software on hardware components to detect faults in perception and dynamic driving tasks. In at least one embodiment, outputs from ADAS system 1638 can be provided to a supervisory MCU. In at least one embodiment, if outputs from a primary computer and a secondary computer conflict, a supervisory MCU decides how to reconcile the conflict to ensure safe operation.
[0260] In at least one embodiment, a host computer can be configured to provide a confidence score to a supervisory MCU to indicate a host computer’s confidence in a selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU can follow the host computer’s instructions regardless of whether the secondary computer provided conflicting or inconsistent results. In at least one embodiment, in cases where a confidence score does not satisfy a threshold, and in cases where the host computer and secondary computer indicate different results (e.g., a conflict), the supervisory MCU can arbitrate between the computers to determine an appropriate result.
[0261] In at least one embodiment, a supervisory MCU can be configured to run a neural network trained and configured to determine conditions under which a secondary computer provides false alarms based at least in part on outputs from a host computer and a secondary computer. In at least one embodiment, a neural network in a supervisory MCU can learn when to trust outputs of a secondary computer, and when not to. For example, in at least one embodiment, when the secondary computer is a RADAR-based FCW system, a neural network in a supervisory MCU can learn when the FCW system identifies metal objects that are not actually dangerous, such as drain grates or manhole covers that would trigger an alert. In at least one embodiment, when the secondary computer is a camera-based LDW system, a neural network in a supervisory MCU can learn to override LDW when there is a bicyclist or pedestrian present and it is actually safest to lane depart. In at least one embodiment, a supervisory MCU can include at least one of a DLA or GPU suitable for running a neural network with associated memory. In at least one embodiment, a supervisory MCU can include and / or be included as a component of one or more SoCs 1604.
[0262] In at least one embodiment, an ADAS system 1638 can include a secondary computer that performs ADAS functions using traditional computer vision rules. In at least one embodiment, the secondary computer can use classic computer vision rules (if-then), and the presence of a neural network in a supervisory MCU can improve reliability, safety, and performance. For example, in at least one embodiment, a diversified implementation and intentional non-identity make the overall system more fault-tolerant, especially to faults caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if there is a software bug or error in software running on a host computer, and non-identical software code running on a secondary computer provides the same overall result, a supervisory MCU can be more confident that the overall result is correct, and that the bug in software or hardware on the host computer did not cause a significant error.
[0263] In at least one embodiment, output of ADAS system 1638 can be input into a perception module of host computer and / or a dynamic driving task module of host computer. For example, in at least one embodiment, if ADAS system 1638 indicates a forward collision warning due to an object directly in front of vehicle, perception block can use this information when identifying the object. In at least one embodiment, as described herein, a secondary computer can have its own neural network that is trained to reduce risk of false positives.
[0264] In at least one embodiment, vehicle 1600 can further include infotainment SoC 1630 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as a SoC, in at least one embodiment, infotainment system SoC 1630 can not be a SoC and can include, without limitation, two or more discrete components. In at least one embodiment, infotainment SoC 1630 can include, without limitation, a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephony (e.g., hands free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, radio data system, vehicle related information such as fuel level, total covered distance, brake fluid level, oil level, doors open / close, air cleaner information, etc.) to vehicle 1600. For example, infotainment SoC 1630 can include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, car, car entertainment system, WiFi, steering wheel audio controls, hands-free voice controls, heads-up display (“HUD”), HMI display 1634, telematics equipment, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, infotainment SoC 1630 can be further used to provide information (e.g., visual and / or audible) to a user of vehicle, such as information from ADAS system 1638, autonomous driving information (such as planned vehicle maneuvers), trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0265] In at least one embodiment, infotainment SoC 1630 can include any number and type of GPU functionality. In at least one embodiment, infotainment SoC 1630 can communicate with other devices, systems, and / or components of vehicle 1600 over bus 1602 (e.g., CAN bus, Ethernet, etc.). In at least one embodiment, infotainment SoC 1630 can be coupled to a supervisory MCU such that a GPU of an infotainment system can perform some autonomous driving functions in the event of a failure of a host controller 1636 (e.g., a primary computer and / or a backup computer of vehicle 1600). In at least one embodiment, infotainment SoC 1630 can cause vehicle 1600 to enter a driver-to-safe-stop mode, as described herein.
[0266] In at least one embodiment, vehicle 1600 can further include an instrument cluster 1632 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument panel, etc.). Instrument cluster 1632 can include, without limitation, a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). In at least one embodiment, instrument cluster 1632 can include, without limitation, any number and combination of gauges such as a speedometer, a fuel level, an oil pressure, a tachometer, an odometer, a turn indicator, a shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, auxiliary restraint system (e.g., airbag) information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between infotainment SoC 1630 and instrument cluster 1632. In at least one embodiment, instrument cluster 1632 can be included as part of infotainment SoC 1630, and vice versa.
[0267] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Inference and / or training logic 1315 can be used in system 1300 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations / architectures or neural network use cases described herein. Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided with respect to FIGS. 1 and 2. Figure 16C In at least one embodiment, inference and / or training logic 1315 can be used in system 1300 to infer or predict operations based, at least in part, on weight parameters calculated using neural network training operations / architectures or neural network use cases described herein.
[0268] The above techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0269] Figure 16DIt is based on at least one embodiment in a cloud-based server and Figure 16A A diagram of a system 1676 for communication between autonomous vehicles 1600 is provided. In at least one embodiment, system 1676 may include, but is not limited to, one or more servers 1678, one or more networks 1690, and any number and type of vehicles, including vehicle 1600. In at least one embodiment, one or more servers 1678 may include, but is not limited to, multiple GPUs 1684(A)-1684(H) (collectively referred to herein as GPU 1684), PCIe switches 1682(A)-1682(D) (collectively referred to herein as PCIe switch 1682), and / or CPUs 1680(A)-1680(B) (collectively referred to herein as CPU 1680). GPU 1684, CPU 1680, and PCIe switch 1682 may be interconnected with high-speed cables, such as, but not limited to, NVLink interface 1688 developed by NVIDIA and / or PCIe connection 1686. In at least one embodiment, the GPU 1684 is connected via NVLink and / or NVSwitchSoC, and the GPU 1684 and PCIe switch 1682 are connected via PCIe interconnect. In at least one embodiment, although eight GPUs 1684, two CPUs 1680, and four PCIe switches 1682 are shown, this is not intended to be limiting. In at least one embodiment, each of one or more servers 1678 may include, but is not limited to, any combination of any number of GPUs 1684, CPUs 1680, and / or PCIe switches 1682. For example, in at least one embodiment, one or more servers 1678 may each include eight, sixteen, thirty-two, and / or more GPUs 1684.
[0270] In at least one embodiment, one or more servers 1678 can receive image data representative of images from vehicles over one or more networks 1690, which show unexpected or changing road conditions, such as road work that has recently started. In at least one embodiment, one or more servers 1678 can transmit neural networks 1692, updated etc. neural networks 1692, and / or map information 1694, including but not limited to information about traffic and road conditions, to vehicles over one or more networks 1690. In at least one embodiment, updates to map information 1694 can include, but are not limited to, updates to HD map 1622, such as information about construction sites, potholes, detours, floods, and / or other obstacles. In at least one embodiment, neural networks 1692, updated etc. neural networks 1692, and / or map information 1694 can be the result of new training and / or experience represented in data received from any number of vehicles in an environment, and / or based at least on training performed at a data center (e.g., using one or more servers 1678 and / or other servers).
[0271] In at least one embodiment, one or more servers 1678 can be used to train machine learning models (e.g., neural networks) based at least in part on training data. In at least one embodiment, training data can be generated by vehicles, and / or can be generated in simulations (e.g., using a game engine). In at least one embodiment, any amount of training data is labeled (e.g., in cases where a related neural network benefits from supervised learning) and / or undergoes other pre-processing. In at least one embodiment, no amount of training data is labeled and / or pre-processed (e.g., in cases where an associated neural network does not require supervised learning). In at least one embodiment, once a machine learning model is trained, a machine learning model can be used by vehicles (e.g., transmitted to vehicles over one or more networks 1690, and / or a machine learning model can be used by one or more servers 1678 to remotely monitor vehicles.
[0272] In at least one embodiment, one or more servers 1678 can receive data from vehicles and apply the data to up-to-date real-time neural networks for real-time intelligent inference. In at least one embodiment, one or more servers 1678 can include deep-learning supercomputers and / or specialized Al computers powered by one or more GPUs 1684, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, one or more servers 1678 can include deep-learning infrastructure of a data center using CPU power.
[0273] In at least one embodiment, deep learning infrastructure of server(s) 1678 can be capable of fast, real-time inferencing and can use this capability to assess and validate health of processors, software, and / or related hardware in vehicle 1600. For example, in at least one embodiment, deep learning infrastructure can receive periodic updates from vehicle 1600, such as sequences of images and / or objects that vehicle 1600 locates in that sequence of images (e.g., through computer vision and / or other machine learning object classification techniques). In at least one embodiment, deep learning infrastructure can run its own neural networks to identify objects and compare them to objects identified by vehicle 1600, and if results do not match and deep learning infrastructure concludes that AI in vehicle 1600 is malfunctioning, server(s) 1678 can send a signal to vehicle 1600 instructing a failsafe computer of vehicle 1600 to take control, notify passengers, and complete a safe parking operation.
[0274] In at least one embodiment, server(s) 1678 can include GPU(s) 1684 and programmable inference accelerator(s) (such as NVIDIA’s TensorRT 3). In at least one embodiment, a combination of GPU-driven servers and inference-accelerated servers can make real-time responses possible. In at least one embodiment, CPU-, FPGA-, and other processor-driven servers can be used for inference, for example, in cases where performance is less critical.
[0275] In at least one embodiment, hardware structure 1315 is used to perform one or more embodiments. Described herein in conjunction with FIG. 13A are, among other things, Figure 13A and / or Figure 13B Details regarding hardware structure 1315 are provided.
[0276] Computer system
[0277] Figure 17 is a block diagram illustrating an exemplary computer system that can be a system with interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor that can include execution units to execute an instruction, according to at least one embodiment. In at least one embodiment, computer system 1700 can include, without limitation, components such as processor 1702 that has execution units with logic to perform an algorithm for process data according to present disclosure, such as embodiments described herein. In at least one embodiment, computer system 1700 can include a processor, such as a family, Xeon TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM microprocessors, although other systems (including PCs having other microprocessors, engineering workstations, set-top boxes, etc.) can also be used. In at least one embodiment, computer system 1700 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash.
[0278] Embodiments can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor (“DSP”), a system on a chip, a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system that can perform one or more instructions in accordance with at least one embodiment.
[0279] In at least one embodiment, computer system 1700 can include, but is not limited to, processor 1702, which can include, but is not limited to, one or more execution units 1708 to perform, e.g., machine learning model training and / or inference, in accordance with the techniques described herein. In at least one embodiment, system 1700 is a single processor desktop or server system, although in another embodiment, system 1700 can be a multiprocessor system. In at least one embodiment, processor 1702 can include, but not be limited to, a complex instruction set computer (“CISC”) microprocessor, a reduced instruction set computing (“RISC”) microprocessor, a very long instruction word (“VLIW”) microprocessor, a processor implementing a combo of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 1702 is coupled to a processor bus 1710 that can transmit data signals between processor 1702 and other components in computer system 1700.
[0280] In at least one embodiment, processor 1702 can include, without limitation, level 1 (“L1”) internal cache memory (“cache”) 1704. In at least one embodiment, processor 1702 can have a single- level internal cache or multi-level internal cache. In at least one embodiment, cache memory can reside in the processor 1702’s outer level. Other embodiments can include a combination of internal and external cache memory depending on the specific implementation and needs. In at least one embodiment, register file 1706 can store different types of data, including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers in various registers.
[0281] In at least one embodiment, execution unit 1708, including, without limitation, logic to perform integer and floating point operations, also resides in processor 1702. Processor 1702 can also include microcode (“ucode”) read-only memory (“ROM”) that stores microcode for certain macro instructions. In at least one embodiment, execution unit 1708 can include logic to handle a packed instruction set 1709. In at least one embodiment, by including the packed instruction set 1709 in an instruction set of a general- purpose processor 1702, along with associated circuitry to execute the instructions, the general purpose processor 1702 can be used to perform the operations for a number of multimedia applications faster and more efficiently than a system that does not so include. In one or more embodiments, by using the full width of a processor’s data bus over one or more clock cycles to send a single data element, many multimedia applications can be accelerated as compared to systems in which smaller units of data are transferred over a full width of a processor’s data bus in multiple clock cycles to execute one or more operations.
[0282] In at least one embodiment, execution unit 1708 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 1700 can include, without limitation, memory 1720. In at least one embodiment, memory 1720 can be implemented as a Dynamic Random Access Memory (“DRAM”) device, a Static Random Access Memory (“SRAM”) device, a flash memory device, or other memory device. Memory 1720 can store data signals expressed as instructions 1719 and / or data 1721 to be executed by processor 1702.
[0283] In at least one embodiment, a system logic chip can be coupled to processor bus 1710 and memory 1720. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 1716 and processor 1702 can communicate with MCH 1716 via processor bus 1710. In at least one embodiment, MCH 1716 can provide a high bandwidth memory path 1718 to memory 1720 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 1716 can direct data signals between processor 1702, memory 1720, and other components in computer system 1700 and communicate data signals between processor bus 1710, memory 1720, and system I / O 1722. In at least one embodiment, system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 1716 can be coupled to memory 1720 through a high bandwidth memory path 1718 and a graphics / video card 1712 can be coupled to MCH 1716 through an Accelerated Graphics Port (“AGP”) interconnect 1714.
[0284] In at least one embodiment, computer system 1700 can use system I / O 1722, which is a proprietary hub interface bus to couple MCH 1716 to I / O controller hub (“ICH”) 1730. In at least one embodiment, ICH 1730 can provide a direct connection to some I / O devices and
[0285] In at least one embodiment, Figure 17 A system including interconnected hardware devices or “chips” is shown, while in other embodiments, Figure 17 An exemplary system on a chip (SoC) can be shown. In at least one embodiment, Figure 17The devices illustrated in FIG. 17 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 1700 are interconnected using a Compute Express Link (CXL) interconnect.
[0286] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In examples in which inference and / or training logic 1315 are used for inferencing, the inference and / or training logic 1315 can be used to implement neural network inference processing. Figure 13A And / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In examples in which inference and / or training logic 1315 are used for inferencing, the inference and / or training logic 1315 can be used to implement neural network inference processing. Figure 17 Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In examples in which inference and / or training logic 1315 are used for inferencing, the inference and / or training logic 1315 can be used to implement neural network inference processing.
[0287] The above-described techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0288] Figure 18 FIG. 18 is a block diagram illustrating an electronic device 1800 for use with a processor 1810, in accordance with at least one embodiment. In at least one embodiment, the electronic device 1800 can be, for example, and without limitation, a notebook, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0289] In at least one embodiment, system 1800 can include, without limitation, a processor 1810 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1810 is coupled using a bus or interface, such as an Industry Standard 2 In at least one embodiment, system 1800 can include, without limitation, a processor 1810 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1810 is coupled using a bus or interface, such as an Industry Standard Figure 18 In at least one embodiment, system 1800 can include, without limitation, a processor 1810 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1810 is coupled using a bus or interface, such as an Industry Standard Figure 18 In at least one embodiment, system 1800 can include, without limitation, a processor 1810 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1810 is coupled using a bus or interface, such as an Industry Standard Figure 18 The devices illustrated in FIG. 17 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 1700 are interconnected using a Compute Express Link (CXL) interconnect.Figure 18 One or more components are interconnected using Computational Fast Link (CXL) interconnects.
[0290] In at least one embodiment, Figure 18 It may include a display 1824, a touch screen 1825, a touchpad 1830, a near field communication unit (“NFC”) 1845, a sensor hub 1840, a thermal sensor 1846, a fast chipset (“EC”) 1835, a trusted platform module (“TPM”) 1838, a BIOS / firmware / flash (“BIOS, FW Flash”) 1822, a DSP 1860, a drive 1820 (e.g., a solid-state drive (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 1850, a Bluetooth unit 1852, a wireless wide area network unit (“WWAN”) 1856, a global positioning system (GPS) 1855, a camera (“USB 3.0 camera”) 1854 (e.g., a USB 3.0 camera), and / or a low-power double data rate (“LPDDR”) memory unit (“LPDDR3”) 1815 implemented in, for example, the LPDDR3 standard. These components can each be implemented in any suitable way.
[0291] In at least one embodiment, other components may be communicatively coupled to processor 1810 via the components described herein. In at least one embodiment, accelerometer 1841, ambient light sensor (“ALS”) 1842, compass 1843, and gyroscope 1844 may be communicatively coupled to sensor hub 1840. In at least one embodiment, thermal sensor 1839, fan 1837, keyboard 1836, and touchpad 1830 may be communicatively coupled to EC 1835. In at least one embodiment, speaker 1863, earphone 1864, and microphone (“mic”) 1865 may be communicatively coupled to audio unit (“audio codec and Class D amplifier”) 1862, which in turn may be communicatively coupled to DSP 1860. In at least one embodiment, audio unit 1862 may include, for example, but not limited to, audio encoder / decoder (“codec”) and Class D amplifier. In at least one embodiment, SIM card (“SIM”) 1857 may be communicatively coupled to WWAN unit 1856. In at least one embodiment, components such as WLAN unit 1850, Bluetooth unit 1852, and WWAN unit 1856 can be implemented as next-generation form factor (NGFF).
[0292] Inference and / or training logic 1315 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 13A and / or Figure 13BDetails regarding inference and / or training logic 1315 are provided. In at least one embodiment, inference and / or training logic 1315 can be used in system Figure 18 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0293] The foregoing techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0294] Figure 19 A computer system 1900 according to at least one embodiment is shown. In at least one embodiment, computer system 1900 is configured to implement various processes and methods described throughout this disclosure.
[0295] In at least one embodiment, computer system 1900 includes, without limitation, at least one central processing unit (“CPU”) 1902 that is connected to a communication bus 1910 implemented using any suitable protocol, such as PCI (“Peripheral Component Interconnect”), peripheral component interconnect express (“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1900 includes, without limitation, a main memory 1904 and control logic (e.g., implemented in hardware, software, or a combination thereof) and data can be stored in the main memory 1904 in the form of random-access memory (“RAM”).
[0296] In at least one embodiment, computer system 1900 includes, without limitation, an input device 1908, parallel processing system 1912, and display device 1906, which can be implemented using a conventional cathode ray tube (“CRT”), liquid crystal display (“LCD”), light emitting diode (“LED”), plasma display, or other suitable display technologies. In at least one embodiment, user input is received from input device 1908, such as a keyboard, mouse, touchpad, microphone, or the like. In at least one embodiment, each of the foregoing modules can be located on a single semiconductor platform.
[0297] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In various embodiments, inference and / or training logic 1315 can be used in system FIG. 12 for inferencing or predicting operations associated with deep learning neural networks, as further described below. Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In various embodiments, inference and / or training logic 1315 can be used in system FIG. 12 for inferencing or predicting operations associated with deep learning neural networks, as further described below. Figure 19 Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In various embodiments, inference and / or training logic 1315 can be used in system FIG. 12 for inferencing or predicting operations associated with deep learning neural networks, as further described below.
[0298] The above techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate a grasp for an object held by a human hand as described above.
[0299] Figure 20 A computer system 2000 according to at least one embodiment is shown. In at least one embodiment, computer system 2000 includes, without limitation, a computer 2010 and a USB stick 2020. In at least one embodiment, computer 2010 can include, without limitation, any number and type of processor(s) (not shown) and memory (not shown). In at least one embodiment, computer 2010 includes, without limitation, a server, a cloud instance, a laptop computer, and a desktop computer.
[0300] In at least one embodiment, USB stick 2020 includes, without limitation, a processing unit 2030, a USB interface 2040, and USB interface logic 2050. In at least one embodiment, processing unit 2030 can be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, processing unit 2030 can include, without limitation, any number and type of processing core(s) (not shown). In at least one embodiment, processing core 2030 includes an application-specific integrated circuit (“ASIC”) optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, processing core 2030 is a tensor processing unit (“TPC”) optimized to perform machine learning inferencing operations. In at least one embodiment, processing core 2030 is a vision processing unit (“VPU”) optimized to perform machine vision and machine learning inferencing operations.
[0301] In at least one embodiment, USB interface 2040 can be any type of USB connector or USB socket. For example, in at least one embodiment, USB interface 2040 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, USB interface 2040 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 2050 can include any amount and type of logic that enables processing unit 2030 to interface with a device (e.g., computer 2010) via USB connector 2040.
[0302] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13 A and / or 13B. In at least one embodiment, inference and / or training logic 1315 can be used in system FIG. 1 related to object detection, image recognition, or other inferencing or training related operations. Figure 20 Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13 A and / or 13B. In at least one embodiment, inference and / or training logic 1315 can be used in system FIG. 1 related to object detection, image recognition, or other inferencing or training related operations.
[0303] The above techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0304] Figure 21A An exemplary architecture is shown in which a plurality of GPUs 2110-2113 are communicatively coupled to a plurality of multi-core processors 2105-2106 via high-speed links 2140-2143 (e.g., buses / point-to-point interconnects, etc.). In one embodiment, high-speed links 2140-2143 support a communication throughput of 4GB / s, 30GB / s, 80GB / s or higher. Various interconnect protocols can be used including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0.
[0305] Further, in one embodiment, two or more of GPUs 2110-2113 are interconnected by high-speed links 2129-2130, which can be implemented using the same or different protocol / links as used for high-speed links 2140-2143. Similarly, two or more of multi-core processors 2105-2106 can be connected by an interprocessor bus 2128, which can be an on-die bus of a system Figure 21Aall communications between the various system components shown in FIG. 1.
[0306] In one embodiment, each multi-core processor 2105-2106 is communicatively coupled to processor memories 2101-2102 via memory interconnects 2126-2127, respectively, and each GPU 2110-2113 is communicatively coupled to GPU memories 2120-2123 by GPU memory interconnects 2150-2153, respectively. In at least one embodiment, memory interconnects 2126-2127 and 2150-2153 can utilize similar or different memory access technologies. By way of example and not limitation, processor memories 2101-2102 and GPU memories 2120-2123 can be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or can be non-volatile memories such as 3D XPoint or Nano-Ram. In one embodiment, certain portions of processor memories 2101-2102 can be volatile memory while another portion can be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0307] As described herein, although various processors 2105-2106 and GPUs 2110-2113 can be physically coupled to particular memories 2101-2102, 2120-2123, respectively, a unified memory architecture can be implemented in which the same virtual system address space (also referred to as an “effective address” space) is distributed across the various physical memories. For example, processor memories 2101-2102 can each contain 64 GB of system memory address space, and in this example, GPU memories 2120-2123 can each contain 32 GB of system memory address space, resulting in a total of 256 GB of addressable memory size.
[0308] Figure 21B Additional details are shown for an interconnect between multi-core processor 2107 and graphics acceleration module 2146, according to one exemplary embodiment. In at least one embodiment, graphics acceleration module 2146 can include one or more GPU chips integrated on a line card that is coupled via a high-speed link 2140 to processor 2107. Alternatively, graphics acceleration module 2146 can be integrated on the same package or chip as processor 2107.
[0309] In at least one embodiment, processor 2107 is shown including multiple cores 2160A-2160D, each with a translation lookaside buffer 2161 A-2161 D and one or more caches 2162A-2162D. In at least one embodiment, cores 2160A-2160D can include various other components not shown for purposes of illustration and explanation, to execute instructions and process data. Caches 2162A-2162D can include Level 1 (LI) and Level 2 (L2) caches. In addition, one or more shared caches 2156 can be included in caches 2162A-2162D and shared by groups of cores 2160A-2160D. For example, one embodiment of processor 2107 includes 24 cores, each with its own LI cache, twelve shared L2 caches, and twelve shared L3 caches. In that embodiment, two adjacent cores share one or more L2 and L3 caches. Processor 2107 and graphics acceleration module 2146 are connected with system memory 2114, which can include Figure 21A processor memories 2101-2102 in FIG. 21.
[0310] Consistency for data and instructions stored in respective caches 2162A-2162D, 2156 and system memory 2114 is maintained through inter-core communication via coherence bus 2164. In at least one embodiment, for example, each cache can have cache coherency logic / circuitry associated therewith to communicate through coherence bus 2164 in response to detecting a read or write to a particular cache line. In one embodiment, a cache snoop protocol is implemented through coherence bus 2164 to snoop cache accesses.
[0311] In at least one embodiment, agent circuit 2125 communicatively couples graphics acceleration module 2146 to coherence bus 2164, allowing graphics acceleration module 2146 to participate in the cache coherence protocol as a peer to cores 2160A-2160D. In particular, interface 2135 provides connectivity to agent circuit 2125 over high-speed link 2140 (e.g., a PCIe bus, NVLink, etc.) and interface 2137 connects graphics acceleration module 2146 to link 2140.
[0312] In one embodiment, the accelerator integration circuit 2136 represents a plurality of graphics processing engines 2131, 2132, N of the graphics processing module 2146 that provide graphics processing services to the graphics processing engines 2131, 2132, N. In at least one embodiment, the graphics processing engines 2131, 2132, N can each include a graphics processing unit (GPU). Alternatively, the graphics processing engines 2131, 2132, N include different types of graphics processing engines, such as graphics execution units, media processing engines, samplers, and blit engines. In at least one embodiment, the graphics acceleration module 2146 can be a GPU with a plurality of graphics processing engines 2131-2132, N or the graphics processing engines 2131-2132, N can be individual GPUs integrated on a common package, line card, or chip.
[0313] In one embodiment, the accelerator integration circuit 2136 includes a memory management unit (MMU) 2139 to provide memory management services for processes executed by the graphics acceleration module 2146. In at least one embodiment, the memory management unit 2139 includes address translation lookaside buffers (TLBs) and global and local cache structures.
[0314] A set of registers 2145 store context data for threads executed by the graphics processing engines 2131-2132, N, and context management circuit 2148 manages thread contexts. For example, the context management circuit 2148 can perform save and restore operations to save and restore the context of individual threads during context switches (e.g., where a first thread is saved and a second thread is stored so that it can be executed by the graphics processing engines). For example, the context management circuit 2148, upon context switch, can store current register values to a designated area in memory (e.g., identified by a context pointer). The register values can then be restored when the context is returned. In one embodiment, the interrupt management circuit 2147 receives and processes interrupts received from system devices.
[0315] In one implementation, the MMU 2139 translates virtual / effective addresses from the graphics processing engines 2131-2132, N to real / physical addresses in system memory 2114. In at least one embodiment, the accelerator integration circuit 2136 supports multiple (e.g., 4, 8, 16) graphics processor modules 2146 and / or other accelerator devices. The graphics processor modules 2146 can be dedicated to a single application or shared between multiple applications. In one embodiment, a virtualized graphics execution environment is presented in which resources of the graphics processing engines 2131-2132, N are shared between multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into “slices” that are assigned to different VMs and / or applications based on processing requirements and priority level associated with the VMs and / or applications.
[0316] In at least one embodiment, the accelerator integration circuit 2136 performs as a bridge to the system for the graphics processor modules 2146 and provides address translation and system memory cache services. In addition, the accelerator integration circuit 2136 can provide virtualization facilities to allow a host processor to manage a virtualization of the graphics processing engines 2131-2132, N, interrupts, and memory management.
[0317] Because the hardware resources of the graphics processing engines 2131-2132, N are explicitly mapped to the real address space seen by the host processor 2107, any host processor can directly address these resources using effective address values. One function of the accelerator integration circuit 2136 is to physically separate the graphics processing engines 2131-2132, N so that they appear as independent units to the system.
[0318] In at least one embodiment, one or more graphics memory 2133-2134, M is coupled to each graphics processing engines 2131-2132, N respectively. Graphics memory 2133-2134, M stores instructions and data used by each graphics processing engines 2131-2132, N. Graphics memory 2133-2134, M can be a volatile memory, such as DRAM (including stacked DRAM 2133-2134, M, GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or can be a non-volatile memory, such as 3D XPoint or Nano-Ram.
[0319] In one embodiment, to reduce data traffic on link 2140, a biasing technique is used to ensure that data stored in graphics memory 2133-2134, M is that which is most frequently used by graphics processing engines 2131-2132, N and that which is least used by cores 2160A-2160D (at least frequently). Similarly, the biasing mechanism attempts to keep data needed by the cores (and preferably not by graphics processing engines 2131-2132, N) in the caches 2162A-2162D, 2156 and system memory 2114 of the cores.
[0320] Figure 21C Another exemplary embodiment is shown in which accelerator integration circuit 2136 is integrated within processor 2107. In this embodiment, graphics processing engines 2131-2132, N communicate directly over high-speed link 2140 to accelerator integration circuit 2136 via interface 2137 and interface 2135 (which can be any form of bus or interface protocol, as desired). Accelerator integration circuit 2136 can perform same operations as those described with regard to Figure 21B But can have higher throughput due to its close proximity to coherence bus 2164 and caches 2162A-2162D, 2156. One embodiment supports different programming models, including a dedicated process programming model (no graphics acceleration module virtualization) and a shared programming model (with virtualization), which can include programming models controlled by accelerator integration circuit 2136 and programming models controlled by graphics acceleration module 2146.
[0321] In at least one embodiment, graphics processing engines 2131-2132, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to graphics processing engines 2131-2132, N, providing virtualization within a VM / partition.
[0322] In at least one embodiment, graphics processing engines 2131-2132, N can be shared by multiple VM / application partitions. In at least one embodiment, a shared model can use a hypervisor to virtualize graphics processing engines 2131-2132, N to allow access by each operating system. In at least one embodiment, for a single-partition system without a hypervisor, an operating system owns graphics processing engines 2131-2132, N. In at least one embodiment, an operating system can virtualize graphics processing engines 2131-2132, N to provide access to each process or application.
[0323] In at least one embodiment, graphics acceleration module 2146 or individual graphics processing engines 2131-2132, N use a process handle to select a process element. In one embodiment, a process element is stored in system memory 2114 and can be addressed using effective to real address translation techniques described herein. In at least one embodiment, a process handle can be an implementation-specific value provided to a host process when it registers its context with a graphics processing engine 2131-2132, N (i.e., calls system software to add a process element to a process element linked list). In at least one embodiment, a lower 16 bits of a process handle can be an offset into a process element linked list for a process element.
[0324] Figure 21D An exemplary accelerator integration slice 2190 is shown. As used herein, a “slice” comprises a specified portion of processing resources of accelerator integration circuit 2136. An application is an effective address space 2182 in system memory 2114 that stores a process element 2183. In at least one embodiment, a process element 2183 is stored in response to a GPU invocation 2181 from an application 2180 executing on processor 2107. In at least one embodiment, a process element 2183 contains process state for a respective application 2180. In one embodiment, a work descriptor (WD) 2184 contained in process element 2183 can be a single job requested by an application or can contain a pointer to a queue of jobs. In at least one embodiment, WD 2184 is a pointer to a job request queue in an application’s address space 2182.
[0325] Graphics acceleration module 2146 and / or individual graphics processing engines 2131-2132, N can be shared by all or a subset of processes in a system. In at least one embodiment, can include infrastructure for setting up process state and sending a WD 2184 to a graphics acceleration module 2146 to start a job in a virtualized environment.
[0326] In at least one embodiment, a dedicated process programming model is implementation specific. In this model, a single process owns a graphics acceleration module 2146 or individual graphics processing engines 2131-2132, N. As the graphics acceleration module 2146 is owned by a single process, a hypervisor initializes the accelerator integration circuit for the owned partition, and an operating system initializes the accelerator integration circuit 2136 for the owned process when the graphics acceleration module 2146 is assigned.
[0327] In operation, a WD fetch unit 2191 in an accelerator integration slice 2190 fetches a next WD 2184, which includes an indication of work to be completed by one or more graphics processing engines of a graphics acceleration module 2146. In at least one embodiment, data from WD 2184 can be stored in registers 2145 and used by MMU 2139, interrupt management circuit 2147, and / or context management circuit 2148, as illustrated. For example, one embodiment of MMU 2139 includes segment / page walk circuitry to access segment / page tables 2186 within an OS virtual address space 2185. In at least one embodiment, interrupt management circuit 2147 can handle interrupt events 2192 received from a graphics acceleration module 2146. In at least one embodiment, effective addresses 2193 generated by graphics processing engines 2131-2132, N when performing graphics operations are translated to real addresses by MMU 2139.
[0328] In one embodiment, a same set of registers 2145 is replicated for each graphics processing engine 2131-2132, N and / or graphics acceleration module 2146, and can be initialized by a hypervisor or operating system. In at least one embodiment, each of these replicated registers can be included in an accelerator integration slice 2190. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.
[0329] Table 1 - Hypervisor-Initialized Registers
[0330]
[0331]
[0332] Exemplary registers that can be initialized by an operating system are shown in Table 2.
[0333] Table 2 - Operating System-Initialized Registers
[0334] 1 Process and thread identification 2 Effective address (EA) context save / restore pointer 3 Virtual address (VA) accelerator utilization record pointer 4 Virtual address (VA) storage segment table pointer 5 Authority mask 6 Work descriptor
[0335] In one embodiment, each WD 2184 is specific to a particular graphics acceleration module 2146 and / or graphics processing engine 2131-2132, N. In at least one embodiment, it contains all information needed for the graphics processing engine 2131-2132, N to complete the work, or it can be a pointer to a memory location where an application has set up a command queue of work to be completed.
[0336] Figure 21E Additional details of one exemplary embodiment of a shared model are shown. This embodiment includes a hypervisor real address space 2198 in which a list of process elements 2199 is stored. The hypervisor real address space 2198 is accessible via a hypervisor 2196 that virtualizes the graphics acceleration module engine for operating systems 2195.
[0337] In at least one embodiment, a shared programming model allows all processes or a subset of processes from all partitions or a subset of partitions in a system to use a graphics acceleration module 2146. There are two programming models in which a graphics acceleration module 2146 is shared by multiple processes and partitions, i.e., time-sliced sharing and graphics-directed sharing.
[0338] In this model, the system hypervisor 2196 owns the graphics acceleration module 2146 and makes its functionality available to all operating systems 2195. In at least one embodiment, for a graphics acceleration module 2146 to support virtualization by the system hypervisor 2196, the graphics acceleration module 2146 can adhere to the following: (1) application job requests must be autonomous (i.e., no state needs to be kept between jobs), or the graphics acceleration module 2146 must provide a context save and restore mechanism, (2) the graphics acceleration module 2146 guarantees that an application’s job request completes within a specified amount of time, including any translation faults, or the graphics acceleration module 2146 provides the ability to preempt processing of a job, and (3) fairness between graphics acceleration module 2146 processes must be ensured when operating in a directed shared programming model.
[0339] In one embodiment, the application 2180 needs to make an operating system 2195 system call using a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the graphics acceleration module type describes a target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 2146 and can take the form of a graphics acceleration module 2146 command, a valid address pointer to a user-defined structure, a valid address pointer to a command queue, or any other data structure describing work to be done by the graphics acceleration module 2146. In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to how an application program sets the AMR. If the accelerator integration circuit 2136 and graphics acceleration module 2146 implementation does not support a user authority mask override register (UAMOR), the operating system can apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. The hypervisor 2196 can selectively apply the current authority mask override register (AMOR) value before placing the AMR in the process element 2183. In at least one embodiment, the CSRP is one of registers 2145 that contains the effective address of a region in the application’s address space 2182 for the graphics acceleration module 2146 to save and restore context state. This pointer is optional if there is no need to save state between jobs or when a job is preempted. In at least one embodiment, the context save / restore region can be a fixed system memory.
[0340] Upon receiving the system call, the operating system 2195 can verify that the application 2180 is registered and has been granted authority to use the graphics acceleration module 2146. The operating system 2195 then invokes the hypervisor 2196 using the information shown in Table 3.
[0341] Table 3 - Operating system to hypervisor call parameters
[0342]
[0343]
[0344] Upon receiving the hypervisor call, the hypervisor 2196 verifies that the operating system 2195 is registered and has been granted authority to use the graphics acceleration module 2146. The hypervisor 2196 then places the process element 2183 in a process element linked list of the corresponding graphics acceleration module 2146 type. The process element can include the information shown in Table 4.
[0345] Table 4 - Process Element Information
[0346] 1 Work descriptor (WD) 2 Authority mask register (AMR) value (possibly masked) 3 Effective address (EA) context save / restore area pointer (CSRP) 4 Process ID (PID) and optional thread ID (TID) 5 Virtual address (VA) accelerator utilization record pointer (AURP) 6 Storage segment table pointer virtual address (SSTP) 7 Logical interrupt service number (LISN) 8 Interrupt vector table derived from hypervisor call parameters 9 State register (SR) value 10 Logical partition ID (LPID) 11 Real address (RA) hypervisor accelerator utilization record pointer 12 Storage descriptor register (SDR)
[0347] In at least one embodiment, hypervisor initializes a plurality of accelerator integration slice 2190 registers 2145.
[0348] As Figure 21F shown in at least one embodiment, a unified memory is used that is addressable via a common virtual memory address space for accessing physical processor memory 2101-2102 and GPU memory 2120-2123. In this implementation, operations executing on GPUs 2110-2113 utilize the same virtual / effective memory address space to access processor memory 2101-2102 and vice versa, simplifying programmability. In at least one embodiment, a first portion of virtual / effective address space is allocated to processor memory 2101, a second portion to second processor memory 2102, a third portion to GPU memory 2120, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as effective address space) is thus distributed among processor memory 2101-2102 and GPU memory 2120-2123, allowing any processor or GPU to access a memory with a virtual address that maps to that memory.
[0349] In one embodiment, bias / coherence management circuitry 2194A-2194E within one or more MMU 2139A-2139E ensures cache coherency between one or more host processors (e.g., 2105) and caches of GPUs 2110-2113, and implements bias techniques that dictate a bias of physical memory where certain types of data should be stored. While multiple instances of bias / coherence management circuitry 2194A-2194E are shown in Figure 21F FIG. 21, bias / coherence circuitry can be implemented within MMU(s) of one or more host processors 2105 and / or within accelerator integration circuit 2136.
[0350] One embodiment allows GPU-attached memory 2120-2123 to be mapped as part of system memory and accessed using shared virtual memory (SVM) techniques, but without suffering the performance penalties associated with full system cache coherency. In at least one embodiment, the ability to access GPU-attached memory 2120-2123 as system memory without the heavy cache coherency overhead provides a favorable operating environment for GPU offload. This arrangement allows software of host processor 2105 to set operands and access computation results without the overhead of traditional I / O DMA data copies. In at least one embodiment, such traditional copies include driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, which are all less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU-attached memory 2120-2123 without cache coherency overhead can be critical to the execution time of offloaded computations. For example, in cases with a large amount of streaming write memory traffic, cache coherency overhead can significantly reduce the effective write bandwidth seen by GPU 2110-2113. In at least one embodiment, the efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation can all play a role in determining the effectiveness of GPU offload.
[0351] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. In at least one embodiment, for example, a bias table can be used, which can be a page-granularity structure (i.e., controlled at the granularity of a memory page) that includes a 1 or 2 bits per GPU-attached memory page. In at least one embodiment, with or without a bias cache in GPU 2110-2113 (e.g., to cache frequently / recently used entries of the bias table), the bias table can be implemented in the stolen memory range of one or more GPU-attached memories 2120-2123. Alternatively, the entire bias table can be maintained within the GPU.
[0352] In at least one embodiment, prior to actually accessing GPU memory, the bias table entry associated with each access to GPU-attached memory 2120-2123 is accessed, resulting in the following operations. First, local requests from GPUs 2110-2113 that find their pages in GPU bias are forwarded directly to the corresponding GPU memory 2120-2123. Local requests from GPUs that find their pages in host bias are forwarded to processor 2105 (e.g., over a high-speed link as described herein). In one embodiment, requests from processor 2105 that find the requested pages in host processor bias complete the request similarly to a normal memory read. Alternatively, requests that point to GPU-biased pages can be forwarded to GPUs 2110-2113. In at least one embodiment, if the page is not currently in use by the GPU, the GPU can then migrate the page to host processor bias. In at least one embodiment, the bias state of a page can be changed by software-based mechanisms, hardware-assisted software-based mechanisms, or in limited cases purely hardware-based mechanisms.
[0353] One mechanism for changing bias state employs an API call (e.g., OpenCL) that in turn invokes a device driver of a GPU, which in turn sends a message (or causes a command descriptor to be enqueued) to the GPU, directing the GPU to change the bias state, and in certain migrations to perform a cache flush operation in the host. In at least one embodiment, the cache flush operation is used for migrations from host processor 2105 bias to GPU bias, but not for the reverse.
[0354] In one embodiment, cache coherency is maintained by temporarily rendering GPU-biased pages that host processor 2105 cannot cache. In at least one embodiment, to access these pages, processor 2105 can request access from GPU 2110, which can or can not grant access immediately. Thus, in at least one embodiment, to reduce communication between processor 2105 and GPU 2110, it is beneficial to ensure that GPU-biased pages are pages that are needed by the GPU and not by host processor 2105, and vice versa.
[0355] One or more hardware structures 1315 are used to perform one or more embodiments. Details regarding one or more hardware structures 1315 can be found in this document in connection with Figure 13A and / or Figure 13B Details regarding one or more hardware structures 1315 are provided.
[0356] Figure 22Exemplary integrated circuits and associated graphics processors in accordance with various embodiments described herein are shown, which can be fabricated using one or more IP cores. In addition to the illustrated, other logic and circuitry can be included in the at least one embodiment, including additional graphics processors / cores, peripheral interface controllers or general purpose processor cores.
[0357] Figure 22 is a block diagram illustrating an exemplary system on a chip integrated circuit 2200 that can be fabricated using one or more IP cores, in accordance with at least one embodiment. In at least one embodiment, integrated circuit 2200 includes one or more application processor(s) 2205 (e.g., CPUs), at least one graphics processor 2210, and can additionally include an image processor 2215 and / or a video processor 2220, any of which can be a modular IP core. In at least one embodiment, integrated circuit 2200 includes peripheral or bus logic including a USB controller 2225, a UART controller 2230, an SPI / SDIO controller 2235, and an I2S / I2C controller 2240. In at least one embodiment, integrated circuit 2200 can include a display device 2245 coupled to one or more of a high-definition multimedia interface (HDMI) controller 2250 and a mobile industry processor interface (MIPI) display interface 2255. In at least one embodiment, storage can be provided by a flash memory subsystem 2260 including flash memory and a flash memory controller. Memory interface can be provided via a memory controller 2265 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 2270. 2 S / I 2 Ccontroller 2240. In at least one embodiment, integrated circuit 2200 can include a display device 2245 coupled to one or more of a high-definition multimedia interface (HDMI) controller 2250 and a mobile industry processor interface (MIPI) display interface 2255. In at least one embodiment, storage can be provided by a flash memory subsystem 2260 including flash memory and a flash memory controller. Memory interface can be provided via a memory controller 2265 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 2270.
[0358] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In examples in which Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In examples in which inference and / or training logic 1315 are used to perform inferencing operations, at least portions of inference and / or training logic 1315 can be considered to be a form of
[0359] The above-described techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0360] Figures 23A-23BExemplary integrated circuits and associated graphics processors according to various embodiments described herein can be fabricated using one or more IP cores. In addition to the illustrated IP cores, other logic and circuits can be included in the integrated circuits in at least one embodiment, including e.g., additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0361] Figures 23A-23B is a block diagram illustrating an exemplary graphics processor used within SoCs according to embodiments described herein. Figure 23A An exemplary graphics processor 2310 of a system on a chip integrated circuit according to at least one embodiment is shown, which can be fabricated using one or more IP cores. Figure 23B An additional exemplary graphics processor 2340 of a system on a chip integrated circuit according to at least one embodiment is shown, which can be fabricated using one or more IP cores. In at least one embodiment, Figure 23A The graphics processor 2310 of is a low power graphics processor core. In at least one embodiment, Figure 23B The graphics processor 2340 of is a higher performance graphics processor core. In at least one embodiment, each graphics processor 2310, 2340 can be Figure 22 a variant of the graphics processor 2210 of
[0362] In at least one embodiment, the graphics processor 2310 includes a vertex processor 2305 and one or more fragment processor(s) 2315A-2315N (e.g., 2315A, 2315B, 2315C, 2315D, through 2315N-1, and 2315N). In at least one embodiment, the graphics processor 2310 can execute different shader programs via separate logic for vertex processing, so vertex processor 2305 is optimized to be a lower power than one or more fragment processers 2315A-2315N, which can be optimized for a higher performance. In at least one embodiment, the vertex processor 2305 executes operations along a single instruction stream to produce vertex data. The vertex channel processes the vertex data to produce finished vertices which are passed to the memory and cache of the one or more fragment processor(s) 2315A-2315N. In at least one embodiment, the graphics multi-processor 2310 can be of any appropriate type such as Intel GMA, AMD Radeon, Nvidia GeForce, or any other graphics processor core of any other company. In at least one embodiment, the graphics processor 2310 includes a ring interconnect 2320. In at least one embodiment, the ring interconnect 2320 is a high-speed interprocessor interconnect that is used to communicate between the vertex processor 2305 and the one or more fragment processor(s) 2315A-2315N. In at least one embodiment, the ring interconnect 2320 is a ring based input queue having descramble logic for writing and scatter logic for reading.
[0363] In at least one embodiment, graphics processor 2310 additionally includes one or more memory management units (MMUs) 2320A-2320B, cache memory 2325A-2325B, and one or more circuit interconnects 2330A-2330B. In at least one embodiment, one or more MMUs 2320A-2320B provide for virtual to physical address mapping for graphics processor 2310, including for vertex processor 2305 and / or fragment processor 2315A-2315N, which can reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more cache memories 2325A-2325B. In at least one embodiment, one or more MMUs 2320A-2320B can be synchronized with one or more MMUs within Figure 22 application processors 2205, image processors 2215, and / or video processors 2220, such that each processor 2205-2220 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 2330A-2330B enable graphics processor 2310 to interface with other IP cores within a SoC, either via an internal bus, or via a direct connection.
[0364] In at least one embodiment, graphics processor 2340 includes one or more MMUs 2320A-2320B, caches 2325A-2325B, and circuit interconnects 2330A-2330B of graphics processor 2310 of FIG. 23. In at least one embodiment, graphics processor 2340 includes one or more shader core(s) 2355A-2355N (e.g., 2355A, 2355B, 2355C, 2355D, 2355E, 2355F through 2355N-1, and 2355N) that provide for a unified shader core architecture in which a single core or type or core can be Figure 23A In at least one embodiment, graphics processor 2340 includes one or more shader core(s) 2355A-2355N (e.g., 2355A, 2355B, 2355C, 2355D, 2355E, 2355F through 2355N-1, and 2355N) that provide for a unified shader core architecture in which a single core or type or core can be
[0365] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are described in more detail below. Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided. In at least one embodiment, inference and / or training logic 1315 can be used in Figure 23A and / or Figure 23B for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions or architectures, or neural network use cases described herein.
[0366] The above techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0367] Figures 24A-24B Additional exemplary graphics processor logic in accordance with the embodiments described herein is shown. In at least one embodiment, graphics processor 2210 can include graphics core(s) 2400 that can be included within graphics processor 2210, and in at least one embodiment, can be unified shader core(s) 2355A-2355N as shown in FIG. 23C. Figure 24A Graphics core(s) 2400 are shown that can be included within graphics processor 2210, and in at least one embodiment, can be unified shader core(s) 2355A-2355N as shown in FIG. 23C. Figure 22 Graphics core(s) 2400 are shown that can be included within graphics processor 2210, and in at least one embodiment, can be unified shader core(s) 2355A-2355N as shown in FIG. 23C. Figure 23B Graphics core(s) 2400 are shown that can be included within graphics processor 2210, and in at least one embodiment, can be unified shader core(s) 2355A-2355N as shown in FIG. 23C. Figure 24B A highly parallel general purpose graphics processing unit (“GPGPU”) 2430 suitable for deployment on a multi-chip module is shown in at least one embodiment.
[0368] In at least one embodiment, graphics core 2400 includes a shared instruction cache 2402, texture unit 2418, and cache / shared memory 2420, which are common to execution resources within graphics core 2400. In at least one embodiment, graphics core 2400 can include multiple slices 2401A-2401N or partitions of each core and graphics processor can include multiple instances of graphics core 2400. In at least one embodiment, slices 2401A-2401N can include support logic including a local instruction cache 2404A-2404N, a thread scheduler 2406A-2406N, a thread dispatcher 2408A-2408N, and a set of registers 2410A-2410N. In at least one embodiment, slices 2401A-2401N can include a set of additional functional units (AFUs 2412A-2412N), floating point units (FPUs 2414A-2414N), integer arithmetic logic units (ALUs 2416A-2416N), address computation units (ACUs 2413A-2413N), double precision floating point units (DPFPUs 2415A-2415N), and matrix processing units (MPUs 2417A-2417N).
[0369] In at least one embodiment, FPUs 2414A-2414N can perform single precision (32-bit) and half precision (16-bit) floating point operations, while DPFPUs 2415A-2415N perform double precision (64-bit) floating point operations. In at least one embodiment, ALUs 2416A-2416N can perform variable precision integer operations in 8-bit, 16-bit, and 32-bit precision, and can be configured to operate in mixed precision formats. In at least one embodiment, MPUs 2417A-2417N can also be configured for mixed precision matrix operations including half precision floating point operations and 8-bit integer operations. In at least one embodiment, MPUs 2417-2417N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general matrix to matrix multiplication (GEMM). In at least one embodiment, AFUs 2412A-2412N can perform additional logical operations not supported by floating point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).
[0370] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In at least one embodiment, inference and / or training logic 1315 can be used in graphics processor 1312 for inferencing or predicting operations based at least in part on weight parameters calculated using one or more neural networks. Figure 13A and / or Figure 13BDetails regarding inference and / or training logic 1315 are provided. In at least one embodiment, inference and / or training logic 1315 can be used in graphics core 2400 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.
[0371] The above-described techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate a grasp of an object held by a human hand as described above.
[0372] Figure 24B A general purpose processing unit (GPGPU) 2430 is shown in at least one embodiment, which can be configured to enable highly parallel computing operations to be performed by a group of graphics processing units. In at least one embodiment, GPGPU 2430 can be linked directly to other instances of GPGPU 2430 to create a multi-GPU cluster to improve speed of training for deep neural networks. In at least one embodiment, GPGPU 2430 includes a host interface 2432 to enable connection to a host processor. In at least one embodiment, host interface 2432 is a PCI Express interface. In at least one embodiment, host interface 2432 can be a proprietary
[0373] In at least one embodiment, GPGPU 2430 includes memory 2444A-2444B coupled with compute clusters 2436A-2436H via a set of memory controllers 2442A-2442B. In at least one embodiment, memory 2444A-2444B can include various types of memory devices including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.
[0374] In at least one embodiment, compute clusters 2436A-2436H each include a group of graphics cores, such as graphics cores 2438A-2438F, which can be configured to perform Figure 24Agraphics core 2400, which can include multiple types of integer and floating point logic units, which can perform compute operations across a range of precisions for a computer, including precisions suitable for machine learning computations. For example, in at least one embodiment, at least a subset of floating point units in each compute cluster 2436A-2436H can be configured to perform 16- or 32-bit floating point operations, while a different subset of floating point units can be configured to perform 64-bit floating point operations.
[0375] In at least one embodiment, multiple instances of GPGPU 2430 can be configured to function as compute clusters. In at least one embodiment, communication of compute clusters 2436A-2436H for synchronization and data exchange varies between embodiments. In at least one embodiment, multiple instances of GPGPU 2430 communicate over host interface 2432. In at least one embodiment, GPGPU 2430 includes an I / O hub 2439 that couples GPGPU 2430 with a GPU link 2440 that enables direct connection to other instances of GPGPU 2430. In at least one embodiment, GPU link 2440 is coupled to a specialized GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGP 2430. In at least one embodiment, GPU link 2440 is coupled with a high speed interconnect to transmit and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 2430 are located in separate data processing systems and communicate over a network device that is accessible through host interface 2432. In at least one embodiment, GPU link 2440 can be configured to enable connection to a host processor in addition to or as an alternative to host interface 2432.
[0376] In at least one embodiment, GPGPU 2430 can be configured to train neural networks. In at least one embodiment, GPGPU 2430 can be used within an inferencing platform. In at least one embodiment, where GPGPU 2430 is used for inferencing, GPGPU 2430 can include fewer compute clusters 2436A-2436H relative to when GPGPU 2430 is used to train neural networks. In at least one embodiment, memory technology associated with memory 2444A-2444B can vary between inferencing and training configurations, with higher bandwidth memory technology dedicated to training configurations. In at least one embodiment, an inferencing configuration of GPGPU 2430 can support inferencing specific instructions. For example, in at least one embodiment, an inferencing configuration can provide support for one or more 8-bit integer dot product instructions, which can be used during inferencing operations for deployed neural networks.
[0377] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In various embodiments, inference and / or training logic 1315 can be used in place of or in conjunction with inference and / or training logic 1215, 1311, and / or 1317 described above. Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In various embodiments, inference and / or training logic 1315 can be used in place of or in conjunction with inference and / or training logic 1215, 1311, and / or 1317 described above.
[0378] The above-described techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0379] Figure 25 A block diagram of a computer system 2500 is shown, in accordance with at least one embodiment. In at least one embodiment, computer system 2500 includes a processing subsystem 2501 with one or more processors 2502 and a system memory 2504 communicating via an interconnection path 2505, which can include a memory hub 2505. In at least one embodiment, memory hub 2505 can be a separate component coupled with one or more processors 2502, or can be integrated into one or more processors 2502. In at least one embodiment, memory hub 2505 couples with an I / O subsystem 2511 via an interconnection path 2506. In at least one embodiment, I / O subsystem 2511 includes an I / O hub 2507 that can enable computing system 2500 to receive input from and provide output to one or more input and output devices 2508. In at least one embodiment, I / O hub 2507 can enable multiple processing subsystems 2501 to communicate with one or more input and output devices 2508 via one or more I / O hubs 2507. In at least one embodiment, one or more processing subsystems 2501 can include additional I / O hubs 2507 to enable communication with one or more input and output devices 2508.
[0380] In at least one embodiment, processing subsystem 2501 includes one or more parallel processors 2512 coupled to memory hub 2505 via a bus or other communication link 2513. In at least one embodiment, communication link 2513 can be any of a number of standard communication links, such as, but not limited to, a PCI Express bus or other bus, and can be of a different communication link depending on particular implementation needs.
[0381] In at least one embodiment, system storage 2514 can connect to memory hub 2507 for use in processing system 2500. In at least one embodiment, I / O switch 2516 can be used to provide an interface mechanism to connect I / O hub 2507 to other components, such as network adapter 2518 and / or wireless network adapter 2519 that can be integrated into a platform, as well as various other devices that can be added via one or more add-in devices 2520. In at least one embodiment, network adapter 2518 can be an Ethernet adapter or another wired
[0382] In at least one embodiment, computer system 2500 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which can also be connected to I / O hub 2507. In at least one embodiment, interconnection of the various components of system 2500 can be implemented using any suitable protocols, including those based on PCI (Peripheral Component Interconnect) or other bus or point-to-point communication interfaces and / or protocols, such as NV-Link high-speed interconnect, or interconnect protocols. Figure 25
[0383] In at least one embodiment, parallel processor(s) 2512 include circuitry optimized for graphics and video processing, including i.e., video output circuitry, and is configured for use in a gaming console, a mobile phone, a personal computer, a workstation, and / or a similar electronic device. In at least one embodiment, parallel processor(s) 2512 include circuitry optimized for general use applications, including i.e., circuitry configured to execute general-purpose computational tasks, and is configurable for use in a gaming console, a mobile phone, a personal computer, a workstation, and / or a similar electronic device. In at least one embodiment, components of computer system 2500 can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, parallel processor(s) 2512, memory hub 2505, processor(s) 2502, and I / O hub 2507 can be integrated into a system on a chip (SoC) integrated circuit. In at least one embodiment, components of computer system 2500 can be integrated into a single package to form a system in a package (SIP) configuration. In at least one embodiment, at least a portion of components of computer system 2500 can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules into a module-in-module (MIM) configuration.
[0384] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In at least one embodiment, inference and / or training logic 1315 can be used in Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In at least one embodiment, inference and / or training logic 1315 can be used in system 2500 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. Figure 25 Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In at least one embodiment, inference and / or training logic 1315 can be used in system 2500 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0385] The above-described techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0386] Processor
[0387] Figure 26A A parallel processor 2600, according to at least one embodiment, is shown in FIG. 26. In at least one embodiment, various components of parallel processor 2600 can be implemented using one or more integrated circuits, which can be programmable integrated circuits, application-specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). Figure 25 Variations of one or more parallel processors 2512 are shown.
[0388] In at least one embodiment, parallel processor 2600 includes a parallel processing unit 2602. In at least one embodiment, parallel processing unit 2602 includes an I / O unit 2604 that enables communication with other devices, including other instances of parallel processing unit 2602. In at least one embodiment, I / O unit 2604 can be directly connected to other devices. In at least one embodiment, I / O unit 2604 connects with other devices via use of a hub or switch, for example, memory hub 2505. In at least one embodiment, connections between memory hub 2505 and I / O unit 2604 form a communication link 2513. In at least one embodiment, I / O unit 2604 connects with a host interface 2606 and a memory crossbar switch 2616, where host interface 2606 receives commands directed to processing operations and memory crossbar switch 2616 receives commands directed to memory operations.
[0389] In at least one embodiment, when host interface 2606 receives a command buffer via I / O unit 2604, host interface 2606 can direct work operations to execute those commands to front end 2608. In at least one embodiment, front end 2608 couples with a scheduler 2610, which is configured to assign commands or other work items to a processing cluster array 2612. In at least one embodiment, scheduler 2610 ensures that processing cluster array 2612 is correctly configured and in an active state before tasks are assigned to processing cluster array 2612. In at least one embodiment, scheduler 2610 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, microcontroller implemented scheduler 2610 is configurable to perform complex scheduling and work allocation operations on a coarse and fine grain level for implementation of a fast thread pre-emption and context switch for threads executing on processing array 2612. In at least one embodiment, host software can prove a workload for scheduling on processing array 2612 through one of a number of graphics processing doorbells. In at least one embodiment, workload can then be automatically allocated on processing array 2612 by scheduler 2610 logic within microcontroller including scheduler 2610.
[0390] In at least one embodiment, processing cluster array 2612 can include up to “N” processing clusters (e.g., cluster 2614A, cluster 2614B, through cluster 2614N). In at least one embodiment, each cluster 2614A-2614N of processing cluster array 2612 can execute a large number of concurrent threads. In at least one embodiment, scheduler 2610 can allocate work to clusters 2614A-2614N of processing cluster array 2612 using various scheduling and / or work distribution algorithms, which can be determined by workload produced by each program or computational type. In at least one embodiment, scheduling can be handled dynamically by scheduler 2610, or can be aided in part by compiler logic during compilation of program logic configured for execution by processing cluster array 2612. In at least one embodiment, different clusters 2614A-2614N of processing cluster array 2612 can be allocated for processing different types of programs or for performing different types of computations.
[0391] In at least one embodiment, processing cluster array 2612 can be configured to perform a variety of types of parallel processing operations. In at least one embodiment, processing cluster array 2612 is configured to perform general-purpose parallel compute operations. For example, in at least one embodiment, processing cluster array 2612 can include logic to perform processing tasks comprising filtering of video and / or audio data, performing modeling operations, including physics operations, and performing data transformations.
[0392] In at least one embodiment, processing cluster array 2612 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 2612 can include additional logic to support the performance of such graphics processing operations including, but not limited to, texture sampling logic to perform texture operations for three-dimensional (3D) graphics, and tessellation logic and other vertex processing logic. In at least one embodiment, processing cluster array 2612 can be configured to execute shader programs associated with the performance of graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing unit 2602 can transfer data to be processed from system memory via I / O unit 2604. In at least one embodiment, during processing, results can be stored back to system memory after processing or processed data can be written directly to graphics memory.
[0393] In at least one embodiment, when parallel processing unit 2602 is used to perform graphics processing, scheduler 2610 can be configured to divide the processing workload into approximately equal sized tasks to better enable distribution of graphics processing operations across multiple clusters 2614A-2614N of processing cluster array 2612. In at least one embodiment, portions of processing cluster array 2612 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to
[0394] In at least one embodiment, processing cluster array 2612 can receive processing tasks to be executed from scheduler 2610, which receives commands defining the processing tasks from front end 2608. In at least one embodiment, processing tasks can comprise indices of data to be processed, e.g., surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands (e.g., what programs to execute) that control how the data is to be processed. In at least one embodiment, scheduler 2610 can be configured to fetch the indices corresponding to a task, or can receive the indices from front end 2608. In at least one embodiment, front end 2608 can be configured to ensure that processing cluster array 2612 is configured in an effective state before a workload initiated by an incoming command buffer (e.g., a batch-buffer, a push buffer, etc.) is launched.
[0395] In at least one embodiment, each of one or more instances of parallel processing unit 2602 can be coupled to a parallel processor memory 2622. In at least one embodiment, parallel processor memory 2622 can be accessed by the memory crossbar 2616, which can receive memory requests from the processing cluster array 2612 and I / O units 2604. In at least one embodiment, memory crossbar 2616 can access parallel processor memory 2622 via a memory interface 2618. In at least one embodiment, memory interface 2618 can include a number of memory ports and can be configured to communicate data with memory component 2610 via one or more memory channels 2626. In at least one embodiment, memory component 2610 can be one or more volatile memory devices, such as synchronous dynamic random access memory (SDRAM), dynamic interactive random access memory (DRAM), or the like. In at least one embodiment, memory component 2610 can be configured as cache memory. In at least one embodiment, memory component 2610 can be non-volatile memory devices, such as flash memory, readonly memory, programmable read-only memory (PROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), early eeprom (E2PROM), video rom (VRom), or the like.
[0396] In at least one embodiment, memory units 2624A-2624N can include various types of memory devices including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 2624A-2624N can also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps can be stored across memory units 2624A-2624N, allowing partition units 2620A-2620N to write portions of each rendering target in parallel to effectively use available bandwidth of parallel processor memory 2622. In at least one embodiment, local instances of parallel processor memory 2622 can be excluded to facilitate a unified memory design that utilizes system memory in combination with local cache memory.
[0397] In at least one embodiment, any of clusters 2614A-2614N of processing cluster array 2612 can process data to be written into any of memory locations 2624A-2624N within parallel processor memory 2622. In at least one embodiment, memory crossbar 2616 can be configured to transmit outputs of each cluster 2614A-2614N to any partition unit 2620A-2620N or another cluster 2614A-2614N, which can perform additional processing operations on the outputs. In at least one embodiment, each cluster 2614A-2614N can communicate with memory interface 2618 through memory crossbar 2616 to read from or write to various external memory devices. In at least one embodiment, memory crossbar 2616 has a connection to memory interface 2618 to communicate with I / O unit 2604, and a local instance of connection to parallel processor memory 2622, to enable processing clusters 2614A-2614N within different processing clusters 2614A-2614N to communicate with system memory or other memories not local to parallel processing unit 2602. In at least one embodiment, memory crossbar 2616 can use virtual channels to separate traffic streams between clusters 2614A-2614N and partition units 2620A-2620N.
[0398] In at least one embodiment, multiple instances of parallel processing unit 2602 can be provided on a single add-in card, or multiple add-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 2602 can be configured to operate in coordination with each other to enable single program multiprocessor (SPM) functionality. In at least one embodiment, different instances of parallel processing unit 2602 can be configured to operate in a locked-step computing mode where the different instances of parallel processing unit 2602 execute the same program but execute different iterations of a program, or execute different programs at different times. In at least one embodiment, different instances of parallel processing unit 2602 can be configured to operate independently of each other.
[0399] Figure 26B is a block diagram of a partition unit 2620 in accordance with at least one embodiment. In at least one embodiment, partition unit 2620 is a Figure 26Aone of the partition units 2620A-2620N of FIG. 26. In at least one embodiment, partition unit 2620 includes an L2 cache 2621, a frame buffer interface 2625, and a ROP 2626 (Raster Operations Unit). L2 cache 2621 is a read / write cache that is configured to perform load and store operations received from memory crossbar 2616 and ROP 2626. In at least one embodiment, L2 cache 2621 outputs read misses and urgent write-back requests to frame buffer interface 2625 for processing. In at least one embodiment, updates can also be sent to a frame buffer via frame buffer interface 2625 for processing. In at least one embodiment, frame buffer interface 2625 interacts with one of memory units 2624A-2624N (e.g., within parallel processor memory 2622) in parallel processor memory. Figure 26A
[0400] In at least one embodiment, ROP 2626 is a processing unit that performs raster operations including, for example, fill, line, polygon, and / or tessellation operations. In at least one embodiment, ROP 2626 is used to perform raster operations on primitive drawings rendered through graphics processing pipeline 2600. In at least one embodiment, ROP 2626 includes, for example, storage for tags (e.g., depth buffers, stencil buffers, and / or the like). In at least one embodiment, ROP 2626 2626 includes a data forwarder that provides retrieved data to graphics processing pipeline 2600. In at least one embodiment, data forwarder provides the data to work distribution crossbar 2613.
[0401] In at least one embodiment, ROP 2626 is included within each processing cluster (e.g., clusters 2614A-2614N) of parallel processor 2600, instead of in partition unit 2620. In at least one embodiment, read and write requests for pixel data are transmitted over memory crossbar 2616 instead of pixel fragment data by memory crossbar 2616. In at least one embodiment, processed graphics data can be displayed on display device(s) 2510, routed by processor 2502 for further processing, or routed to one of processing entities within parallel processor 2600 for further processing. Figure 26A Figure 25 Figure 26A
[0402] Figure 26C is a block diagram of a processing cluster 2614 within a parallel processing unit in accordance with at least one embodiment. In at least one embodiment, processing cluster is a Figure 26A one of the processing clusters 2614A-2614N. In at least one embodiment, processing cluster 2614 can be configured to perform a number of threads in parallel, where a “thread” is an instance of a particular program executing on a particular set of input data. In at least one embodiment, Single-Instruction, Multiple-Data (SIMD) instruction issue techniques are used to support parallel execution of a single program thread by multiple execution units. In at least one embodiment, Single-Instruction, Multiple Thread (SIMT) techniques are used to support parallel execution of multiple program threads, using a common instruction pool and per-thread context-specific data stores.
[0403] In at least one embodiment, operation of processing cluster 2614 can be controlled via a pipeline manager 2632 that is assigned to processing tasks by scheduler 2610. In at least one embodiment, pipeline manager 2632 manages execution of instructions associated with a thread on a group of processing engines within processing cluster 2614, in at least one embodiment, pipeline manager 2632 manages fetch, decode, and exception handling for those instructions, and issues the instructions for execution. Figure 26A In at least one embodiment, processing cluster 2614 can include a graphics multiprocessor 2634 and / or a texture unit 2636 that provide graphics processing, including three dimensional (3D) graphics processing. In at least one embodiment, graphics multiprocessor 2634 executes graphics processing programs such as video
[0404] In at least one embodiment, each graphics multiprocessor 2634 within processing cluster 2614 can include an identical set of functional execution logic (e.g., arithmetic logic, load store units, etc.). In at least one embodiment, functional execution logic can be configured in a pipelined manner in which new instructions can be issued before previous instructions are complete. In at least one embodiment, functional execution logic supports a variety of operations including integer and floating point arithmetic, comparison operations, Boolean operations, shift operations, and a multitude of algebraic functions. In at least one embodiment, same functional-unit hardware can be leveraged to perform different operations using different settings of control bits in those instructions.
[0405] In at least one embodiment, instructions delivered to processing cluster 2614 constitute a thread. In at least one embodiment, a set of threads executing across a set of parallel processing engines constitutes a warp. In at least one embodiment, a thread group is a group of threads executing the same program, although each thread within a thread group can be at different instruction points within the program. In at least one embodiment, a thread group is associated with a same set of instruction boundaries. In at least one embodiment, a thread group includes fewer threads than are available processing engines within graphics multiprocessor 2634. In at least one embodiment, when a thread group includes fewer threads than the number of processing engines within graphics multiprocessor 2634, one or more processing engines can be idle during a cycle when a thread group is not available. In at least one embodiment, a thread group can include more threads than the number of processing engines within graphics multiprocessor 2634. In at least one embodiment, when a thread group includes more threads than the number of processing engines within graphics multiprocessor 2634, multiple threadside can be executed in a single cycle. In at least one embodiment, thread groups of different sizes can be supported.
[0406] In at least one embodiment, graphics multiprocessor 2634 includes internal cache memory, to perform load and store operations. In at least one embodiment, graphics multiprocessor 2634 can discard internal cache and use cache memory within processing cluster 2614 (e.g., LI cache 2648). In at least one embodiment, each graphics multiprocessor 2634 can also have access to L2 Cache within a partition unit (e.g., partition units 2620A-2620N) that is shared among all processing clusters 2614 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 2634 can also have access to off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processor 2602 can be used as global memory. In at least one embodiment, processing cluster 2614 includes multiple instances of graphics multiprocessor 2634, which can share common instructions and data stored in LI cache 2648. Figure 26A
[0407] In at least one embodiment, each processing cluster 2614 can include a memory management unit (MMU) 2645 to map virtual addresses into physical addresses, as is known to those skilled in the art. In at least one embodiment, one or more instances of MMU 2645 can reside within graphics multiprocessor 2634. In at least one embodiment, MMU 2645 can include address translation lookaside buffers (TLBs) to improve translation of virtual addresses into physical addresses. Figure 26A In at least one embodiment, MMU 2645 includes a set of page table entries (PTEs) used to map virtual addresses into physical addresses for task computation and / or graphics operations. In at least one embodiment, MMU 2645 can include address translation lookaside buffer (TLB) or can reside in graphics multiprocessor 2634 or within level 1 cache memory 2630. In at least one embodiment, processing physical addresses enables access locality to be determined for surface data accesses to efficiently alternate requests between partitions.
[0408] In at least one embodiment, processing clusters 2614 can be configured such that each graphics multiprocessor 2634 is coupled to a texture unit 2636 for performing texture mapping operations, e.g., determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within graphics multiprocessor 2634 as needed, and texture data is fetched from an L2 cache, local parallel processor memory, or system memory, as needed. In at least one embodiment, each graphics multiprocessor 2634 outputs processed tasks to data crossbar 2640 to provide processed tasks to another processing cluster 2614 for further processing or to store processed task data in an L2 cache, local parallel processor memory, or system memory via memory crossbar 2616. Figure 26A In at least one embodiment, preROP 2642 (Raster Operations Pre-unit) is configured to receive data from graphics multiprocessor 2634, direct data to ROP unit which can be located within or outside of processing cluster 2614, and perform optimizations related to color blending and pixel ordering. In at least one embodiment, ROP unit 2644 is configured to take processed tasks from preROP 2642 and generate image data based on those tasks. In at least one embodiment, image data is transmitted from ROP unit 2644 via memory crossbar 2616 to be stored in system memory, encoded on to a display, or both.
[0409] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 7 and 8. Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 7 and 8. In at least one embodiment, inference and / or training logic 1315 can be used in graphics processing cluster 2614 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0410] The above techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate a grasp for an object held by a human hand as described above.
[0411] Figure 26D A graphics processing unit 2634 according to at least one embodiment is shown. In at least one embodiment, graphics processing unit 2634 is coupled with a pipeline manager 2632 of processing cluster 2614. In at least one embodiment, graphics processing unit 2634 has an execution pipeline that includes, without limitation, an instruction cache 2652, an instruction unit 2654, an address mapping unit 2656, a register file 2658, one or more general-purpose graphics processing unit (GPGPU) cores 2662, and one or more load / store units 2666. In at least one embodiment, GPGPU cores 2662 and load / store units 2666 are coupled with cache memory 2672 and shared memory 2670 via a memory and cache interconnect 2668.
[0412] In at least one embodiment, instruction cache 2652 receives a stream of instructions 2650 to be executed by graphics processing unit 2634 from pipeline manager 2632. In at least one embodiment, instructions are cached in instruction cache 2652 and dispatched for execution by instruction unit 2654. In one embodiment, instruction unit 2654 can dispatch instructions as a thread group (e.g., a warp), assigning each thread of the thread group to a different execution unit within GPGPU cores 2662. In at least one embodiment, instructions can access any of a number of different address spaces through the use of addresses that are specified in a unified address space. In at least one embodiment, address mapping unit 2656 can be used to translate an address in the unified address space into a different address that can be accessed by load / store units 2666.
[0413] In at least one embodiment, register file 2658 provides a set of registers for functional units of graphics processing unit 2634. In at least one embodiment, register file 2658 provides temporary storage for operands of the data paths connected to the functional units (e.g., GPGPU cores 2662, load / store units 2666) of graphics processing unit 2634. In at least one embodiment, register file 2658 is partitioned between different thread groups being executed by graphics processing unit 2634 such that each thread group is allocated a dedicated section of register file 2658. In at least one embodiment, register file 2658 is partitioned between different warps being executed by graphics processing unit 2634.
[0414] In at least one embodiment, GPGPU cores 2662 can each include floating point units (FPUs) and / or integer arithmetic logic units (ALUs) that are capable of performing instructions for processing graphics primitives and / or performing general purpose computing tasks. In at least one embodiment, GPGPU cores 2662 can be similar to the GPGPU cores 2650 described herein and / or can be configured similarly to the GPGPU cores 2650 described herein. In at least one embodiment, first portion of GPGPU cores 2662 includes single precision floating point
[0415] In at least one embodiment, GPGPU cores 2662 include SIMD logic capable of performing a single instruction on multiple sets of data. In at least one embodiment, GPGPU cores 2662 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for GPGPU cores can be generated at compile time by a shader compiler or automatically generated when executing programs written and compiled for single program multiple data (SPMD) or SIMT architectures.
[0416] In at least one embodiment, memory and cache interconnect 2668 is an interconnect network that connects each functional unit of graphics multiprocessor 2634 to register file 2658 and shared memory 2670. In at least one embodiment, memory and cache interconnect 2668 is a crossbar interconnect that allows load / store units 2666 to implement load and store operations between shared memory 2670 and register file 2658. In at least one embodiment, register file 2658 can operate at same frequency as GPGPU cores 2662, such that latency of data transfers between GPGPU cores 2662 and register file 2658 is very low. In at least one embodiment, shared memory 2670 can be used to enable communication between threads executing on functional units within graphics multiprocessor 2634. In at least one embodiment, cache memory 2672 can be used to cache texture data communicated between texture unit 2636 and functional units. In at least one embodiment, shared memory 2670 can also be used as a program managed cache. In at least one embodiment, in addition to automatically cached data stored in cache memory 2672, threads executing on GPGPU cores 2662 can store data in shared memory in a programmed manner.
[0417] In at least one embodiment, parallel processor or GPGPU as described herein is communicatively coupled to host / processor cores to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general purpose GPU (GPGPU) functions. In at least one embodiment, GPU can be communicatively coupled to host processor / cores by a bus or other interconnect (e.g., a high speed
[0418] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In at least one embodiment, inference and / or training logic 1315 is used in conjunction with components of system 1300, for example, to perform inferences, predictions, calculations, and / or training operations associated with one or more embodiments described herein. Figure 13A and / or Figure 13BDetails regarding inference and / or training logic 1315 are provided. In at least one embodiment, inference and / or training logic 1315 can be used in graphics multiprocessor 2634 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0419] The above-described techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate a grasp of an object held by a human hand as described above.
[0420] Figure 27 A multi-GPU computing system 2700 according to at least one embodiment is shown. In at least one embodiment, multi-GPU computing system 2700 can include a processor 2702 coupled to a plurality of general purpose graphics processing units (GPGPUs) 2706A-D via a host interface switch 2704. In at least one embodiment, host interface switch 2704 is a PCI Express switch device that couples processor 2702 to a PCI Express bus over which processor 2702 can communicate with GPGPUs 2706A-D. GPGPUs 2706A-D can be interconnected via a set of high-speed P2P GPU-to-GPU links 2716. In at least one embodiment, GPU-to-GPU links 2716 connect to each of GPGPUs 2706A-D via a dedicated GPU link. In at least one embodiment, P2P GPU links 2716 enable direct communication between each GPGPU 2706A-D without having to communicate through host interface bus 2704 to which processor 2702 is connected. In at least one embodiment, host interface bus 2704 remains available for system memory access or communication with other instances of multi-GPU computing system 2700, for example, via one or more network devices, in event that GPU-to-GPU traffic is directed to P2P GPU links 2716. While, in at least one embodiment, GPGPUs 2706A-D are connected to processor 2702 via host interface switch 2704, in at least one embodiment, processor 2702 includes direct support for P2P GPU links 2716 and can be directly connected to GPGPUs 2706A-D.
[0421] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, inference and / or training logic 1315 used to perform inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. Figure 13A and / or Figure 13BDetails regarding inference and / or training logic 1315 are provided. In at least one embodiment, inference and / or training logic 1315 can be used in multi-GPU computing system 2700 for performing inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0422] The above-described techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate a grasp for an object held by a human hand as described above.
[0423] Figure 28 FIG. 28 is a block diagram of a graphics processor 2800 according to at least one embodiment. In at least one embodiment, graphics processor 2800 includes ring interconnect 2802, front-end pipeline 2804, media engine 2837, and graphics cores 2880A-2880N. In at least one embodiment, ring interconnect 2802 couples graphics processor 2800 to other processing units including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, graphics processor 2800 is one of a number of processors integrated within a multi-core processing system.
[0424] In at least one embodiment, graphics processor 2800 receives batches of commands via ring interconnect 2802. In at least one embodiment, incoming commands are interpreted by a command streamer 2803 in pipeline front-end 2804. In at least one embodiment, graphics processor 2800 includes scalable execution logic to perform 3D geometry processing and media processing via the graphics cores 2880A-2880N. In at least one embodiment, for 3D geometry processing commands, command streamer 2803 supplies commands to geometry pipeline 2836. In at least one embodiment, for at least some media processing commands, command streamer 2803 supplies commands to a video front end 2834, which couples with a media engine 2837. In at least one embodiment, media engine 2837 includes a video quality engine (VQE) 2830 for video and image post-processing, and a multi-format encode / decode (MFX) 2833 engine to provide hardware-accelerated media
[0425] In at least one embodiment, graphics processor 2800 includes a scalable thread execution resource including a graphics core 2880A-2880N (which sometimes referred to as a core slice) featuring multiple sub-cores 2850A-2850N, 2860A-2860N (sometimes referred to as a core sub-slice). In at least one embodiment, graphics processor 2800 can have any number of graphics cores 2880A-2880N. In at least one embodiment, graphics processor 2800 includes graphics core 2880A featuring at least a first sub-core 2850A and a second sub-core 2860A. In at least one embodiment, graphics processor 2800 is a low power processor with a single sub-core (e.g., 2850A). In at least one embodiment, graphics processor 2800 includes multiple graphics cores 2880A-2880N each including a set of first sub-cores 2850A-2850N and a set of second sub-cores 2860A-2860N. In at least one embodiment, each sub-core in first sub-cores 2850A-2850N includes at least a first set of execution units 2852A-2852N and a media / texture samplers 2854A-2854N. In at least one embodiment, each sub-core in second sub-cores 2860A-2860N includes at least a second set of execution units 2862A-2862N and samplers 2864A-2864N. In at least one embodiment, each sub-core 2850A-2850N, 2860A-2860N shares a set of shared resources 2870A-2870N. In at least one embodiment, shared resources include shared cache memory and pixel operation logic.
[0426] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In examples in which inference and / or training logic 1315 are used for inferencing, the inference and / or training logic 1315 can be used to implement neural network inference operations, including operations performed by neural network 1302. Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In examples in which inference and / or training logic 1315 are used for inferencing, the inference and / or training logic 1315 can be used to implement neural network inference operations, including operations performed by neural network 1302.
[0427] The techniques described above can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0428] Figure 29is a block diagram illustrating microarchitecture for a processor 2900 according to at least one embodiment, which can include logic circuits to execute instructions. In at least one embodiment, processor 2900 can execute instructions including x86 instructions, ARM instructions, specialized instructions for application specific integrated circuits (ASICs), and the like. In at least one embodiment, processor 2900 can include registers to store packed data, such as 64-bit wide MMX® registers enabled by Intel® MMX Technology in microprocessors by Intel Corporation, Santa Clara, CA, as well as SIMD TM registers available for integer and floating point number formats can operate with packed data elements accompanying single instruction multiple data (“SIMD”) and streaming SIMD extensions (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or higher (generically referred to as “SSEx”) technology can hold such packed data operands. In at least one embodiment, processor 2900 can execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0429] In at least one embodiment, processor 2900 includes an in-order front-end (“front-end”) 2901 to fetch instructions to be executed and to prepare instructions for execution by other pipelines. In at least one embodiment, front-end 2901 can include several units. In at least one embodiment, instruction prefetcher 2926 fetches instructions from memory and provides pre-fetched instructions to instruction decoder 2928 which, in turn, decodes or interprets instructions. For example, in at least one embodiment, instruction decoder 2928 decodes a received instruction into one or more operations called “micro-instructions” or “micro-operations” (also called “micro-op” or “uops”) that the machine can execute. In at least one embodiment, instruction decoder 2928 parses instruction headers to operational codes and corresponding data and control fields that can be used by micro-architecture to perform operations according to at least one embodiment. In at least one embodiment, trace cache 2930 can assemble decoded micro-instructions into program ordered sequences or traces in micro-instruction queue 2934 for execution. In at least one embodiment, when trace cache 2930 encounters a complex instruction, microcode ROM 2932 provides uops needed to complete operation.
[0430] In at least one embodiment, some instructions can be converted into a single micro- operation, while others can require several micro-operations to complete. In at least one embodiment, if more than four micro-instructions are needed to complete a single instruction, then instruction decoder 2928 can access microcode ROM 2932 to perform that instruction. In at least one embodiment, instructions can be decoded into a small number of micro-instructions to handle at instruction decoder 2928. In at least one embodiment, if multiple micro-instructions are needed to complete an operation, then an instruction can be stored in microcode ROM 2932. In at least one embodiment, a trace cache 2930 references an entry point programmable logic array (“PLA”) to determine a correct micro-instruction pointer for reading a microcode sequence from microcode ROM 2932 to complete one or more instructions, in accordance with at least one embodiment. In at least one embodiment, after microcode ROM 2932 completes sequencing of micro-operations for an instruction, a front end 2901 of a machine can resume fetching micro-operations from trace cache 2930.
[0431] In at least one embodiment, out-of-order execution engine (“out-of-order engine”) 2903 can prepare instructions for execution. In at least one embodiment, out-of-order execution logic has multiple buffers to smooth and reorder instruction flow to optimize performance as instructions are pipelined down and dispatched for execution. In at least one embodiment, out-of-order execution engine 2903 includes, without limitation, an allocator / register renamer 2940, a memory micro instruction queue 2942, an integer / float micro instruction queue 2944, a memory scheduler 2946, a fast scheduler 2902, a slow / general floating point scheduler (“slow / general FP scheduler”) 2904, and a simple floating point scheduler (“simple FP scheduler”) 2906. In at least one embodiment, fast scheduler 2902, slow / general floating point scheduler 2904, and simple floating point scheduler 2906 are also collectively referred to as “micro instruction schedulers 2902, 2904, 2906.” Allocator / register renamer 2940 allocates machine buffers and resources needed for each micro instruction to execute in sequence. In at least one embodiment, allocator / register renamer 2940 renames logical registers to entries in a register file. In at least one embodiment, allocator / register renamer 2940 also allocates entries for each micro instruction in one of two micro instruction queues, memory micro instruction queue 2942 for memory operations and integer / float micro instruction queue 2944 for non-memory operations, in front of memory scheduler 2946 and micro instruction schedulers 2902, 2904, 2906. In at least one embodiment, micro instruction schedulers 2902, 2904, 2906 determine when micro instructions are ready to execute based on readiness of their dependent input register operand sources and availability of execution resource micro instructions needed to complete. Fast scheduler 2902 of at least one embodiment can schedule on each half of a main clock cycle, while slow / general floating point scheduler 2904 and simple floating point scheduler 2906 can schedule once per main processor clock cycle. In at least one embodiment, micro instruction schedulers 2902, 2904, 2906 arbitrate for a dispatch port to dispatch micro instructions for execution.
[0432] In at least one embodiment, execution block 2911 includes, without limitation, integer register file / bypass network 2908, floating point register file / bypass network (“FP register file / bypass network”) 2910, address generation units (“AGUs”) 2912 and 2914, fast ALUs 2916 and 2918, slow ALUs 2920, floating point ALUs (“FPs”) 2922, and floating point move unit (“FP move”) 2924. In at least one embodiment, integer register file / bypass network 2908 and floating point register file / bypass network 2910 are also referred to herein as “register files 2908, 2910.” In at least one embodiment, AGUs 2912 and 2914, fast ALUs 2916 and 2918, slow ALUs 2920, floating point ALUs 2922, and floating point move unit 2924 are also referred to herein as “execution units 2912, 2914, 2916, 2918, 2920, 2922, and 2924.” In at least one embodiment, execution block 2911 can include, without limitation, any number (including zero) and type of register files, bypass networks, address generation units, and execution units (in any combination).
[0433] In at least one embodiment, register files 2908, 2910 can be arranged between micro-instruction schedulers 2902, 2904, 2906 and execution units 2912, 2914, 2916, 2918, 2920, 2922, and 2924. In at least one embodiment, integer register file / bypass network 2908 performs integer operations. In at least one embodiment, floating point register file / bypass network 2910 performs floating point operations. In at least one embodiment, each of register files 2908, 2910 can include, without limitation, a bypass network that can bypass or forward a result just completed that has not yet been written into a register file to a new dependee. In at least one embodiment, register files 2908, 2910 can communicate data to each other. In at least one embodiment, integer register file / bypass network 2908 can include, without limitation, two separate register files, one for low order 32 bits data, a second for high order 32 bits data. In at least one embodiment, floating point register file / bypass network 2910 can include, without limitation, 128 bit wide entries, as floating point instructions typically have operands that are 64 to 128 bits wide.
[0434] In at least one embodiment, execution units 2912, 2914, 2916, 2918, 2920, 2922, 2924 can execute instructions. In at least one embodiment, register files 2908, 2910 store integer and floating point data operand values that microinstructions need to execute. In at least one embodiment, processor 2900 can include, without limitation, any number and combination of execution units 2912, 2914, 2916, 2918, 2920, 2922, 2924. In at least one embodiment, floating point ALU 2922 and floating point move unit 2924 can execute floating point, MMX, SIMD, AVX and SSE, or other operations, including specialized machine learning instructions. In at least one embodiment, floating point ALU 2922 can include, without limitation, a 64 bit by 64 bit floating point divider to execute divide, square root, and remainder micro-ops. In at least one embodiment, instructions for floating point
[0435] In at least one embodiment, micro-instruction schedulers 2902, 2904, 2906 schedule dependent operations prior to completion of parent load execution. In at least one embodiment, because micro-instructions can be speculatively scheduled and executed in processor 2900, processor 2900 can also include logic to handle memory misses. In at least one embodiment, if a data load in a data cache misses, there can be a dependent operation running in a pipeline that causes the scheduler to temporarily have incorrect data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, dependent operations can need to be replayed and independent operations can be allowed to complete. In at least one embodiment, a scheduler and replay mechanism of at least one embodiment of a processor can also be designed to capture instruction sequences for text string compare operations.
[0436] In at least one embodiment, the term “register” can refer to an on-board processor storage location that can be used as part of an instruction that identifies an operand. In at least one embodiment, a register can be one that can be used from outside of a processor (from a programmer’s perspective). In at least one embodiment, a register can not be limited to a particular type of circuit. Rather, in at least one embodiment, a register can store data, provide data, and perform functions described herein. In at least one embodiment, registers described herein can be implemented by circuitry within a processor using a variety of different techniques, such as dedicated physical registers, physical registers allocated dynamically with register renaming, a combination of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, an integer register stores 32-bit integer data. A register file of at least one embodiment also contains eight multimedia SIMD registers for packing data.
[0437] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In various embodiments, inference and / or training logic 1315 can be used in place of or in conjunction with inference and / or training logic 1205 described above. Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In various embodiments, inference and / or training logic 1315 can be used in place of or in conjunction with inference and / or training logic 1205 described above. For example, in at least one embodiment, training and / or inference techniques described herein can use one or more ALUs illustrated in execution block 2911. Further, weight parameters can be stored in on-chip or off-chip memory and / or registers (illustrated or not) that configure ALUs of execution block 2911 to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0438] The above techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0439] Figure 30 A deep learning application processor 3000 according to at least one embodiment is shown. In at least one embodiment, deep learning application processor 3000 uses instructions that, if executed by deep learning application processor 3000, cause deep learning application processor 3000 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, deep learning application processor 3000 is an application specific integrated circuit (ASIC). In at least one embodiment, application processor 3000 performs matrix multiplication operations or is “hardwired” into hardware as a result of executing one or more instructions or both. In at least one embodiment, deep learning application processor 3000 includes, without limitation, processing clusters 3010(1)-3010(12), inter-chip links (“ICLs”) 3020(1)-3020(12), inter-chip controllers (“ICCs”) 3030(1)-3030(2), second generation high bandwidth memory (“HBM2”) 3040(1)-3040(4), memory controllers (“MemCtrlrs”) 3042(1)-3042(4), high bandwidth memory physical layers (“HBM PHYs”) 3044(1)-3044(4), management controller central processing units (“management controller CPUs”) 3050, serial peripheral interface, internal integrated circuit, and general purpose input / output blocks (“SPI, I2C, GPIO”) 3060, peripheral component interconnect express controllers and direct memory access blocks (“PCIe controllers and DMA”) 3070, and a peripheral component interconnect express x 16 port (“PCI Express x 16”) 3080.
[0440] In at least one embodiment, processing clusters 3010 can perform deep learning operations, including inference or prediction operations based on weight parameters calculated based on one or more training techniques, including those described herein. In at least one embodiment, each processing cluster 3010 can include, without limitation, any number and type of processor. In at least one embodiment, deep learning application processor 3000 can include any number and type of processing clusters 3000. In at least one embodiment, inter-chip links 3020 are bidirectional. In at least one embodiment, inter-chip links 3020 and inter-chip controllers 3030 enable multiple deep learning application processors 3000 to exchange information, including activation information resulting from execution of one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, deep learning application processor 3000 can include any number (including zero) and type of ICLs 3020 and ICCs 3030.
[0441] In at least one embodiment, HBM2 3040 provides a total of 32 GB of memory. HBM2 3040(i) is associated with both a memory controller 3042(i) and an HBM PHY 3044(i). In at least one embodiment, any number of HBM2s 3040 can provide any type and total amount of high bandwidth memory and can be associated with any number (including zero) and type of memory controllers 3042 and HBM PHYs 3044. In at least one embodiment, SPI, I2C, GPIO 3360, PCIe controller and DMA 3070 and / or PCIe 3080 can be replaced with any number and type of block to implement any number and type of communication standard in any technically feasible fashion.
[0442] Inference and / or training logic 1315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 1315 are provided below in conjunction with FIGS. 13A and / or 13B. In at least one embodiment, inference and / or training logic 1315 can be used in Figure 13A and / or Figure 13B Details regarding inference and / or training logic 1315 are provided below. In at least one embodiment, deep learning application processor is used to train a machine learning model (e.g., neural network) to predict or infer information provided to deep learning application processor 3000. In at least one embodiment, deep learning application processor 3000 is used to infer or predict information based on a trained machine learning model (e.g., neural network) that has been trained by another processor or system or by deep learning application processor 3000. In at least one embodiment, processor 3000 can be used to perform one or more neural network use cases described herein.
[0443] The above techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0444] Figure 31 is a block diagram of a neuromorphic processor 3100, in accordance with at least one embodiment. In at least one embodiment, neuromorphic processor 3100 can receive one or more inputs from a source external to neuromorphic processor 3100. In at least one embodiment, these inputs can be transmitted to one or more neurons 3102 within neuromorphic processor 3100. In at least one embodiment, neurons 3102 and components thereof can be implemented using circuitry or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, neuromorphic processor 3100 can include, without limitation, thousands or millions of instances of neurons 3102, but any suitable number of neurons 3102 can be used. In at least one embodiment, each instance of neurons 3102 can include neuron inputs 3104 and neuron outputs 3106. In at least one embodiment, neurons 3102 can generate outputs that can be transmitted to inputs of other instances of neurons 3102. In at least one embodiment, neuron inputs 3104 and neuron outputs 3106 can be interconnected via synapses 3108.
[0445] In at least one embodiment, neurons 3102 and synapses 3108 can be interconnected such that neuromorphic processor 3100 operates to process or analyze information received by neuromorphic processor 3100. In at least one embodiment, a neuron 3102 can send out a pulse of output (or a “spike” or “peak”) when input received through neuron input 3104 exceeds a threshold value. In at least one embodiment, neuron 3102 can sum or integrate signals received at neuron input 3104. For example, in at least one embodiment, neuron 3102 can be implemented as a leaky integrate-and-fire neuron, where neuron 3102 can produce an output (or “spike”) using a transfer function such as a sigmoid or threshold function if a sum (called “membrane potential”) exceeds a threshold value. In at least one embodiment, a leaky integrate-and-fire neuron can sum signals received at neuron input 3104 into a membrane potential, and can apply a program decay factor (or leak) to reduce the membrane potential. In at least one embodiment, a leaky integrate-and-fire neuron can spike if multiple input signals are received at neuron input 3104 fast enough to exceed a threshold value (i.e., before the membrane potential decays too low to spike). In at least one embodiment, neuron 3102 can be implemented using circuitry or logic that receives input, integrates input into a membrane potential, and decays the membrane potential. In at least one embodiment, input can be averaged, or any other suitable transfer function can be used. Moreover, in at least one embodiment, neuron 3102 can include, without limitation, comparator circuitry or logic that produces an output spike at neuron output 3106 when a result of applying a transfer function to neuron input 3104 exceeds a threshold value. In at least one embodiment, once neuron 3102 spikes, it can ignore previously received input information by, for example, resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, neuron 3102 can resume normal operation after a suitable period of time (or refractory period).
[0446] In at least one embodiment, neurons 3102 can be interconnected by synapses 3108. In at least one embodiment, synapses 3108 can operate to transmit a signal from an output of a first neuron 3102 to an input of a second neuron 3102. In at least one embodiment, a neuron 3102 can transmit information over more than one instance of a synapse 3108. In at least one embodiment, one or more instances of neuron output 3106 can be connected through an instance of synapse 3108 to an instance of neuron input 3104 in the same neuron 3102. In at least one embodiment, an instance of neuron 3102 that produces an output to be transmitted over an instance of synapse 3108 can be referred to as a “presynaptic neuron” with respect to that instance of synapse 3108. In at least one embodiment, an instance of neuron 3102 that receives input transmitted through an instance of synapse 3108 can be referred to as a “postsynaptic neuron” with respect to that instance of synapse 3108. In at least one embodiment, with respect to various instances of synapse 3108, because an instance of neuron 3102 can receive input from one or more instances of synapse 3108 and can also transmit output through one or more instances of synapse 3108, a single instance of neuron 3102 can be both a “presynaptic neuron” and a “postsynaptic neuron”.
[0447] In at least one embodiment, neurons 3102 can be organized into one or more layers. Each instance of neuron 3102 can have one neuron output 3106 that can fan out to one or more neuron inputs 3104 through one or more synapses 3108. In at least one embodiment, neuron outputs 3106 of neurons 3102 in a first layer 3110 can be connected to neuron inputs 3104 of neurons 3102 in a second layer 3112. In at least one embodiment, layer 3110 can be referred to as a “feedforward layer”. In at least one embodiment, each instance of neuron 3102 in an instance of first layer 3110 can fan out to each instance of neuron 3102 in a second layer 3112. In at least one embodiment, first layer 3110 can be referred to as a “fully connected feedforward layer”. In at least one embodiment, each instance of neuron 3102 in each instance of second layer 3112 fans out to fewer than all instances of neuron 3102 in a third layer 3114. In at least one embodiment, second layer 3112 can be referred to as a “sparsely connected feedforward layer”. In at least one embodiment, neurons 3102 in second layer 3112 can fan out to neurons 3102 in multiple other layers, including to neurons 3102 in second layer 3112. In at least one embodiment, second layer 3112 can be referred to as a “recurrent layer”. Neuromorphic processor 3100 can include, without limitation, any suitable combination of recurrent and feedforward layers, including without limitation sparsely connected feedforward layers and fully connected feedforward layers.
[0448] In at least one embodiment, neuromorphic processor 3100 can include, without limitation, a reconfigurable interconnect architecture or dedicated hardwired interconnects to connect synapses 3108 to neurons 3102. In at least one embodiment, neuromorphic processor 3100 can include, without limitation, circuitry or logic that, depending on a neural network topology and neuron fan-in / fan-out, allows synapses to be allocated to different neurons 3102 as needed. For example, in at least one embodiment, synapses 3108 can be connected to neurons 3102 using an interconnect structure such as a network-on-chip or through dedicated connections. In at least one embodiment, synapse interconnects and components thereof can be implemented using circuitry or logic.
[0449] The above-described techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate a grasp on an object held by a human hand as described above.
[0450] Figure 32A processing system is shown in accordance with at least one embodiment. In at least one embodiment, system 3200 includes one or more processor(s) 3202 and one or more graphics processor(s) 3208, and can be a single processor desktop system, a multiprocessor workstation system, or a server system having many processor(s) 3202 or processor core(s) 3207. In at least one embodiment, system 3200 is a processing platform incorporated within a system on a chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.
[0451] In at least one embodiment, system 3200 can include or be incorporated within a server-based gaming platform, including a game console, a mobile gaming console, a handheld gaming console, or an online gaming console that includes game and media processing consoles. In at least one embodiment, system 3200 is a mobile phone, a smart phone, a tablet device, or a mobile internet device. In at least one embodiment, processing system 3200 can also include a wearable device coupled to or integrated within the wearable device, such as a smart watch wearable device, smart glasses device, augmented reality device, or virtual reality device. In at least one embodiment, processing system 3200 is a television or set-top box device having one or more processor(s) 3202 and graphics interface generated by one or more graphics processor(s) 3208.
[0452] In at least one embodiment, one or more processor(s) 3202 each include one or more processor core(s) 3207 to process instructions which, when executed, implement the operations for system and user software. In at least one embodiment, each of the one or more processor core(s) 3207 is configured to process a specific instruction set 3209. In at least one embodiment, instruction set 3209 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via a very long instruction word (VLIW). In at least one embodiment, processor core(s) 3207 can each process a different instruction set 3209, which can include instructions to facilitate emulation of other instruction sets. In at least one embodiment, processor core(s) 3207 can include other processing devices, such as a digital signal processor (DSP).
[0453] In at least one embodiment, processor 3202 includes cache memory 3204. In at least one embodiment, processor 3202 can have a single level of internal cache or multiple levels of internal caches. In at least one embodiment, cache memory is shared among multiple components of processor 3202. In at least one embodiment, processor 3202 also uses an external cache (e.g., a level three (L3) cache, or last level cache (LLC)) (not shown), which can be shared between processor cores 3207 using known cache coherency techniques. In at least one embodiment, register file 3206 is additionally included in processor 3202, which processor can include different types of registers to store different kinds of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, register file 3206 can include general registers or other registers.
[0454] In at least one embodiment, one or more processor(s) 3202 are coupled with one or more interface bus(es) 3210 for communicating data between processor 3202 and other components in system 3200, such as address, data, or control signals. In at least one embodiment, interface bus 3210 can be a version of a processor bus, such as a direct media interface (DMI) bus, in at least one embodiment. In at least one embodiment, interface 3210 is not limited to DMI bus, and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, processor 3202 includes an integrated memory controller 3216 and platform controller hub 3230. In at least one embodiment, memory controller 3216 facilitates communication between memory devices and other components of processing system 3200, while platform controller hub 3230 provides connections to input / output (I / O) devices via local I / O bus.
[0455] In at least one embodiment, memory device 3220 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, flash memory device, or a phase change memory device, among others. In at least one embodiment, memory device 3220 can be a system memory of processing system 3200, to store data 3222 and instructions 3221 for use when one or more processors 3202 executes an application or process. In at least one embodiment, memory controller 3216 also couples with an optional external graphics processor 3212, which can communicate with one or more graphics processors 3208 in processors 3202 to perform graphics and media operations.
[0456] In at least one embodiment, the platform controller hub 3230 enables peripheral devices to connect to the storage device 3220 and the processor 3202 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 3246, a network controller 3234, a firmware interface 3228, a wireless transceiver 3226, a touch sensor 3225, and a data storage device 3224 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 3224 may be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 3225 may include a touchscreen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 3226 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or LTE transceiver. In at least one embodiment, the firmware interface 3228 enables communication with the system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, network controller 3234 may enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to interface bus 3210. In at least one embodiment, audio controller 3246 is a multi-channel high-definition audio controller. In at least one embodiment, processing system 3200 includes an optional legacy I / O controller 3240 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. In at least one embodiment, platform controller hub 3230 may also be connected to one or more Universal Serial Bus (USB) controllers 3242 that connect input devices, such as a keyboard and mouse combination 3243, a camera 3244, or other USB input devices.
[0457] In at least one embodiment, instances of the memory controller 3216 and platform controller hub 3230 may be integrated into a discrete external graphics processor, such as external graphics processor 3212. In at least one embodiment, the platform controller hub 3230 and / or the memory controller 3216 may be external to one or more processors 3202. For example, in at least one embodiment, system 3200 may include an external memory controller 3216 and a platform controller hub 3230, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset communicating with processor 3202.
[0458] Inference and / or training logic 1315 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 13A and / or Figure 13BDetails regarding the inference and / or training logic 1315 are provided. In at least one embodiment, some or all of inference and / or training logic 1315 can be incorporated with graphics processor 3200. For example, in at least one embodiment, the training and / or inference techniques described herein can use one or more ALUs embodied in 3D pipeline 3212. Additionally, in at least one embodiment, the inference and / or training operations described herein can be accomplished with logic other than that shown. Figure 13A or Figure 13B In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not) that configure ALUs of graphics processor 3200 to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0459] The above-described techniques can be used, for example, to implement a system for performing human-object handoff. Some examples use inference and / or training logic to create a neural network trained to generate grasps for objects held by a human hand as described above.
[0460] Figure 33 is a block diagram of a processor 3300 having one or more processor cores 3302A-3302N, an integrated memory controller 3314, and an integrated graphics processor 3308, according to at least one embodiment. In at least one embodiment, processor 3300 can include additional cores, up to and including an additional core 3302N represented by a dashed lined in at least one embodiment. In at least one embodiment, each processor core 3302A-3302N includes one or more internal cache units 3304A-3304N. In at least one embodiment, each processor core can also include access to one or more shared cache units 3306.
[0461] In at least one embodiment, internal cache units 3304A-3304N and shared cache unit 3306 represent a cache memory hierarchy within processor 3300. In at least one embodiment, cache memory units 3304A-3304N can include at least one level of cache memory such as a level one instruction cache and a level one data cache within each processor core, and one or more shared level two cache unit...
Claims
1. A processor comprising one or more computers including one or more processors to: generate a point cloud from an image of a hand holding an object; identify a first portion of the point cloud representing the object; determine a pose of the object from the first portion of the point cloud; identify a second portion of the point cloud representing the hand; classify the second portion of the point cloud as one of a plurality of types of hand poses; determine a pose of the hand from the type of the second portion of the point cloud; generate a set of grasping poses that allow a robot to grasp the object; select, based at least in part on the pose of the hand, a target grasping pose from the set of grasping poses that does not interfere with the hand; and cause the robot to perform the target grasping pose.
2. The processor of claim 1, wherein the pose of the hand identifies a plurality of segments and joint angles.
3. The processor of claim 1, wherein: the image is a three-dimensional image acquired from a depth camera; and the one or more processors produce the point cloud from the three-dimensional image.
4. The processor of claim 1, wherein: the set of grasping poses are poses of a robotic gripper of the robot; and the robotic gripper has two opposing fingers that perform the grasping.
5. The processor of claim 1, wherein the pose of the object includes three angles indicating an orientation of the object and information identifying a location of the object.
6. The processor of claim 1, wherein the robot takes the object from the hand.
7. A system comprising: one or more processors coupled to computer-readable media; computer-readable media storing executable instructions that, as a result of execution by the one or more processors, cause the system to: identify an object from a three-dimensional image of an appendage holding the object; determine a pose of the object; identify the appendage from the three-dimensional image; classify the identified appendage as one of a plurality of types of appendage poses; determine a pose of the appendage from the classification of the appendage; determine a set of grasping poses that allow a robotic gripper to grasp the object; select, from the set of grasping poses, a grasping pose that does not interfere with the appendage; and perform the grasping pose.
8. The system of claim 7, wherein the three-dimensional image is generated from a depth camera, a radar image, a LIDAR image, or a three-dimensional medical imaging device.
9. The system of claim 7, wherein the executable instructions, as a result of execution by the one or more processors, further cause the system to: generate a point cloud from the three-dimensional image; identify the appendage from the three-dimensional image includes identifying a first portion of the point cloud representing the appendage; and wherein identifying the object from the three-dimensional image includes identifying a second portion of the point cloud representing the object. wherein 10. The system of claim 7, wherein the executable instructions, as a result of execution by the one or more processors, further cause the system to: generate a point cloud from the three-dimensional image; and providing the point cloud to a trained model that outputs the grasp pose.
11. The system of claim 10, wherein the training model is trained by providing ground truth data comprising point cloud information and corresponding grasp poses.
12. The system of claim 7, wherein the grasp pose is determined to interfere with the appendage when it is predicted that the robotic hand will contact the appendage during execution of the grasp pose.
13. A machine-readable medium having stored thereon a set of instructions, which, if executed by one or more processors, cause the one or more processors to at least: obtain a three-dimensional image of an appendage holding an object; generate a 3D model from the three-dimensional image; identify a first portion of the 3D model that represents the object; determine a pose of the object from the first portion of the 3D model; identify a second portion of the 3D model that represents the appendage; classify the second portion of the 3D model as one of a plurality of types of appendage poses; determine a pose of the appendage from the type of the second portion of the 3D model; provide the determined pose of the object and the determined pose of the appendage to a trained network that produces a grasp pose for a robotic hand that is capable of grasping the object without contacting the appendage; and cause the robotic hand to execute the grasp pose.
14. The machine-readable medium of claim 13, wherein the trained network is trained at least in part by providing training data to a network, the training data comprising color 3D models of appendages holding objects and suggested grasps for the robotic hand that enable the robotic hand to receive the object from the appendage without contacting the appendage.
15. The machine-readable storage medium of claim 14, wherein the training data is generated at least by: generating a dataset of human-object handoffs; and annotating the dataset with ground truth hand poses and ground truth object poses.
16. The machine-readable storage medium of claim 13, wherein: the grasp pose is based at least in part on the determined pose of the object and the determined pose of the appendage; and the grasp pose is determined to be a grasp that successfully grasps the object without contacting the appendage.
17. The machine-readable storage medium of claim 13, wherein the network is trained using images of the appendage holding different object types. one or more arithmetic logic units (ALUs) to train one or more neural networks to produce a grasp pose for a robotic hand that enables the robotic hand to receive an object from an appendage without contacting the appendage, at least in part, by providing training data to a network, wherein the processor:
18. A processor comprising: generates a 3D model from a three-dimensional image of an appendage holding an object; identifies a first portion of the 3D model that represents the object; determines a pose of the object from the first portion of the 3D model; identifies a second portion of the 3D model that represents the appendage; and determines a pose of the appendage from the type of the second portion of the 3D model. classifying a second portion of the 3D model into one of a plurality of types of appendage pose; determining the pose of the appendage from the type of the second portion of the 3D model; providing the determined pose of the object and the determined pose of the appendage to the one or more neural networks to produce a grasp pose for a robotic hand; and causing the robotic hand to perform the grasp pose.
19. The processor of claim 18, wherein the training data is generated at least by: generating a dataset of human-object handoffs; and annotating the dataset with ground truth hand poses and ground truth object poses.
20. The processor of claim 18, wherein the 3D model is a colorized 3D model.
21. The processor of claim 18, wherein the 3D model is a point cloud.
22. The processor of claim 18, wherein the appendage is a human hand, a robotic hand or paw, a portion of an animal, or a human leg or arm.
Citation Information
Patent Citations
Noncontact type joint angle measuring system
JP2005091085A
Image processing device, image processing method, image processing program and recording medium readable by computer as well as equipment with the same recorded
JP2018144167A
Information processing device, image recognition method and non-transitory computer readable medium
US20190087976A1