Reactive interactions for robotic applications and other automation systems

By using a fast-response predictive robot control system and model predictive control, the robot's motion path is optimized at multiple future time steps, solving the problems of unsmooth and unsafe object transfer in existing technologies, and achieving smooth and reliable object transfer.

CN116787423BActive Publication Date: 2026-05-12NVIDIA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NVIDIA CORP
Filing Date
2023-02-24
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing robots, when performing object transfer tasks, move in a way that is not smooth, intuitive, or reliable, which may lead to contact with people, obstruction of cameras, or other undesirable actions, increasing safety risks.

Method used

Employing a fast, responsive, and safe predictive robot control system, which combines model predictive control (MPC) and machine learning models, optimizes the robot's motion path across multiple future time steps, taking into account changes in the environment and human hand, to ensure smooth and reliable object transfer.

Benefits of technology

This enables robots to safely and smoothly transfer objects to humans in different environments, avoiding pinching injuries and accidental contact, and improving the reliability and efficiency of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116787423B_ABST
    Figure CN116787423B_ABST
Patent Text Reader

Abstract

The present disclosure relates to reactive interaction for robotic applications and other automated systems. The methods presented herein provide predictive control of a robot or automated assembly in performing a particular task. The task to be performed can depend on the position and orientation of the robot performing the task. A predictive control system can determine a state of a physical environment at each of a sequence of time steps, and can select an appropriate position and orientation at each of these time steps. At various time steps, an optimization process can determine a sequence of future motions or acceleration profiles that conform to one or more constraints of the motion. For example, at various time steps, a respective action in the sequence can be performed, and then another sequence of motions is predicted for the next time step, which can help drive robot motion based on predicted future motion, and allow for fast reaction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Patent No. 63 / 321,755, filed March 20, 2022, entitled "Reactive Handovers with Fast Joint-Space Model-Predictive Behavior," the entire contents of which are incorporated herein by reference for all purposes. Background Technology

[0003] Robots and other automated devices are increasingly being used to assist in performing a variety of tasks. At least some of these tasks involve interaction with humans or other entities, such as performing handover actions where a robot grasps and removes an object from a person's hand. For such actions, it is important that the robot grasps the object in a way that does not pinch or otherwise contact the person or entity from which it is removing the object. In existing systems, the robot's movements when performing such tasks may be unsmooth, unintuitive, or unreliable, potentially leading to rapid or accidental actions by humans. Such movements may increase the likelihood of contact with the robot, or may result in a human hand obstructing a camera used to provide the robot with a view of its environment, among other undesirable actions. Attached Figure Description

[0004] Various embodiments of this disclosure will be described with reference to the accompanying drawings, in which:

[0005] Figure 1A Figures 1B, 1C, and 1D show images of a robot performing a handover operation according to at least one embodiment;

[0006] Figure 2A , Figure 2B , Figure 2C and Figure 2D A method for a robot to grasp an object during a handover operation according to at least one embodiment is shown;

[0007] Figure 3A , Figure 3B and Figure 3C The diagram illustrates the reactive actions that a robot, according to at least one embodiment, can take in response to changes in environmental conditions, at least in part.

[0008] Figure 4 An example system for enabling a robot to perform one or more actions in an environment, according to at least one embodiment, is shown;

[0009] Figure 5A and Figure 5BAn example process for moving a robot or automated assembly to a determined position and orientation to perform a task, according to at least one embodiment, is shown;

[0010] Figure 6 Components of a distributed system, according to at least one embodiment, are shown that can be used to enable a robot to perform one or more tasks;

[0011] Figure 7A The inference and / or training logic according to at least one embodiment is illustrated;

[0012] Figure 7B The inference and / or training logic according to at least one embodiment is illustrated;

[0013] Figure 8 An example data center system according to at least one embodiment is shown;

[0014] Figure 9 A computer system according to at least one embodiment is shown;

[0015] Figure 10 A computer system according to at least one embodiment is shown;

[0016] Figure 11 At least a portion of a graphics processor according to one or more embodiments is shown;

[0017] Figure 12 At least a portion of a graphics processor according to one or more embodiments is shown;

[0018] Figure 13 This is an example data flow diagram of an advanced computing pipeline according to at least one embodiment;

[0019] Figure 14 This is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment; and

[0020] Figure 15A and Figure 15B A data flow diagram of the process for training a machine learning model according to at least one embodiment is shown, as well as a client-server architecture for enhancing annotation tools using a pre-trained annotation model. Detailed Implementation

[0021] In the following description, various embodiments will be described. Specific configurations and details are set forth for illustrative purposes to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may also be practiced without specific details. Furthermore, well-known features may be omitted or simplified to avoid obscuring the described embodiments.

[0022] The systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), driving and non-driving robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a variety of purposes, such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational artificial intelligence (AI), optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable applications.

[0023] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, marine systems, smart area monitoring systems, systems performing deep learning operations, systems performing simulation operations, systems performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing optical transport simulations, systems for performing collaborative content creation for 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0024] Methods according to various embodiments can provide optimized path planning for automated and partially automated devices and systems, such as object transfer between humans and robots. Information about the environment (e.g., images or sensor data) can be obtained, which can be used to determine a set of options for the robot to perform a task, such as a set of potential grasping options for the robot to grasp an object from a human's hand. A set of potential grasping options, along with the robot's current state information and a set of constraints for the transfer (or other actions to be performed), can be fed into the path planning and optimization process. Such a process can attempt to determine, at least in part, the motion (e.g., acceleration) to be taken in some future time steps—e.g., the next 10-15 time steps—or the path of motion over a future period, based at least in part on the current robot and environmental state information. Such a planning process can also attempt to optimize the path or sequence of motion to try to minimize one or more cost factors or functions while adhering to one or more constraints on the motion, which may include general constraints (e.g., favoring linear motion, limiting acceleration or variation, avoiding collisions (e.g., with objects or human hands), and avoiding camera obstruction) as well as constraints that may be specific to a particular task. Multiple potential grasping options and optimizations can be determined for each individual time step, but only the first action (or a subset of actions) from the determined motion or acceleration sequence may be executed at that time step, because other motions or accelerations in the predicted sequence can instead be used to attempt improvements such as motion smoothness based on expected future motion. This optimization can be performed using a model predictive control (MPC) system or other such predictive or optimization framework. During optimization, MPC can determine, select, or modify the best or appropriate grasping option based on current state information (e.g., at each time step). Contact detection can be used to let the robot know when to close the gripper or otherwise grasp a target object (or perform a specific action), which may be in a different position than expected, for example, due to the movement of a human hand.

[0025] In view of the teachings and suggestions contained herein, various other such functions that will be apparent to those skilled in the art may also be used within the scope of various embodiments.

[0026] Figure 1AThe illustration depicts a first image 100 of an example environment in which a robot 102 or at least a partially automated assembly can perform various tasks. This environment for operating the robot can include any suitable setting, such as a laboratory, home, warehouse, workbench, vehicle, hospital, or factory. The environment may be located in the same location as the robot's user or operator, or it may be geographically different. In at least some embodiments, the robot can be operated at least partially using human-provided commands, which can be interactive, recorded, or synthesized by an algorithm. Furthermore, while blocks are illustrated as example objects with which objects can interact, it should be understood that such objects can be any physical object that might be located in one of these operating environments. Image 100 may be captured by one of one or more cameras used to capture image or video data about the environment to help guide the robot 102 in performing tasks or actions within that environment, for example, by using computer vision analysis to determine the position and orientation of the robot and one or more objects for interaction, as well as objects for collision avoidance. Although described as using a camera to generate sensor data or image data, this is not limiting, and the sensor data can be generated using one or more sensors of any or more types, such as LiDAR sensors, RADAR sensors, ultrasonic sensors, IMU sensors, and / or other sensor types. Furthermore, the sensor data can represent various data representations, such as images, point clouds, projected images, and / or other sensor data representation types.

[0027] In this example, the human is holding block 104 in their hand. The tasks to be performed by robot 102 may include performing a handover task, which could involve causing the gripper or other such part of robot 102 to perform a grasping action 122, as shown in image 120 of Figure 1B, and to hold block 104 in such a way that the human hand can release the block, and block 104 will remain held by robot 102. The robot can then perform various actions on the block or on another such object acquired during the handover action, such as moving block 104, as shown in image 140 of Figure 1C, and placing the block at a target location.

[0028] As previously mentioned, during tasks such as handover, it may be desirable to avoid any contact between the robot and the hands or other parts of a person holding an object for the handover action, unless such contact is necessary for performing the task. Avoiding contact prevents pinching or any injury to the person, as well as preventing the object from being dropped or damaged by accidental contact. It is also desirable for the robot to move smoothly, reliably, and intuitively toward the object held by the person, so that the person is not startled by the movement or triggered to drop or damage the object by the robot's "brutal" movements (e.g., sudden changes in speed and / or direction) or unexpected movements. In some embodiments, accidental movement of a person may not damage the object, but may result in obstruction of cameras used to guide the robot or other such adverse consequences, or at least may delay the execution of the handover as the robot will need to adapt to the new position of the object.

[0029] Methods according to various embodiments can attempt to overcome at least some of these and other such deficiencies associated with previous and current methods of operating robots or at least semi-automated assemblies or mechanisms to perform one or more tasks. Such methods can utilize fast, reactive, and safe predictive robot control systems, useful for tasks such as human-robot handover, which are easily extendable to include new, additional, or alternative constraints. Such systems can allow humans to hold any objects in any manner or orientation that the robot may not have previously known or encountered, as long as these objects are within the robot's operability. This ability to transfer objects to or from humans in a variety of environments and conditions is crucial for collaborative robots in diverse environments such as factories and homes. In at least one embodiment, such a robot control framework can provide fast, joint spatial model predictive control for tasks such as human-to-robot (H2R) and robot-to-human (R2H) handover. Such reactive systems can accelerate performance, for example, by configuring control functions to implement scheme execution using parallel processors, such as multiple graphics processing units (GPUs), parallel processing units (PPUs), data processing units (DPUs), and / or processor cores. Such systems allow for the integration of constraints that were difficult to adjust in previous systems, which helps to provide safer and faster handovers and other such actions. For example, these constraints can help prevent the robot from pinching an object when it attempts to grasp it, and the task model can be appropriately developed and modified to ensure that the robot's intentions are clear to the humans involved in the handover.

[0030] In the example system, tasks and actions can be performed on any object in a planned manner to ensure that each action taken by the robot is reliable, intuitive, and obvious to a human, so that the human can understand its intentions and will not be surprised by any movement. This can be achieved in various embodiments by optimizing the sequence of movements performed by the robot at each time step. As an example, figure 160 in Figure 1D illustrates the approach of robot 102's gripper to grasp block 104 as part of a handover task. In many existing systems, robot 102 will determine a single next movement or action to take at each time step. Determining each individual movement or action individually and (at least to some extent independently) may result in jitter or irregular trajectories in the robot's movement toward the block. Instead, the method according to various embodiments can optimize the movement 162 or action over multiple future time steps. In this way, the control system can optimize the current movement not only based on current conditions but also based on future conditions that exist as a result of the current movement. In this way, the current movement may be selected that differs from the movement selected when considering the current movement alone, where the current movement will place the robot in a better position for future actions to produce smooth and predictable motion. This prediction of future motion can also be updated at each time step, so the system remains reactive and can quickly update motion or action based on any changing conditions in the environment (e.g., the motion or rotation of the object to be grasped), while still providing improved motion smoothness and predictability. This forward projection also helps enforce various constraints, such as ensuring the robot never partially occludes itself in front of cameras or other sensors that capture image data or other sensor data used to guide the robot, and that it doesn't have to reverse motion or change its trajectory to avoid such occlusion that might be necessary in a previous system where each motion was determined independently. In at least one embodiment, a smoothness reward (or jitter penalty) may be applied when determining motion at these future time steps to help optimize these predicted future motions as part of a model predictive control (MPC) scheme that incorporates human comfort as part of the optimization process. An example MPC framework can integrate perception and complex domain-specific constraints into the optimization problem. Other optimization software or techniques or predictive frameworks may also be used in other embodiments, as different cost terms or criteria can be used for optimization. A learning-based grasping accessibility model can be used to select candidate grasps, maximizing the robot's maneuverability and providing the robot with greater degrees of freedom to satisfy these constraints. Force (e.g., torque) classifiers based on machine learning models (e.g., neural networks) can also be used to detect contact events from noisy data.

[0031] In at least one embodiment, a system for performing actions such as the handover of arbitrary objects may use one or more learning-based grasping planners. Grasping can be selected based on factors such as grasping stability and one or more task constraints, and the robot can be driven to the posture of the gripper or other end effector to complete the handover process. To determine the optimal, desired, or suitable grasping position, and the associated motion of the robot to that position, a machine learning model (e.g., a neural network) trained to predict the maneuverability of the grasping posture can be used for motion-aware grasping selection. Along with grasping stability, this approach allows the control system to select stable grasping positions and actions that are also reachable and can be manipulated in a smooth and predictable manner. Furthermore, this approach allows for the parallel evaluation of the reachability and maneuverability of a batch of grasps compared to sequentially checking the reachability of each grasp using inverse kinematics, which is more efficient for real-time robotic tasks such as human-to-robot handover.

[0032] Beyond the grasping action, another challenge involves planning smooth and natural motion from the robot's current position and orientation to a determined grasping position. Given a chosen grasp, robots in existing systems are typically driven to the grasping end-effector pose by local (or non-global) policies, such as Riemannian motion policies or visual servoing. However, natural motion involves more than just acceleration toward a specific end-effector pose; it considers how to achieve that pose while adhering to complex task constraints. Furthermore, without well-planned motion, the robot might choose a more distant grasping approach, leading to slower or less efficient task execution. Systems using frameworks such as highly parallelized stochastic MPC frameworks can provide real-time motion planning while also allowing for the incorporation of a range of complex task-specific constraints, such as preventing robot-hand pinching. To bridge the gap between grasp selection and motion planning, MPC frameworks can integrate the grasp selection process into the optimization given a set of candidate grasps. This unified MPC framework allows the system to simultaneously select grasps and prescribe motion, taking into account both grasping reachability and domain-specific constraints. Furthermore, during the physical handover phase, such as from the moment the receiver's hand first contacts the object until its release, force feedback can be used to determine when to release the object (R2H) or when to grasp it (H2R). In some cases, raw force / torque sensor readings may be noisy and / or difficult to interpret directly. Therefore, data representing how humans interact with mobile robots can be collected and used as training data to train machine learning models (e.g., neural networks) to detect when contact occurs between the hand and the gripper (or other human-robot interactions). This detection can be combined with vision during the physical handover phase to trigger appropriate robotic grasping or releasing actions.

[0033] Figures 2A to 2DThe images in the illustrations depict different stages of the human-to-robot object transfer according to at least one embodiment. Figure 2A Image 200 shows a robot 202 in a waiting phase, where the robot 202 may remain idle until the robot or the system at least partially responsible for guiding or controlling the robot detects a human hand holding an object 204 that will be taken by the robot 202. During the waiting phase, the robot may move to its original position and not otherwise move until the robot detects a person holding an object in the workspace, or otherwise receives or determines a movement instruction or trigger. Once the object is detected as being held, the MPC framework can analyze various gripping and approach positions 206 that the robot can use to approach and grasp the object 204. As previously described, this may include determining motions on a sequence of future time steps, even if only the next motion can be used at each time step before these motions are redefined. In the approach phase, as Figure 2B As shown in image 220, the robot can move toward the object along a proximity vector using a determined optimal, desired, or suitable motion. In one embodiment, the standoff posture is 15 cm from the final grasping position of the object, where the optimal, desired, or suitable path to this standoff position is determined by the MPC system. Once the robot reaches the determined standoff posture, it can grasp the object during the grasping phase using blocking strategies, such as... Figure 2C Image 240 is shown. A blocking strategy can be used because estimations of body tracking, segmentation, and / or grasping may become less reliable as the robot gets closer to the object. After moving the end effector forward to the grasping position, the robot can close its gripper or end effector and return to a standoff posture. Missed grasping can be detected by observing the distance between the gripper fingers after closure. After a failed grasp, the robot can return to the approach phase and attempt to grasp the object again. Instead of using vision, or as a supplement to vision, a classifier can be trained to detect contact with the gripper given a raw force or torque signal. If a human pushes an object into the robot's gripper, the grasping motion can be terminated early and the gripper will close. The robot can then perform actions on the object at least after the human hand has released it, such as placing the object in a preset position during the descent phase, as... Figure 2D Image 260 is shown.

[0034] As previously mentioned, camera system 302 or other monitoring systems can be used to capture useful information about the state of one or more objects in a defined environment, such as a hand holding object 306 within the camera's field of view. Figure 3A As illustrated in example view 300. In addition to adjusting gripping options and motion at each time step in response to changes in object position, the predictive control system can also attempt to make necessary rapid adjustments to impose constraints on the motion in order to optimize the motion path. For example, consider... Figure 3B In scenario 320, the user accidentally moves object 306 toward the robot's end effector 322 in direction 324, for example, when a person attempts to help place an object into the end effector. Since this could lead to accidental contact between the robot and the person, violating constraints, the system can quickly determine the robot's corresponding motion 326 or acceleration in response to this state change. Figure 3C As shown in Example 340, in addition to avoiding collisions or contact caused by accidental movement 342, this can also lead to adjustments in movement for an updated path 344. However, in at least some embodiments, if path planning does not provide sufficient time to avoid collisions or accidental human contact, the robot may pull away from accidental movements in the corresponding direction without updating the path planning.

[0035] A system according to at least one embodiment can generate a set of candidate grips for human-robot handover by tracking the poses of at least one relevant part of a robot and at least one relevant part of a human. This can include, for example, tracking the robot's pose using Dense Articulated Real-Time (DART) tracking and tracking the human body using, for example, a body tracking software development kit (SDK). The tracked body data can be used to perform tracking of objects and hands, for example, by segmenting the hand and the object in the hand using a pre-trained segmentation model. Given a segmentation, a time-consistent version of a six-degree-of-freedom (DoF) GraspNet or other suitable neural network can be used to generate potential gripping positions because it may correspond to a segmented object point cloud inferred using that segmentation model.

[0036] Neural networks such as GraspNet can estimate a score for each predicted grasp to measure grasp stability. Because considering all potential stable grasps can be challenging, and to reduce the number of grasps considered, a ranking or cost function that can predict the maneuverability of the grasp pose can be used. Maneuverability of the grasp pose can provide a measure of the robot's distance from a singular configuration. Such a metric can be particularly valuable when attempting to reach a moving target, as a highly maneuverable grasp is likely to remain reachable as the object moves. Unlike existing methods that require interpolation using simple heuristics, surrogate representations, or pre-computed data, the method according to at least one embodiment can learn maneuverability metrics directly from the data to quickly predict scores.

[0037] In one example, the manipulability index Given as |J TJ|, where J is the Jacobian determinant of the inverse kinematics (IK) of the grasp. If no IK solution exists, the closest realizable end-effector pose and the negative torsional distance between the target can be used, which may result in:

[0038]

[0039] Where |·| is the determinant operator, J is the Jacobian determinant of the IK solution to g, and dist twist (g,e) is the torsional distance between the gripping g and the closest achievable pose of the end effector pose e.

[0040] The robot's joint constraints can be further incorporated into the maneuverability score, for example, by forming a weighted matrix using joint constraint performance metrics. In at least one embodiment, the metric for each joint can be given by the following formula:

[0041] θ={θ i :i=1,2,…,DoF}:

[0042]

[0043] This index can have a minimum value (0) in the middle of the joint range and a maximum value at both ends of the joint range. The weighting matrix can be a diagonal matrix, with entries on the diagonal corresponding to each joint, and can be given by the following formula:

[0044]

[0045] Then the joint-aware maneuverability can be calculated as... This is used for grasping poses with IK solutions. The results can be used to generate datasets for training a joint constraint-aware reachability model. For ease of training, a dataset of grasping poses and corresponding manipulability scores can be generated, and a multilayer aware model can be trained using mean squared error loss. The input to the MLP can be a 6-DOF grasping pose. The output can be a manipulability score. During the experiment, the median inference time for this type of trained model was 0.0013 seconds.

[0046] To better coordinate the timing of the physical handover phase, feedforward neural networks can be trained to detect contact events between a hand or object and a robot gripper, end effector, or other relevant parts or mechanisms. Such networks can take raw sensor data as input, potentially including data on joint velocities, force points, forces, and torques from the last T=5 steps, and can predict the probability of contact events. An example model is a temporal convolutional neural network that encodes each individual time step using two layers of MLPs, then performs two one-dimensional (1D) convolutions on historical data, and then another MLP (with missing information) predicts a single output indicating whether a force event was detected within the relevant window.

[0047] Frameworks such as the MPC framework can encode heuristics to facilitate smooth human-robot handover in reactive motion generation. For example, these can include heuristics that encourage the robot's gripper to move in a straight line, avoid collisions between the robot and the human (or other objects), ensure the human hand remains within the field of view of at least one camera or display associated with the robot (to the extent possible), and reduce jitter or promote smoother robot motion during operation. Given a set of grips G, the robot can be moved to one of the gripping poses X. g ∈G, and then grab objects from the human body, such as Figure 2C As shown. Since some heuristic algorithms may correspond to operations in joint space, while others may correspond to Cartesian space with one or more links, this problem can be formulated as a kinematic joint space trajectory optimization problem, which can then be solved in real time or near real time, for example, using stochastic MPC.

[0048] Given the joint position θ0 and velocity at time step 0 Robot's joint acceleration θ t∈[0,H-1] It can be calculated across H time steps, which are determined to minimize the cost term C(·) while satisfying one or more constraints. Example constraints can be given as follows:

[0049]

[0050] stS e (θ t )<0.0

[0051] S r (θ t )<0.0

[0052]

[0053] θ min ≤θ t ≤θ max

[0054]

[0055] The first equation contains the cost terms defined for human-robot handover. An equation is also illustrated, listing the maneuverability and stopping costs that help the robot avoid local minima and overshoot, respectively. Collision avoidance constraints that help prevent collisions between the robot and its environment, humans, and itself are also explained. Euler integral equations are shown for obtaining joint positions, and joint velocities are obtained from joint accelerations. The last three equations provide the limits for robot joint positions, velocities, and accelerations.

[0056] Given target pose X g And the current gripping posture X t =FK(θ) t ), using the robot to configure θ at the current joint t Forward kinematics calculations at a point can define the pose distance index dist(X). t ,X g ), as given by the following formula:

[0057]

[0058] Among them wR g and w R t These are postures X g and X t The rotation matrix of X. g and X t The translation vectors are respectively from w R g and w R t express.

[0059] In at least one embodiment, the aforementioned distance metric can be used as the target cost C. g (·) = dist(·) to achieve the given gripper posture. The optimization process can also try writing the target cost as the current gripper posture X. t The distance between the nearest pose in the target set and the target pose is used to optimize the attainment of a pose from a set of grasping poses G, and can be given as follows:

[0060] C(X t ,G)=min(dist(X t ,X g∈G ))

[0061] This formula for achieving the target set allows MPC inference to optimize various grasping postures internally, while also taking into account other heuristics and constraints. This approach allows MPC-based systems to smoothly move from a currently determined grasp to another, given one or more new constraints; for example, a person's hand holding an object moves to a position at least partially within the current grasping region.

[0062] In some robotic systems, the robot may follow a very circular trajectory in Cartesian space, at least in part due to the presence of numerous rotary joints in the manipulator. These circular trajectories are difficult for users to predict or anticipate, especially those with limited domain knowledge. To address this issue, at least in part, a cost term can be used to penalize the linear velocity of the gripper or other end effectors of the robot in a direction not parallel to the vector connecting the current gripper position and the target position, as given by:

[0063]

[0064] The current velocity of the gripper can be determined by configuring θ using the current joint. t Kinematics Jacobi J(θ) t ) and current joint velocity The calculation is as follows:

[0065]

[0066] This vector can then be normalized to calculate the current gripper movement direction. Therefore, the cost can be written as:

[0067]

[0068] This cost is zero when the two vectors are parallel.

[0069] During the gripper's movement, efforts can be made to avoid collisions between the robot and nearby objects (such as a table or floor supporting the robot, or a human hand). Furthermore, efforts can be made to prevent the hand from being gripped by the robot during its movement. In at least one embodiment, the table (or other underlying surface or support) can be represented as a cuboid (or other shape), and the human hand can be represented as a sphere (or other shape). Lines connecting the origin of the monitoring camera and the position of the human hand can be used to construct a capsule or other spatial representation. The robot's links can also be represented using a sphere or similar geometry, and then analyzed using an analytical function S. e (θ tThe signed distance between the robot and its environment is calculated. Furthermore, self-collisions of the robot can be avoided, at least in part, by using a machine learning model (e.g., a neural network) trained with a loss or cost function as presented above. The machine learning model can be configured with θ at a given joint. t In the case of , output the signed distance between the two nearest links, and then you can combine it with the S above. r (θ t Use them together. This method can help robots avoid collisions with human hands and prevent pinching injuries.

[0070] To allow for responsiveness, the optimization problem can be solved, at least in part, using stochastic model predictive control (also known as sampled model predictive control), which can work with a number of cost terms. Stochastic model predictive control optimizes by sampling sequences of actions from a distribution of many particles, deriving these actions, calculating the cost generated by each particle, and then using these costs to update the distribution. By iterating this process, given a large number of particles, stochastic MPC can generate motions, for example, at 50-100 Hz, sufficient to maintain responsiveness in human-robot interactions. At least in some embodiments, stochastic MPC cannot handle constraints, so these constraints can be set as cost terms with large weights in the optimization problem. This optimization problem can also be used for motion generation in other stages, where the objective cost can be changed to an L2 loss over joint configuration.

[0071] Figure 4 The illustration shows components of an example system 400 according to at least one embodiment, which can be used to perform actions such as automated handover. In this example, one or more cameras 402 may be used to capture image data 414 (e.g., sequences, images, or videos) representing a physical environment, including a robot 406 and at least one object 404 interacting with the robot. The cameras may include cameras for capturing two-dimensional images, cameras for capturing stereo images, or cameras for capturing images including color and depth information (e.g., RGB-D images). Other sensors or devices may also be used to capture a representation of the environment, such as depth sensors, LiDAR scanners, etc. The captured environmental data may be fed as input to a state monitor 424 of a control system 420. This input may be updated and provided appropriately, for example, for each captured frame of a video stream or in response to any detected motion in the environment.

[0072] In this example system 400, the state monitor can use this (and any other relevant) input to generate a representation of the current state of the environment. For example, this could include the current position and orientation of robot 406 or a portion of robot 406 (e.g., end effector 408 or gripper), and the current position and orientation of at least one object 404 that robot 406 is to contact. The state monitor can generate any suitable type of representation, such as a 3D model or point cloud representing the current state of the environment. This determined or predicted state information can be provided to various components of the control system 420, such as gripping predictor 426 and motion predictor 428. As discussed herein, at each time step of the process, or at least at periodic time steps, gripping predictor 426 (or other position and orientation predictors for other types of tasks) for handover operations can determine one or more target grips, or other target positions and orientations of at least one end effector 408 (or other portion) of robot 406, to at least initiate the determined task, action, or other interaction with object 404. This can include updating the determined grip in response to any changes in the environment, such as changes in the position or orientation of the robot 406, object 404, hand or other support holding the object, or another object or entity in the environment, among other such options. Updating the target grip helps ensure that the end effector of the robot 406 does not perform undesirable actions, such as contacting a human hand or obstructing the view of the camera 402, and ensures that the grip on the object 404 is sufficient to hold the object 404 in the face of any potential changes in its position or orientation without the risk of dropping or damaging the object 404.

[0073] At each time step, motion predictor 428 can also analyze current state data, as well as any updates to the target grasp determination. As previously described, in some embodiments, grasp predictor 426 can provide a full set of potential grasps, and may include a motion path predictor or MPC. Motion predictor 428 can select the current target grasp to use. In some embodiments, all potential grasps in a set can be treated equally as long as these grasps meet one or more minimum criteria, while in other embodiments, relevant grasps may be weighted or ranked in some way, for example, based on factors such as distance or intensity of the grasp location. Motion predictor 428 can take into account any constraints on motion and may use motion planner or optimizer applications, algorithms, or neural networks, for example, to predict the best, desired, or suitable path between the robot's current position and orientation and the object's current position and orientation, and the hand grasping the object. As discussed herein, optimization may also take into account other factors or constraints, such as a preference for smooth or linear motion, or avoidance of camera obstacles. In at least some embodiments, state data may be received from the robot more frequently than state data is updated based on camera data, such that in some embodiments motion data may be updated more frequently than grasp data. Motion predictor 428 can then determine a series of motions, or motion paths to be executed over multiple future time steps. Motion predictor 428 can then provide information about this sequence or future motions, a first subset of these future motions, or only the first predicted motions to robot controller 422 of control system 422. In this way, optimizing the motion by observing the range of future motions, rather than just the current motion, helps improve factors such as the smoothness and predictability of the motion. If motion predictor 428 has not yet provided instructions, robot controller 422 can then determine instructions that cause robot 406 to perform actions corresponding to that motion, and can send those instructions to robot 406, where drive system 412 or other mechanical controls of robot 406 can cause robot 406 to perform that motion, for example, performing the determined motion for the current time step to bring robot 406's end effector 408 to a target position relative to object 404. In many cases, this will result in the end effector moving toward the target grasping position and orientation, but due to factors such as the movement of the hand or object, it may involve moving away from the object to avoid collisions or obstructions relative to the hand or object. In this example, the end effector 408 may have one or more contact sensors that provide contact data to a contact detection system 410, which in turn provides this information to the drive system 412 and / or robot controller 422 of the control system 420. In some embodiments, the drive system 412 may have a built-in safety mechanism that stops the robot 406 if contact is detected.In some embodiments, the robot controller 422 may alternatively or additionally update state information to include contact information, which can then be provided to the motion predictor 428 to determine any subsequent actions or movements to be taken, such as moving away if the contact is not the desired contact. In some embodiments, if the end effector 408 is in the target grasping position and orientation and detects the expected contact, the motion predictor 428 may provide information to the robot controller 422 to cause the robot 406 to remain in that position for at least a defined period of time to allow the hand to be removed and the handover action completed. In other embodiments, the robot controller 422 may make such a decision. Such a system can also be used to control the robot 406 to perform other tasks, such as determining position or orientation with similar or alternative constraints applied. In some embodiments, the grasping predictor 426, the motion predictor 428, or the robot controller 422 may need to determine applicable constraints for a given task or action, which may be obtained from the constraint repository 430 or other such locations.

[0074] Use such as Figure 4 The system shown was tested using the Franka Emika Panda arm and an externally mounted 1280×720 resolution Azure Kinect RGBD camera. The sensing pipeline, such as... Figure 4 As shown, this system is used to generate potential grabs. The system is distributed across three computers, with an additional real-time desktop for Franka control. Four NVIDIA RTX 2080ti GPUs are used for perception and MPC. The performance of this example system was evaluated using three objects of different shapes and sizes: a banana, a cookie box, and a pepper shaker. Handover trials with these objects were repeated until three successful handovers were recorded, after which system performance was measured using metrics including success rate, approach time, and total time for successful handover. Other metrics observed during the experiments included velocity, acceleration, and jitter during a given handover.

[0075] This method uses a selected baseline approach, along with the crawling selection criteria and the learning accessibility metric R(x) described elsewhere in this paper. appr This determines which grabs are reachable. This set of potential grabs can then be further restricted to include only highly manipulable grabs, as given below:

[0076] C = w s min(ss min ,0)+w prev d(x appr ,x prev )+w home d(x appr ,x home )

[0077] +w R R(x appr )

[0078] Where s and s min It refers to the score to be captured and the minimum acceptable score, x. appr and x prev These represent the current and previously selected fetches, respectively, x home This indicates the orientation of the end effector in its original position, w. s w prev w home ,wR is the weight. Above, d(x1,x2) is a distance metric with both position and rotation components. The above baseline is compared with the following variants of the proposed method, including: 1) MPC / MPC-R: using MPC for motion planning for a selected grasp (with or without an accessibility metric); and 2) MPC-GoalSet: using MPC to generate motion given a set of grasps. It can be noted that the accessibility metric is inherently embedded in the cost function of multi-objective MPC.

[0079] The efficiency of this method is evaluated in response to the handover location and the different ways of holding the objects during H2R handovers at different locations. The process involves handing over three objects at three locations, typically relative to the robot's left, center, and right sides. For each location, the object is handed over three times using three different gripping methods. Compared to previous methods, MPC significantly reduced approach time (11.5 s to 7.1 s) and total time (13.3 s to 7.9 s), and also found an improvement in success rate from 85.7% to 88.3%. The experiments also investigated the reachability model, where the reachability metric was observed to allow for improved gripping selection compared to previous methods, reducing approach time by two seconds and increasing the MPC success rate by 7.5%. This is at least partly attributable to the fact that, using the reachability model, the robot tends to choose more reachable and maneuverable gripping postures. Compared to MPC / MPC-R, the MPC-GoalSet variant was observed to have a slightly longer approach time and a slightly lower success rate, but it was observed to achieve smoother motion while significantly reducing jitter.

[0080] After the robot's motion begins, performance is further investigated by rotating each object approximately 45 degrees along its vertical axis. All MPC-based systems exhibit shorter approach times and overall total time to successful handover because these systems can react quickly and adapt to changes in object orientation. The reachability model was observed to reduce approach times while maintaining a similar success rate to baseline methods and improving the success rate of MPC variants. While MPC-GoalSet achieved the shortest approach time for successful handover in the tests, it was observed to have a lower success rate when handling orientation changes, which may be at least partly due to inconsistencies in grasping over time. To better understand motion patterns, the experiments recorded metrics such as robot position, velocity, acceleration, and jitter during motion. Generally, the tested MPC-based methods were observed to have more consistent velocity and acceleration compared to baseline methods. The proposed method was observed to have fewer sudden accelerations, resulting in less jitter. In particular, the MPC-GoalSet system was observed to achieve minimal jitter because, in this formula, MPC optimization is performed to reach a single grasp from the grasp set while taking into account the robot's current position, velocity, and acceleration.

[0081] Another experiment was conducted to verify that the system or method proposed in this paper allows for fluid H2R transfer, particularly relative to the proposed previous baseline method. In this experiment, four human participants were recruited, each participating in two rounds of transfer and interacting with two systems. Participants were asked to transfer ten items from a set of household items to the robot at a time. The humans were then asked to rate the systems immediately after each round of interaction using Likert scale questions. This experiment negated the order in which participants interacted with the two systems. After completing all interactions, participants were asked to rank the two systems across different dimensions and share their opinions and comments through open-ended questions.

[0082] Metrics such as approach time and success rate were reported. The proposed method was observed to take over objects with significantly reduced time, except for two objects, in this case scissors and toothpaste, and the overall approach time was significantly shorter than the previous baseline system. Success rates were similar across systems. From the perspective of human test participants, most preferred the proposed system because it was more predictable, less abrupt, less aggressive, safer, and more comfortable to participate in.

[0083] Figure 5AAn example process 500 for causing a robot to perform a handover action is illustrated, which can be performed according to various embodiments. It should be understood that, for this and other processes discussed herein, unless specifically stated otherwise, additional, fewer, or alternative steps may be performed in a similar or alternative order or at least partially in parallel within the various embodiments. Furthermore, while this process is described with respect to a robot and a handover action, aspects of this process can be used with other types of automation and for other types of tasks within the various embodiments. In this example, image information representing the environment, along with possibly other sensor or imaging data, may be captured 502 or received. In this example, the environment includes at least a portion of the robot or automated assembly, a grasped object, and other potential objects. The image information may be analyzed 504 to determine the current state of the environment. This may include, for example, the robot's current position and orientation, the robot's end effector or gripper, the object to be handed over and the hand grasping the object, and the position of cameras or sensors. In some embodiments, at least for any object of interest or objects that may create obstacles or collisions, a virtual representation of the current state of the environment may be generated and maintained, and may be updated as the state changes. For this current state, a set of potential grips (e.g., positions and orientations where the end effector can safely and reliably grip the object without contacting the hand holding it) 506 can be determined for the robot relative to the object, which will not contact the hand or violate any associated gripping constraints. In some embodiments, a grip can be determined only after manipulating the robot to a position sufficiently close to the target handover or other such action, in a standoff or other such position. From this set of potential grips, a target grip 508 can be selected for the robot to grip the object. As previously described, this is selected as the optimal, desired, or at least suitable current state for the environment at the current time, and can be updated at any subsequent time step for any future changes (or other related changes) to the state. In at least one embodiment, a complete set of potential grips can be provided to an optimization framework, which can then automatically select the optimal, desired, or suitable target end effector pose as the target grip, and this can be performed or updated for each time step or state evaluation.

[0084] Based at least in part on the current state and the selected target grasp, a sequence of motions (or a series of motion paths over future time steps) for the robot to reach the selected target grasp can be determined. This may include taking into account any relevant motion constraints, such as collision avoidance, camera obstacle avoidance, favoring smooth or linear motion, etc. This can be determined for a defined or predetermined number of time steps or future time periods, or it may be determined at least in part based on the current state or changes in the state of the environment. For example, for a relatively stable environment, the sequence of future motions to be optimized may not need to be that long, perhaps only about five future motions or time steps, but if the environment changes significantly, such as a human placing an object in a significantly different position and orientation with their other hand, then a longer time or sequence may be predicted to attempt to provide the robot with smoother and more intuitive (at least from a human perspective) motion. Once the sequence or future path is determined, and any optimizations have been applied as discussed and suggested herein, the robot can be made to perform a first motion in the determined sequence. In some embodiments, the robot can be made to perform a first subset of the motions in the sequence, for example, where it may not be possible to predict every individual time step, or where the motion path is tracked until changes in the path are determined. In this example, the process can continue until contact is detected at 514, indicating contact with an object near the target grasping position. If not, the process can continue by capturing image information for the next time step and updating the state and motion data. If contact with a different object is detected, or at a completely different position or orientation from the grasp, the process can continue to attempt to recover from this potentially accidental or undesirable contact. However, if it is determined at 514 that contact between the end effector and the object occurred in a position and orientation close to the target grasp, at least within the range of variation allowed due to safety or other constraints or requirements, then it can be determined that the robot has been safely guided to its nearest target grasping position, and unless other factors prevent this determination, it can be determined that the end effector has correctly grasped the object. Then, at 516, the robot in this example can hold the object still for a period of time to allow the human hand to release the object before performing one or more subsequent actions using the held object. During this handover process, the robot can also properly support the object to ensure that it does not fall or become damaged, or cause accidental contact with the human hand after release. In some embodiments, the robot may keep the object stationary for a period of time to allow the hand to be removed, or updated state information may be analyzed to determine when the hand has at least a minimum distance before performing a subsequent action. In some embodiments, the control system may also analyze changes in motion and state over time to attempt, for example, to predict or infer the future position or state of one or more objects in the environment using one or more neural networks, which can further help optimize motion and path planning.

[0085] Figure 5B Another example process 550 is illustrated, in which a similar method can be used to move any automated (or partially automated) assembly, device, component, or system to a position and orientation for performing any action, which has at least one spatial component or requirement. In this example, 552 the action to be performed by such an automated assembly can be determined. The current state of the environment in which the action is to be performed can also be determined, for example by analyzing image or sensor data as described above. Based at least in part on the current state of the environment and the spatial component of the action to be performed, 554 the optimal or suitable position and orientation (or at least the position and orientation with the highest score, ranking, or confidence value) is determined from a set of possible positions and orientations. 556 The optimal (highest rank or score) or suitable sequence of motion (or a motion path over a future period) can be predicted or calculated to guide the assembly or relevant parts of the assembly to the optimal or suitable position and orientation. As described above, predicting the sequence of motion or future path can help produce smoother and more intuitive motions and can help the motion better satisfy one or more constraints. Once determined, 558 at least a first motion or a subset of first motions can be performed by the assembly. As described above, the motion to be performed can be updated at each time step or for each state change; for example, only the current motion at the current time step can be executed before the predicted sequence can be updated. It can be determined whether at least the relevant part of assembly 560 has reached its target position and orientation; if not, the process can continue to the next time step. If it is determined that at least the relevant part of the robot has reached the target position and orientation, at least within an acceptable deviation, robot 562 can perform a determined action at that position. As previously mentioned, this approach can help improve actions, such as human-to-robot handover, by incorporating more knowledge about robot motion planning problems—specifically, using MPC to find smooth, consistent paths. This approach also allows various tasks to be reformulated as model predictive control problems. A learned reachability model can also be used, allowing the model to prioritize positions where the robot has higher maneuverability.

[0086] As discussed above, the various methods presented in this paper are lightweight and capable of being executed in real time on client devices such as personal computers or game consoles. Such processing can be performed on the client device to generate content received by that client device or from an external source (e.g., streaming sensor data or other content received via at least one network). In some cases, the processing and / or determination of this content can be performed by one of these other devices, systems, or entities and then provided to the client device (or another such receiver) for presentation or other similar purposes.

[0087] As an example, Figure 6An example network configuration 600 is shown that can be used to provide, generate, modify, encode, and / or transmit content. In at least one embodiment, client device 602 can generate or receive data for a session using components of control application 604 on client device 602 and data locally stored on the client device. In at least one embodiment, control application 624 (e.g., an image generation or editing application) executing on control server 620 (e.g., a cloud server or edge server) can initiate a session associated with at least client device 602, such as by utilizing a session manager and user data stored in user database 634, and can cause content 632 to be determined by content manager 626. Control application 630 can acquire image data of a scene or environment and work with instruction module 626, planning module 628, control module 630, or other such components to generate a sequence of motions or accelerations corresponding to the instructions to be executed by robot or automation system 670 in that environment. In this example, the content may relate to conditions of the environment, discrete motion sequences, machine-executable instructions, or other such data discussed or suggested herein. At least a portion of the content can be transmitted to client device 602 using a suitable transmission manager 622 for transmission via download, streaming, or another such transmission channel. An encoder can be used to encode and / or compress at least some of this data before transmission to client device 602. In at least one embodiment, content may alternatively be provided to automation system 670 for execution while monitoring or instruction data is received from client device 602, either directly or via at least one network 640. In at least one embodiment, client device 602 receiving such content may provide it to a corresponding control application 604, which may also or alternatively include a graphical user interface 610, a planning module 612, and a control module 614 or process. A decoder can also be used to decode data received via network 640 for presentation by client device 602, such as image or video content on display 606 and audio such as sound and music via at least one audio playback device 608 such as a speaker or headphones. In at least one embodiment, at least some of the content may have already been stored on, rendered on, or made accessible to the client device 602, such that at least this portion of the content does not need to be transmitted over the network 640, for example, in cases where the content may have been previously downloaded or locally stored on a hard drive or optical disc. In at least one embodiment, a transmission mechanism such as data streaming may be used to transmit the content from the server 620 or the user database 634 to the client device 602.In at least one embodiment, at least a portion of the content may be obtained or streamed from another source, such as a third-party content service 660 or other client device 650, which may also include a content application 662 for generating or providing the content. In at least one embodiment, a portion of the function may be performed using multiple computing devices or multiple processors within one or more computing devices, such as a combination of CPU and GPU.

[0088] In this example, these client devices can include any suitable computing device, such as desktop computers, laptops, set-top boxes, streaming devices, game consoles, smartphones, tablets, VR headsets, AR goggles, wearable computers, or smart TVs. Each client device can submit requests across at least one wired or wireless network, which can include the Internet, Ethernet, a local area network (LAN), or a cellular network, among other such options. In this example, these requests can be submitted to an address associated with a cloud provider that operates or controls one or more electronic resources within a cloud provider environment, such as a data center or server cluster. In at least one embodiment, the request can be received or processed by at least one edge server located at the network edge and outside at least one security layer associated with the cloud provider environment. In this way, latency can be reduced by enabling client devices to interact with servers in closer proximity, while also improving the security of resources within the cloud provider environment.

[0089] In at least one embodiment, such a system can be used to perform graphics rendering operations. In other embodiments, such a system can be used for other purposes, such as providing image or video content to test or validate autonomous machine applications, or for performing deep learning operations. In at least one embodiment, such a system can be implemented using edge devices, or can be combined with one or more virtual machines (VMs). In at least one embodiment, such a system can be implemented at least partially in a data center or at least partially using cloud computing resources.

[0090] Reasoning and training logic

[0091] Figure 7A Inference and / or training logic 715 is shown for performing inference and / or training operations associated with one or more embodiments. The following is in conjunction with... Figure 7A and / or Figure 7B Provide details about reasoning and / or training logic 715.

[0092] In at least one embodiment, inference and / or training logic 715 may include, but is not limited to, code and / or data storage 701 for storing forward and / or output weights and / or input / output data, and / or other parameters configuring neurons or layers of a neural network trained for and / or used for inference in one or more embodiments. In at least one embodiment, training logic 715 may include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, code and / or data storage 701 stores weight parameters and / or input / output data of each layer of a neural network trained or used in one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 701 may be included within other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0093] In at least one embodiment, any portion of the code and / or data storage 701 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 701 may be a cache memory, dynamic random access memory (“DRAM”), static random access memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 701 is internal or external to the processor, for example, or composed of DRAM, SRAM, flash memory, or some other storage type, may depend on the available on-chip or off-chip storage space, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.

[0094] In at least one embodiment, inference and / or training logic 715 may include, but is not limited to, code and / or data storage 705 for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in one or more embodiments. In at least one embodiment, during training and / or inference using one or more embodiments, code and / or data storage 705 stores weight parameters and / or input / output data for each layer of a neural network trained or used in one or more embodiments during backpropagation of input / output data and / or weight parameters. In at least one embodiment, training logic 715 may include or be coupled to code and / or data storage 705 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic including integer and / or floating-point units (collectively, an arithmetic logic unit (ALU)). In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data storage 705 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 705 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 705 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice between the code and / or data storage 705 being internal or external to the processor, for example, whether it consists of DRAM, SRAM, flash memory, or some other type of storage, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.

[0095] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be the same storage structure. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be partially identical and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0096] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating-point units) for performing logical and / or mathematical operations at least in part based on or instructed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values ​​from layers or neurons within a neural network) stored in activation storage 720, which are functions of input / output and / or weight parameter data stored in code and / or data storage 701 and / or code and / or data storage 705. In at least one embodiment, activation is activated in response to execution instructions or other code, and linear algebraic and / or matrix-based mathematical generation performed by ALU 710 is stored in activation storage 720, wherein weight values ​​stored in code and / or data storage 705 and / or code and / or data storage 701 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, and any or all of these can be stored in code and / or data storage 705 or code and / or data storage 701 or other on-chip or off-chip storage.

[0097] In at least one embodiment, one or more processors or other hardware logic devices or circuits include one or more ALUs 710, while in another embodiment, one or more ALUs 710 may be located outside the processor or other hardware logic device or the circuitry using them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 710 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and activation storage 720 may be on the same processor or other hardware logic device or circuitry, while in another embodiment, they may be on different processors or other hardware logic devices or circuitries, or in some combination of the same and different processors or other hardware logic devices or circuitries. In at least one embodiment, any portion of activation storage 720 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0098] In at least one embodiment, the active memory 720 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 720 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 720 is internal to or external to the processor may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or other memory types. In at least one embodiment, Figure 7A The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM Inference processing unit (IPU) or from Intel Corp. (e.g., "LakeCrest") processor. In at least one embodiment, Figure 7A The inference and / or training logic 715 shown can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as field programmable gate array (“FPGA”)

[0099] Figure 7B Inference and / or training logic 715 according to at least one or more embodiments is illustrated. In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, hardware logic, wherein computational resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 7B The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM Inference processing unit (IPU) or from Intel Corp. (e.g., "LakeCrest") processor. In at least one embodiment, Figure 7BThe inference and / or training logic 715 shown can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field-programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 715 includes, but is not limited to, code and / or data storage 701 and code and / or data storage 705, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 7B In at least one embodiment shown, each of code and / or data storage 701 and code and / or data storage 705 is associated with dedicated computing resources (e.g., computing hardware 702 and computing hardware 706), respectively. In at least one embodiment, each of computing hardware 702 and computing hardware 706 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) only on the information stored in code and / or data storage 701 and code and / or data storage 705, respectively, and the results of the function execution are stored in activation storage 720.

[0100] In at least one embodiment, each of the code and / or data storage 701 and 705 and the corresponding computing hardware 702 and 706 corresponds to a different layer of the neural network, such that activation obtained from one "store / computation pair 701 / 702" of the code and / or data storage 701 and computing hardware 702 provides input as input to the next "store / computation pair 705 / 706" of the code and / or data storage 705 and computing hardware 706, in order to reflect the conceptual organization of the neural network. In at least one embodiment, each store / computation pair 701 / 702 and 705 / 706 may correspond to more than one neural network layer. In at least one embodiment, additional store / computation pairs (not shown) may be included in the inference and / or training logic 715 after or in parallel with the store / computation pairs 701 / 702 and 705 / 706.

[0101] Data Center

[0102] Figure 8 An example data center 800 that can be used with at least one embodiment is shown. In at least one embodiment, the data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840.

[0103] In at least one embodiment, such as Figure 8As shown, the data center infrastructure layer 810 may include a resource coordinator 812, grouped computing resources 814, and node computing resources (“nodes CR”) 816(1)-816(N), where “N” represents any positive integer. In at least one embodiment, nodes CR 816(1)-816(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (“NWI / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR 816(1)-816(N) may be servers having one or more of the aforementioned computing resources.

[0104] In at least one embodiment, the grouped computing resources 814 may include individual groups (not shown) of node CRs housed in one or more racks, or a plurality of racks (also not shown) housed in data centers in various geographic locations. Individual groups of node CRs within the grouped computing resources 814 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.

[0105] In at least one embodiment, resource coordinator 812 may configure or otherwise control one or more nodes CR816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource coordinator 812 may include a Software Design Infrastructure (“SDI”) management entity for data center 800. In at least one embodiment, resource coordinator 108 may include hardware, software, or some combination thereof.

[0106] In at least one embodiment, such as Figure 8As shown, framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, framework layer 820 may include a framework of software 832 supporting software layer 830 and / or one or more applications 842 supporting application layer 840. In at least one embodiment, software 832 or application 842 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 820 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark") which can leverage distributed file system 828 for large-scale data processing (e.g., "big data"). In at least one embodiment, job scheduler 832 may include Spark drivers to facilitate the scheduling of workloads supported by the various layers of data center 800. In at least one embodiment, configuration manager 824 may be able to configure different layers, such as software layer 830 and framework layer 820 including Spark and distributed file system 828 for supporting large-scale data processing. In at least one embodiment, resource manager 826 is capable of managing cluster or group computing resources mapped to or allocated to support distributed file system 828 and job scheduler 822. In at least one embodiment, cluster or group computing resources may include group computing resources 814 on data center infrastructure layer 810. In at least one embodiment, resource manager 826 may coordinate with resource coordinator 812 to manage these mapped or allocated computing resources.

[0107] In at least one embodiment, the software 832 included in the software layer 830 may include software used by at least a portion of the nodes CR816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0108] In at least one embodiment, the application layer 840 may include one or more applications 842 that can be used by at least a portion of nodes CR816(1)-816(N), grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0109] In at least one embodiment, any of the configuration manager 824, resource manager 826, and resource coordinator 812 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. In at least one embodiment, self-modification actions can mitigate potentially poor configuration decisions by data center operators of data center 800 and can prevent underutilization and / or poor performance of the data center.

[0110] In at least one embodiment, data center 800 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above with respect to data center 800. In at least one embodiment, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to data center 800 by using weight parameters calculated through one or more training techniques described herein.

[0111] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to utilize the aforementioned resources to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0112] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7BDetails are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 can be implemented in the system. Figure 8 Used in systems for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0113] Such components can be used to determine and optimize the motion of automated systems to perform defined tasks.

[0114] Computer System

[0115] Figure 9 This is a block diagram illustrating an exemplary computer system according to at least one embodiment. The exemplary computer system may be a system of interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof formed with a processor, which may include an execution unit to execute instructions. In at least one embodiment, according to this disclosure, such as the embodiments described herein, computer system 900 may include, but is not limited to, components such as processor 902, whose execution unit includes logic to execute algorithms for process data. In at least one embodiment, computer system 900 may include a processor, such as those available from Intel Corporation of Santa Clara, California. Processor family, Xeon™ XScale™ and / or StrongARM™ Core TM or Nervana TM A microprocessor may be used, although other systems (including PCs, engineering workstations, set-top boxes, etc.) with other microprocessors may also be used. In at least one embodiment, computer system 900 may execute a version of the Windows operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0116] The embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (“DSP”), a system-on-a-chip (SoC), a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system that can execute one or more instructions according to at least one embodiment.

[0117] In at least one embodiment, the computer system 900 may include, but is not limited to, a processor 902, which may include, but is not limited to, one or more execution units 908, to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 900 is a single-processor desktop or server system, but in another embodiment, the computer system 900 may be a multiprocessor system. In at least one embodiment, the processor 902 may include, but is not limited to, a Complex Instruction Set Computer (“CISC”) microprocessor, a Reduced Instruction Set Computing (“RISC”) microprocessor, a Very Long Instruction Word (“VLIW”) microprocessor, a processor implementing instruction set combinations, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 902 may be coupled to a processor bus 910, which can transmit data signals between the processor 902 and other components in the computer system 900.

[0118] In at least one embodiment, processor 902 may include, but is not limited to, a Level 1 (“L1”) internal cache memory (“cache”) 904. In at least one embodiment, processor 902 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 902. Depending on specific implementation and requirements, other embodiments may also include a combination of internal and external caches. In at least one embodiment, register file 906 may store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.

[0119] In at least one embodiment, a logic execution unit 908, including but not limited to performing integer and floating-point operations, is also located within the processor 902. In at least one embodiment, the processor 902 may further include a microcode (“ucode”) read-only memory (“ROM”) for storing microcode of certain macro instructions. In at least one embodiment, the execution unit 908 may include logic for processing a packaged instruction set 909. In at least one embodiment, by including the packaged instruction set 909 in the instruction set of a general-purpose processor, along with the associated circuitry for executing the instructions, packaged data in the processor 902 can be used to perform operations used by numerous multimedia applications. In one or more embodiments, many multimedia applications can be executed more quickly and efficiently by using the full width of the processor’s data bus to perform operations on the packaged data, which may eliminate the need to transfer smaller data units on the processor’s data bus to perform one or more operations on one data element at a time.

[0120] In at least one embodiment, the execution unit 908 may also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuitry. In at least one embodiment, the computer system 900 may include, but is not limited to, memory 920. In at least one embodiment, memory 920 may be implemented as a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, or other storage device. In at least one embodiment, memory 920 may store instructions 919 and / or data 921 represented by data signals that can be executed by processor 902.

[0121] In at least one embodiment, the system logic chip may be coupled to the processor bus 910 and the memory 920. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (“MCH”) 916, and the processor 902 may communicate with the MCH 916 via the processor bus 910. In at least one embodiment, the MCH 916 may provide a high-bandwidth memory path 918 to the memory 920 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 916 may initiate data signals between the processor 902, the memory 920, and other components in the computer system 900, and bridge data signals between the processor bus 910, the memory 920, and the system I / O 922. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 916 may be coupled to the memory 920 via the high-bandwidth memory path 918, and the graphics / video card 912 may be coupled to the MCH 916 via an Accelerated Graphics Port (“AGP”) interconnect 914.

[0122] In at least one embodiment, computer system 900 may use system I / O 922, which is a proprietary hub interface bus, to couple MCH 916 to I / O controller hub (“ICH”) 930. In at least one embodiment, ICH 930 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to memory 920, chipset, and processor 902. Examples may include, but are not limited to, audio controller 929, firmware hub (“FlashBIOS”) 928, wireless transceiver 926, data storage 924, a conventional I / O controller 923 including user input and keyboard interfaces, serial expansion port 927 (e.g., a Universal Serial Bus (USB) port), and network controller 934. Data storage 924 may include hard disk drives, floppy disk drives, CD-ROM devices, flash memory devices, or other mass storage devices.

[0123] In at least one embodiment, Figure 9 The illustration shows a system comprising interconnected hardware devices or "chips," while in other embodiments, Figure 9An exemplary system-on-a-chip (SoC) may be illustrated. In at least one embodiment, the device may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the computer system 900 are interconnected using a compute fast link (CXL) interconnect.

[0124] The inference and / or training logic 715 is used to perform inference and / or training operations related to one or more embodiments. (The following is in conjunction with...) Figure 7A and / or Figure 7B Details are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 can be... Figure 9 Used in systems for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0125] Such components can be used to determine and optimize the motion of automated systems to perform defined tasks.

[0126] Figure 10 This is a block diagram illustrating an electronic device 1000 for utilizing a processor 1010 according to at least one embodiment. In at least one embodiment, the electronic device 1000 may be, for example, but not limited to, a laptop computer, tower server, rack server, blade server, desktop computer, tablet computer, mobile device, telephone, embedded computer, or any other suitable electronic device.

[0127] In at least one embodiment, system 1000 may include, but is not limited to, processor 1010 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1010 uses a bus or interface coupling, such as an I2C bus, system management bus (“SMBus”), low pin count (LPC) bus, serial peripheral interface (“SPI”), high-definition audio (“HAD”) bus, serial advanced technology accessory (“SATA”) bus, universal serial bus (“USB”) (versions 1, 2, and 3), or universal asynchronous receiver / transmitter (“UART”) bus. In at least one embodiment, Figure 10 The system shown includes interconnected hardware devices or "chips," while in other embodiments, Figure 10 An exemplary system-on-a-chip (SoC) can be illustrated. In at least one embodiment, Figure 10 The device shown can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 10 One or more components are interconnected using Computational Fast Link (CXL) interconnects.

[0128] In at least one embodiment, Figure 10 It may include a display 1024, a touch screen 1025, a touchpad 1030, a near field communication unit (“NFC”) 1045, a sensor hub 1040, a thermal sensor 1046, a fast chipset (“EC”) 1035, a trusted platform module (“TPM”) 1038, a BIOS / firmware / flash (“BIOS, FWFlash”) 1022, a DSP 1060, a drive 1020 (e.g., a solid-state drive (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 1050, a Bluetooth unit 1052, a wireless wide area network unit (“WWAN”) 1056, a global positioning system (GPS) 1055, a camera (“USB 3.0 camera”) 1054 (e.g., a USB 3.0 camera), and / or a low-power double data rate (“LPDDR”) memory unit (“LPDDR3”) 1015 implemented in, for example, the LPDDR3 standard. These components can each be implemented in any suitable way.

[0129] In at least one embodiment, other components may be communicatively coupled to processor 1010 via the components described above. In at least one embodiment, accelerometer 1041, ambient light sensor (“ALS”) 1042, compass 1043, and gyroscope 1044 may be communicatively coupled to sensor hub 1040. In at least one embodiment, thermal sensor 1039, fan 1037, keyboard 1036, and touchpad 1030 may be communicatively coupled to EC 1035. In at least one embodiment, speaker 1063, earphone 1064, and microphone (“mic”) 1065 may be communicatively coupled to audio unit (“audio codec and Class D amplifier”) 1062, which in turn may be communicatively coupled to DSP 1060. In at least one embodiment, audio unit 1062 may include, for example, but not limited to, audio encoder / decoder (“codec”) and Class D amplifier. In at least one embodiment, SIM card (“SIM”) 1057 may be communicatively coupled to WWAN unit 1056. In at least one embodiment, components such as WLAN unit 1050, Bluetooth unit 1052, and WWAN unit 1056 can be implemented as next-generation form factor (NGFF).

[0130] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. The following is combined with... Figure 7A and / or Figure 7B Details are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 can be... Figure 10The system is used to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0131] Such components can be used to determine and optimize the motion of automated systems to perform defined tasks.

[0132] Figure 11 This is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 1100 includes one or more processors 1102 and one or more graphics processors 1108, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 1102 or processor cores 1107. In at least one embodiment, system 1100 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.

[0133] In at least one embodiment, system 1100 may include or be integrated into a server-based gaming platform, including a game console, mobile game console, handheld game console, or online game console, which are game and media consoles. In at least one embodiment, system 1100 is a mobile phone, smartphone, tablet computing device, or mobile internet device. In at least one embodiment, processing system 1100 may also include components coupled to or integrated into a wearable device, such as a smartwatch, smart glasses, augmented reality, or virtual reality device. In at least one embodiment, processing system 1100 is a television or set-top box device having one or more processors 1102 and a graphical interface generated by one or more graphics processors 1108.

[0134] In at least one embodiment, one or more processors 1102 each include one or more processor cores 1107 for processing instructions that, when executed, perform operations against the system and user software. In at least one embodiment, each of the one or more processor cores 1107 is configured to process a specific instruction set 1109. In at least one embodiment, the instruction set 1109 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation via Very Long Instruction Word (VLIW). In at least one embodiment, the processor cores 1107 may each process different instruction sets 1109, which may include instructions that facilitate the emulation of other instruction sets. In at least one embodiment, the processor cores 1107 may also include other processing devices, such as digital signal processors (DSPs).

[0135] In at least one embodiment, processor 1102 includes cache memory 1104. In at least one embodiment, processor 1102 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of processor 1102. In at least one embodiment, processor 1102 also uses an external cache (e.g., a Level 3 (L3) cache or a last-level cache (LLC)) (not shown), which can be shared among processor cores 1107 using known cache coherence techniques. In at least one embodiment, processor 1102 further includes a register file 1106, which may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). In at least one embodiment, register file 1106 may include general-purpose registers or other registers.

[0136] In at least one embodiment, one or more processors 1102 are coupled to one or more interface buses 1110 to transmit communication signals, such as address, data, or control signals, between the processors 1102 and other components in the system 1100. In at least one embodiment, the interface bus 1110 may be a processor bus, such as a version of the Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 1110 is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 1102 includes an integrated memory controller 1116 and a platform controller hub 1130. In at least one embodiment, the memory controller 1116 facilitates communication between memory devices and other components of the processing system 1100, while the platform controller hub (PCH) 1130 provides connectivity to I / O devices via a local I / O bus.

[0137] In at least one embodiment, memory device 1120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or a device with suitable performance for use as processor memory. In at least one embodiment, memory device 1120 may be used as system memory of processing system 1100 to store data 1122 and instructions 1121 for use when one or more processors 1102 execute an application or process. In at least one embodiment, memory controller 1116 is also coupled to an optional external graphics processor 1112, which may communicate with one or more graphics processors 1108 of processor 1102 to perform graphics and media operations. In at least one embodiment, display device 1111 may be connected to processor 1102. In at least one embodiment, display device 1111 may include one or more internal display devices, such as in mobile electronic devices or laptop devices, or external display devices connected via a display interface (e.g., DisplayPort). In at least one embodiment, the display device 1111 may include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) or augmented reality (AR) applications.

[0138] In at least one embodiment, the platform controller hub 1130 allows peripheral devices to connect to the storage device 1120 and the processor 1102 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 1146, a network controller 1134, a firmware interface 1128, a wireless transceiver 1126, a touch sensor 1125, and a data storage device 1124 (e.g., a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 1124 may be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1125 may include a touchscreen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1126 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or LTE transceiver. In at least one embodiment, the firmware interface 1128 allows communication with the system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, network controller 1134 may allow network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to interface bus 1110. In at least one embodiment, audio controller 1146 is a multi-channel high-definition audio controller. In at least one embodiment, processing system 1100 includes an optional legacy I / O controller 1140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to system 1100. In at least one embodiment, platform controller hub 1130 may also be connected to one or more Universal Serial Bus (USB) controllers 1142 that connect input devices, such as a keyboard and mouse combination 1143, a camera 1144, or other USB input devices.

[0139] In at least one embodiment, instances of the memory controller 1116 and platform controller hub 1130 may be integrated into a discrete external graphics processor, such as external graphics processor 1112. In at least one embodiment, the platform controller hub 1130 and / or the memory controller 1116 may be external to one or more processors 1102. For example, in at least one embodiment, system 1100 may include external memory controller 1116 and platform controller hub 1130, which may be configured as a memory controller hub and peripheral controller hub in a system chipset communicating with processor 1102.

[0140] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. The following is combined with... Figure 7A and / or Figure 7B Details regarding the inference and / or training logic 715 are provided. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into the graphics processor 1100. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in the graphics processor. Furthermore, in at least one embodiment, the inference and / or training operations described herein may use, in addition to Figure 7A or Figure 7B The logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0141] Such components can be used to determine and optimize the motion of automated systems to perform defined tasks.

[0142] Figure 12 This is a block diagram of a processor 1200 having one or more processor cores 1202A-1202N, an integrated memory controller 1214, and an integrated graphics processor 1208 according to at least one embodiment. In at least one embodiment, the processor 1200 may include additional cores, up to and including additional cores 1202N indicated by dashed boxes. In at least one embodiment, each processor core 1202A-1202N includes one or more internal cache units 1204A-1204N. In at least one embodiment, each processor core may also access one or more shared cache units 1206.

[0143] In at least one embodiment, internal cache units 1204A-1204N and shared cache unit 1206 represent a cache memory hierarchy within processor 1200. In at least one embodiment, internal cache units 1204A-1204N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared intermediate cache, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, wherein the highest level of cache preceding external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1206 and 1204A-1204N.

[0144] In at least one embodiment, the processor 1200 may further include a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, one or more bus controller units 1216 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1210 provides management functions for various processor components. In at least one embodiment, the system agent core 1210 includes one or more integrated memory controllers 1214 to manage access to various external memory devices (not shown).

[0145] In at least one embodiment, one or more processor cores 1202A-1202N include support for multi-threaded concurrent processing. In at least one embodiment, system agent core 1210 includes components for coordinating and operating cores 1202A-1202N during multi-threaded processing. In at least one embodiment, system agent core 1210 may additionally include a power control unit (PCU) including logic and components for regulating one or more power states of processor cores 1202A-1202N and graphics processor 1208.

[0146] In at least one embodiment, processor 1200 further includes a graphics processor 1208 for performing graph processing operations. In at least one embodiment, graphics processor 1208 is coupled to a shared cache unit 1206 and a system proxy core 1210 including one or more integrated memory controllers 1214. In at least one embodiment, system proxy core 1210 further includes a display controller 1211 for driving graphics processor outputs to one or more coupled displays. In at least one embodiment, display controller 1211 may also be a separate module coupled to graphics processor 1208 via at least one interconnect, or it may be integrated within graphics processor 1208.

[0147] In at least one embodiment, ring-based interconnect unit 1212 is used to couple internal components of processor 1200. In at least one embodiment, alternative interconnect units, such as point-to-point interconnects, switched interconnects, or other technologies, may be used. In at least one embodiment, graphics processor 1208 is coupled to ring interconnect 1212 via I / O link 1213.

[0148] In at least one embodiment, I / O link 1213 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory module 1218 (e.g., eDRAM module). In at least one embodiment, each of processor cores 1202A-1202N and graphics processor 1208 uses embedded memory module 1218 as a shared last-level cache.

[0149] In at least one embodiment, processor cores 1202A-1202N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, processor cores 1202A-1202N are heterogeneous in terms of instruction set architecture (ISA), with one or more processor cores 1202A-1202N executing a common instruction set, while one or more other processor cores 1202A-1202N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, processor cores 1202A-1202N are heterogeneous in terms of microarchitecture, with one or more cores having relatively high power consumption coupled to one or more power cores having lower power consumption. In at least one embodiment, processor 1200 may be implemented on one or more chips or implemented as a SoC integrated circuit.

[0150] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. The following is combined with... Figure 7A and / or Figure 7B Details regarding the inference and / or training logic 715 are provided. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into the processor 1200. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in... Figure 12 The graphics processor 1512, graphics core 1202A-1202N, or other components are used. Furthermore, in at least one embodiment, the inference and / or training operations described herein can use, except... Figure 7A or Figure 7B The logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 1200 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0151] Such components can be used to determine and optimize the motion of automated systems to perform defined tasks.

[0152] Virtualization computing platform

[0153] Figure 13 This is an example data flow diagram of process 1300 for generating and deploying an image processing and inference pipeline according to at least one embodiment. In at least one embodiment, process 1300 may be deployed for use with imaging devices, processing devices, and / or other device types at one or more facilities 1302. Process 1300 may be executed within training system 1304 and / or deployment system 1306. In at least one embodiment, training system 1304 may be used to train, deploy, and implement machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for use in deployment system 1306. In at least one embodiment, deployment system 1306 may be configured to offload processing and computing resources in a distributed computing environment to reduce the infrastructure requirements of facility 1302. In at least one embodiment, one or more applications in the pipeline may use or invoke services of deployment system 1306 (e.g., inference, visualization, computation, AI, etc.) during application execution.

[0154] In at least one embodiment, some applications used in the advanced processing and inference pipeline may use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, a machine learning model may be trained at facility 1302 using data 1308 (e.g., imaging data) generated at facility 1302 (and stored on one or more Picture Archiving and Communication System (PACS) servers at facility 1302), imaging or sequencing data 1308 from another or more facilities, or a combination thereof. In at least one embodiment, training system 1304 may be used to provide applications, services, and / or other resources to generate a deployable machine learning model for the work of deploying system 1306.

[0155] In at least one embodiment, the model registry 1324 may be supported by an object storage system that supports version control and object metadata. In at least one embodiment, it may be available from within a cloud platform via, for example, cloud storage (e.g., Figure 14 The system uses a cloud-compatible application programming interface (API) (1426) to access object storage. In at least one embodiment, machine learning models within the model registry 1324 can be uploaded, listed, modified, or deleted by the developer or partner of the system interacting with the API. In at least one embodiment, the API can provide access to methods that allow users with appropriate credentials to associate models with applications, enabling the models to be executed as part of the containerized instantiation of the application.

[0156] In at least one embodiment, training pipeline 1404 ( Figure 14This can include situations where facility 1302 is training its own machine learning model or has an existing machine learning model that needs optimization or updating. In at least one embodiment, imaging data 1308 generated by imaging devices, sequencing devices, and / or other types of devices can be received. In at least one embodiment, once the imaging data 1308 is received, AI-assisted annotation 1310 can be used to help generate annotations corresponding to the imaging data 1308 for use as ground-based data for the machine learning model. In at least one embodiment, AI-assisted annotation 1310 can include one or more machine learning models (e.g., convolutional neural networks (CNNs)) that can be trained to generate annotations corresponding to certain types of imaging data 1308 (e.g., from certain devices). In at least one embodiment, AI-assisted annotation 1310 can then be used directly or adjusted or fine-tuned using annotation tools to generate ground-based data. In at least one embodiment, AI-assisted annotation 1310, labeled clinical data 1312, or a combination thereof can be used as ground-based data for training the machine learning model. In at least one embodiment, the trained machine learning model may be referred to as output model 1316 and may be used by deployment system 1306 as described herein.

[0157] In at least one embodiment, training pipeline 1404 ( Figure 14This may include situations where facility 1302 requires a machine learning model to perform one or more processing tasks for deploying one or more applications in system 1306, but facility 1302 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for this purpose). In at least one embodiment, an existing machine learning model may be selected from model registry 1324. In at least one embodiment, model registry 1324 may include machine learning models trained to perform various inference tasks on imaging data. In at least one embodiment, the machine learning model in model registry 1324 may be trained on imaging data from a different facility (e.g., a remote facility) instead of facility 1302. In at least one embodiment, the machine learning model may have already been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when training on imaging data from a specific location, training may be performed at that location, or at least in a manner that protects the confidentiality of the imaging data or restricts the transfer of the imaging data from off-site locations. In at least one embodiment, once a model has been trained or partially trained at a location, a machine learning model may be added to model registry 1324. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be used in model registry 1324. In at least one embodiment, a machine learning model (and referred to as output model 1316) can then be selected from model registry 1324, and can be executed in deployment system 1306 for one or more processing tasks for one or more applications of the deployment system.

[0158] In at least one embodiment, in training pipeline 1404 ( Figure 14In this scenario, the scenario may include facility 1302, which requires a machine learning model to perform one or more processing tasks for deploying one or more applications in system 1306, but facility 1302 may not currently have such a machine learning model (or may not have an optimized, efficient, or effective model). In at least one embodiment, the machine learning model selected from model registry 1324 may not be fine-tuned or optimized for the imaging data 1308 generated at facility 1302 due to population variability, robustness, anomalous diversity of training data, and / or other problems with the training data used to train the machine learning model. In at least one embodiment, AI-assisted annotation 1310 may be used to help generate annotations corresponding to the imaging data 1308 for use as ground-based data for training or updating the machine learning model. In at least one embodiment, labeled clinical data 1312 may be used as ground-based data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model may be referred to as model training 1314. In at least one embodiment, model training 1314 (e.g., AI-assisted annotation 1310, labeled clinical data 1312, or a combination thereof) can be used as ground-based data to retrain or update the machine learning model. In at least one embodiment, the trained machine learning model can be referred to as output model 1316 and can be used by deployment system 1306, as described herein.

[0159] In at least one embodiment, deployment system 1306 may include software 1318, service 1320, hardware 1322, and / or other components, features, and functions. In at least one embodiment, deployment system 1306 may include a software "stack" such that software 1318 can be built on top of service 1320 and can be used to perform some or all of the processing tasks, and service 1320 and software 1318 can be built on top of hardware 1322 and use hardware 1322 to perform the deployment system's processing, storage, and / or other computational tasks. In at least one embodiment, software 1318 may include any number of different containers, each of which can perform an instantiation of an application. In at least one embodiment, each application can perform one or more processing tasks (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.) in a high-level processing and inference pipeline. In at least one embodiment, in addition to receiving and configuring imaging data for use by each container and / or by facility 1302 after processing through the pipeline, advanced processing and inference pipelines (e.g., to convert output back to available data types) can be defined based on the selection of different containers desired or required for processing imaging data 1308. In at least one embodiment, a combination of containers within software 1318 (e.g., constituting a pipeline) may be referred to as a virtual instrument (as described in more detail herein), and the virtual instrument may utilize service 1320 and hardware 1322 to perform some or all of the processing tasks of an application instantiated within the container.

[0160] In at least one embodiment, the data processing pipeline may receive input data (e.g., imaging data 1308) in a specific format in response to an inference request (e.g., a request from a user of deployment system 1306). In at least one embodiment, the input data may represent one or more images, videos, and / or other data representations generated by one or more imaging devices. In at least one embodiment, the data may be preprocessed as part of the data processing pipeline to prepare it for processing by one or more applications. In at least one embodiment, post-processing may be performed on the output of one or more inference tasks or other processing tasks of the pipeline to prepare output data for the next application and / or to prepare output data for user transmission and / or use (e.g., as a response to an inference request). In at least one embodiment, the inference task may be performed by one or more machine learning models, such as trained or deployed neural networks, which may include the output model 1316 of training system 1304.

[0161] In at least one embodiment, the tasks of the data processing pipeline can be encapsulated in containers, each container representing a discrete, fully functional instantiation of an application and a virtualized computing environment capable of referencing a machine learning model. In at least one embodiment, containers or applications can be published to a private (e.g., limited access) area of ​​a container registry (described in more detail herein), and trained or deployed models can be stored in a model registry 1324 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., a container image) can be used in the container registry, and once a user selects an image from the container registry for deployment in the pipeline, that image can be used to generate containers for instantiation of the application for use by the user's system.

[0162] In at least one embodiment, a developer (e.g., a software developer, clinician, physician, etc.) can develop, publish, and store an application (e.g., as a container) for performing image processing and / or inference on provided data. In at least one embodiment, a software development kit (SDK) associated with the system can be used to perform development, publication, and / or storage (e.g., to ensure that the developed application and / or container conforms to or is compatible with the system). In at least one embodiment, the developed application can be tested locally using the SDK (e.g., at a first facility, testing data from a first facility), the SDK serving as a system (e.g.,...). Figure 14 System 1400 may support at least some services 1320. In at least one embodiment, since DICOM objects may contain one to hundreds of images or other data types, and due to variations in the data, the developer may be responsible for managing (e.g., setting up constructions for preprocessing built into the application, etc.) the extraction and preparation of incoming data. In at least one embodiment, once verified by system 1400 (e.g., for accuracy), the application becomes available in the container registry for user selection and / or implementation to perform one or more processing tasks on data at the user's facility (e.g., a second facility).

[0163] In at least one embodiment, the developer can then share the application or container over a network for the system (e.g., Figure 14The system 1400 allows for user access and use. In at least one embodiment, completed and validated applications or containers may be stored in a container registry, and associated machine learning models may be stored in a model registry 1324. In at least one embodiment, a requesting entity (which provides an inference or image processing request) may browse the container registry and / or model registry 1324 to obtain applications, containers, datasets, machine learning models, etc., select desired combinations of elements to include in the data processing pipeline, and submit an image processing request. In at least one embodiment, the request may include input data necessary to execute the request (and, in some examples, patient-related data), and / or may include selections of applications and / or machine learning models to be executed when the request is processed. In at least one embodiment, the request may then be passed to one or more components of the deployment system 1306 (e.g., the cloud) to perform processing in the data processing pipeline. In at least one embodiment, processing performed by the deployment system 1306 may include referencing elements (e.g., applications, containers, models, etc.) selected from the container registry and / or model registry 1324. In at least one embodiment, once the results are generated through the pipeline, the results can be returned to the user for reference (e.g., for viewing in a suite of viewing applications executed locally, on a local workstation, or on a terminal).

[0164] In at least one embodiment, service 1320 may be utilized to assist in processing or executing applications or containers in the pipeline. In at least one embodiment, service 1320 may include computing services, artificial intelligence (AI) services, visualization services, and / or other service types. In at least one embodiment, service 1320 may provide functionality common to one or more applications in software 1318, thus abstracting functionality into services that can be invoked or utilized by applications. In at least one embodiment, the functionality provided by service 1320 can operate dynamically and more efficiently, while also allowing applications to process data in parallel (e.g., using...). Figure 14The parallel computing platform 1430 in the system can be scaled well. In at least one embodiment, it is not required that each application providing the same functionality as service 1320 must have a corresponding instance of service 1320, but service 1320 can be shared between and among various applications. In at least one embodiment, as a non-limiting example, the service may include an inference server or engine that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service may be included, which can provide the ability to train and / or retrain machine learning models. In at least one embodiment, a data augmentation service may be further included, which can provide GPU-accelerated data (e.g., DICOM, RIS, CIS, conforming to REST, RPC, raw, etc.) extraction, resizing, scaling, and / or other enhancements. In at least one embodiment, a visualization service may be used, which can add image rendering effects (e.g., ray tracing, rasterization, denoising, sharpening, etc.) to add realism to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual instrument service may be included, which provides beamforming, segmentation, inference, imaging, and / or support for other applications within the virtual instrument pipeline.

[0165] In at least one embodiment, where service 1320 includes an AI service (e.g., an inference service), as part of application execution, one or more machine learning models can be executed by invoking (e.g., as an API call) the inference service (e.g., an inference server) to execute one or more machine learning models or their processing. In at least one embodiment, where another application includes one or more machine learning models for a segmentation task, the application can invoke the inference service to execute the machine learning models for performing one or more processing operations associated with the segmentation task. In at least one embodiment, software 1318 implementing advanced processing and inference pipelines, including a segmentation application and an anomaly detection application, can be pipelined because each application can invoke the same inference service to execute one or more inference tasks.

[0166] In at least one embodiment, hardware 1322 may include a GPU, CPU, graphics card, AI / deep learning system (e.g., an AI supercomputer, such as NVIDIA's DGX), cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 1322 may be used to provide efficient, specially built support for software 1318 and services 1320 in deployment system 1306. In at least one embodiment, GPU processing may be used to perform local processing (e.g., at facility 1302) within the AI / deep learning system, in the cloud system, and / or other processing components of deployment system 1306 to improve the efficiency, accuracy, and performance of image processing and generation. In at least one embodiment, as a non-limiting example, software 1318 and / or services 1320 may be optimized for GPU processing in relation to deep learning, machine learning, and / or high-performance computing. In at least one embodiment, at least some of the computing environment of deployment system 1306 and / or training system 1304 may be executed in a data center, one or more supercomputers, or high-performance computing systems with GPU-optimized software (e.g., a hardware and software combination of an NVIDIA DGX system). In at least one embodiment, as described herein, hardware 1322 may include any number of GPUs that can be invoked to perform data processing in parallel. In at least one embodiment, the cloud platform may also include GPU-optimized execution for deep learning tasks, GPU processing for machine learning tasks, or other computational tasks. In at least one embodiment, an AI / deep learning supercomputer and / or GPU-optimized software (e.g., as provided on NVIDIA's DGX systems) may be used as a hardware abstraction and scaling platform to execute the cloud platform (e.g., NVIDIA's NGC). In at least one embodiment, the cloud platform may integrate application container cluster systems or coordination systems (e.g., Kubernetes) across multiple GPUs to allow for seamless scaling and load balancing.

[0167] Figure 14 This is a system diagram of an example system 1400 for generating and deploying an imaging deployment pipeline according to at least one embodiment. In at least one embodiment, system 1400 can be used to implement Figure 13 The process 1300 and / or other processes include advanced processing and inference pipelines. In at least one embodiment, system 1400 may include training system 1304 and deployment system 1306. In at least one embodiment, training system 1304 and deployment system 1306 may be implemented using software 1318, service 1320 and / or hardware 1322, as described herein.

[0168] In at least one embodiment, system 1400 (e.g., training system 1304 and / or deployment system 1306) may be implemented in a cloud computing environment (e.g., using cloud 1426). In at least one embodiment, system 1400 may be implemented locally (in relation to a healthcare facility) or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to the API in cloud 1426 may be restricted to authorized users by establishing security measures or protocols. In at least one embodiment, the security protocol may include a network token, which may be signed by an authentication service (e.g., AuthN, AuthZ, Gluecon, etc.) and may carry appropriate authorization. In at least one embodiment, the API of the virtual instrument (described herein) or other instances of system 1400 may be restricted to a set of public IPs that have been audited or authorized for interaction.

[0169] In at least one embodiment, the various components of system 1400 may communicate with each other using any of a variety of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between facilities and components of system 1400 (e.g., for sending inference requests, for receiving the results of inference requests, etc.) may be transmitted via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.

[0170] In at least one embodiment, similar to the description herein. Figure 13 As described, training system 1304 can execute training pipeline 1404. In at least one embodiment, where deployment system 1306 uses one or more machine learning models in deployment pipeline 1410, training pipeline 1404 can be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1406 (e.g., without retraining or updating). In at least one embodiment, as a result of training pipeline 1404, output model 1316 can be generated. In at least one embodiment, training pipeline 1404 can include any number of processing steps, such as, but not limited to, transformation or adaptation of imaging data (or other input data). In at least one embodiment, different training pipelines 1404 can be used for different machine learning models used by deployment system 1306. In at least one embodiment, similar to the description of... Figure 13 The training pipeline 1404 described in the first example can be used for the first machine learning model, similar to the one described above. Figure 13 The training pipeline 1404 described in the second example can be used for a second machine learning model, similar to the one described above. Figure 13The training pipeline 1404 of the third example described can be used for a third machine learning model. In at least one embodiment, any combination of tasks within the training system 1304 can be used according to the requirements of each corresponding machine learning model. In at least one embodiment, one or more machine learning models may have already been trained and are ready for deployment, so the training system 1304 may not perform any processing on the machine learning models, and one or more machine learning models may be implemented by the deployment system 1306.

[0171] In at least one embodiment, depending on the implementation or embodiment, the output model 1316 and / or the pre-trained model 1406 may include any type of machine learning model. In at least one embodiment, and not limited thereto, the machine learning model used by system 1400 may include models using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutions, recursion, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid state machines, etc.), and / or other types of machine learning models.

[0172] In at least one embodiment, the training pipeline 1404 may include AI-assisted annotations, as described herein regarding at least Figure 15BMore specifically, in at least one embodiment, labeled clinical data 1312 can be generated using any number of techniques (e.g., conventional annotation). In at least one embodiment, in some examples, labels or other annotations can be generated by drawing programs (e.g., annotation programs), computer-aided design (CAD) programs, tagging programs, another type of application suitable for generating annotations or labels for ground reality, and / or can be hand-drawn. In at least one embodiment, ground reality data can be synthetically generated (e.g., generated from computer models or renderings), realistically generated (e.g., designed and generated from real-world data), machine-generated (e.g., extracting features from data using feature analysis and learning, and then generating labels), human-annotated (e.g., taggers or annotation experts, defining the placement of labels), and / or combinations thereof. In at least one embodiment, for each instance of imaging data 1308 (or other data types used by machine learning models), there may be corresponding ground reality data generated by training system 1304. In at least one embodiment, AI-assisted annotation can be performed as part of deployment pipeline 1410; supplementing or replacing AI-assisted annotation included in training pipeline 1404. In at least one embodiment, system 1400 may include a multi-layer platform, which may include a software layer (e.g., software 1318) of a diagnostic application (or other application type) capable of performing one or more medical imaging and diagnostic functions. In at least one embodiment, system 1400 may be communicatively coupled (e.g., via an encrypted link) to a network of PACS servers in one or more facilities. In at least one embodiment, system 1400 may be configured to access and reference data from PACS servers to perform operations such as training machine learning models, deploying machine learning models, image processing, inference, and / or other operations.

[0173] In at least one embodiment, the software layer may be implemented as a secure, encrypted, and / or certified API that can invoke (e.g., call) an application or container from an external environment (e.g., facility 1302). In at least one embodiment, the application may then invoke or execute one or more services 1320 to perform computational, AI, or visualization tasks associated with their respective applications, and the software 1318 and / or service 1320 may utilize the hardware 1322 to perform processing tasks efficiently and effectively.

[0174] In at least one embodiment, deployment system 1306 may execute deployment pipeline 1410. In at least one embodiment, deployment pipeline 1410 may include any number of applications, which may be sequential, non-sequential, or otherwise applied to imaging data (and / or other data types) – including AI-assisted annotation, the imaging data being generated by imaging devices, sequencing devices, genomics devices, etc., as described above. In at least one embodiment, as described herein, deployment pipeline 1410 for an individual device may be referred to as a virtual instrument for the device (e.g., a virtual ultrasound instrument, a virtual CT scanner, a virtual sequencing instrument, etc.). In at least one embodiment, for a single device, more than one deployment pipeline 1410 may exist, depending on the desired information from the data generated from the device. In at least one embodiment, a first deployment pipeline 1410 may exist if it is desired to detect an anomaly from an MRI machine, and a second deployment pipeline 1410 may exist if it is desired to perform image enhancement from the output of the MRI machine.

[0175] In at least one embodiment, the image generation application may include processing tasks that utilize machine learning models. In at least one embodiment, a user may wish to use their own machine learning model or select a machine learning model from the model registry 1324. In at least one embodiment, a user may implement their own machine learning model or select a machine learning model to be included in the application performing the processing tasks. In at least one embodiment, the application may be optional and customizable, and by defining the application's construction, the deployment and implementation of the application for a specific user is presented as a more seamless user experience. In at least one embodiment, by leveraging other features of system 1400 (e.g., service 1320 and hardware 1322), the deployment pipeline 1410 can be more user-friendly, provide easier integration, and produce more accurate, efficient, and timely results.

[0176] In at least one embodiment, deployment system 1306 may include user interface 1414 (e.g., graphical user interface, web interface, etc.) which may be used to select applications to be included in deployment pipeline 1410, deploy applications, modify or change applications or their parameters or configurations, use and interact with deployment pipeline 1410 during setup and / or deployment, and / or otherwise interact with deployment system 1306. In at least one embodiment, although not shown with respect to training system 1304, user interface 1414 (or different user interfaces) may be used to select models to be used in deployment system 1306, to select models to be trained or retrained in training system 1304, and / or to otherwise interact with training system 1304.

[0177] In at least one embodiment, in addition to the application coordination system 1428, a pipeline manager 1412 may also be used to manage interactions between applications or containers deploying pipeline 1410 and services 1320 and / or hardware 1322. In at least one embodiment, the pipeline manager 1412 may be configured to facilitate interactions from application to application, from application to service 1320, and / or from application or service to hardware 1322. In at least one embodiment, although shown as included in software 1318, this is not intended to be limiting, and in some examples (e.g., as...) Figure 12 As shown, pipeline manager 1412 may be included in service 1320. In at least one embodiment, application coordination system 1428 (e.g., Kubernetes, DOCKER, etc.) may include container coordination system that can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications (e.g., rebuilding applications, splitting applications, etc.) from deployment pipeline 1410 with individual containers, each application can execute in a self-contained environment (e.g., at the core level) to improve speed and efficiency.

[0178] In at least one embodiment, each application and / or container (or its image) can be developed, modified, and deployed independently (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separate from the first user or developer). This allows focus on the tasks of a single application and / or container without being hindered by the tasks of another application or container. In at least one embodiment, the pipeline manager 1412 and the application coordination system 1428 can facilitate communication and collaboration between different containers or applications. In at least one embodiment, the application coordination system 1428 and / or the pipeline manager 1412 can facilitate communication and resource sharing between and within each application or container, provided that the expected inputs and / or outputs of each container or application are known to the system (e.g., based on the construction of the application or container). In at least one embodiment, since one or more applications or containers in the deployment pipeline 1410 can share the same services and resources, the application coordination system 1428 can coordinate, load balance, and determine the sharing of services or resources between and within the various applications or containers. In at least one embodiment, the scheduler can be used to track the resource requirements of applications or containers, the current or planned use of these resources, and resource availability. Therefore, in at least one embodiment, the scheduler can allocate resources to different applications and distribute resources between and among applications, taking into account the system's needs and availability. In some examples, the scheduler (and / or other components of the application coordination system 1428) can determine resource availability and distribution based on constraints imposed on the system (e.g., user constraints), such as Quality of Service (QoS), the urgency of data output (e.g., to determine whether to perform real-time processing or delayed processing), etc.

[0179] In at least one embodiment, service 1320, utilized and shared by applications or containers in deployment system 1306, may include computing service 1416, AI service 1418, visualization service 1420, and / or other service types. In at least one embodiment, an application may invoke (e.g., execute) one or more services 1320 to perform processing operations for the application. In at least one embodiment, an application may utilize computing service 1416 to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, one or more computing services 1416 may be utilized to perform parallel processing (e.g., using parallel computing platform 1430) to process data substantially simultaneously through one or more applications and / or one or more tasks of a single application. In at least one embodiment, parallel computing platform 1430 (e.g., NVIDIA's CUDA) may allow general-purpose computing on a GPU (GPGPU) (e.g., GPU 1422). In at least one embodiment, the software layer of parallel computing platform 1430 may provide access to the GPU's virtual instruction set and parallel computing elements to execute computing cores. In at least one embodiment, the parallel computing platform 1430 may include memory, and in some embodiments, memory may be shared between and within multiple containers, and / or between and within different processing tasks within a single container. In at least one embodiment, inter-process communication (IPC) calls may be generated for multiple containers and / or multiple processes within containers to enable the use of the same data (e.g., multiple different stages of one or more applications processing the same information) from a shared memory segment of the parallel computing platform 1430. In at least one embodiment, instead of copying data and moving it to different locations in memory (e.g., read / write operations), the same data in the same memory location can be used for any number of processing tasks (e.g., at the same time, at different times, etc.). In at least one embodiment, this information about the new location of the data can be stored and shared between applications because the resulting data from processing is used to generate new data. In at least one embodiment, the location of the data, and the location of the updated or modified data, may be part of the definition of how the payload in the container is understood.

[0180] In at least one embodiment, AI service 1418 may be used to perform an inference service for executing a machine learning model associated with the application (e.g., a task to perform one or more processing tasks of the application). In at least one embodiment, AI service 1418 may utilize AI system 1424 to execute a machine learning model (e.g., a neural network such as a CNN) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, the application deploying pipeline 1410 may use one or more output models 1316 of self-training system 1304 and / or other models of the application to perform inference on imaging data. In at least one embodiment, two or more examples of using application coordination system 1428 (e.g., a scheduler) for inference may be available. In at least one embodiment, a first category may include a high-priority / low-latency path that can implement a higher service level protocol, such as for performing inference on urgent requests in emergency situations or for radiologists during diagnostic procedures. In at least one embodiment, a second category may include a standard priority path that can be used for requests that may not be urgent or for situations where analysis can be performed at a later time. In at least one embodiment, the application coordination system 1428 may allocate resources (e.g., services 1320 and / or hardware 1322) based on priority paths for different inference tasks of the AI ​​service 1418.

[0181] In at least one embodiment, shared memory may be installed into AI service 1418 in system 1400. In at least one embodiment, shared memory may operate as a cache (or other storage device type) and may be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of deployment system 1306 may receive the request and may select one or more instances (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request may be fed into a database, and if not already in the cache, a machine learning model may be located from model registry 1324. A verification step may ensure that an appropriate machine learning model is loaded into the cache (e.g., shared memory), and / or a copy of the model may be saved to the cache. In at least one embodiment, if the application is not already running or there are not enough instances of the application, a scheduler (e.g., the scheduler of pipeline manager 1412) may be used to start the application referenced in the request. In at least one embodiment, if an inference server has not yet been started to execute the model, an inference server may be started. Any number of inference servers may be started for each model. In at least one embodiment, in a pull model that clusters inference servers, the model can be cached whenever load balancing is favorable. In at least one embodiment, the inference servers can be statically loaded into the corresponding distributed servers.

[0182] In at least one embodiment, an inference server running in a container can be used to perform inference. In at least one embodiment, an instance of the inference server can be associated with a model (and optionally multiple versions of the model). In at least one embodiment, if an instance of the inference server does not exist when a request to perform inference on the model is received, a new instance can be loaded. In at least one embodiment, when the inference server is started, a model can be passed to the inference server, allowing the same container to be used to serve different models, as long as the inference server runs as different instances.

[0183] In at least one embodiment, during application execution, an inference request for a given application can be received, and a container (e.g., an instance of a hosted inference server) can be loaded (if not already loaded), and a launcher can be invoked. In at least one embodiment, preprocessing logic within the container can (e.g., using a CPU and / or GPU) load, decode, and / or perform any additional preprocessing on the incoming data. In at least one embodiment, once the data is ready for inference, the container can infer the data as needed. In at least one embodiment, this can include a single inference call for an image (e.g., a hand X-ray) or can request inference for hundreds of images (e.g., a chest CT scan). In at least one embodiment, the application can summarize the results before completion, which may include, but is not limited to, a single confidence score, pixel-level segmentation, voxel-level segmentation, generating visualizations, or generating text to summarize the results. In at least one embodiment, different priorities can be assigned to different models or applications. For example, some models may have a real-time (TAT less than 1 minute) priority, while other models may have a lower priority (e.g., TAT less than 10 minutes). In at least one embodiment, model execution time can be measured from the requesting agency or entity, and may include cooperative network traversal time and inference service execution time.

[0184] In at least one embodiment, the transfer of requests between service 1320 and the inference application can be hidden behind a software development kit (SDK) and robust transfer can be provided via queues. In at least one embodiment, requests are placed in queues via an API for individual application / tenant ID combinations, and the SDK pulls requests from the queues and provides them to the application. In at least one embodiment, the name of the queue can be provided in the environment where the SDK picks up the queue. In at least one embodiment, asynchronous communication via queues may be useful because it allows any instance of the application to pick up work when it becomes available. Results can be sent back via queues to ensure no data loss. In at least one embodiment, queues can also provide the ability to partition work, as the highest priority work can go into a queue connected to a majority of instances of the application, while the lowest priority work can go into a queue connected to a single instance that processes tasks in the order they are received. In at least one embodiment, the application can run on a GPU-accelerated instance generated in cloud 1426, and the inference service can perform inference on the GPU.

[0185] In at least one embodiment, visualization service 1420 can be used to generate visualizations for viewing the output of application and / or deployment pipeline 1410. In at least one embodiment, visualization service 1420 can utilize GPU 1422 to generate visualizations. In at least one embodiment, visualization service 1420 can implement rendering effects such as ray tracing to generate higher quality visualizations. In at least one embodiment, visualizations can include, but are not limited to, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slicing, virtual reality display, augmented reality display, etc. In at least one embodiment, a virtualized environment can be used to generate virtual interactive displays or environments (e.g., virtual environments) for system users (e.g., doctors, nurses, radiologists, etc.) to interact with. In at least one embodiment, visualization service 1420 can include an internal visualizer, cinematic and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).

[0186] In at least one embodiment, hardware 1322 may include GPU 1422, AI system 1424, cloud 1426, and / or any other hardware for performing training system 1304 and / or deployment system 1306. In at least one embodiment, GPU 1422 (e.g., NVIDIA's TESLA and / or QUADRO GPUs) may include any number of GPUs that can be used to perform processing tasks for any feature or function of computing service 1416, AI service 1418, visualization service 1420, other services, and / or software 1318. For example, for AI service 1418, GPU 1422 may be used to perform preprocessing on imaging data (or other data types used by machine learning models), postprocessing on the output of machine learning models, and / or perform inference (e.g., to execute machine learning models). In at least one embodiment, cloud 1426, AI system 1424, and / or other components of system 1400 may use GPU 1422. In at least one embodiment, cloud 1426 may include a GPU-optimized platform for deep learning tasks. In at least one embodiment, AI system 1424 may use a GPU, and one or more AI systems 1424 may be used to perform cloud 1426 (or at least part of a task for deep learning or inference). Similarly, although hardware 1322 is shown as a discrete component, this is not intended to be limiting, and any component of hardware 1322 may be combined with or utilized by any other component of hardware 1322.

[0187] In at least one embodiment, AI system 1424 may include a specially built computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, in addition to CPU, RAM, memory, and / or other components, features, or functions, AI system 1424 (e.g., NVIDIA's DGX) may also include GPU-optimized software (e.g., a software stack) that can be executed using multiple GPUs 1422. In at least one embodiment, one or more AI systems 1424 may be implemented in a cloud 1426 (e.g., in a data center) to perform some or all of the AI-based processing tasks of system 1400.

[0188] In at least one embodiment, cloud 1426 may include GPU-accelerated infrastructure (e.g., NVIDIA's NGC) that can provide a GPU-optimized platform for performing processing tasks of system 1400. In at least one embodiment, cloud 1426 may include AI system 1424 for performing one or more AI-based tasks of system 1400 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, cloud 1426 may be integrated with application coordination system 1428 utilizing multiple GPUs to allow seamless scaling and load balancing between and within applications and services 1320. In at least one embodiment, as described herein, cloud 1426 may be responsible for performing at least some of the services 1320 of system 1400, including computing service 1416, AI service 1418, and / or visualization service 1420. In at least one embodiment, cloud 1426 may perform large and small batch inference (e.g., perform NVIDIA's TENSORRT), provide accelerated parallel computing APIs and platform 1430 (e.g., NVIDIA's CUDA), perform application coordination system 1428 (e.g., KUBERNETES), provide graphics rendering APIs and platform (e.g., for ray tracing, 2D graphics, 3D graphics and / or other rendering techniques to produce higher quality cinematic effects), and / or provide other functionalities for system 1400.

[0189] Figure 15A A data flow diagram of a process 1500 for training, retraining, or updating a machine learning model according to at least one embodiment is shown. In at least one embodiment, a non-limiting example can be used. Figure 14System 1400 executes process 1500. In at least one embodiment, process 1500 may utilize services 1320 and / or hardware 1322 of system 1400, as described herein. In at least one embodiment, the refined model 1512 generated by process 1500 may be executed by deployment system 1306 for one or more containerized applications in deployment pipeline 1410.

[0190] In at least one embodiment, model training 1314 may include retraining or updating the initial model 1504 (e.g., a pre-trained model) using new training data (e.g., new input data, such as customer dataset 1506, and / or new ground reality data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, the output or loss layer of the initial model 1504 may be reset or deleted, and / or replaced with an updated or new output or loss layer. In at least one embodiment, the initial model 1504 may have previously finely tuned parameters (e.g., weights and / or biases) retained from previous training, so training or retraining 1314 may not require as much time or processing as training the model from scratch. In at least one embodiment, during model training 1314, by resetting or replacing the output or loss layer of the initial model 1504, on a new customer dataset 1506 (e.g., new input data, such as customer dataset 1506, and / or new ground reality data associated with the input data), the initial model 1504 may be retrained or updated. Figure 13 When generating predictions on image data (1308), the parameters of the new dataset can be updated and readjusted based on the loss calculation associated with the accuracy of the output or loss layer.

[0191] In at least one embodiment, the pre-trained model 1406 may be stored in a data storage or registry (e.g., Figure 13(Model registry 1324). In at least one embodiment, the pre-trained model 1406 may have been trained at least partially at one or more facilities other than the facility executing process 1500. In at least one embodiment, to protect the privacy and rights of patients, subjects, or customers at different facilities, the pre-trained model 1406 may have been trained locally using locally generated customer or patient data. In at least one embodiment, the pre-trained model 1406 may be trained using cloud 1426 and / or other hardware 1322, but confidential, privacy-protected patient data may not be transferred to, used by, or accessed by any component of cloud 1426 (or other non-local hardware). In at least one embodiment, if the pre-trained model 1406 is trained using patient data from more than one facility, the pre-trained model 1406 may have been trained separately for each facility before training on patient or customer data from another facility. In at least one embodiment, such as when customer or patient data has been published for privacy reasons (e.g., by abandonment, for experimental purposes, etc.), or where customer or patient data is included in a public dataset, customer or patient data from any number of facilities can be used to train a pre-trained model 1406 locally and / or externally, such as in a data center or other cloud computing infrastructure.

[0192] In at least one embodiment, when selecting an application for use in deployment pipeline 1410, the user may also select a machine learning model for a specific application. In at least one embodiment, the user may not have a model available, so the user may select a pre-trained model 1406 to use with the application. In at least one embodiment, the pre-trained model 1406 may not be optimized to generate accurate results on the user facility's customer dataset 1506 (e.g., based on patient diversity, demographics, type of medical imaging equipment used, etc.). In at least one embodiment, the pre-trained model 1406 may be updated, retrained, and / or fine-tuned for use at various facilities before being deployed to deployment pipeline 1410 for use with one or more applications.

[0193] In at least one embodiment, a user may select a pre-trained model 1406 to be updated, retrained, and / or fine-tuned, and the pre-trained model 1406 may be referred to as the initial model 1504 of the training system 1304 in process 1500. In at least one embodiment, a client dataset 1506 (e.g., imaging data, genomic data, sequencing data, or other data types generated by equipment at the facility) may be used to perform model training 1314 (which may include, but is not limited to, transfer learning) on ​​the initial model 1504 to generate a refined model 1512. In at least one embodiment, ground-based data corresponding to the client dataset 1506 may be generated by the training system 1304. In at least one embodiment, ground-based data (e.g., such as...) may be generated at the facility at least in part by clinicians, scientists, physicians, practitioners, etc. Figure 13 Clinical data marked in 1312).

[0194] In at least one embodiment, AI-assisted annotation 1310 may be used in some examples to generate ground reality data. In at least one embodiment, AI-assisted annotation 1310 (e.g., implemented using an AI-assisted annotation SDK) may leverage machine learning models (e.g., neural networks) to generate suggested or predicted ground reality data for a customer dataset. In at least one embodiment, user 1510 may use the annotation tool within a user interface (graphical user interface (GUI)) on computing device 1508.

[0195] In at least one embodiment, user 1510 can interact with the GUI via computing device 1508 to edit or fine-tune annotations or automatic annotations. In at least one embodiment, polygon editing features can be used to move the vertices of a polygon to more precise or fine-tuned positions.

[0196] In at least one embodiment, once the customer dataset 1506 has associated ground-based data, the ground-based data (e.g., from AI-assisted annotations, manual labeling, etc.) can be used to generate a refined model 1512 during model training 1314. In at least one embodiment, the customer dataset 1506 can be applied to the initial model 1504 an arbitrary number of times, and the ground-based data can be used to update the parameters of the initial model 1504 until an acceptable level of accuracy is achieved for the refined model 1512. In at least one embodiment, once the refined model 1512 is generated, it can be deployed within one or more deployment pipelines 1410 at the facility to perform one or more processing tasks related to medical imaging data.

[0197] In at least one embodiment, the refined model 1512 can be uploaded to the pre-trained model 1406 in the model registry 1324 for selection by another facility. In at least one embodiment, this process can be completed at any number of facilities, allowing the refined model 1512 to be further refined any number of times on a new dataset to generate a more general model.

[0198] Figure 15B This is an example illustration of a client-server architecture 1532 for enhancing an annotation tool using a pre-trained annotation model, according to at least one embodiment. In at least one embodiment, an AI-assisted annotation tool 1536 may be instantiated based on the client-server architecture 1532. In at least one embodiment, the annotation tool 1536 in an imaging application can assist ray surgeons, for example, in identifying organs and abnormalities. In at least one embodiment, the imaging application may include software tools, as a non-limiting example, that help user 1510 identify several extreme points on a specific organ of interest in a raw image 1534 (e.g., in a 3D MRI or CT scan) and receive automatic annotation results for all 2D slices of that specific organ. In at least one embodiment, the results may be stored in a data store as training data 1538 and used as (e.g., but not limited to) ground-based data for training. In at least one embodiment, when computing device 1508 sends extreme points for AI-assisted annotation 1310, for example, a deep learning model may receive this data as input and return inference results for segmenting organs or abnormalities. In at least one embodiment, a pre-instantiated annotation tool (e.g., Figure 15B The AI-assisted annotation tool 1536B can be enhanced by making API calls (e.g., API call 1544) to a server (such as annotation assistant server 1540), which may include a set of pre-trained models 1542 stored, for example, in an annotation model registry. In at least one embodiment, the annotation model registry may store pre-trained models 1542 (e.g., machine learning models, such as deep learning models) that have been pre-trained to perform AI-assisted annotation on specific organs or abnormalities. In at least one embodiment, these models can be further updated using a training pipeline 1404. In at least one embodiment, the pre-installed annotation tool can be improved over time as newly labeled clinical data 1312 is added.

[0199] Such components can be used to determine and optimize the motion of automated systems to perform defined tasks.

[0200] Other variations are within the spirit of this disclosure. Therefore, while the disclosed technology is readily adaptable to various modifications and alternative constructions, certain embodiments thereof are illustrated in the accompanying drawings and have been described in detail above. However, it should be understood that the disclosure is not intended to be limited to one or more specific forms disclosed, but rather, it is intended to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of this disclosure as defined in the appended claims.

[0201] Unless otherwise stated or obviously contradicted by the context, the terms “a,” “an,” and “the,” and similar references, used in the context of describing the disclosed embodiments (particularly in the context of the appended claims), should be interpreted as encompassing both singular and plural forms, rather than as definitions of the terms. Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including, but not limited to”). The term “connection” (referring to a physical connection where not modified) should be interpreted as partially or wholly contained, attached to, or joined together, even with some intervention. Unless otherwise indicated herein, references to numerical ranges herein are intended only as a way of abbreviating each individual value falling within that range, and each individual value is incorporated into the specification as if it were separately described herein. Unless otherwise indicated or contradicted by the context, the use of the terms “set” (e.g., “item set”) or “subset” should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise indicated or contradicted by the context, the term "subset" of a corresponding set does not necessarily refer to an appropriate subset of the corresponding set, but rather the subset and the corresponding set can be equal.

[0202] Unless otherwise explicitly stated or clearly contradicted by the context, connective phrases such as “at least one of A, B, and C” or “at least one of A, B, and C” are understood in the context to generally refer to items, terms, etc., which can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in an illustrative example of a set with three members, the connective phrases “at least one of A, B, and C” and “at least one of A, B, and C” refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Therefore, such connective language is generally not intended to imply that some embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by the context, the term “multiple” indicates a plural state (e.g., “multiple items” means multiple items). The number of items in a multiple item is at least two, but may be more if explicitly indicated or indicated by the context. Furthermore, unless otherwise stated or clearly understood from the context, the phrase “based on” means “at least partially based on” rather than “based on only”.

[0203] Unless otherwise indicated herein or clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations thereof and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that are executed jointly on one or more processors via hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media comprises multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media lack all the code, but the multiple non-transitory computer-readable storage media collectively store all the code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors; for example, the non-transitory computer-readable storage media store the instructions, and the main central processing unit (“CPU”) executes some instructions while the graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and the different processors execute different subsets of the instructions.

[0204] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that perform the operations of the processes described herein, either individually or collectively, and such a computer system is configured with suitable hardware and / or software to enable the implementation of the operations. Furthermore, the computer system implementing at least one embodiment of this disclosure is a single device, and in another embodiment it is a distributed computer system comprising multiple devices operating in different ways, such that the distributed computer system performs the operations described herein, and that a single device does not perform all the operations.

[0205] The use of any and all examples or exemplary language (e.g., “such as”) provided herein is intended only to better illustrate embodiments of this disclosure and does not constitute a limitation on the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of the disclosure.

[0206] All references cited in this article, including publications, patent applications and patents, are incorporated herein by reference as if each reference were individually and specifically indicated to be incorporated herein by reference and the entire contents of which are described herein.

[0207] The terms “coupled” and “connected”, and their derivatives, may be used in the specification and claims. It should be understood that these terms may not be intended to be synonyms with each other. Rather, in certain examples, “connected” or “coupled” may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0208] Unless otherwise expressly stated, it will be understood that throughout this specification, terms such as “processing,” “computing,” “determining,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that process and / or convert data represented as physical quantities (e.g., electrons) in the registers and / or memory of the computing system into other data represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.

[0209] In a similar manner, the term "processor" can refer to any device or part of memory that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process can refer to multiple processes that execute instructions sequentially or intermittently, sequentially, or in parallel. The terms "system" and "method" are used interchangeably herein, provided that a system can embody one or more methods, and a method can be considered a system.

[0210] This document refers to the process of acquiring, obtaining, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. Analog and digital data can be acquired, obtained, received, or input in various ways, such as by receiving data as a parameter to a function call or a call to an application programming interface (API). In some implementations, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference can also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter to a function call, an API call, or an inter-process communication mechanism.

[0211] While the discussion above illustrates example implementations of the described technologies, other architectures can be used to implement the described functionality and are intended to fall within the scope of this disclosure. Furthermore, although specific assignments of responsibilities have been defined above for discussion purposes, various functions and responsibilities can be assigned and divided in different ways depending on the circumstances.

[0212] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims.

Claims

1. A method for controlling a robot, comprising: Receive data representing the environment, which includes objects held by a person's hand; The target grasping option is determined from one or more potential grasping options, at least in part, based on an evaluation of one or more potential grasping options corresponding to the robot. The motion sequence corresponding to the robot's initial and final positions is determined based on the target grasping options, and the motion sequence is determined at least in part based on minimizing one or more cost functions and satisfying one or more motion constraints; The robot is made to perform the corresponding motion in the motion sequence at each time step corresponding to one or more motions in the motion sequence; Detecting contact between the robot and the object corresponding to the target grasping position; and The robot's end effector grasps the object.

2. The method of claim 1, wherein the one or more motion constraints include at least one of the following: constraints limiting acceleration, constraints favoring linear motion, constraints avoiding collisions, or constraints avoiding occlusion of sensors used to capture the data.

3. The method of claim 1, wherein determining the motion sequence comprises using a model predictive control (MPC) system to execute at least one optimization algorithm having the one or more motion constraints.

4. The method of claim 3, wherein determining the target crawling option is performed using the MPC system.

5. The method of claim 3, wherein the MPC system is used to optimize the motion sequence for each of the one or more potential grasping options.

6. The method of claim 1, further comprising: The one or more motion constraints are modified, at least in part, based on one or more user inputs.

7. The method of claim 1, further comprising: During each of the time steps of the movement, monitor whether the end effector comes into contact with the human hand; as well as Once it is determined that the end effector has made contact with the human hand, one or more operations are performed.

8. The method of claim 1, wherein each motion in the motion sequence is determined using one or more joint accelerations optimized for each time step.

9. A method for controlling a robot, comprising: During the time step sequence, a set of positions is determined where the robot will perform actions; Determine a motion sequence between the current position and the final position in the set of positions, which satisfies one or more motion constraints, the motion sequence allowing the target position in the set of positions to be changed at each time step in the time step sequence; The robot is made to perform the corresponding motion in the motion sequence at each time step in the time step sequence, so that at least a portion of the robot moves relative to the target position; as well as The robot performs the action when at least a portion of it is determined to be within a threshold distance from the target location.

10. The method of claim 9, wherein one or more of the set of locations includes location and orientation information.

11. The method of claim 9, wherein the motion sequence is determined using a prediction model and one or more optimization criteria.

12. The method of claim 11, wherein the prediction model optimizes the motion sequence at the set of locations and evaluates the target location at each time step in the time step sequence.

13. The method of claim 9, wherein the one or more motion constraints include at least one of the following: constraints that limit acceleration, constraints that facilitate linear motion, constraints that prevent collisions, or constraints that prevent sensor obstruction.

14. The method of claim 9, further comprising: During the time step sequence, sensor data representing the physical environment in which the robot will perform the actions is captured, and The motion sequence is determined at least in part based on the determined changes in the physical environment.

15. A system for controlling a robot, comprising: One or more processing units are used for: During the time step sequence, a set of positions is determined where the robot will perform actions; Determine a motion sequence between the current position and the final position in the set of positions, which satisfies one or more motion constraints, the motion sequence allowing the target position in the set of positions to be changed at each time step in the time step sequence; The robot is made to perform the corresponding motion in the motion sequence at each time step in the time step sequence, so that at least a portion of the robot moves relative to the target position; as well as When at least a portion of the robot is determined to be within a threshold distance from the target location, the robot performs the action.

16. The system of claim 15, wherein the one or more processing units are further configured to: A predictive model is used to determine the motion sequence based at least in part on one or more optimization criteria.

17. The system of claim 16, wherein the one or more processing units are further configured to: The prediction model is used to optimize the motion sequence at the set of locations, and the target location is evaluated at each time step in the time step sequence.

18. The system of claim 15, wherein the one or more processing units are further configured to: During the time step sequence, sensor data representing the physical environment in which the robot will perform the actions is captured, and The change in the target location is determined at least in part based on the determined changes in the physical environment.

19. The system of claim 15, wherein the one or more motion constraints include at least one of the following: constraints limiting acceleration, constraints favoring linear motion, constraints avoiding collisions, or constraints avoiding sensor obstruction.

20. The system of claim 15, wherein the system comprises at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for collaborative content creation of 3D assets; A system used to perform deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing conversational AI operations; A system for generating synthetic data; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.