STRATEGY LEVELS FOR MACHINE CONTROL

DE112022001174B4Active Publication Date: 2025-07-10NVIDIA CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
DE112022001174
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-27
Filing Date
2022-04-26
Publication Date
2025-07-10
Estimated Expiration
2042-04-26

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer system comprising one or more processors and a computer-readable memory storing instructions executable by the one or more processors to cause the computer system to at least: Identifying a strategy (404) to cause a machine (106, 206) to perform at least one movement, the strategy including at least a plurality of strategy levels comprising: a first strategy level (108) for causing the machine (106, 206) to execute a first movement sequence that reaches a neutral state, the first movement sequence being limited by at least a first parameter associated with the machine and a second parameter associated with a range in which the machine is to operate, and a second strategy level (110) for causing the machine (106, 206) to execute a second movement sequence without affecting the neutral state associated with the first strategy level; and Executing the strategy (404) to cause the machine to perform the at least one movement, wherein the at least one movement comprises at least the first movement sequence and the second movement sequence.
Need to check novelty before this filing date? Find Prior Art

Description

FIELD OF TECHNOLOGY

[0001] At least one embodiment relates to the control of a machine. For example, at least one embodiment relates to a machine, such as a robot, that is controlled based on a strategy. STATE OF THE ART

[0002] Many machines are programmed to use one or more end effectors to manipulate one or more objects. For example, a machine, such as a robot, may use a gripper or other device to apply force to an object and thereby cause that object to move. For example, the machine may use the gripper or other device to translate an object without necessarily grasping it. In addition, the machine may use the gripper to pick up an object from a first location, move the object to a second location, and place the object at the second location. The complexity of the tasks performed by machines is increasing. The programming associated with these machines is evolving in response to this increasing complexity to ensure that the machines perform their tasks correctly and safely.Furthermore, reference is made to US 2012 / 0 173 021 A1, US 8 977 394 B2, US 2019 / 0 130 312 A1, US 2020 / 0 114 506 A1, US 10 926 408 B1, and EP 1 195 231 A1. It is an object of the present invention to further improve the control of a machine.

[0003] This object is achieved by the features of the independent claims. Preferred embodiments are described in the dependent claims. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 illustrates an example system environment for controlling a machine according to at least one embodiment; Fig. 2 illustrates an example system environment for controlling a robot to initiate engagement with an object, according to at least one embodiment; Fig. 3 illustrates an example system environment for controlling a robot to engage an object, according to at least one embodiment; Fig. 4 illustrates another example of a system environment for controlling a robot according to at least one embodiment; Fig. 5 illustrates additional details of the system environment for controlling the robot according to at least one embodiment; Fig. 6 illustrates additional details of the system environment for controlling the robot according to at least one embodiment; Fig. 7 illustrates an example of a process for causing a computer-implemented action, such as causing movement of a machine, according to one embodiment; Fig. 8 illustrates another example of a process for causing a computer-implemented action, such as causing movement of a machine, according to one embodiment; Fig. 9A illustrates inference and / or training logic according to at least one embodiment; Fig. 9B illustrates inference and / or training logic according to at least one embodiment; Fig. 10 illustrates the training and deployment of a neural network according to at least one embodiment; Fig. 11 illustrates an exemplary data center system according to at least one embodiment; Fig. 12A illustrates an exemplary autonomous vehicle according to at least one embodiment; Fig. Figure 12B illustrates an example of camera locations and fields of view for the autonomous vehicle from Fig. 12A according to at least one embodiment; Fig. Figure 12C is a block diagram illustrating an example system architecture for the autonomous vehicle of Fig. 12A illustrates, according to at least one embodiment; Fig. 12D is a diagram illustrating a system for communication between cloud-based server(s) and the autonomous vehicle. Fig. 12A illustrates, according to at least one embodiment; Fig. 13 is a block diagram illustrating a computer system, according to at least one embodiment; Fig. 14 is a block diagram illustrating a computer system, according to at least one embodiment; Fig. 15 illustrates a computer system according to at least one embodiment; Fig. 16 illustrates a computer system according to at least one embodiment; Fig. 17A illustrates a computer system according to at least one embodiment; Fig. 17B illustrates a computer system according to at least one embodiment; Fig. 17C illustrates a computer system according to at least one embodiment; Fig. 17D illustrates a computer system according to at least one embodiment; Fig. 17E and Fig. 17F illustrate a shared programming model according to at least one embodiment; Fig. 18 illustrates example integrated circuits and associated graphics processors according to at least one embodiment; Fig. 19A and Fig. 19B illustrate example integrated circuits and associated graphics processors according to at least one embodiment; Fig. 20A and Fig. 20B illustrate additional example graphics processor logic according to at least one embodiment; Fig. 21 illustrates a computer system according to at least one embodiment; Fig. 22A illustrates a parallel processor according to at least one embodiment; Fig. 22B illustrates a partition unit according to at least one embodiment; Fig. 22C illustrates a processing cluster according to at least one embodiment; Fig. 22D illustrates a graphics multiprocessor according to at least one embodiment; Fig. 23 illustrates a system having multiple graphics processing units (GPUs) according to at least one embodiment; Fig. 24 illustrates a graphics processor according to at least one embodiment; Fig. 25 is a block diagram illustrating a processor microarchitecture for a processor, according to at least one embodiment; Fig. 26 illustrates a deep learning application processor according to at least one embodiment; Fig. 27 is a block diagram illustrating an exemplary neuromorphic processor, according to at least one embodiment; Fig. 28 illustrates at least portions of a graphics processor according to one or more embodiments; Fig. 29 illustrates at least portions of a graphics processor according to one or more embodiments; Fig. 30 illustrates at least portions of a graphics processor according to one or more embodiments; Fig. 31 is a block diagram of a graphics processing engine of a graphics processor according to at least one embodiment; Fig. 32 is a block diagram of at least portions of a graphics processor core according to at least one embodiment; Fig. 33A and Fig. 33B illustrate thread execution logic including an array of processing elements of a graphics processor core, according to at least one embodiment; Fig. 34 illustrates a parallel processing unit (“PPU”) according to at least one embodiment; Fig. 35 illustrates a general processing cluster (“GPC”) according to at least one embodiment; Fig. 36 illustrates a memory partition unit of a parallel processing unit ("PPU") according to at least one embodiment; Fig. 37 illustrates a streaming multiprocessor according to at least one embodiment; Fig. 38 is an example data flow diagram for an enhanced compute pipeline according to at least one embodiment; Fig. 39 is a system diagram for an example system for training, adapting, instantiating, and deploying machine learning models in an advanced compute pipeline, according to at least one embodiment; Fig. 40 includes an exemplary illustration of an advanced compute pipeline 3910A for processing imaging data in accordance with at least one embodiment; Fig. 41A includes an example data flow diagram of a virtual instrument supporting an ultrasound device, according to at least one embodiment; Fig. 41B includes an example data flow diagram of a virtual instrument supporting a CT scanner, according to at least one embodiment; Fig. 42A illustrates a data flow diagram for a process for training a machine learning model according to at least one embodiment; and Fig. 42B is an example illustration of a client-server architecture for extending annotation tools with pre-trained annotation models, according to at least one embodiment. DETAILED DESCRIPTION

[0004] Programming a machine, such as a robot or vehicle, to perform a desired movement is an important task. Specifically, programming the machine should ensure that the machine is capable of performing a desired movement. Furthermore, programming should ensure that the desired movement is performed in a predictable and safe manner.

[0005] In at least one embodiment, a strategy for controlling a machine, for example, for causing the machine to move, is designed using a level-based approach. In particular, the strategy is generated from a plurality of strategy levels. Each strategy level in the strategy can control a desired movement of the machine. In at least one embodiment, a base or first strategy level of the strategy is used to control the movement of the machine from a rest state to a destination associated with a work or task space or area of the machine. The base strategy level is designed to cause the machine to reach a static or neutral state at the destination. In particular, the base strategy level is designed to cause the machine to come to a standstill at the destination.In at least one embodiment, a subsequent or second level is added to the strategy that includes the first level. The second level controls movement of the machine from the neutral state or the goal reached using the first level of the strategy. Alternatively, the second level may be added to the strategy that includes the first level to cause the machine to avoid at least one obstacle, such as an object. In at least one embodiment, the second level is intended to cause movement of the machine without affecting the distortion of the machine caused by the first level. For example, in at least one embodiment, the second level of the strategy may cause movement of the machine without affecting a neutral state that the machine has reached based on the first level of the strategy.A strategy implemented using one or more strategy levels may be referred to herein as a geometric structure. In at least one embodiment, the geometric structure is intended to cause the machine to perform a shaped movement to accomplish an assigned task. In at least one embodiment, each strategy level causes a machine to perform at least one movement sequence or at least one movement, and each subsequent strategy level builds upon one or more movement sequences or movements caused by one or more previously executed strategy levels.

[0006] As indicated, the generation of motion sequences is important for the control of machines, such as robots and vehicles. In particular, fast, reactive motion sequences are important for most modern tasks, especially in highly dynamic and uncertain collaborative environments. The geometric structures described herein provide a direct construction of stable robot behavior in modular parts. Geometric structures define nominal behavior for machines independent of a particular task, incorporating cross-task commonalities such as joint limit avoidance, obstacle avoidance, and redundancy resolution. In at least one embodiment, the described geometric structures cause a machine to execute one or more motions through a strategy-level approach implemented by a geometric structure.

[0007] In at least one embodiment, a machine is expected to reach a target point. For example, a robot may be expected to reach a target point using its end effector, such as a gripper. The robot's movement to reach the target point should avoid obstacles and limits of the robot's joints. The geometric structure that causes the robot's movement should intelligently resolve redundancies, implement global navigation heuristics, and the geometric structure may need to shape the gripper's path to approach a target from a specific direction. In at least one embodiment, the geometric structure includes a plurality of strategy layers, each layer having an associated differential equation used to implement a movement the robot is expected to perform.In at least one embodiment, the differential equation associated with each level of the geometric structure is a second-order differential equation. The second-order differential equation of each level may be a second-order differential equation with an associated metric tensor that optimizes under constraints. In at least one embodiment, the levels of the geometric structure are based on a Lagrangian and / or geometric formulations. Lagrangians include associated equations of motion, and general nonlinear geometries may include a special class of Finsler geometries. At a higher level, a nonlinear geometry is a differential equation whose solutions define a set of paths (not velocity-dependent trajectories, but velocity-independent paths).An example is a Riemannian geometry, a type of Finsler geometry whose paths are length-minimizing under the Riemann length measure.

[0008] In at least one embodiment, at least one level of a strategy, for example, a level in a geometric structure, can be generated by a nonlinear geometry and an energy input to that geometry to find a representation of the geometry that includes a set of trajectories along paths that are energy-conserving under a given Lagrangian energy (e.g., an energy transformation to the geometry). The nonlinear geometry defines the basic nominal behavior of the machine (i.e., the paths it follows when unperturbed), and the Lagrangian energy defines the nature of the forces acting to push the machine away from those nominal paths. In particular, the energy has an associated energy tensor that acts as a generalized mass matrix that defines how the machine accelerates away from the geometric paths when forced.

[0009] In at least one embodiment, a computer system may be provided that generates a strategy or geometric structure that includes strategy levels. The computer system may include at least one processor and memory, and a user may operate the computer system to cause the computer system to generate the strategy. In at least one embodiment, the computer system includes a user interface (UI) that the user uses to generate the strategy levels included in the strategy. For example, the user may use the UI to generate a first strategy level to cause a machine, such as a robot, to generate a first motion sequence that reaches a neutral state.For example, the first strategy level may cause a machine to reach a steady state or to come to a stop within a range in which the machine is intended to operate. The first strategy level may cause a machine to reach the neutral data based on one or more parameters. In at least one embodiment, a first parameter may be associated with a joint of the machine. The first parameter may prevent the machine from exceeding a joint limit of the machine. The one or more parameters may also include a second parameter, where the second parameter is a coordinate within the range in which the machine is intended to operate. In at least one embodiment, the first strategy level causes the machine to come to a stop at the coordinate within the range in which the machine is intended to operate.

[0010] The user may use the UI to generate a second strategy level included in the strategy. In at least one embodiment, the second strategy level is to cause the machine to generate a second motion sequence. The second motion sequence may build on the first motion sequence that the machine executes based on the first strategy level. In at least one embodiment, the second motion sequence associated with the second strategy level does not affect the distortion associated with the first motion sequence caused by the first strategy level. For example, the second motion sequence caused by the second strategy level does not affect the first motion sequence reaching the neutral state.In at least one embodiment, the acceleration and / or velocity associated with the second motion sequence does not affect the acceleration and / or velocity associated with the first motion sequence provided by the first strategy level.

[0011] Fig. 1 illustrates an example of a system environment according to at least one embodiment. In at least one embodiment, the system environment includes a computer system 100. The computer system 100 may include one or more processors 102 and a computer memory 104. The one or more processors 102 may include a graphics processing unit described below. In at least one embodiment, the computer memory 104 may include memory storing executable instructions that, as a result of execution by the one or more processors 102, cause the computer system 100 to control a machine 106. In at least one embodiment, the computer system 100 may cause the machine 106 to perform one or more motion sequences or movements.In at least one embodiment, computer system 100 may cause machine 106 to perform one or more motion sequences or movements based on one or more geometric structures stored in memory 104. In at least one embodiment, the one or more geometric structures stored in memory 104 include one or more strategy levels 108 and 110. Multiple strategy levels may be stored in memory 104.

[0012] In at least one embodiment, machine 106 is a robot. The robot may be an articulated robot that includes one or more arms. In at least one embodiment, machine 106 is a vehicle, such as an automobile, that can be controlled by a user. In at least one embodiment, one or more movements of the vehicle are controlled by strategy layers of computer system 100. In at least one embodiment, one or more of the strategy layers may be combined in memory 104 to provide a strategy or geometric structure that causes machine 106 to perform one or more motion sequences or movements.

[0013] In at least one embodiment, the strategy layer 108 includes a differential equation to cause the machine 106 to perform a desired movement or sequence of movements. The differential equation may be a second-order differential equation. In at least one embodiment, the differential equation is homogeneous of degree two. More specifically, in at least one embodiment, the differential equation includes one or more trajectories that cause the machine 106 to perform a sequence of movements. Furthermore, the differential equation includes a path consistency property that ensures that one or more integral curves emanating from a particular position and having a particular velocity follow a desired path of movement.In at least one embodiment, strategy level 110 also includes a differential equation to cause machine 106 to perform a desired movement or movement sequence following the movement or movement sequence caused by strategy level 108. The differential equation may be a second-order differential equation. In at least one embodiment, the differential equation is homogeneous of degree two. More specifically, in at least one embodiment, the differential equation includes one or more trajectories that cause machine 106 to perform a movement sequence that builds on the movement sequence caused by strategy level 108. Furthermore, the differential equation includes a path consistency property that ensures that one or more integral curves emanating from a particular position and having a particular velocity follow a desired movement path.

[0014] The strategy layers 108 and / or 110 may be implemented via a UI 112. In particular, a user may interface with the computer system 100 via the UI 112 to create the strategy layers 108 and / or 110. In one embodiment, the strategy layers 108 and / or 110 may be constructed via the UI 112 in portions distributed across a transformed tree of a relevant task space or domain in which the machine 106 is to operate. In at least one embodiment, each of the strategy layers 108 and 110 provides stable machine movement to one or more neutral goals due to their construction as nonlinear geometries of one or more paths. In at least one embodiment, the strategy layer 110 and the machine movements prompted thereby build upon the strategy layer 108 and the machine movements prompted thereby.The level-based design and implementation of strategy levels 108 and 110 reduces design complexity and allows independent control of execution speed by accelerating machine 106 along a motion direction without compromising the overall quality of motion behavior of machine 106.

[0015] In at least one embodiment, one or more strategies of memory 104 are designed by the user and combined to create an overall strategy or geometric structure that is communicated to machine 106. This strategy communicated to machine 106, when processed by machine 106, causes the machine 106 to move from any initial configuration to a desired target position. In at least one embodiment, the strategy or geometric structure may cause machine 106 to move from any of a plurality of initial or starting states to a target state, such as a neutral state. During movement, the strategy or geometric structure may cause machine 106 to adjust its movement pattern based on one or more environmental conditions, such as one or more obstacles.Furthermore, the strategy or geometric structure may cause the machine 106 to perform one or more movements or motion sequences based on joint limits, stiffness, and / or other parameters of the machine 106. In at least one embodiment, the generation of the strategy layers 108 and / or 110 may be assisted by one or more learning technologies, such as one or more neural networks.

[0016] Fig. 2 illustrates an example of a system environment according to at least one embodiment. In at least one embodiment, the system environment includes a computer system 200. The computer system 200 may include one or more processors 202 and a computer memory 204. The one or more processors 202 may include a graphics processing unit described below. In at least one embodiment, the computer memory 204 may include memory storing executable instructions that, as a result of execution by the one or more processors 202, cause the computer system 200 to control a robot 206. In at least one embodiment, a 7-DoF robotic arm, Franka Emika Panda, is used for a grasping task. In at least one embodiment, the computer system 200 may cause the robot 206 to perform one or more motion sequences or movements.In at least one embodiment, computer system 200 may cause robot 206 to perform one or more motion sequences or movements based on one or more geometric structures stored in memory 204. In at least one embodiment, the one or more geometric structures stored in memory 204 comprise one or more strategy layers 208. Multiple strategy layers may be stored in memory 204. Computer system 200 may also include a UI 212 that can be used by a user to generate or create strategy layer 208, which may cause robot 206 to perform one or more motion sequences or movements.

[0017] In at least one embodiment, robot 206 is an articulated robot, such as the Franka robot mentioned above. In at least one embodiment, robot 206 is controlled at least in part by a control computer system implementing a neural network. The control computer system may be implemented by computer system 200. In at least one embodiment, the control computer system includes memory storing executable instructions that, as a result of execution by the one or more processors, cause the system to grasp an object. In at least one embodiment, the control computer system implements strategies and a neural network trained to generate control signals that cause the articulated robot to grasp an object.

[0018] In the Fig. 2, the robot 206 includes an arm 214 connected to a base. In at least one embodiment, the arm 214 is connected to a wrist 216. In at least one embodiment, a gripper 218 is mounted on the wrist 216. In at least one embodiment, an object 220 is to be grasped by the gripper 218 under the control of the computer system 200 and one / or more strategies of the computer system 200. In at least one embodiment, a camera is mounted on the wrist 216, and the camera is mounted such that the view of the camera is directed along the axis of the gripper 218 toward the object 220 to be grasped. In at least one embodiment, the system 200 and / or the robot 206 includes one or more additional cameras overlooking the workspace, e.g., the task space, in which the robot 206 is to operate.

[0019] In at least one embodiment, computer system 200 executes strategy layer 208 to cause robot 206 to move arm 214. In at least one embodiment, strategy layer 208, when executed by computer system 200, causes robot 206 to move arm 214 in a straight line or a substantially straight line. Strategy layer 208 implements one or more parameters to at least ensure that joint limits, such as a joint limit associated with arm 214 at wrist 216, are met. In at least one embodiment, strategy layer 208 causes robot 206 to move arm 214 to a coordinate point within the workspace of robot 206. Arm 214 comes to a stop or a neutral state at the coordinate point within the workspace of robot 206 according to strategy layer 208.In at least one embodiment, according to strategy level 208, arm 214 performs a sequence of motion to cause gripper 218 to come to a stop near object 220. In at least one embodiment, object 220 may be located in a box, compartment, or other container. Strategy level 208, when executed alone, may cause arm 214 and / or gripper 218 to contact the box, compartment, or other container. Additional one or more strategy levels built upon strategy level 208 are configured to prevent arm 214 and / or gripper 218 from contacting the box, compartment, or other container that may contain object 220.

[0020] Fig. 3 illustrates an example system environment according to at least one embodiment. In Fig. 3, the computer system 200 is illustrated as including a strategy layer 308 in a computer memory 204. The strategy layer 308 may be created or generated by a user to build upon the movement of the robot 206 prompted by the strategy layer 208. In at least one embodiment, the strategy layer 308, when executed by the computer system 200 to cause the robot 206 to perform one or more movements, builds upon the movements of the robot 206 according to the strategy layer 208. In at least one embodiment, the strategy layer 308 has no influence on the distortion prompted by the strategy layer 208.

[0021] In at least one embodiment, as in Fig. 3, the strategy layer 306 causes the wrist 216 of the arm 214 to pivot or rotate toward the object 220. This pivoting or rotation of the wrist 216 may enable the gripper 218 to grasp the object 220.

[0022] In at least one embodiment, the described computer systems may be integrated into the described machines and robots. Alternatively, in at least one embodiment, the computer systems may be separate systems configured to control the described machines and robots.

[0023] Fig. 4 illustrates an example of a system environment according to at least one embodiment. In at least one embodiment, the system environment includes a computer system 400 and a robot 402. In at least one embodiment, the computer system 400 and the robot 402 are integrated with each other. Alternatively, the computer system 400 and the robot 402 may each be a separate computer-based system including one or more processors and one or more computer memories including computer-executable instructions. In at least one embodiment, the computer system 400 may include one or more of the system elements described in detail herein.

[0024] In at least one embodiment, computer system 400 includes a strategy 404. Strategy 404 may also be referred to as a geometric structure according to the disclosed techniques and methods. In at least one embodiment, strategy 404 is associated with one or more computer memories of computer system 400. Strategy 404 may include a plurality of strategy levels 406-410. One or more of strategy levels 406-410 may be generated by the techniques and methods disclosed herein. In at least one embodiment, one or more of strategy levels 406-410 are implemented by a user of computer system 400. Strategy 404 may be used to control the movements performed by robot 402 and / or to cause robot 402 to perform one or more movements.In at least one embodiment, strategy 404 is executed based on a layered approach, with each of strategy layers 406-410 building upon one or more moves initiated by a previously executed layer associated with strategy 404.

[0025] In at least one embodiment, with reference to Fig. 4, the level 406 of the strategy 404 is intended to cause a robot arm 412 to move generally in the direction of movement of the arm indicated by the dashed line in Fig. 4. In at least one embodiment, plane 406 causes arm 412 to move in a manner that avoids joint limiting parameters of arm 412. Furthermore, plane 406 may cause arm 412 to move based at least on trajectory, acceleration, and / or velocity parameters, taking into account redundancy resolution parameters, posture control parameters associated with arm 412, and / or cause an end effector of arm 412 to move to one or more desired positions. In at least one embodiment, plane 406 causes arm 412 to move from or to a box 414, to a box 416, and then to a box 418.As described herein, subsequent movements of arm 412 prompted by other levels of strategy 404 are configured to cause arm 412 and its associated end effector to grasp or otherwise contact one or more objects associated with one or more of boxes 414-418. A goal of strategy 404 is to cause arm 412 to grasp the one or more objects associated with one or more of boxes 414-418 while avoiding contact with the box and any other obstacle that may be near a task space of robot 402. The task space of robot 402 may include a task space surface, such as a table 420, upon which boxes 414-418 rest.

[0026] As in Fig. 5, the computer system 400 may cause the robot arm 412 to perform one or more movements based on the plane 408. In at least one embodiment, the one or more movements caused by the plane 408 build upon and cooperate with the movements caused by the plane 406. In at least one embodiment, the plane 408 includes one or more trajectory, acceleration, and / or velocity parameters that cause the end effector, such as a gripper, of the arm 412 to move in and out of one or more of the boxes 414, 416, and 418. However, in at least one embodiment, the combined movements caused by the combination of the planes 406 and 408 may still result in the arm 412 contacting obstacles, such as one or more of the boxes 414-418.

[0027] As in Fig. 6, the computer system 400 may cause the robot arm 412 to perform one or more movements based on the plane 410. In at least one embodiment, the one or more movements caused by the plane 410 build upon and cooperate with the movements caused by the planes 406 and 408. In at least one embodiment, the plane 410 includes one or more trajectory, acceleration, and / or velocity parameters to cause the end effector, such as a gripper, of the arm 412 to move into and out of one or more of the boxes 414, 416, and 418 without contacting one or more of the surfaces associated with one or more of the boxes 414, 416, and 418.

[0028] The following description provides extensive technical details on generating geometric structures that can be used to provide the types of strategies and strategy levels described above. In general, a strategy level of a structure is a neutral spectral half-spray, i.e., a second-order differential equation with an associated metric tensor that optimizes under forcing. A structure can be derived based on a Lagrangian and geometric formulations. The description also covers the general class of spectral semi-sprays (Specs), Lagrangians and their equations of motion, and general nonlinear geometries, including the special class of Finsler geometries.At a higher level, a nonlinear geometry is a differential equation whose solutions define a set of paths (not velocity-dependent trajectories, but velocity-independent paths). An example is a Riemannian geometry, a type of Finsler geometry whose paths are length-minimizing under the Riemann length measure.

[0029] Many Lagrangians have energies associated with them, defined by their Hamiltonians; the main result of this revelation is that a structure having at least one level of strategy can be defined by starting from a nonlinear geometry and energizing it by finding a representation of the geometry consisting of a set of trajectories along paths that are energy-conserving under a given Lagrangian energy (e.g., an energy transformation of the geometry). The nonlinear geometry defines the fundamental nominal behavior of the system (the paths it follows when unperturbed), and the Lagrangian energy defines the nature of the forces acting to push the system away from these nominal paths.In particular, the energy has an associated energy tensor, which acts as a generalized mass matrix that defines how the system accelerates when forced away from the geometric paths. Since the resulting system is a structure, the potential is ensured to be optimized by applying a potential (and damping). The following text also provides the necessary conditions for Lagrangian equations of motion to define a structure, also referred to herein as a Lagrangian structure. The term system, as used herein, may include at least one computationally implemented system comprising at least one geometric structure and a machine, such as a robot, that is controlled based at least in part on the geometric structure and the one or more strategy levels of that geometric structure.

[0030] Furthermore, the energization transformation commutes with pullbacks across differentiable maps, showing that it is possible to either energize in the codomain of the differentiable map and pullback the resulting system, or pullback the energy and geometry independently and then energize the resulting pulled-back geometry in the domain, and both lead to the same geometric structure.This result, in conjunction with an analysis of the minimum damping conditions required to ensure optimization when following alternative velocity profiles during optimization, shows that nonlinear geometries and associated energies can be placed on a transformation tree, and that the geometries can be used as acceleration strategies to design behaviors and the energies to design spectral priority weights that define what the behaviors pay attention to and how they are combined with each other.

[0031] In the context of Riemannian Motion Policies (RMPs), this result means that geometric structures can be used as formally provably stable tools for the design of flexible RMPs that separate the design of the acceleration strategy (nonlinear geometry) from the priority specification (energy design). Furthermore, it is shown that these geometric formulation methods allow the velocity to be modulated independently of the choice of metric, leading to smooth and consistent, provably stable behavior.Furthermore, this formal separation of behavior into a task-independent geometric structure and a task-specific constraint potential serves as a way to filter out common behavioral elements that span many tasks, so that the geometric structure encodes a well-informed behavioral prior that can be reused for many tasks, thereby improving the generalization of learned potentials in line with recent findings on learning RMPs.

[0032] The concepts and notations surrounding differential geometry are complex. This description uses a notation that avoids the typical coordinate-free or tensor-based notations of differential geometry and instead uses an advanced calculus notation. The manifolds on which the equations are derived are well-defined with respect to the standard constructions of differential geometry, and the equations are simply derived with respect to a coordinate system, as is common in physics.

[0033] This description introduces a variety of terms, so the following text summarizes the definitions in one place for quick reference. In addition, the text provides a list of the different types of structures that appear throughout the description, along with a concise taxonomy of their closure status under the operations of the Spec-algebra defined herein.

[0034] Below is a list of terms defined in this description, each with a brief contextual note: 1. Spectral half-spray, or Spec for short: A pair (M(x, ẋ), f(x, ẋ)) representing a differential equation Mẍ + f = 0. Pullback and combination operations define an associated Spec algebra over transformation trees. 2. Equations of motion: The equation ∂x˙x˙2L+∂x˙xLx˙−∂x⋅L=0, which results from the application of the Euler-Lagrange equation to a stationary Lagrange function L(x,x˙) results. 3. Forcing a system: Adding the gradient of a potential function, often together with a damper, to a Spec. If the original Spec is Mẍ + f = 0, the forced system is Mẍ + f = - ∂ x ψ - Bẋ. 4. A nonlinear geometry: A geometrically consistent system of velocity-independent paths, defined by a differential equation ẍ + h2 (x. ẋ) = 0, where h2 is homogeneous of degree 2 of the velocity, is called a (geometry) generator. Each generator is associated with a geometric shape Px˙⊥[x¨+h2(x,x˙)]=0 which is called the corresponding geometric equation. 5. Finsler structure: A Lagrangian that is positive if ẋ ≠ 0 and homogeneous of degree 1 in velocity. Such Lagrangians define action integrals that are like "path length" in that they are a positive measure that is invariant to temporal reparameterization of the trajectory, so that all trajectories following the same path give them the same action measure. A Finsler structure can also be considered a geometric Lagrangian, as it is a special form of the Lagrangian that incorporates the notion of a positive, velocity-independent path measure with a locally unique minimum. 6. A Finsler geometry is a nonlinear geometry defined by the equations of motion of a Finsler structure. Its generator is called the Finsler generator and is given by the equations of motion of the corresponding Finsler energy. 7. Lagrange energy: The Hamiltonian of a general Lagrange function. This energy is HL=∂x˙LTx˙−L and can generally be different from the Lagrangian. Energy should not be confused with the Lagrangian itself. This is only the case when the Hamiltonian coincides with the Lagrangian (for example, in the Finsler energies below). When the context of the Lagrangian is clear, energy is often used. 8. Finsler Energy: The form of energy Le=12Lg2 a Finsler structure Kind regards In this particular case, the Lagrange energy (Hamiltonian function) that The is assigned, even Le. If the context of the Finsler structure Lg Finsler energy is often referred to as energy or energy form of Lg The Finsler energy is always homogeneous of degree 2 in ẋ and can be used to describe the Finsler geometry as Lg=2Le to be defined if the defining properties for the derived Finsler structure Lg apply. 9. Bending a structure: Adding an energy-conserving geometric term (homogeneous of degree 2) to a geometric structure generator. The resulting generator generates a specific geometry, but the resulting generator still retains the same energy and remains a structure. "Bending" is different from "forcing" (see above). 10. Metric tensor of a Lagrangian: Defined as the Hessian of the Lagrangian L, which defines energy M=∂x˙x˙2L is used. 11. Energization or energization transformation: Given a nonlinear geometry and a Lagrangian (which defines an energy), the energization operation transforms the geometry into a geometric or semi-geometric structure, given by the metric tensor of the Lagrangian and a geometry generator that represents the given nonlinear geometry but preserves the given notion of energy. The resulting structure is either a geometric or semi-geometric structure, depending on whether the energy is a Finsler energy or a more general Lagrangian, respectively. 12. Energized structure: The energy-conserving structure resulting from an energizing transformation. If the energy is a Finsler energy, the structure is a curved Finsler structure, called a curved Finsler representation.

[0035] The following is a list of the classes of structures defined in this description: 1. Optimization structure or short structure: A Spec (M, f) that optimizes when forced, i.e. M e ẍ + f + ∂ x ψ + Bẋ = 0 optimizes psi. 2. Conservative structure: A structure that is characterized by an energy-conserving equation of the form M e ẍ + f e + f f = 0, where (M e ,f e ) from the energy Lagrange function Le(x,x˙) comes from and f f represents a zero labor contribution. 3. Lagrangian structure: A class of conservative structures defined by the equations of motion of an energy Lagrangian function. If the structure is more precisely defined by a Finsler energy, it is called a Finsler structure. 4. Energized structure: A structure formed by energizing a differential equation. If a Finsler energy is used to energize the differential equation, it is called a Finsler-energized structure. 5. Geometric structure: A structure created by energizing a geometry generator with a Finsler energy.

[0036] x and ẋ denote a position and velocity in a task space X, which is assumed to be represented by some selected coordinates. Optimization creates instances of systems of the form Mẍ + f = 0, where M(x, ẋ) is symmetric and invertible and f(x, ẋ) is a function of both position and velocity, with the property that they optimize when forced with a potential, which is discussed in detail in the following description. First, however, a brief overview of the broader class of differential equations is provided.

[0037] A (non-spectral) half-spray itself is a differential equation of the form ẍ + h(x, ẋ) = 0], and the above equation can be written as ẍ + M -1 f = 0 to conform to this form, but it is crucial to represent M explicitly to track how it transforms between spaces, as discussed below.

[0038] For a differentiable map ϕ:Q→X With the action notation x = ϕ(q) an explicit expression for the covariant transformation of a spec (M,f)X on the codomain X into a spec (M˜,f˜)Q on the domain Q. If the Jacobian function of the map is denoted as J=∂ q ϕ and noting ẍ = Jq̈ + J̇q̇, the covariant transformation of the left-hand side of the differential equation Mẍ + f = 0, which is represented by the Spec, is: JT(Mx¨+f)=JT(M(Jq¨+J˙q˙)+f) =(JTMJ)q¨+JT(f=J˙q˙) =M˜q¨+f˜. where M̃ = J T MJ and f̃ = J T (f + J̇q̇). This means that the following covariant pullback operation can be derived: pullϕ(M,f)X=(JTMJ,JT(f+J˙q˙))Q.

[0039] Similarly, since M1ẍ + f1 + (M2ẍ + f2) = (M1 + M2) + (f1 + f2), an associative and commutative summation operation of the form can be derived: (M1,f1)X+(M2,f2)X=(M1+M2,f1+f2)X

[0040] Note that the above operations are the natural form of specs. They can also be viewed in canonical form, which in the current setting, where M is completely invertible, (M,−M−1f)Xc The expression -M -1 f defines the acceleration, since the differential equation represented by the Spec can be solved to give ẍ = -M -1 f. With respect to this canonical acceleration form, the summation operation calculates the combined acceleration as a metric-weighted average of the individual accelerations: (M1,a1)Xc+(M2,a2)Xc=(M1+M2,(M1+M2)−1(M1a1+M2a2))Xc

[0041] These pullback and summation operations define the Spec algebra.

[0042] Differentiable maps can be composed to create a tree of spaces that, in configuration space, and is called a transformation tree, where directed edges denote the differentiable maps and the nodes are given by the spaces resulting for these differentiable maps.

[0043] It can be shown that in the context of a transformation tree, the space of specs becomes a compatible linear structure over the tree under the spec algebra defined above. This means that the tree can be used to represent a composite spec at the root by placing specs at the tree's nodes, pulling them back, and recursively combining them until a single resulting spec is at the root.

[0044] Since the specs form a compatible linear structure on the tree, this computation is independent of the computational path, which means that, at least for the purposes of the theoretical analysis here, the transformation tree can be considered star-shaped, where each node has a single independent map leading directly from the root to the node. Such a star-shaped tree can be represented by a collection of n differentiable maps ϕi:Q→Xi,i=1,…,n The withdrawn and combined spec, which is defined on Q and is a collection of specs {(Mi,fi)}i=1n which are defined on the n spaces is: ∑i=1npullϕ,(Mi,fi)Xi=(∑iJiRMiJi,∑iJiT((fi+Jiq˙)))Q which can be considered as a metric weighted average of the individual withdrawn specs in canonical form.

[0045] A class of specs is said to be closed under tree operations, or in short, closed, if the set of operations is unique if the application of the operations to elements of the class results in an element of the same class.

[0046] In particular, if a given class of specs is closed under tree operations, if the transformation tree is filled with specs from that class, the spec resulting at the root from pullback and combination is of the same class.

[0047] A Spec can be enforced by a position-dependent potential function f (x), using Mx¨+Ef=∂xψ where the gradient - ∂ xψ defines the force added to the system. In most cases, forcing an arbitrary spec does not lead to a system guaranteed to converge to a local minimum of ψ. However, if it does, the spec is optimizing and forms an optimization structure, or structure for short. In this section, the class of specs that form structures is characterized by definitions and results of increasing specificity. These results are used to define geometric structures, which represent a concrete set of tools for designing structures.

[0048] Note that the accelerations of a forced system ẍ = -M -1 f - M -1 ∂ x ψ into nominal accelerations of the system -M -1 f and forced accelerations -M -1 ∂ xψ. The spectrum of M therefore plays a key role in defining how the potential force -∂ x ψ acts to push the system away from the nominal path. It may be easy for potentials to push in some directions but difficult to push in others. The metric M dictates the profile of how the potential function can push away from the system's nominal paths.

[0049] Definition 4.1 X be a manifold with a boundary. Its boundary is denoted by ∂X and its interior is called int (X)=X\∂X designated.

[0050] Note that in all the following cases, if you focus on the edge ∂X of a manifold, assumes that ∂X could be empty unless explicitly stated otherwise.

[0051] Definition 4.2 X denote a smooth manifold of dimension n. The space of all velocities at a point x is denoted by TxX and is known in the theory of manifolds as the tangent space at x. It is often convenient to consider the set of all available positions and velocities on a manifold. This space is known as the tangent bundle and is denoted by TX=∐x∈∂XTx∂X where ∐ is called the disjoint union. That is, (x,x˙)∈TX if and only if x˙∈TxX for some x∈X. The edge of a manifold ∂X is a separate smooth manifold of dimension n - 1 with its own tangent bundle of lower dimension T∂X=∐x∈∂XTx∂X Likewise, the manifold of dimension n consisting of all interior points, where the tangent bundle of consistent dimension n with Tint(X)=∐x∈∂int(X)TxX With these definitions, the complete manifold with a boundary is defined as the disjoint union of the separated interior and boundary manifolds X=int(X)∐∂X understood, and its tangent bundle is the disjoint union of the separated tangent bundles TX=Tint(X)∐T∂X.

[0052] Definition 4.3 Let X be a manifold with boundary ∂X (possibly empty). A spec (M,f)X is considered internal if for each internal starting point (x0,x˙0)∈Tint(X) the integral curve x(t) is always inside x(t)∈int(X) for all t ≥ 0.

[0053] Definition 4.4 X be a manifold with a (possibly empty) boundary. A spec S=(M,f)x is considered rough if all its integral curves x(t) converge: lim t → ∞ x(t) = x ∞ with x∞∈X (including the possibility x∞∈∂X If S is not rough, but each of its muted variants SB=(M,f+Bx˙) , where B(x, ẋ) is smooth and positive definite, the spec is said to be frictionless. The damped variants of a frictionless spec are also called rough variants of the spec.

[0054] Definition 4.5 Let ψ(x) be a smooth potential function with gradient ∂ x ψ and (M,f)X be a Spec. Then (M, f) + a x ψ is the enforced variant of the spec and enforces the spec with the potential ψ. Furthermore, ψ is finite if ||∂ x ψ || < ∞ everywhere on X lies.

[0055] Definition 4.6 A spec S forms a rough structure if it is rough, if it is forced by a finite potential ψ(x) and if every convergent point x ∞ a Karush-Kuhn-Tucker (KKT) solution of the constrained optimization problem min x∈X ψ(x). A forced spec is a frictionless structure, its rough variants form rough structures.

[0056] Lemma 4.7. If a spec forms a rough (or frictionless) structure, then it is (unforcedly) a rough (or frictionless) spec.

[0057] Definition 4.8 A spec S=(M,f)X is edge-compliant if the following conditions apply: 1. is inside. 2. M(x, ẋ) and f(x, ẋ) are finite for all (x,x˙)∈TX (more precisely for all (x,x˙)∈Tint(X) and (x,x˙)∈T∂X). 3. The inverse metric has a finite limit M−1→M∞−1 with ‖M∞−1‖<∞ along each boundary trajectory x→x∞∈∂X.

[0058] A boundary-conforming metric is a metric that satisfies conditions (2) and (3) of this definition. Furthermore, f is said to be boundary-conforming with respect to M if M is a boundary-conforming metric and (M,f)X forms an edge-conforming spec.

[0059] Note that this definition of boundary conformity implies that M either approaches a finite matrix along trajectories that bound the boundary, or that it approaches a matrix that is finite along eigendirections parallel to the tangent space of the boundary, but explodes to infinity along the direction orthogonal to the tangent space. This means that M∞−1 is either of full rank or of reduced rank and its column space spans the tangent space of the boundary Tx∞∂X.

[0060] Definition 4.9 A boundary-conforming spec (M,f)X is neutral if for every convergent trajectory x(t) with x → x ∞ , the V∞Tf(x,x˙)→0 provides, where V ∞ is a matrix whose columns are a basis V∞ for Tx∞X If the spec is not neutral, it is biased. The term f alone is either biased or neutral if the context of M is clear. Note that all neutral specs must also be boundary-conforming by definition. Therefore, the spec is neutral if one assumes that it is also boundary-conforming by definition.

[0061] Remark 4.10 If in the above definition x∞∈int(X) holds, the base V contains ∞ a complete set of n linearly independent vectors such that the condition V∞Tf(x,x˙)→0f(x,x˙)→0 This is implied in not the case.

[0062] Since the definition of neutral requires that the Spec is boundary-conforming, the spectrum of the metric M is always finite in the relevant directions (all directions for interior points and directions parallel to the boundary for boundary points). The property of being neutral is therefore linked to zero acceleration in the relevant subspaces. This property is used in the following theorem to characterize general structures.

[0063] Theorem 4.11 (General Structures). Suppose S=(M,f)X is a boundary-conformal Spec. Then S forms a rough structure if and only if it is neutral and converges when it is defined by a potential ψ(x) with ‖∂ψ‖<∞ on x∈X∂ is enforced.

[0064] Proof. The enforced Spec defines the equation Mx¨+f=−∂xψ, where the damping term in f is absorbed if it is the rough variant of a frictionless spec (this has no influence on the hypotheses about f).

[0065] First, assume that f is neutral. Since x converges, ẋ → 0 also implies ẍ → 0.

[0066] If x(t) approaches an interior point converges, M is finite, so Mẍ → 0, since ẍ → 0. And since the Spec is neutral, f(x, ẋ) → 0 as ẋ → 0. Therefore, the left-hand side of equation 9 approaches the value 0, so that ∂ x ψ → 0 satisfies the (unconditional) KKT conditions.

[0067] Alternatively, if x(t) approaches a boundary point converges, then analyze the expression: x¨=−M−1(f+∂xψ)→0. since ẍ → 0. Since M is boundary conformal, the inverse metric limit is M∞−1 finite, and since ∂ψ also on is finite, the term M converges -1 ∂ψ to the finite vector M∞−1∂ψ∞ Therefore, according to equation 10M−1f→M∞−1f∞=M∞−1∂ψ∞. At the limit M∞−1 the full rank over therefore the above boundary equality implies f∞ / / =−∂ψ∞ / / , where f∞ / / and ∂ψ∞ / / the components of f ∞ or ∂ψ ∞ which lie in the tangent space of the boundary Tx∞∂X. Since f is neutral and the edge-parallel component f∞ / / =0 must also ∂ψ∞ / / =0 Therefore, ∂ψ ∞ either orthogonal to Tx∞∂X or zero. If it is zero, the KKT conditions are automatically satisfied. If it is nonzero, since x(t) lies inside, - f(x. ẋ) must lie inside near the boundary, so = -∂ψ ∞ be directed outwards and -f ∞This alignment, in addition to orthogonality, implies that the boundary point satisfies the (conditional) KKT conditions.

[0068] Finally, to prove the opposite, we assume that f is distorted. Then there exists a point for which f(x*, 0) ≠ 0. An objective potential with a unique global minimum at x* can be constructed. The enforced system cannot stop at x* in this case, since f is nonzero there, so no optimization is guaranteed for (M, f) and it is not a structure.

[0069] The above theorem characterizes the most general class of structures and shows that all structures are necessarily neutral in the sense of Definition 4.9. The theorem relies on the hypothesis that the system always converges if forced, and is therefore more of a template for proving that a given system forms a structure than a direct characterization. Proving convergence is generally not trivial. For the structures presented below, convergence is proven using energy conservation and boundedness properties.

[0070] Note that the above theorem does not impose any restrictions on whether the metric or damping in eigendirections orthogonal to the tangent space of the surface is finite or not. In practice, it may be convenient to let these metrics increase to infinity in these directions, so that the effects of forces orthogonal to the surface of the boundary are increasingly damped by the large mass. Such metrics can lead to smoother optimization behavior when optimizing for local minima of the interface.

[0071] Definition 4.12 A stationary Lagrange function L(x,x˙) is edge-compliant (on ) if their induced equations of motion ∂x˙x˙2Lx¨+∂x˙xLx˙−∂xL=0 under the Euler-Lagrange equation a boundary-conforming spec SL=(ML,fL)X form, whereby ML=∂x˙x˙2L and fL=∂x˙xLx˙−∂xL This spec is considered the L associated Lagrange function-Spec. L is additionally neutral if SL is neutral. Since neutral specs are boundary-conforming by definition, a neutral Lagrangian L implicitly also boundary-compliant.

[0072] Definition 4.13 Le(x,x˙) be a stationary Lagrange function with Hamilton function He(x,x˙)=∂x˙LeTx−⋅Le.Le is an energy Lagrange function if Hey is nontrivial (not zero everywhere) and He(x,x˙) on is finite. The equations of motion of an energy Lagrangian function M e ẍ + f e = 0 are often referred to as their energy equations, where Spec is Se=(Me,fe)X is called.

[0073] Definition 4.14 (Energy boundary condition). The be an energy Lagrangian function with energy Hey. If for every trajectory x(t) for which there exists a t0 < ∞ such that x(t0)∈∂X and x∉T˙x(t0)∂X (a trajectory intersecting the edge) limt→t0He(x,x˙)=∞ ∞ is provided and Hey the boundary condition is met.

[0074] Remark 4.15. When designing boundary-conforming energy Lagrangians, it must be ensured that the spec of the Lagrangian lies inside. To achieve an inside spec, it is often helpful to design energies that prevent energy-conserving trajectories from It can be shown that the Spec of an energy Lagrangian is interior if and only if its energy is boundary, i.e., the energy approaches infinity for any trajectory that intersects the boundary.

[0075] Lemma 4.16. The be an energy Lagrangian function with energy He=∂x˙LeTx˙−Le. The energy-time derivative is: H˙e=x˙T(Mex¨+fe), where M e and f e from the Lagrange equations of motion M e ẍ + f e = 0.

[0076] Proof. The calculation is a simple time derivative of the Hamiltonian function: HLe=ddt[∂x˙Le−Le](12)=(∂x˙x˙Lex¨+∂x˙x˙Lex˙)Tx˙+∂x˙LeTx¨−(∂x˙Lex¨+ ∂xLeTx˙)(13)=x˙T(∂x˙x˙Lex¨+∂x˙xLex˙−∂˙x˙LeTx˙)(14)x˙T(Mex¨+fe)(15)

[0077] Lemma 4.17. The be an energy Lagrangian function. Then M e ẍ + f e + f f = 0 energy conserving and only if ẋ T f f = 0. Such a spec S=(Me,fe+ff) is considered as a conservative spec under energy Lagrangian The designated.

[0078] Proof. This energy is conserved if its time derivative is zero. Inserting x¨=−Me−1(fe+fr) ẍ into equation 11 of Lemma 4.16 and setting it to zero gives: HLe=x˙T(Me(−Me−1(fe+ff))+fe)(16)=x˙T(−fe−ff+fe)(17)=x˙Tff=0.(18)

[0079] The energy is therefore only conserved if the last condition is fulfilled.

[0080] Proposition 4.18 (Conservative Structures). Suppose S=(Me,fe,ff)X is a conservative neutral spec under energy Lagrangian The with a work term of zero f f . The S forms a frictionless structure.

[0081] Proof. Let ψ(x) be a lower bounded potential function with finite ∂ψ on X ∂ and an expression of how the total amount Heψ=He+ψ(x) changes over time: Heψ=H˙e+ψ˙(19)=x˙T(Mex¨+fe)+∂xψTx˙(20)=x˙T(Mex¨+fe+∂xψ).(21)

[0082] With damping matrix B(x, ẋ) let M e ẍ + f e + f f + ∂ x ψ + Bẋ =0 is the forced and damped variant of the conservative Spec system. Inserting this system into equation 21, one obtains: H˙eψ=x˙T(Me(−Me−1(fe+ff+∂xψ+Bx˙))+fe+∂xψ)(22)=x˙T(−fe−∂xψ−Bx˙+fe+∂xψ)−x˙Tff(23)=x˙TBx˙(24) since all terms except the damping term cancel. If B is strictly positive definite, the rate of change for x ≠ 0 is strictly negative. Since Heψ=He+ψ(x) is limited downwards and H˙eψ≤0 and H˙eψ→0, which implies x → 0, the convergence of the system is guaranteed.

[0083] Since it is also boundary-conformal and neutral, it forms a structure according to Theorem 4.11.

[0084] Finally, if B = 0, Equation 22 shows that the total energy is conserved, a system starting with nonzero energy cannot converge. Therefore, the undamped system is a frictionless structure with rough variants defined by the added damping term.

[0085] Corollary 4.19 (Lagrangian and Finsler structures). If Le(x,x˙) is a neutral Lagrange function, then Se=(Me,fe)X a frictionless structure called a Lagrangian structure. If The is a Finsler energy, this structure is more accurately called a Finsler structure.

[0086] Proof. M e ẍ + f e = 0 is conservative by Lemma 4.18 with f f = 0. Since it is also neutral, it forms a frictionless structure according to Proposition 4.18.

[0087] The following lemma collects some results around common matrices and operators that arise in the analysis of energy conservation.

[0088] Lemma 4.20. The be an energy Lagrangian function. Then with p e = M e ẋ Rpe=Me−1−x˙x˙Tx˙TMex˙ has a null space bounded by p e is spanned, and Rx˙=Me−pepeTpeTMe−1pe has a null space spanned by x. These matrices are defined by R ẋ = M e R pe M e linked and the matrix Me−1Rx˙=MeRpe=Pe M e is a projection operator of the form Pe=Me12[I−v^v^T]Me−12 where v=Me12x˙ and v^=v‖v‖ is the normalized vector. In addition, ẋ T P e f = 0 for all f(x, ẋ)

[0089] Proof: The right multiplication of R pe with p eresults in: Rpepe=(Me−1−x˙x˙Tx˙TMex˙)Mex˙(28)=x˙−x˙(x˙TMex˙x˙TMex˙)=0(29) and the null space is not larger than this value, since every matrix is formed by subtracting a rank-1 term from a full-rank matrix.

[0090] The relationship between R pe and R x can be represented algebraically: MeRpeMe=Me(Me−1−x˙x˙Tx˙TMex˙)Me=Me−Mex˙x˙TMex˙TMex˙(30)=Me−pepeTpeTMe−1pe=Rx˙(31)

[0091] Therefore, R ẋ the same rank as R pe and its null space must be spanned by ẋ, since Rx˙x˙=MeRpeMex˙˙=MeRpe=0.

[0092] With a little algebraic manipulation, the following is provided: MeRpe=Me(Me−1−x˙x˙Tx˙TMex˙˙)(32)=Me12(I−Me12x˙x˙TMe12x˙TMe12Me12x˙)Me−12(33)=Me12(I−vvTvTv)Me−12(34)=Pe(35) there vvTvTv=V^V^T. Over and beyond, PePe=Me12(I−vvTvTv)Me−12Me12(I−vvTvTv)Me−12(36)=Me12(I−vvTvTv)(I−vvTvTv)Me−12(37)=Me12(I−v^v^T)Me−12=Pe,(38) da P ⊥ = I - v̂v̂ T is an orthogonal projection operator. Therefore, Pe2=Pe, that it is also a projection operator.

[0093] Finally, the following: x˙TPe=x˙TMeRpe=x˙TMe(Me−1−x˙x˙Tx˙TMex˙)(39)=[x˙T−(x˙TMex˙x˙TMex˙)x˙T]=0(40)

[0094] Therefore, for every f, M e ẍ + f e + f f = 0.

[0095] Corollary 4.21. The be an energy Lagrangian function. Then for each f̃ f (x, ẋ) under the constraint term f f = P e f f the equations of motion: Mex¨+fe+ff=0 are energy conserving.

[0096] Proof. By Lemma 4.20, ẋ T f f = ẋ T P ef = 0, so according to Lemma 4.17, equation 41 is energy-conserving.

[0097] Proposition 4.22 (System energization). Let ẍ + h(x, ẍ) = 0 be a differential equation and assume that The any energy Lagrange function with equations of motion Mex¨+fe=0 and energy Hey is. Then x¨+h(x,x˙)+αHex˙=0 where αHe=−(x˙TMex¨˙)−1x˙T[Meh−fe] is energy-conserving and differs from the original system only by an acceleration along the direction of motion. The new system can be expressed as follows: Mex˙+fe+Pe[Meh−fe]=0

[0098] This modification of the system is called energy transformation and using the Spec notation is denoted as ShLe=(Me,fe+Pe[Meh−fe])x=energizeLe{Se} where Sh=(I,h)X is the Spec that represents ẍ + h(x, ẋ) = 0.

[0099] Proof. Equation 11 of Lemma 4.16 gives the time derivative of energy. Substituting a system of the form x¨+h(x,x˙)=αHex˙, ẍ setting to zero and solving for αHe results H˙e=x˙T[Me(−h−αHex˙)+f]=0(45)⇒−xTMeh−αHex˙TMex˙+x˙Tfe=0(46)⇒αHe=−x˙TMeh−x˙Tfex˙TMex˙(47)=(−x˙TMex˙)−1x˙T[Meh−fe].(48)

[0100] If you replace this solution for αHe back, you get: x¨=−h+(x˙Tx˙TMex˙[Meh−fe])x˙(49)=−h[x˙x˙Tx˙TMex˙](Meh−fe).(50)

[0101] Algebraically it helps, h=Me−1[ff+fe] to initially calculate the result as the difference of f e away. If you do this and move all terms to the left side of the equation, you get x¨+h−[x˙x˙Tx˙TMex˙](Meh−fe)=0(51)⇒x¨+Me−1[ff+fe]−[x˙x˙Tx˙TMex˙](MeMe−1[ff+fe]−fe)=0(52)⇒Mex¨+ff+f e−Me[x˙x˙Tx˙TMex˙](ff+(fe−fe))(53)⇒Mex¨+fe+Me[Me−1−x˙x˙Tx˙TMex˙]ff(54)⇒Mex¨+fe+MeRpe(Meh−fe),(55) where f f = M e f f - f e is substituted back into back into. By Lemma 4.20, M e R pe = P e to obtain equation 43.

[0102] Definition 4.23 A metric M(x, ẋ) is called boundary-aligned if for every convergent x(t) with the limit limt→∞M−1(x,x˙)=M∞−1 exists and is finite, Tx∞∂X of a subset of the eigenbasis of M∞−1 is spanned.

[0103] Lemma 4.24. Suppose M e be boundary-conforming. If then (M e, f) and (M e, g) both are neutral, then (M e , αf, βg) \neutral.

[0104] Proof. Suppose x(t) is a convergent trajectory with x → x ∞ . If x ∞ ∈ then f → 0 and g → 0, so αf + βg → 0. If x ∞ ∈ ∂X, then f = f1 + f2 with f1 → 0 and f2→f⊥⊥Tx∞∂X, and similarly with g = g1 + g2 with g1 = 0 and g2→g⊥⊥Tx∞∂X. Then αf + βg = (αf1 + βf2) + (αg1 + βg2) → αg1 + βg2 is orthogonal to Tx∞∂X, since each component in the linear combination is orthogonal. Therefore, αf + βg is neutral.

[0105] Lemma 4.25. For every matrix A there exists a constant c > 0 such that ||Au|| ≤ c||u|| for every vector u, where ||·|| can be an arbitrary norm.

[0106] For the sake of concreteness, it is assumed in the following that the matrix norm is the Frobenius norm.

[0107] Lemma 4.26. M e be boundary-conforming. Then, if f is neutral with respect to M e , P e f is neutral, where Pe=Me12[I−v^v^T]Me−12 p e is the projection operator defined in equation 27.

[0108] Proof: Let X(t) converge with x → x ∞ as t → ∞. If then exists, since P e finally on is, according to Lemma 4.25, a constant c > 0, which ||P e f|| ≤ c||f||. Since ||f|| → 0| as t → 0, it must be true that ||P e f|| → 0.

[0109] Let us now consider the case in which Da P e is a projection operator, for each z a decomposition into linearly independent components z = z1 + z2 with z2 in the kernel must be performed, so that P e z = P e z1 + P e z2 = z1 (so that Pe2z=Pez1z1+Pez ). Therefore, if z has a property if and only if all elements of a decomposition have this property, then P ez = z1 also have this property. In particular, if z lies in a subspace (or is orthogonal to a subspace), then P e z = z1 also in this subspace (or is orthogonal to a subspace).

[0110] The property of being neutral is determined by the behavior of the decomposition f=f / / + f ⊥ where these components are parallel or perpendicular to the limit value Tx∞∂X According to the index convention described above to characterize the behavior of the projection, the following is provided: Pef=Pe(f / / +f⊥)=Pef / / +Pef⊥(56)=f1 / / +f1⊥→f1⊥(57) since f is neutral, what f / / → 0 and therefore f1 / / →0 implied.

[0111] Therefore, P satisfies e f in both cases satisfies the conditions to be neutral if f does so.

[0112] Lemma 4.27. Suppose M e(x,ẋ), is boundary-conforming and boundary-aligned. If (I, h) is neutral, then (M e , M e h) neutral.

[0113] Proof: Let x(t) converge with x → x ∞ as t → ∞. If then exists, since M e finally on Tint(X) is, according to Lemma 4.25, a constant c > 0 such that ||M e h|| ≤ c||h||. Since ||h|| → 0 as t → 0, it must be that ||M e h|| → 0.

[0114] Da M e is boundary-aligned, there is a subset of eigenvectors that define the tangent space in the limit as span. V / / be a matrix containing the eigenvectors restricted to the spanning of the tangent space, and D / / a diagonal matrix containing the corresponding eigenvalues. Likewise, let V ⊥ the remaining eigenvectors (vertical in the limit) with the eigenvalues D ⊥ The metric decomposes as Me=Me / / +Me⊥ with Me / / =V / / D / / V / / T. M e Since h is neutral, express h = h / / + h / / = V / / y / / + V ⊥y⊥ where y / / and y ⊥ Coefficients are and y / / → 0 as t → ∞. Therefore Meh=(Me / / +Me⊥)(h / / +h⊥)(58)=(V / / D / / V / / T+V⊥D⊥V⊥T)(V / / y / / +V⊥y⊥)(59)=V / / D / / y / / +V⊥D⊥y⊥(60)→V⊥z⊥Tx∞∂X where z=D⊥y⊥(61)

[0115] Therefore, if x ∞ ∈ ∂X, the component parallel to the tangent space in the limit.

[0116] Lemma 4.28. Suppose Sh=(I,h) is a neutral (acceleration) spec and The is a neutral energy Lagrangian with boundary-aligned metric M e . Then ShLe=energizeLe{Sh} neutral.

[0117] Proof. M e is boundary-conform according to the hypothesis to The Furthermore, the energized equation has the form Mex¨+fe+Pe[Meh−fe]=0 as shown in Proposition 4.22. By Lemma 4.27, M e h neutral, since h is neutral, and likewise M e h - f e neutral by Lemma 4.24, since f e based on the hypothesis about The is neutral. This means that P e [M e h - f e ] is also neutral by Lemma 4.26. Finally, by applying Lemma 4.24 again, it is shown that the entire set of f e + P e [M e h - f e ] is neutral.

[0118] Theorem 4.29 (Energized structures). The be a neutral energy Lagrangian with boundary-aligned Me=∂x˙x˙2Le M e and lower-limited energy Hey, and (I, h) be a neutral Spec. Then the energized Spec ShLe=energizeLe{Sh} according to Proposition 4.22 a frictionless structure.

[0119] In the following, it is shown that Finsler geometries form conservative structures if the underlying system is energy-conserving, and in particular, that a broad class of neutral systems can be energized to transform them into structures. In general, however, the energy transformation affects the behavior of the system, since systems are generally not invariant to the traversal velocity. In particular, the system follows different paths if it is forced to decelerate or accelerate. An intuitive example of such a system is a particle in a gravitational field. If it moves around a mass at the right speed, the particle can describe an orbit. However, if it is accelerated or decelerated along the direction of motion, it will either break out of the orbit or spiral into the mass; in both cases, the path necessarily changes.

[0120] This section introduces a special class of differential equations, called geometry generators, which maintain the consistency of geometric paths despite the traversal speed. For this special class of equations, the path behavior of a system is invariant under energization, so that an energized structure formed by energizing a geometry generator follows the same path as the original system and is itself a geometry generator. These energized structures are called geometric structures.

[0121] Definition 5.1 (Geometry generator). A geometry generator, or generator for short, is a second-order differential equation of the form: x¨+h(x,x˙)=0 where h2 (x, ẋ) is a smooth, covariant map h2 : ℝ d × ℝ d → No. dwhich is positively homogeneous of degree 2 in velocities in the sense of h2 (x, αẋ) = α 2 h2(x, ẋ) for α > 0. Solutions to the generator are called generating solutions or trajectories, and if these trajectories are guaranteed to conserve a known amount of energy, they are called energy levels.

[0122] As a second-order differential equation, the solutions of a generator are unique for certain initial values (x0, x0). The homogeneity of a generator of degree 2 means that all solutions of initial value problems of the form (x0, αv̂0), where v̂0 is an arbitrary unit vector defining a direction in space, follow the same path. In other words, all generating trajectories that start from a given point x0 and whose initial velocity points in the same direction x̂0 = αv̂0, follow the same path.

[0123] Generators are often referred to as sprays in differential geometry, although the term generator more clearly describes its role in generating the geometry of a geometric equation, as defined below.

[0124] Definition 5.2 (Geometric equation). The geometric equation corresponding to a generator of the form ẍ + h2(x, ẋ) = 0 is an equation of the form Px˙⊥[x¨+h2(x,x˙)]=0 where Px˙⊥ is a projector that projects orthogonally to x. Solutions to this geometric equation are called geometric solutions or trajectories.

[0125] Every matrix R x with a null space spanned by x would suffice in this definition. The projector is chosen for clarity of its role.

[0126] While the generator's solutions are unique but follow the same path if the velocities of the initial conditions point in the same direction, the geometric equation is a redundant equation (the redundancy arises from the reduced rank matrix Px˙⊥), whose solutions are the set of all trajectories following this single path. It can be shown that every solution to the generating equation is a smooth time reparameterization of any generated trajectory.

[0127] Furthermore, for a given geometric trajectory, there exists a generating trajectory whose instantaneous velocity at a particular point x coincides, so the geometric trajectory can be viewed as accelerating and decelerating along the direction of motion to move smoothly between generating solutions, hence the name generating solution. As these generating solutions form energy levels, the geometric solution accelerates and decelerates along the direction of motion to move smoothly between the energy levels of the system.

[0128] The geometric equation completely characterizes a geometry of paths. The equivalence class of solutions to the problems x0, αv̂0, and α > 0 is (locally) the set of reparameterizations of a one-dimensional smooth subbifurcation of space. The collection of these velocity-independent paths defines a nonlinear geometry of space.

[0129] All solutions of the geometric null space equation given in equation 64 can be expressed as follows: x¨=h2(x,x˙)+γ(t)x˙ where γ(t) is any smooth function of time. This is an explicit expression showing that geometric solutions are formed by acceleration and deceleration along the direction of motion ẋ.

[0130] Finsler geometry is the study of nonlinear geometries whose geometric equation is defined by the equations of motion of the Finsler structure.

[0131] Definition 5.3 (Finsler structure). A Finsler structure is a stationary Lagrangian function Lg(x,x˙) with the following properties: 1. Positivity: Lg(x,x˙)>0 for all ẋ≠ 0. 2. Homogeneity: Lg is positively homogeneous of degree 1 in velocities in the sense of Lg(x˙,αx˙)=αLg(x,x˙) for α > 0. 3. Energy tensor invertibility: ∂x˙x˙2Le is invertible everywhere, where Le=12Lg2. The is used as an energy form of Lg and is sometimes also called Finsler energy.

[0132] Note that the first two conditions together mean that L(x,0)=0.

[0133] Finsler structures could be more descriptively called geometric Lagrangians, since their geometric properties (invariance to time reparameterization) stem directly from the conditions of the Lagrangian and its effects on the resulting action.

[0134] Proposition 5.4 (Energy of a Finsler geometry). Le=12Lg2 be the Finsler energy of the Finsler structure Kind regards The Hamiltonian function (conserved set) of Le is Le. In particular He=∂x˙LeTx˙−Le=Le.

[0135] Proof. According to Euler's theorem on homogeneous functions, if f(y) is homogeneous of degree k, then ∂ y f T y = kf (y). This means for Lagrange functions thus ∂x˙LTx˙=kL, what means H=∂x˙LTx˙−L=kL−L=(k−1)L

[0136] There Lg homogeneous of degree 1 in x, is The homogeneous of degree 2 in x, so for Lg the above analysis He=(2−1)Le=Le.

[0137] Equivalent forms of higher order energy can also be Lek=12Lgk be defined, and all the following results also hold, but for the sake of clarity only k = 2 is considered here.

[0138] Lemma 5.5 (Homogeneity of the Finsler energy tensor). Lg Let be a Finsler structure with the energy form Le=12Lg2. Then the energy tensor is Me=∂x˙x˙2Le homogeneous of degree 0.

[0139] The above lemma means that M e only through its norm x˙^ depends on the speed ẋ, i.e. rescaling ẋ has no influence on the energy tensor.

[0140] Theorem 5.6 (Generation of Finsler geometry). Lg be a Finsler structure with energy Le=12Lg2. The equations of motion of The define a geometry generator whose geometric equation is determined by the equations of motion of Lg is given.

[0141] These results mean that Finsler structures Lg define nonlinear geometries of paths whose generators are energy levels of the energy The define.

[0142] The expression in equation 43 shows that energization can be considered as a modification of the energy-motion equations with zero work. If the original differential equation ẍ + h = 0 is a geometry generator (i.e., h is homogeneous of degree 2) and The Finsler, then the Finsler equations of motion, the original differential equation, and the energized equation are all geometry generators, and most importantly, the energized equation in 43 generates a geometry that corresponds to the generated geometry of the original equation. Therefore, this zero-work modification can be considered as a bending of the Finsler geometry that corresponds to the desired geometry without affecting the system energy. This result is summarized in the following proposition.

[0143] Corollary 5.7 (Curved Finsler representation). Suppose that h2(x,ẋ) is homogeneous of degree 2 such that ẍ + h2 (x, ẋ) = 0 is a geometry generator, and The be a Finsler structure (and therefore also a Finsler energy). Then the energized system M e ẍ + f e + P e [M e h2 - f e ] = 0 is a geometry generator whose geometry coincides with the geometry of the original system. Since the Finsler system M e ẍ + f e = 0 is also a geometry generator, the energized system is to be considered as a geometric modification of the Finsler geometry with zero work, which is referred to herein as a bending of the geometric system.

[0144] Proof. The energized system takes the form ẍ + ĥ2(x, ẋ) = 0, where h˜2=Me−1fe+Re[Meh2−fe] da P e = M e R pe . The energy The is Finsler, so f e homogeneous of grade 2 and M e is homogeneous of degree 0, which means that the first term in the combination is homogeneous of degree 2. Furthermore, Rpe=Me−1−x˙x˙Tx˙TMex˙ R pe homogeneous of degree 0, since the numerator and denominator scalars in the second term would cancel when ẋ is scaled. Thus, the energized system as a whole forms a geometry generator.

[0145] Since ẍ + h2 (x, ẋ) = 0 is a geometry generator, instantaneous accelerations along the direction of motion ẋ do not change the paths taken by the system. ẍ + ĥ2 (x, ẋ) = 0 thus forms a generator whose geometry coincides with the original geometry defined by ẍ + h2 (x, ẋ) = 0.

[0146] Corollary 5.7 shows that geometries are invariant under energization transformations performed with respect to Finsler energies.

[0147] Note that for the generated system to be a generator, the original system must be a generator, and the energy must also be Finsler. If the energy is not Finsler, the resulting energized system will still follow the same paths as the original geometry (since it is, by definition, formed by acceleration along the direction of motion), but it will not be a generator itself (i.e., the paths will not be invariant under reparameterization).

[0148] Corollary 5.8 (Geometric Structures). Suppose h2(x, ẋ) is homogeneous of degree 2 such that ẍ + h2(x, ẋ) = 0 is a geometry generator, and suppose The is Finsler. Then the energized system is a structure defined by a generator whose geometry matches the geometry of the original generator. Such a structure is called a geometric structure.

[0149] Proof. By Corollary 5.7, the energized system is a generator with matching geometry, and by Theorem 4.29, this energized system forms a structure.

[0150] Each class of structures described above is closed under the Spec-algebra. In particular, Specs that form a structure of a given class remain structures of the same class under the Spec-algebra operations of combination and pullback. This result is summarized in the following theorem.

[0151] Theorem 6.1. The following classes of structures are closed under Spec-algebra: optimization structures, conservative structures, Finsler-energized structures, and geometric structures. A structure of any of these types remains a structure of the same type under Spec-algebra operations in regions where the differentiable transformations remain finite.

[0152] In Theorem 6.1 above, it is shown that structures behave naturally under spec-algebra operations in the sense that each of the classes described above is closed under these operations. Here, it is additionally shown that the energization operation described in Theorem 4.22 commutes with the pullback operator as long as the pullback is performed with respect to the metric defined by the energy used for energization.

[0153] As is known from Lagrangian mechanics that an energy Lagrangian function Le(x,x˙) is defined, the Euler-Lagrange equation M e ẍ + f e = 0 in χ and after withdrawn to (J T M e J)ẍ + J T (f e -J̇ q̇ ) = 0, or the Lagrange function according to withdrawn to first Le(q,q˙)=Le(ϕ(q),Jq˙) and the direction of the Euler-Lagrange equation is applied to this pullback Lagrange function to obtain M̃ e q + f e = 0, and the resulting equations of motion are M̃ e (J T M e J) and f e = J T (f e - J̇q̇). This standard result shows that the operation of deriving the Euler-Lagrange equation from a Lagrangian commutes with the pullback transformation. Thus, the system can either apply the Euler-Lagrange equation in the codomain (ambient space) and pull back the resulting equations of motion, or pull the Lagrangian back into the domain and apply the Euler-Lagrange equation directly there. The resulting equations will agree.

[0154] The result presented in Theorem 7.1 shows that the same kind of commutativity also holds for the operation of energization. This result is specific to the case where the differentiable map defines an embedding, and the intuition also follows from understanding this case. For example, if the differentiable map defines a full-valued mapping from a d-dimensional space with generalized coordinates to a higher-dimensional neighborhood space χ of dimension n > d. The embedded manifold can be considered as a condition, and the coordinates define generalized coordinates for this condition. The theorem states that if an energy Lagrangian is defined on the ambient space together with a differential equation ẍ + h(x,ẋ) = 0, the system can either energize the ambient space and pullback the resulting equations, or first pullback the differential equation with respect to the metric of the energy Lagrangian and energize the equation there. The theorem shows that when the operation of energization is used to define a structure, the most important element is the energy metric. The energy defines how the pullback of the differential equation must be performed to remain consistent with the operation of energization.

[0155] In practice, the system does not need to be explicitly energized. Instead, the system can simply reduce the differential equation with respect to the energy metric to its square root and track the energy itself to define a lower bound on the damping required to maintain stability and instead control an alternative execution energy. The main role of the energy Lagrangian in formulating the equations of motion is to define the metric. Theorem 7.2 shows that this basic theorem allows us to view energization as metric-weighted averages of acceleration strategies in different task spaces.

[0156] Theorem 7.1. The be an energy Lagrange function and ẍ+h(x,ẋ) = 0 be a second-order differential equation with associated Spec (M e , f) natural form under the metric Me=∂x˙x˙2Le, where f = M e h. Suppose x = ϕ(q) is a differentiable map for which the pullback metric J T M e J has the full rank. Then energizepullLe(pull∅(Me,f2))=pullϕ(energizeLe(Me,f2)).

[0157] The energization operation commutes with the pullback transformation.

[0158] Proof. The computational equivalence is shown. The energization of ẍ + h = 0 in the form of force M e ẍ + f = 0 with f=Meh is Mex¨+feh, where: feh+fe+Me[Me−1−x˙x˙Tx˙TMex˙] where fe=∂x˙xLex˙−∂xLe, f e so that Mex¨+feh is the energy equation. Let J = ∂ x ϕ. The pullback of the energized geometry generator is: JTMe(Jq¨+J˙q˙)+JTfeh=0(72)⇒(JTMeJ)q¨+J˙T(feh+MeJ˙q˙)=0(73)⇒(JTMeJ)q¨+JTfe+JTMe[Me −1−x˙x˙Tx˙TMex˙](f−fe)+JTMeJ˙q˙=0(74)⇒M˜eq¨+f˜e+JTMe[Me−1−x˙x˙Tx˙TMex˙](f−fe),(75)

[0159] Where M̃ e + J T M e J and f e =J T (f e + M e Jq̇) the standard pullback of (M e , f e ) form.

[0160] Now calculate the geometry pullback with respect to the energy metric M e by taking the metric weighted force form of geometry M e ẍ + f = 0, where again f=M e h applies. The pullback is: JTMe(Jq¨+J˙q˙)+JTf=0(76)⇒(JTMeJ)q¨+JT(f+M2J˙q˙)(77)⇔M˜eq¨+f¯=0(78) where M e =J T M e J as before and f̃ = J T (f + M e Jq). Le=Le(∅(q),Jq˙) be the pullback of the energy function Le,T, the Euler-Lagrange equation commutes with the pullback, so that the application of the Euler-Lagrange equation to this pullback energy L˜e equivalent to withdrawing the Euler-Lagrange equation from The So calculate the Euler-Lagrange equation of The as: (JTMeJ)q¨+JT(fe+MeJ˙q˙)=0 (79)⇔M˜eq¨+f˜e=0 (80) with M e = J T M e J and f e = J T (f e + M e J̇q̇) (both as defined above). Therefore, the energization of 78 with L˜e M˜eq¨+f˜e+Me˜[M˜e−1−q˙q˙Tq˙TM˜eq˙](f˜−f˜e)=0 (81)⇒Me˜q¨+f˜e+(JTMeJ)[M˜e−1−q˙qTq˙TJTMeJq˙](JT(f=MeJ˙q˙)−JT(fe+MeJ˙q˙))=0 (82)⇒M˜eq¨+f˜e+(JTMe)J[(JTMeJ)−1−q˙qTx˙TMex˙]JT(f−fe)=0 (83)⇒M˜eq¨+f˜e+JTMe[J(JTMeJ)−1JT−x˙x˙Tx˙TMex˙](f−fe)=0 (84)

[0161] There JTMeJ(JTMeJ)−1JT=JT=JTMe(Me−1), Equation 84 can be expressed by M˜eq¨+f˜e+JTMe[Me−1−x˙x˙Tx˙TMex˙](f−fe)=0 which is consistent with the expression for the energized geometry pullback in Equation 75.

[0162] Theorem 7.1 shows that a concise method for calculating the energized geometry in the root is to first energize the leaves and then perform standard pullbacks. However, it is equally possible to simply pull back the geometries with respect to the energy metrics and energize the result. The following proposition shows that the system can consider such a pullback geometry as a metric-weighted average of geometries.

[0163] Proposition 7.2 (Metrically weighted average of geometries). x i = Ø(q) denotes for i = 1,..., m the star-shaped reduction of an arbitrary transformation tree, and assume that the leaves are with geometries ẍi + h 2,i = 0 with Finsler energies Lei with energy tensors Mi=∂x˙x˙2Lei Then the metric weighted pullback of the complete leaf geometry is q̈ + h̃2 = 0, with h˜2=(∑i=1mM˜i)−1∑i=1mM˜ih˜2,i where M i = J T M i J and h˜2,i=M˜i†JTMi(h2,i−J˙q˙) the standard pullback components are in acceleration form.

[0164] Proof. The standard algebra on (M i f i ), where f i = M i h 2,i , pulling back gives (M̃ i f i ) and summing gives Σ i (M i f i ) = (Σ i M i , Σ i f i ). Expressing this result in canonical form yields (∑iM˜i)(∑iM˜i)1−∑iM˜ih˜2,i) where h 2,iis the acceleration form of the individual pullbacks. Expanding yields the formula.

[0165] Since Lei Finsler energies, the pullback metrics M̃ i homogeneous of degree 0 in speed (ie they depend only on the normalized speed x˙^ Therefore, h̃2 is homogeneous of degree 2 and the pullback forms a geometry generator.

[0166] In general, once a geometry has been transformed with respect to a certain Finsler energy The energized, does not have a constant Euclidean speed ||ẋ||, since it tries to The Since the underlying geometric structure is defined by a tempo-independent speed, this is relatively easy to achieve. To maintain a constant Euclidean speed, the behavioral Finsler energy would have to The increase or decrease as needed. Intuitively, this energy can be increased by adding energy to the system using the potential function, and the energy can be decreased by damping. These two operations give us scope to regulate execution energies (for example, Euclidean energy) while working within the framework outlined by the geometric structure optimization theorems to ensure convergence and stability. The key tools that enable such regulation of execution energy are provided by the following proposition.

[0167] Proposition 8.1. Suppose that ẍ + h2(x,ẋ)=0 is a geometry generator, The a system energy with metric M e and h a forcing potential. αLe be such that x¨=−h2+αLex˙ the constant The (the energization coefficient), αalt0 be such that x¨=−h2+ αalt0x˙ the constant execution energy Lealt maintains, and αaltψ be such that x¨=−h2− Me−1∂xψ+αaltψx˙ which maintains constant execution energy. α alt be an interpolation between αalt0 and αaltψ. The system x¨=−h2−Me−1∂xψ+αaltx˙−βx˙ optimized when β>αalt−αLe

[0168] Proof. By definition, αLe an energization coefficient, x¨=−h2−Me−1∂xψ+αaltx˙−β˜x˙ is optimized if β > 0. This means: x¨=−h2−Me−1∂xψ+(αLe+αalt−αalt)x˙−β˜x˙ (91)=−h2−Me−1∂xψ+αaltx˙−(αalt−αLe−β˜)x˙ (92)

[0169] Optimized when β̃ > 0 . Using β=αalt−αLe+β˜ thus represents β−(αalt−αLe)+β˜>0 ready what β>αalt−αLe implied.

[0170] The basic requirement from the analysis, the constant The to maintain is β=αold−αLe; To optimize, the system must ensure that β is strictly larger than that (in particular, to converge to a local minimum). The system can use the constraint β ≥ 0 to maintain the semantics of a damper, which satisfies the following inequality β≥max{0,αalt−αLe} To converge, this inequality must be strict.

[0171] Note that α alt is the standard energy transformation for the geometric term -h2 alone, while αaltψ ensures that the constraint term is also included in the transformation. Thus, if β = 0, the following is provided:

[0172] The system under α alt x¨1=−h2−Me−1∂xψ+αaltψx˙ maintains constant execution energy Lealt while it continues to be enforced.

[0173] Likewise, αalt0 it and zero potential ψ = 0 the system x¨2=−h2+αalt0x˙ constant during the forced system Lealt hold x¨3=−h2−Me−1∂xψ+αalt0x˙ will also force the system to switch between energy levels. The difference between equations 95 and 93 must therefore include the additional component of −Me−1∂xh that accelerates the system in terms of this execution energy. This component is x¨3−x¨1=(−h2−Me−1∂xψ+αalt0x˙)−(−h2−Me−1∂xψ+αaltψx˙)(96)=(αalt0−αaltψ)x˙(97)

[0174] An interpolation between equations 95 and 93 results in a scaling of this component: Let η ∈ [0, 1], then: ηx¨3+(1+η)x¨1=η(−h2−Me−1∂xψ+αalt0x˙)−(1−η)(−h2−Me−1∂eψ+αaltψx˙)(98)=−h2−Me−1∂xψ+(ηαalt0+(1−η)αaltψ)x˙(99)

[0175] The added component is now: (ηx¨3+(1−η)x¨1)−x¨1=−h2−Me−1∂xψ+(ηαalt0+(1−η)αaltψ)x˙−(−h2−Me−1∂ xψ+αaltψx˙)(100)=[ηαalt0+(1−η)αaltψ]x˙(101)=η(αalt0−αaltψ)x˙(102) if η = 0 , take the entire system ẍ1, which is the totality of −h2−Me−1∂xψ projected. Accordingly, η(αalt0−αaltψ)x˙=0 if η = 0. Similarly, if η = 1, take the entire ẍ3, which gives the complete Me−1∂xψ remains. Here the entire component η(αalt0−αaltψ)x˙ obtained when η = 1.

[0176] Thus, the interpolated coefficient αalt−ηαalt0+(1−η)αaltψ to which the theorem refers, to modulate the component η(αalt0−αaltψ)x˙, which the amount of −Me−1∂xψ that must be passed to move the system between execution energy levels. A typical speed control strategy could then be: 1. Choose an execution energy Lealt for modulating. 2. Calculate for each cycle αalt0,αaltψ,αLe. 3. Choose η ∈ [0,1] to increase the energy as needed, using the extremes of η = 0 to maintain the execution energy and η = 1 to increase the execution energy completely. 4. Choose the damper under the condition β≥max{0,αalt−αLe} with αold=ηold0+(1−η)αaltψ. Use a strict inequality to remove energy from the system to ensure convergence. Note that the bound also adjusts accordingly at η = 1 (fully active potential), ensuring convergence under strict inequality.

[0177] The entire system executed in each cycle is x¨=v⊥+ηv / / −βx˙, (where η and β are defined as above), where v⊥=PLealt[−h2−Me−1∂xψ]=−h2−Me−1∂xψ+αaltψx˙ is the component of the system that receives the execution energy, Lealt and v / / the remaining (execution energy changing) component, so that v⊥+v / / =−h2−Me−1∂xψ the original system was reconstructed.

[0178] if β+αalt−αLe, the underlying behavioral energy The strictly maintained. If the damping is strictly greater than the lower bound β>αalt−αLe , the energy of the system decreases. As long as this strict inequality is satisfied, convergence is guaranteed.

[0179] The projection equation for α is calculated in the same way as the energization alpha used to energize geometries: α=−(x˙TMex˙)−1x˙T[Mex¨d−fe], where the Spec (M e , f e ) the energy equation M e ẍ + f e = 0 for the underlying behavioral energy The defined and ẍ d in this case a x¨d0=−h2 for αalt0 or x¨dψ=−h2−Me−1∂xψ for αaltψ is.

[0180] Note that for Euclidean Lealt=12‖x˙‖2 assuming that M e = I and fe = 0, so αalt=−(x˙Tx˙)−1x˙T[x¨d−0]=−x˙Tx¨dx˙Tx¨, which results in: x¨=x¨d+αaltx¨,(106)=x¨d−−x˙Tx¨dx˙Tx˙x˙=[I−x˙^x˙^T]x¨d,(107)=Px˙⊥[x¨d],(108) where Px˙⊥ is the projection matrix that projects orthogonally to ẋ. The above analysis is therefore only a generalization of this form of orthogonal projection to arbitrary execution energies.

[0181] Although theory provides that any potential function can be optimized over an arbitrary geometric structure, the shape of the potential function can lead to more desirable system behavior and timely convergence. In particular, an acceleration-based potential design is easy to use and tune and consists of a basis potential ψ1 (x) and an energy tensor M ψ , from the Finsler energy Le,ψ=xTG(x)x. The gradient of ψ1(x) is denoted by Mψ , which gives a gradient of the total potential, ψ(x), as ∂xψ(x)=Mψ∂xψ1(x)

[0182] It is important that Le,ψ should be added to the energy of a system, so that the constraint term in the acceleration of the system is given by ∂ x ψ1(x) can be approximated if M ψ , is large (high priority), whereby the mass of the system, M e , is dominated. That is, Me−1Mψ∂xψ1(x)≈∂xψ1(x),

[0183] Under this condition, the forced acceleration profile x˙=−h2−Me−1Mψ∂xψ1(x)=+αx˙,(111)=−h2−∂xψ1(x)+αx˙,(112) therefore acceleration-based design. Thus ∂ x ψ1(x) comes from a valid scalar potential function, ψ1(x) and M e be chosen so that ∂xx2ψ(x) is symmetric. This property is achieved by designing M ψ= ω(|| x || 2 )I, where ω(·)εℝ + is a scalar function that M ψ , radially symmetric. Likewise, let ψ1(x) = l(|| x || 2 ) is a radially symmetric potential function. The symmetry of ∂xx2ψ(x) can then be represented as ∂xx2ψ(x)=∂x[(ω(‖x˙‖2)I)(2I'(‖x‖2)x)],(113)=∂x[r(‖x‖2)x],(114)=r(‖x‖2)I+2r',(‖x‖2)xxT(115) where r = 2ω(|| x || 2 )l'(|| x || 2 ). The final expression is a sum of symmetric terms and therefore ∂xx2ψ(x) symmetrical.

[0184] As indicated, strategies or geometric structures that induce machines to perform one or more motion sequences can be based on strategy levels and structures that induce the machines to perform the one or more motion sequences using a level-wise or structure-wise approach. These strategies or geometric structures are a special type of structure that expresses their neutral nominal behavior as a generalized nonlinear geometry in the machine's configuration space; they represent the most concrete embodiment of optimization structures and incorporate many of the intuitive properties that make RMPs so powerful, such as acceleration-based strategy design and independent specification of priority metrics. Geometric structures can be conveniently constructed in parts distributed across a transformation tree of relevant task spaces.Importantly, however, they inherit important theoretical properties from the theory of structures, including stability and their neutral behavior. Furthermore, due to their construction as nonlinear geometries of paths, geometric structures exhibit a characteristic geometric consistency that enables their level-based construction to reduce design complexity. Each level independently controls the execution speed by accelerating the machine along a motion direction without compromising the overall quality of the machine's motion behavior.

[0185] Furthermore, geometric structures and strategies, as described, build on the theory of spectral half-sprays (Specs), which generalize the concept of second-order modular differential equations derived and used as RMPs. That is, let C be the d-dimensional configuration space of the machine. A vector notation describes elements of a space in coordinates. Associated task spaces x = ϕ(q) are defined in coordinates that q∈C⊂ℝd and x∈X⊂ℝn with the Jacobian matrix J = ∂ x ϕ, which is used in the relations ẋ = Jq̇ and ẍ = Jq̈ + J̇q̇.

[0186] Specs in natural form (M, f) x represent equations of the form M(x, ẋ)ẍ + f(x, ̇ẋ) = 0, and their algebra is derived from how these equations are summed and transformed under ẍ = Jq̈ + J̈q̇ (see [1]). Specs in canonical form (M,h)XC express the standard acceleration form equation ẍ + h(x, ẋ) = 0, where h - M -1 ẍ . For robotics applications, it is useful to additionally define a strategy shape spec [M,π] x to denote the solved strategy term ẍ = -h(x, ẋ)= π(x. ẋ) to emphasize that π = -h is an acceleration strategy.

[0187] For task spaces in which the specs are located, a transformation tree can be constructed and used. Each directed edge of the tree represents the differentiable map leading from the parent (domain) to the child (co-domain). Specs that populate a transformation tree collectively represent a complete second-order differential equation in parts by linking a particular spec to the root via the chain of differentiable maps that occur along the single path to the root. Denoting this composite map as above by ẋ = ϕ(q), one can use the expressions ẋ = Jq̇ and x¨˙=Jq¨˙+Jq˙, Relate the velocities and accelerations in the task space to velocities and accelerations in the root, derive a spec algebra that defines both how specs are combined in a single space and how they are transformed backward across edges from parent to child. The tree implicitly represents a complete differential equation at the root as a sum of the parts, computed by recursive application of the spec algebra.

[0188] The theory of generalized nonlinear geometry and Finsler energy are important for the derivation of geometric structures. For example, a generalized nonlinear geometry is an acceleration strategy ẍ = π(x, ẋ) for which π has a special homogeneity property, e.g., positive homogeneous of degree two, , which means that for any λ ≥ 0, π(x, λẋ) = λ 2π(x, ẋ). It can be shown that the second-order homogeneity (HD2) property ensures that the differential equation is more than just a collection of trajectories (its integral curves); it also has a path consistency property, where any integral curve starting at a given position x0 and whose velocity ẋ0 = ηn̂ points in a given direction n̂ (here η > 0) follows the same path. In particular, any variant of the differential equation of the form ẍ = π(x, ẋ) + α(t, x, ẋ)ẋ, where α ∈ ℝ, will have integral curves that trace the same paths as π. This geometric consistency property transforms π into a geometry of paths.

[0189] A Finsler energy Le(xx˙) is a generalization of the classical kinetic energy from classical mechanics (the classical kinetic energy K=12xTG(x).x˙ is a form of Finsler energy). Analogous to the classical case, the Euler-Lagrange equation applied to a Finsler energy defines an equation of motion M e (x, ẋ)ẍ = f e (x, ẋ)= 0, in the Me=∂x˙x˙2Le is the energy (or metric) tensor and fe=∂x˙xLex˙=∂xLe the curvature terms (Coriolis and centripetal forces in classical mechanics). This equation agrees with the equations of motion of classical mechanics if Le=K, for the M e (x, ̇ẋ) = G(x).

[0190] In geometric structures, the energy tensor defines the priority metric of the strategy, and the curvature terms f e are used for stability (see Section VI-C). Finsler energies are Lagrange functions, Le(x,x˙), which meet the following requirements: 1) Positivity: Le(x,x˙)≥0, with equality only for ẋ = 0 2) Homogeneity: Le(x,x˙) is positively homogeneous of degree 2 in ẋ; ie for λ ≥ 0, provided Le(x,λx˙)=λ2Le(x,x˙) 3) Energy tensor invertibility: Me=∂x˙x˙2Le is invertible everywhere.

[0191] The metric tensor M e (x, ẋ) is in general a function of both velocity and position, although the above homogeneity requirement forces that M e only on the direction of the speed (Me(x,x˙)=Me(x,x˙^) forx˙≠0) for ẋ ≠ 0) and does not depend on the order of magnitude (i.e., it is homogeneous of degree 0). This directionality allows the design of direction-dependent priority matrices as metric tensors of Finsler energies.

[0192] Geometric structures are a form of optimization structure that is designed as a special type of differential equation to induce behavior by influencing the optimization path of a differential optimizer. This disclosure pragmatically describes how the structure can be effectively designed to encode a desired behavior.

[0193] A forced geometric structure is a collection of structural terms that are represented as pairs (Le,π)x a Finsler energy Le(xx˙) and an acceleration strategy ẍ = π(x, ̇ẋ). Geometric terms define the structure, while constraint terms define the goal. A geometric term is a term (Le,π2)X, for which π2 is an HD2 geometry. A constraint term is a term (Le,−Me−1∂xψ)X, who derives his strategy from a potential function. Structural terms can be added to spaces of a transformation tree to modularize compound behaviors.

[0194] Each structural term defines a triple (M e , f e , π) x , where M e ẍ + f e = 0 from the The applied Euler-Lagrange equation, which can be viewed as two specs, a strategy spec [M e , π ]x and a natural energy spec (M e , f e ) x . Geometric structure summation and pullback is accordingly defined in terms of the algebra of these two constituent specs.

[0195] Geometric structures are neutral and therefore never prevent a system from reaching a local minimum of the goal. The goal therefore encodes concrete task goals independent of the structure's behavior. Furthermore, geometric strategies are geometrically consistent and rate-invariant in the geometry of the paths, which both simplifies the intuition for their summation and enables behavior-invariant control of execution speed.

[0196] The design of a geometric structure follows the intuition of the design of RMPs. Strategy specs [M e , π2] x model both a desired behavior π2 and a priority matrix for this behavior M e , which defines how it combines with other strategies as a metric weighted average of parts. The spectrum of M ecan assign different weights to different directions in space, and both π2 π(x, ̇ẋ) and M e (x, ̇ẋ) have the flexibility to depend on both the position x and the velocity ẋ. Since M e HD0, geometric terms remain geometric upon summation and pullback (e.g., the strategy resulting from a metric-weighted average of geometries is itself a geometry). The energy spec (M e . f e ) x Each structural term is used only to ensure stability during execution. Therefore, users can simply focus on the design of the behavioral strategy specs [M e , π2] x focus.

[0197] Once the constrained geometric structure is defined, it can be transformed by acceleration and deceleration along the motion direction to maintain a certain level of execution energy (e.g., the end-effector velocity or the joint velocity in configuration space). The geometric consistency of the structure means that the behavior remains consistent despite these velocity modulations. Many numerical integrators are suitable for integrating the final differential equation. For example, fourth-order Euler (1 ms time step) or Runge-Kutta methods (10 ms time step) offer a good compromise between speed and accuracy.

[0198] In one embodiment, the motion sequence of a machine can be designed in three parts (using a transformation tree of task spaces): (1) designing the underlying structure, (2) adding a driving potential to define task goals, (3) designing an execution energy for velocity control. In some embodiments, the transformation tree can be designed / constructed by (1) forward traversal: filling the nodes with the current state from the root to the leaves. (2) backward traversal: evaluating the specs and pulling to the root in separate channels, an energy channel for the energy specs of the geometric terms, a strategy channel for the strategy specs of the geometric terms, an execution energy channel for the execution energy specs, and a constraint strategy channel for the strategy specs of the constraint terms. Adding the energy specs of the constraint term to the energy channel.(3) Using the root results of the four channels to calculate the final desired acceleration using the velocity control.

[0199] Geometric structures follow acceleration-based design principles that were captured in canonical form in the original RMPs. A geometric structure is a pair (Le,π2)X, which characterizes two specs, an energy spec and a geometry spec. The energy spec captures stability information, while the geometry spec captures the behavior. Behavioral engineering focuses on the construction of the latter, using the class of HD2 geometries to model π2 and derive M e as the energy tensor of a Finsler geometry The is used.

[0200] When geometric structures are summed ∑i(Le(i),π2(i))X, Σ iis the geometry spec of the combined structure (which captures its behavior) ∑i(Me(i),π2(i))X=(M˜e,π˜2) a metric weighted average of the contributing geometries π˜2=(∑iMe(i))−1∑iMe(i)π2(i), prioritized according to the overall metric Me=∑iMe(i). M e When populating a transformation tree, this intuitive combination rule is applied recursively to each node. Designers only need to focus on intuitively creating modular acceleration strategies (as HD2 geometries) in the different spaces and prioritizing them with metric tensors (from Finsler energies).

[0201] An HD2 geometry is a differential equation ẍ + h2π(x, ẋ), where h2 is HD2, which is usually denoted in the strategy form ẍ = -h2 (x, ẋ) = π2 π(x, ẋ). The construction of an HD2 geometry is straightforward given the following rules for homogeneous functions: (1) a sum of HD2 functions is HD2; (2) multiplication of homogeneous functions adds their degree (an HDk function is denoted by f k denoted, examples are f2 f0 = f2, f1f1 = f2 etc.). A simple way to design an HD2 geometry is, for example, to choose an HD0 strategy π0 (x) that depends only on position and shape π2 (x, ̇ẋ) = ||ẋ|| 2 π0(x) depends; by scaling by ||ẋ|| 2 . π0 (x) can be written as the negative gradient of a potential π0 (x) = - ∂ x ψ(x) can be chosen.

[0202] In at least one embodiment, it may be intuitive to define the constraint potential of a geometric structure as a constraint spec F=[Mf,πf]X in strategy form, so that it is intuitively treated as another acceleration strategy averaged into the final weighted mean. This strategy π f can therefore implicitly be a forcing potential ψ f (x) whose negative gradient is given by - ∂ x ψ f (x) = M f π f When designing F is to enforce that M f and π f remain theoretically compatible in this sense. Choose π f = ∇ x ψ acc (x), where ψ acc is a potential function that is spherically symmetric about its global minimum and expresses the acceleration strategy directly as its negative gradient. Each metric M f ( X) is theoretically compatible, even if it is spherically symmetric about the same global minimal point. Note that pure position metrics are Riemannian metrics and consist of Finsler energies of the form Le=12x˙TMf(x)x˙ derive.

[0203] For speed control, in one embodiment αreg=αexη−βreg(x,x˙)+αboost in x¨=−Me−1∂xψ(x)+π0(x,x˙)+αregx˙ with βreg=sβ(x)B+B_+max{0,αexη−αLe} used. B > 0 is a (small) baseline attenuation, B > B is a larger attenuation coefficient for concise convergence. The switch s β (x) turns on near the target: sβ(x)=12(tanh(−αβ(‖x‖−r))+1) where ap α β ∈ ℝ + is a gain that defines the switching speed, and r ∈ ℝ +the radius at which the switch is half engaged. Designating the desired execution energy as Leex,d use the following strategy for η. η=12(tanh(−αη(Leex−Leex,d)−αshift)+1) where α η , α shift ∈ ℝ + the speed or offset of the switch is adjusted as an affine function of the speed error (execution energy). Finally, a boost as αboost=kη(1−sβ(x))1‖x˙‖+ε modeled, where k ∈ ℝ + is a gain that directly determines the desired acceleration level, η (from above) a boost = 0 when the desired speed is reached, 1 - s β (x) a boost = 0 when the system is in the higher damping range. The normalization by ||ẋ|| + ε ensures that α boost directly together with x˙^ is applied, with a very small positive value for ε to ensure numerical stability. This overall design introduces more energy into the system when −αboost<αexη−αLe. Since this entry occurs for a finite time, the total system energy is still limited. The additional switches ensure that the system continues to be subject to positive damping, thus ensuring convergence.

[0204] In at least one embodiment, the effectiveness of level-based construction of a geometric structure allows the robot to reach into and out of three sets of compartments or crate bins. In one example, six reachable compartments or crate bins are located at the front and two on each side of a Franka arm. The length, width, and depth of each compartment / crate is 0.3 m. The following explains the levels of the structure, with each new level resting on the previous, fixed levels, reducing the complexity of design and tuning.

[0205] Each strategy is represented as an HD2 geometry of the form ẍ = -||ẋ|| 2 ∂ x ψ(x) is defined. The strategies are denoted by M e (x, ̇ẋ) weighted from the Finsler energy, Le=xTG(x,x˙)x˙.ψ(x) and G(x, ̇ẋ) are defined as follows.

[0206] The first level creates a geometric starting structure designed for global, cross-body point-to-point navigation with end-effector without obstacles.

[0207] End-effector attraction: For the attraction to a target, the task map y = Φ att (x) = x t - x used, where x, x t ∈ ℝ 3 are the current and target positions of the end effector in Euclidean space. The metric is simply an identity matrix scaled by s(||y||), G att = s(||y||)I, where s(||y||) = 40 if ||y|| < 0.5, and otherwise s(||y||) = 1. The acceleration-based potential gradient ∂ x ψ1(y) = M a tt(y) ∂ x ψ1(y) is used: ψ1(y)=k(‖y‖+1αψlog(1+e−2αψ‖y‖)) where k ∈ ℝ + controls the overall gradient strength, α ψ ∈ ℝ + the transition rate of ψ1 (y) from a positive constant to 0, and here α ψ = 10, k = 10.

[0208] Joint boundary avoidance: This behavior uses two 1D task cards per joint, xju=ϕu(qj)=q¯j−qj and xjl=ϕl(qj)=qj−q_j, where q j and q j the upper and lower limits of the j th joint. If both are denoted generically as x, the metric G l as Gl(x,x˙)=s(x˙)=s(x˙)λx where s(ẋ) = 0 if x > 0 and s(ẋ) = 1 x otherwise (1D normalization) and λ = 10. This cancels out the effect of the coordinate boundary geometry as soon as the motion is orthogonal or away from the boundary. The acceleration-based potential gradient ∂ q ψ(x) = M l (x) ∂ q ψ 1,1 (x)used: ψ1,l(x)=a1x2+α2log(e−α3(x−α4)+1) and M l (x) comes from the Finsler energy Le(x)=Gl(x)x˙2, where G l (x) G l(x, ̇ẋ) which drops the s(ẋ) term, α1,α2 ∈ ℝ + controls the significance and mutual balance of the first and second terms. α3 ∈ ℝ + controls the sharpness of the smooth rectified linear unit (SmoothReLU), while α4 ∈ ℝ + the SmoothReLU. Overall, ψ1(x) → ∞ is x → 0 and ψ1 (x) → 0 is x → ∞, which hinders the movement towards a boundary. In this experiment, α1 = 0.4, α2 = 0.5, α3 = 20 and α4=π6.

[0209] Standard configuration: The task card for this behavior is x = ϕ dc (q) = q0 - q , where q0 is a standard configuration. The metric G dc is simply an identity matrix scaled by a constant λ dc , G dc = λ dc I , where λ dc = 0.5 . The acceleration-based potential gradient is defined in the same way as Eq. 3, where α ψ= 6.75 and k = 100. Five separate default configurations are created to cover different areas of the robot. This behavior controls the robot posture and eliminates manipulator redundancy.

[0210] The next level allows the end effector to pick up and enter any compartment / box (collisions are ignored). Compartment removal: Compartment removal uses two task cards that represent the distance between: 1) the end effector and the front level of the compartment, y1 = ϕ1(x) = ||x f - x||, if the end effector is inside the compartment, y1 = ϕ1(x) = ||x f - x|| otherwise, where x, x f ∈ ℝ 3 is the position of the end effector or its orthogonal projection onto the front plane; and 2) the end effector and a line centered and orthogonal to the front plane of the target compartment, y2 = ϕ2(x) = ||x c - x||,, where x, x c ∈ ℝ 3x is the position of the end-effector and the point closest to the end-effector on the line, where the line is orthogonal and centered to the front plane. The priority metric is as follows: Gw(x)=s(y1)((m¯−m_)s(y2)+m_)I. where s(y1) = 1 if y1 < 0.1yl and otherwise s(y1) = 0. s(y2) = 0.5(tanh(α m (y2 - r)) + 1), m̅, m ∈ ℝ + are the upper and lower isotropic masses, respectively, and α m ∈ ℝ + defines the transition rate between 0 and 1, while r ∈ ℝ + the transition. For this experiment, m̅ = 5, m = 0, a m = 100 and r = 0.15. Overall, the priority disappears when the end effector is either more than 0.1 m from the front plane (outside the compartment) or within 0.15 m of the centerline of the target compartment. The potential function is the same as defined in Eq. 4, where α1 = 0, α2 = 15, a3 = 100, a4 = 0.05.

[0211] Target-bin attraction: An additional target attraction strategy is used to help pull the end-effector into a target bin. It is equivalent to the end-effector attraction defined in Level 1, but with the priority metric defined as in Eq. 5, where s(y1) = 1 ∀ y1, m̅ = 5, m = 0, α m = -100, and r = 0.15 and α ψ = 10, k = 40 in Eq. 3. Overall, the priority remains zero until it enters a cylindrical region aligned with the target compartment / box. The increasing priority directs the movement into the compartment / box.

[0212] Waypoint attraction: This term guides the arm to the compartment / box opening by pulling it toward a point 0.15 m in front of the front plane. This term is generally equivalent to the attraction term for the end effector above, but with a different target and a toggle function for the metric that deactivates it once it is in the column of the target compartment / box. Specifically, make the following changes: 1) Replace x t by x w in the task card, where x w ∈ ℝ 3 is the position of the waypoint; 2) Define the switching function using the task map y2 = Ø2 (x) = ||x2 - x||, as described in the compartment / box picking strategy; 3) k = 20 is used. Note that this term guides the arm but, as a geometric term, does not affect convergence to the target.

[0213] The next level enables complete collision avoidance with the compartments / boxes. Compartment collision: The task map for this behavior is y = Ø2(x), which captures the minimum distance between a point on the arm and the compartment, where x is in ℝ 3 is a specified collision point on the robot. Denote the nearest point on the shelf as x c and use y2 = Ø2(x)= ||x c - x|| (for simplicity, ignore the dependencies of x c of x). The metric is a function of the position Gb(y)=kby2 defined, where k b ∈ ℝ + is a barrier reinforcement. The acceleration-based potential gradient ∂ q ψ b (y)= M b (y)∂ q ψ 1,b (y) used ψ1,b(y)=αby8, where α b ∈ ℝ + is the barrier reinforcement. Here κ b = 1 and α3 = 0.1.

[0214] The optimization potential is an acceleration-based design and is for the task space y = Ø(x) = x - x t where x, x t ∈ ℝ 3 is the current and target end-effector position in Euclidean space. The acceleration-based potential gradient ∂ y ψ(y)= M ψ (y)∂ y ψ1(y) uses ψ1(y) as in Eq. 3 with α ψ = 10, k = 20. Gψ(y)=((m¯−m_)s(||y||)+m_)I, where s(||y||) = 0.5(tanh(α m (||y||- r)) + 1), m̅, m ∈ ℝ + is the upper or lower isotropic mass and α m ∈ ℝ + defines the rate of transition between 0 and 1, while r ∈ ℝ + the transition. For this experiment, m̅ = 40, m = 0.1, α m= 100, and r = 0.15. This design allows for a low attraction potential at great distances from the target, while the priority increases smoothly with increasing proximity to the target location.

[0215] Damping values are B = 17.5, B = 0. In Eq. 1, α β = 50 and r = 0.15. In Eq. 2, α eta = 10, a shift = 2, Leex=yTy and Leex,d=1, The gain k in a boost is defined as k=−5|‖y˙‖−Leex,d|, where is the current end effector speed.

[0216] Using the above level-based approach to strategy design, the robot intelligently navigates the compartments / bins. This global behavior has been progressively sequenced by adding levels of complexity to the underlying geometric structure, a technique facilitated by geometric consistency and acceleration-based design. Specifically, in at least one embodiment, the first level causes the robot to move directly to the goal, completely ignoring the compartments / bins. The next level enhances compartment / bin navigation. The next level allows the robot to completely avoid collisions.

[0217] Fig. Figure 7 illustrates an example of a process 700 for initiating a computer-implemented action, according to one embodiment. In at least one embodiment, one or more computer systems, such as one or more of the Fig. 1-6 and 8A-41B execute instructions stored in computer-readable memory that cause the computer system to perform process 700. In at least one embodiment, the computer system is a robotic system, such as one or more of the robotic systems described herein. The robotic system may include the one or more of the computer systems described and illustrated herein.

[0218] At 702, a strategy is identified that causes a machine to perform at least one movement. In at least one embodiment, the machine is a robot. For example, the machine may be Fig. 2-3 illustrated robot 206, or the one in Fig. 4-6 illustrated robot 402. Alternatively, the machine can also be the one shown in Fig. 1. In at least one embodiment, the strategy may be similar to the Fig. 1. Similarly, the strategy may be associated with the Fig. 2-6. In at least one embodiment, the strategy comprises at least one strategy level. In one example, the strategy comprises a plurality of strategy levels. The one or more strategy levels may comprise at least one differential equation. In at least one embodiment, the at least one differential equation is a second-order differential equation. The second-order differential equation may be homogeneous of degree two. In at least one embodiment, the one or more strategy levels are energized by one or more Finsler energies. The one or more Finsler energies may be homogeneous of degree two.

[0219] At 704, a first strategy level associated with the plurality of strategy levels is executed to cause the machine to perform a first motion sequence that reaches a neutral state. In at least one embodiment, the first motion sequence caused by the first strategy level is limited by at least a first parameter associated with the machine and a second parameter associated with a region in which the machine operates. In at least one embodiment, the first parameter is a joint limit of the machine. In particular, the joint limit may be associated with a wrist or other joint of a robot. In at least one embodiment, the second parameter comprises data indicating a target position to be reached by the machine.In at least one embodiment, the target position is within a task space or area in which the operator operates. In at least one embodiment, the target position is a coordinate or coordinates within a Euclidean space in which the operator operates.

[0220] At 706, a second strategy level associated with the plurality of strategy levels is executed to cause the machine to perform a second motion sequence. In at least one embodiment, the second motion sequence performed by the machine builds upon the first motion sequence performed by the machine as prompted by the first strategy level. In at least one embodiment, the second motion sequence prompted by the second strategy level does not impact allowing the machine to reach the neutral state based on the first strategy level. In at least one embodiment, one or more motion sequence parameters of the second strategy level may be energized independently of the one or more motion sequence parameters of the first strategy level.The one or more motion parameters may include trajectory motion parameters, acceleration motion parameters, velocity motion parameters, joint limit parameters, redundancy parameters, direction and / or coordinate parameters, and so on.

[0221] Fig. Figure 8 illustrates an example of a process 800 for initiating a computer-implemented action, according to one embodiment. In at least one embodiment, one or more computer systems, such as one or more of the Fig. 1-6 and 8A-41B execute instructions stored in computer-readable memory that cause the computer system to perform process 800. In at least one embodiment, the computer system is a robotic system, such as one or more of the robotic systems described herein. The robotic system may include the one or more of the computer systems described and illustrated herein.

[0222] At 802, a first strategy level associated with a plurality of strategy levels is generated. The first strategy level is intended to cause a machine to perform a first motion sequence that reaches a neutral state. In at least one embodiment, the first motion sequence caused by the first strategy level is limited by at least a first parameter associated with the machine and a second parameter associated with a region in which the machine operates. In at least one embodiment, the first parameter is a joint limit of the machine. In particular, the joint limit may be associated with a wrist or other joint of a robot. In at least one embodiment, the second parameter comprises data indicating a target position to be reached by the machine.In at least one embodiment, the target position is within a task space or area in which the operator operates. In at least one embodiment, the target position is a coordinate or coordinates within a Euclidean space in which the operator operates.

[0223] In at least one embodiment, the machine is a robot. For example, the machine may be Fig. 2-3 illustrated robot 206, or the one in Fig. 4-6 illustrated robot 402. Alternatively, the machine can also be the one shown in Fig. 1. In at least one embodiment, the strategy may be similar to the Fig. 1. Similarly, the strategy may be associated with the Fig. 2-6. In at least one embodiment, the strategy comprises at least one strategy level. In one example, the strategy comprises a plurality of strategy levels. The one or more strategy levels may comprise at least one differential equation. In at least one embodiment, the at least one differential equation is a second-order differential equation. The second-order differential equation may be homogeneous of degree two. In at least one embodiment, the one or more strategy levels are energized by one or more Finsler energies. The one or more Finsler energies may be homogeneous of degree two.

[0224] At 804, a second strategy level associated with the plurality of strategy levels is generated. The second strategy level is to cause the machine to perform a second motion sequence. In at least one embodiment, the second motion sequence to be performed by the machine builds upon the first motion sequence to be performed by the machine caused by the first strategy level. In at least one embodiment, the second motion sequence caused by the second strategy level has no impact on allowing the machine to reach the neutral state based on the first strategy level. In at least one embodiment, one or more motion sequence parameters of the second strategy level may be energized independently of the one or more motion sequence parameters of the first strategy level.The one or more motion parameters may include trajectory motion parameters, acceleration motion parameters, velocity motion parameters, joint limit parameters, redundancy parameters, direction and / or coordinate parameters, and so on.

[0225] At 806, the first strategy level and the second strategy level are executed to cause the machine to execute the first motion sequence and the second motion sequence. In at least one embodiment, the first strategy level and the second strategy level are part of a geometric structure. The second motion sequence caused by the second strategy level can build on the first motion sequence caused by the first strategy level.

[0226] Fig. 9A illustrates inference and / or training logic 915 used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 915 are described below in connection with Fig. 9A and / or 9B provided.

[0227] In at least one embodiment, the inference and / or training logic 915 may include, without limitation, code and / or data storage 901 to store feedforward and / or output weighting and / or input / output data and / or other parameters to configure neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 915 may include or be coupled to code and / or data storage 901 to store graph code or other software to control the timing and / or order in which weighting and / or other parameter information is to be loaded to configure logic, including integer and / or floating-point units (collectively referred to as arithmetic logic units (ALUs)).In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on a neural network architecture to which that code corresponds. In at least one embodiment, code and / or data storage 901 stores weight parameters and / or input / output data of each layer of a neural network being trained or used in connection with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 901 may be included in other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.

[0228] In at least one embodiment, any portion of code and / or data storage 901 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 901 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, a choice of whether the code and / or code and / or data storage 901 is, for example, internal or external to a processor or comprises DRAM, SRAM, flash, or another memory type may depend on the available on-chip or off-chip memory, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inferring and / or training a neural network, or a combination of these factors.

[0229] In at least one embodiment, the inference and / or training logic 915 may include, without limitation, a code and / or data storage 905 to store backward and / or output weighting and / or input / output data corresponding to neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the code and / or data storage 905 stores weighting parameters and / or input / output data of each layer of a neural network being trained or used in connection with one or more embodiments during backpropagation of input / output data and / or weighting parameters during training and / or inference using aspects of one or more embodiments.In at least one embodiment, training logic 915 may include or be coupled to code and / or data storage 905 to store graph code or other software for controlling the timing and / or order in which weight and / or other parameter information is to be loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).

[0230] In at least one embodiment, code, such as graph code, causes weight or other parameter information to be loaded into processor ALUs based on a neural network architecture to which such code corresponds. In at least one embodiment, any portion of code and / or data storage 905 may be included in other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 905 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 905 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, a choice of whether the code and / or data storage 905 is, for example, internal or external to a processor or comprises DRAM, SRAM, flash memory, or another type of memory may depend on the available on-chip or off-chip memory, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inferring and / or training a neural network, or a combination of these factors.

[0231] In at least one embodiment, code and / or data memory 901 and code and / or data memory 905 may be separate memory structures. In at least one embodiment, code and / or data memory 901 and code and / or data memory 905 may be a combined memory structure. In at least one embodiment, code and / or data memory 901 and code and / or data memory 905 may be partially combined and partially separate. In at least one embodiment, any portion of code and / or data memory 901 and code and / or data memory 905 may be included in other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.

[0232] In at least one embodiment, the inference and / or training logic 915 may include, without limitation, one or more arithmetic logic units (“ALU(s)”) 910, including integer and / or floating point units, to perform logical and / or mathematical operations based at least in part on or specified by training and / or inference code (e.g., graph code), a result of which may produce activations (e.g., output values of layers or neurons within a neural network) stored in an activation memory 920 that are functions of input / output and / or weighting parameter data stored in the code and / or data memory 901 and / or the code and / or data memory 905.In at least one embodiment, activations stored in activation memory 920 are generated according to linear algebraic and / or matrix-based mathematics performed by ALU(s) 910 in response to the execution of instructions or other code, using weight values stored in code and / or data memory 905 and / or data memory 901 as operands along with other values, such as distortion values, gradient information, moment values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data memory 905 or code and / or data memory 901 or other on-chip or off-chip memory.

[0233] In at least one embodiment, the ALU(s) 910 are included within one or more processors or other hardware logic devices or circuits, while in another embodiment, the ALU(s) 910 may be external to a processor or other hardware logic device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, the ALU(s) 910 may be included within the execution units of a processor or otherwise within a bank of ALUs accessible by the execution units of a processor, either within the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.).In at least one embodiment, code and / or data memory 901, code and / or data memory 905, and enablement memory 920 may share a processor or other hardware logic device or circuit, while in another embodiment, they may be located in different processors or other hardware logic devices or circuits, or in a combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of enablement memory 920 may be included in other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.In addition, the inference and / or training code may be stored along with other code accessible by a processor or other hardware logic or circuitry and retrieved and / or processed using a processor's fetch, decode, scheduling, execution, elimination, and / or other logic circuitry.

[0234] In at least one embodiment, the activation memory 920 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the activation memory 920 may be located, in whole or in part, within or external to one or more processors or other logic circuitry. In at least one embodiment, a choice of whether the activation memory 920 is, for example, internal or external to a processor or comprises DRAM, SRAM, flash memory, or another type of memory may depend on the available on-chip or off-chip memory, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inferring and / or training a neural network, or a combination of these factors.

[0235] In at least one embodiment, the Fig. 9A may be used in conjunction with an application-specific integrated circuit (“ASIC”), such as a TensorFlow® Processing Unit from Google, an inference processing unit (IPU) from Graphcore™, or a Nervana® processor (e.g., “Lake Crest”) from Intel Corp. In at least one embodiment, the inference and / or training logic 915 illustrated in Fig. 9A may be used in conjunction with central processing unit (“CPU”), graphics processing unit (“GPU”), or other hardware, such as field programmable gate arrays (“FPGAs”).

[0236] Fig. 9B illustrates inference and / or training logic 915 according to at least one embodiment. In at least one embodiment, the inference and / or training logic 915 may include, without limitation, hardware logic in which computational resources are dedicated or otherwise used exclusively in connection with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, the inference and / or training logic 915 may Fig. 9B may be used in conjunction with an application-specific integrated circuit (ASIC), such as Google's TensorFlow® Processing Unit, a Graphcore™ inference processing unit (IPU), or a Nervana® processor (e.g., "Lake Crest") from Intel Corp. In at least one embodiment, the inference and / or training logic 915 illustrated in Fig. 9B may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware, such as field programmable gate arrays (FPGAs). In at least one embodiment, the inference and / or training logic 915 includes, without limitation, code and / or data storage 901 and code and / or data storage 905, which may be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, moment values, and / or other parameter or hyperparameter information. In at least one embodiment, Fig. 9B, each of the code and / or data memory 901 and the code and / or data memory 905 is associated with a dedicated computing resource, such as the computing hardware 902 and the computing hardware 906, respectively. In at least one embodiment, each of the computing hardware 902 and the computing hardware 906 includes one or more ALUs that perform mathematical functions, such as linear algebraic functions, only on information stored in the code and / or data memory 901 and the code and / or data memory 905, respectively, the result of which is stored in the activation memory 920.

[0237] In at least one embodiment, each of the code and / or data memories 901 and 905 and the corresponding computational hardware 902 and 906 correspond to different layers of a neural network, such that the resulting activation from one memory / compute pair 901 / 902 of the code and / or data memory 901 and the computational hardware 902 is provided as input to the next memory / compute pair 905 / 906 of the code and / or data memory 905 and the computational hardware 906 to mirror the conceptual organization of a neural network. In at least one embodiment, each of the memory / compute pairs 901 / 902 and 905 / 906 may correspond to more than one layer of a neural network. In at least one embodiment, additional memory / compute pairs (not shown) may be included subsequent to or in parallel with memory / compute pairs 901 / 902 and 905 / 906 in the inference and / or training logic 915.

[0238] Fig. 10 illustrates training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, the untrained neural network 1006 is trained using a training dataset 1002. In at least one embodiment, the training framework 1004 is a PyTorch framework, whereas in other embodiments, the training framework 1004 is a TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, the training framework 1004 trains an untrained neural network 1006 and allows it to be trained using the processing resources described herein to produce a trained neural network 1008. In at least one embodiment, the weights may be chosen randomly or by pre-training using a deep belief network.In at least one embodiment, the training may be performed in either a supervised, semi-supervised, or unsupervised manner.

[0239] In at least one embodiment, the untrained neural network 1006 is trained using supervised learning, where the training data set 1002 includes an input paired with a desired output for an input, or where the training data set 1002 includes an input with a known output and an output of the neural network 1006 is manually ranked. In at least one embodiment, the untrained neural network 1006 is trained in a supervised manner and processes inputs from the training data set 1002 and compares the resulting outputs to a set of expected or desired outputs. In at least one embodiment, errors are then backpropagated through the untrained neural network 1006. In at least one embodiment, the training framework 1004 sets weights that control the untrained neural network 1006.In at least one embodiment, the training framework 1004 includes tools to monitor how well the untrained neural network 1006 converges to a model, such as the trained neural network 1008, capable of generating correct answers, such as in the result 1014, based on input data, such as a new dataset 1012. In at least one embodiment, the training framework 1004 repeatedly trains the untrained neural network 1006 while adjusting weights to refine an output of the untrained neural network 1006 using a loss function and a tuning algorithm, such as stochastic gradient descent. In at least one embodiment, the training framework 1004 trains the untrained neural network 1006 until the untrained neural network 1006 achieves a desired accuracy.In at least one embodiment, the trained neural network 1008 may then be used to implement any number of machine learning operations.

[0240] In at least one embodiment, the untrained neural network 1006 is trained using unsupervised learning, where the untrained neural network 1006 attempts to train itself using unlabeled data. In at least one embodiment, the training dataset 1002 for unsupervised learning includes input data without associated output data or ground truth data. In at least one embodiment, the untrained neural network 1006 can learn groupings within the training dataset 1002 and determine how individual inputs relate to the untrained dataset 1002. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in the trained neural network 1008 capable of performing operations useful in reducing the dimensionality of the new dataset 1012.In at least one embodiment, unsupervised training may also be used to perform anomaly detection, enabling the identification of data points in the new data set 1012 that deviate from normal patterns of the new data set 1012.

[0241] In at least one embodiment, semi-supervised learning may be used, which is a technique in which the training dataset 1002 includes a mixture of labeled and unlabeled data. In at least one embodiment, the training framework 1004 may be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning allows the trained neural network 1008 to adapt to the new dataset 1012 without forgetting the knowledge taught to the trained neural network 1008 during initial training. DATA CENTER

[0242] Fig. Figure 11 illustrates an example data center 1100 in which at least one embodiment may be used. In at least one embodiment, data center 1100 includes a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and an application layer 1140.

[0243] In at least one embodiment, a data center infrastructure layer 1110, as shown in Fig. 11, include a resource orchestrator 1112, clustered computing resources 1114, and node computing resources (“node CRs”) 1116(1)-1116(N), where “N” represents a positive integer (which may be a different integer “N” than used in other figures). In at least one embodiment, the node CRs 1116(1)-1116(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), storage devices 1118(1)-1118(N) (e.g., dynamic read-only memory, solid-state storage, or hard disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more node CRss from the node CRs 1116(1)-1116(N) may be a server having one or more of the computing resources mentioned above.

[0244] In at least one embodiment, the grouped computing resources 1114 may include separate groupings of node CRs housed in one or more racks (not shown) or in many racks in data centers in different geographic locations (also not shown). Separate groupings of node CRs within the grouped computing resources 1114 may, in at least one embodiment, include grouped computing, networking, memory, or storage resources that may be configured or assigned to support one or more workloads. In at least one embodiment, multiple node CRs including CPUs or processors may be grouped in one or more racks to provide computing resources to support one or more workloads.In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.

[0245] In at least one embodiment, resource orchestrator 1112 may configure or otherwise control one or more node CRs 1116(1)-1116(N) and / or clustered computing resources 1114. In at least one embodiment, resource orchestrator 1112 may include a software design infrastructure (“SDI”) management entity for data center 1100. In at least one embodiment, resource orchestrator 1112 may include hardware, software, or a combination thereof.

[0246] In at least one embodiment, as in Fig. 11, the framework layer 1120 includes a task scheduler 1122, a configuration manager 1124, a resource manager 1126, and a distributed file system 1128. In at least one embodiment, the framework layer 1120 may include a framework for supporting software 1132 of the software layer 1130 and / or one or more applications 1142 of the application layer 1140. In at least one embodiment, the software 1132 or the application(s) 1142 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1120 may be some type of free and open source software web application framework such as, but not limited to, Apache Spark™ (hereinafter "Spark"), which may utilize the distributed file system 1128 for processing large amounts of data (e.g., "Big Data").In at least one embodiment, the task scheduler 1122 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 1100. In at least one embodiment, the configuration manager 1124 may be capable of configuring various layers, such as the software layer 1130 and the framework layer 1120, which includes Spark and a distributed file system 1128 to support large-scale data processing. In at least one embodiment, the resource manager 1126 may be capable of managing clustered or grouped compute resources mapped or allocated to support the distributed file system 1128 and the task scheduler 1122. In at least one embodiment, clustered or grouped compute resources may include grouped compute resources 1114 in the data center infrastructure layer 1110.In at least one embodiment, the resource manager 1126 may coordinate with the resource orchestrator 1112 to manage these mapped or allocated compute resources.

[0247] In at least one embodiment, the software 1132 included in software layer 1130 may include software used by at least portions of node CRs 1116(1)-1116(N), clustered computing resources 1114, and / or distributed file system 1128 of framework layer 1120. One or more types of software may include, but are not limited to, Internet web page crawling software, email virus scanning software, database software, and streaming video content software, in at least one embodiment.

[0248] In at least one embodiment, the applications 1142 included in application layer 1140 may include one or more types of applications used by at least portions of node CRs 1116(1)-1116(N), clustered compute resources 1114, and / or distributed file systems 1128 of framework layer 1120. One or more types of applications may include, in at least one embodiment, any number of a genomics application, a cognitive computation application, and a machine learning application, including, but not limited to, training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in connection with one or more embodiments.

[0249] In at least one embodiment, any of the configuration manager 1124, the resource manager 1126, and the resource orchestrator 1112 may implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible manner. In at least one embodiment, self-modifying actions may relieve a data center operator of the data center 1100 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly performing portions of a data center.

[0250] In at least one embodiment, data center 1100 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weighting parameters according to a neural network architecture using software and computing resources described above with respect to data center 1100.In at least one embodiment, trained machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1100 by using weighting parameters calculated by one or more training techniques described herein.

[0251] In at least one embodiment, the data center may use CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to enable users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0252] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the system of Fig. 11 for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. AUTONOMOUS VEHICLE

[0253] Fig. 12A illustrates an exemplary autonomous vehicle 1200 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1200 (alternatively referred to herein as "vehicle 1200") may be, without limitation, a passenger vehicle, such as a car, a truck, a bus, and / or another type of vehicle that accommodates one or more passengers. In at least one embodiment, the vehicle 1200 may be a semi-trailer truck used to haul cargo. In at least one embodiment, the vehicle 1200 may be an aircraft, a robotic vehicle, or another type of vehicle.

[0254] Autonomous vehicles may be described in terms of automation levels defined by the National Highway Traffic Safety Administration ("NHTSA"), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers ("SAE") "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and prior and future versions of this standard). In at least one embodiment, the vehicle 1200 may be capable of functionality according to one or more of Level 1 through Level 5 of the autonomous driving levels.For example, in at least one embodiment, the vehicle 1200 may be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.

[0255] In at least one embodiment, vehicle 1200 may include, without limitation, components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. In at least one embodiment, vehicle 1200 may include, without limitation, a propulsion system 1250, such as an internal combustion engine, a hybrid electric powerplant, an all-electric motor, and / or another type of propulsion system. In at least one embodiment, propulsion system 1250 may be connected to a drivetrain of vehicle 1200, which may include, without limitation, a transmission, to facilitate propulsion of vehicle 1200. In at least one embodiment, propulsion system 1250 may be controlled in response to receiving signals from throttle / accelerator pedal(s) 1252.

[0256] In at least one embodiment, a steering system 1254, which may include, without limitation, a steering wheel, is used to steer the vehicle 1200 (e.g., along a desired path or route) when the propulsion system 1250 is operating (e.g., when the vehicle 1200 is in motion). In at least one embodiment, the steering system 1254 may receive signals from steering actuator(s) 1256. In at least one embodiment, a steering wheel may be optional for full automation (Level 5) functionality. In at least one embodiment, a brake sensor system 1246 may be used to operate vehicle brakes in response to receiving signals from brake actuator(s) 1248 and / or brake sensors.

[0257] In at least one embodiment, controller(s) 1236, including without limitation one or more systems on chips (“SoCs”) (in Fig. 12A not shown) and / or graphics processing unit(s) ("GPU(s)"), provide signals (e.g., representative of commands) to one or more components and / or systems of the vehicle 1200. For example, in at least one embodiment, the controller(s) 1236 may send signals to operate vehicle brakes via the brake actuator(s) 1248, to operate the steering system 1254 via the steering actuator(s) 1256, to operate the propulsion system 1250 via the throttle / accelerator pedal(s) 1252. In at least one embodiment, the controller(s) 1236 may include one or more on-board (e.g., integrated) computing devices that process sensor signals and issue operational commands (e.g., signals representative of commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 1200.In at least one embodiment, the controller(s) 1236 may include a first controller for autonomous driving functions, a second controller for functional safety functions, a third controller for artificial intelligence functionality (e.g., machine vision), a fourth controller for infotainment functionality, a fifth controller for emergency redundancy, and / or other controllers. In at least one embodiment, a single controller may handle two or more of the above functionalities, two or more controllers may handle a single functionality, and / or any combination thereof.

[0258] In at least one embodiment, the controller(s) 1236 provide signals to control one or more components and / or systems of the vehicle 1200 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data may be received, for example and without limitation, from global navigation satellite systems (“GNSS”) sensor(s) 1258 (e.g., global positioning system sensor(s), RADAR sensor(s) 1260, ultrasonic sensor(s) 1262, LIDAR sensor(s) 1264, inertial measurement unit (“IMU”) sensor(s) 1266 (e.g., accelerometer(s), gyroscope(s), magnetic compass(es), magnetometer(s), etc.), microphone(s) 1296, stereo camera(s) 1268, wide-view camera(s) 1270 (e.g., fisheye cameras), infrared camera(s) 1272, surround camera(s) 1274 (e.g.,360-degree cameras), long-range cameras (in . Fig. 12A not shown), medium-range camera(s) (in Fig. 12A not shown), speed sensor(s) 1244 (e.g., for measuring the speed of the vehicle 1200), vibration sensor(s) 1242, steering sensor(s) 1240, brake sensor(s) (e.g., as part of the brake sensor system 1246), and / or other types of sensors.

[0259] In at least one embodiment, one or more controllers 1236 may receive inputs (e.g., in the form of input data) from an instrument cluster 1232 of a vehicle 1200 and provide outputs (e.g., in the form of output data, display data, etc.) via a human-machine interface (“HMI”) display 1234, an audible annunciator, a speaker, and / or via other components of a vehicle 1200. In at least one embodiment, the outputs may include information such as vehicle speed, velocity, time, map data (e.g., a high-resolution map (in Fig. 12A not shown)), location data (e.g., location of vehicle 1200, e.g., on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and status of objects as perceived by controllers 1236, etc. For example, in at least one embodiment, HMI display 1234 may display information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about maneuvers the vehicle has performed, is currently performing, or will perform (e.g., change lanes now, take exit 34B in two miles, etc.).

[0260] In one embodiment, vehicle 1200 further includes a network interface 1224 that may utilize wireless antennas 1226 and / or modems to communicate over one or more networks. For example, in at least one embodiment, network interface 1224 may be capable of communicating over Long-Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communication ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000"), etc. networks. In at least one embodiment, wireless antenna(s) 1226 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area network(s), such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, etc., and / or low power wide-area networks (“LPWANs”), such as LoRaWAN protocols, SigFox protocols, etc.

[0261] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the system of Fig. 12A for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0262] Fig. 12B illustrates an example of camera locations and fields of view for the autonomous vehicle 1200 of Fig. 12A according to at least one embodiment. In at least one embodiment, the cameras and respective fields of view represent an example embodiment and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located at different locations on the vehicle 1200.

[0263] In at least one embodiment, camera types may include, but are not limited to, digital cameras configured for use with components and / or systems of the vehicle 1200. In at least one embodiment, the camera(s) may operate at automotive safety integrity level (“ASIL”) B and / or another ASIL. In at least one embodiment, the camera types may achieve any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the cameras may use rolling shutter, global shutter, another shutter type, or a combination thereof.In at least one embodiment, the color filter array may include a red-clear-clear-clear color filter array ("RCCC"), a red-clear-clear-blue color filter array ("RCCB"), a red-blue-green-clear color filter array ("RBGC"), a Foveon X3 color filter array, a Bayer sensor color filter array ("RGGB"), a monochrome sensor color filter array, and / or another type of color filter array. In at least one embodiment, clear-pixel cameras, such as cameras with an RCCC, an RCCB, and / or an RBGC color filter array, may be used to increase light sensitivity.

[0264] In at least one embodiment, one or more cameras may be used to perform Advanced Driver Assistance Systems (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function mono camera may be installed, providing functions such as lane departure warning, traffic sign assist, and intelligent headlight control. In at least one embodiment, one or more of the camera(s) (e.g., all cameras) may simultaneously capture and provide image data (e.g., video).

[0265] In at least one embodiment, one or more cameras may be mounted in a mounting assembly, such as a custom-designed (three-dimensionally ("3D") printed) assembly, to cut out stray light and reflections from inside the vehicle 1200 (e.g., reflections from the dashboard reflected in the windshield mirrors) that may impair the cameras' image data collection capabilities. With reference to side view mirror mounting assemblies, in at least one embodiment, the side view mirror assemblies may be custom 3D printed so that a camera mounting plate conforms to a shape of a side view mirror. In at least one embodiment, the camera(s) may be integrated into the outside view mirrors. In at least one embodiment, for side view cameras, the camera(s) may also be integrated within four pillars at each corner of a cab.

[0266] In at least one embodiment, cameras with a field of view that includes portions of an environment in front of a vehicle 1200 (e.g., forward-facing cameras) may be used for the environmental view to help identify forward paths and obstacles, and to provide information critical to establishing an occupancy grid and / or determining preferred vehicle paths with the aid of one or more controllers 1236 and / or control SoCs. In at least one embodiment, forward-facing cameras may be used to perform many similar ADAS functions as LIDAR, including, without limitation, emergency braking, pedestrian detection, and collision avoidance.In at least one embodiment, forward-facing cameras may also be used for ADAS features and systems such as, without limitation, Lane Departure Warnings (“LDW”), Autonomous Cruise Control (“ACC”), and / or other features such as traffic sign recognition.

[0267] In at least one embodiment, a variety of cameras may be used in a forward-facing configuration, including, for example, a monocular camera platform incorporating a CMOS (complementary metal oxide semiconductor) color image sensor. In at least one embodiment, a wide-view camera 1270 may be used to perceive objects coming into view from a periphery (e.g., pedestrians, crossing traffic, or bicycles). Although in Fig. 12B illustrates only one wide-view camera 1270, in other embodiments, any number (including zero) of wide-view cameras may be present on the vehicle 1200. In at least one embodiment, any number of long-range camera(s) 1298 (e.g., a wide-view stereo camera pair) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. In at least one embodiment, the long-range camera(s) 1298 may also be used for object detection and classification, as well as basic object tracking.

[0268] In at least one embodiment, any number of the stereo camera(s) 1268 may also be included in a forward-facing configuration. In at least one embodiment, one or more of the stereo camera(s) 1268 may include an integrated controller unit comprising a scalable processing unit that may provide a field-programmable logic (“FPGA”) and a multi-core microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of an environment of the vehicle 1200, including a distance estimate for all points in an image.In at least one embodiment, one or more of the stereo camera(s) 1268 may include, without limitation, compact stereo vision sensors, which may include, without limitation, two camera lenses (one each on the left and right) and an image processing chip that can measure the distance from the vehicle 1200 to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo camera(s) 1268 may be used in addition to or alternatively to those described herein.

[0269] In at least one embodiment, cameras with a field of view that includes portions of the environment to the side of the vehicle 1200 (e.g., side view cameras) may be used for surround view, thereby providing information used to create and update an occupancy grid and to generate side impact collision warnings. For example, in at least one embodiment, the surround camera(s) 1274 (e.g., four surround cameras, as in Fig. 12B) may be positioned on the vehicle 1200. In at least one embodiment, the surround-view camera(s) 1274 may include, without limitation, any number and combination of wide-view cameras, fisheye cameras, 360-degree cameras, and / or similar cameras. For example, in at least one embodiment, four fisheye cameras may be positioned on a front, rear, and sides of the vehicle 1200. In at least one embodiment, the vehicle 1200 may utilize three surround-view cameras 1274 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.

[0270] In at least one embodiment, cameras with a field of view that includes portions of an environment behind the vehicle 1200 (e.g., rearview cameras) may be used for parking assistance, surround view, rear collision warnings, and for creating and updating an occupancy grid. In at least one embodiment, a wide variety of cameras may be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., long-range cameras 1298 and / or mid-range camera(s) 1276, stereo camera(s) 1268, infrared camera(s) 1272, etc.), as described herein.

[0271] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the system of Fig. 12B for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0272] Fig. 12C is a block diagram illustrating an example system architecture for the autonomous vehicle 1200 of Fig. 12A, according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 1200 is Fig. 12C as connected via a bus 1202. In at least one embodiment, bus 1202 may include, without limitation, a CAN data interface (alternatively referred to herein as a "CAN bus"). In at least one embodiment, a CAN may be a network within vehicle 1200 used to assist in controlling various features and functions of vehicle 1200, such as application of brakes, acceleration, braking, steering, windshield wipers, etc. In at least one embodiment, bus 1202 may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). In at least one embodiment, bus 1202 may be read to determine steering wheel angle, ground speed, engine revolutions per minute (RPMs), button positions, and / or other vehicle status indicators.In at least one embodiment, bus 1202 may be a CAN bus that is ASIL B compliant.

[0273] In at least one embodiment, FlexRay and / or Ethernet protocols may be used in addition to or alternatively to CAN. In at least one embodiment, any number of buses may be present that comprise bus 1202, which may include, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using different protocols. In at least one embodiment, two or more buses may be used to perform different functions and / or used for redundancy. For example, a first bus may be used for collision avoidance functionality and a second bus may be used for actuation control.In at least one embodiment, each bus 1202 may communicate with any component of a vehicle 1200, and two or more buses 1202 may communicate with the same components. In at least one embodiment, each of any number of system(s) on chip(s) ("SoC(s)") 1204 (such as SoC 1204(A) and SoC 1204(B)), each of the controller(s) 1236, and / or each computer within the vehicle may have access to the same input data (e.g., inputs from sensors of the vehicle 1200) and be connected to a common bus, such as the CAN bus.

[0274] In at least one embodiment, the vehicle 1200 may include one or more controllers 1236, such as those described herein with respect to Fig. 12A. In at least one embodiment, the controller(s) 1236 may be used for a variety of functions. In at least one embodiment, the controller(s) 1236 may be coupled to any of various other components and systems of the vehicle 1200 and used to control the vehicle 1200, the artificial intelligence of the vehicle 1200, the infotainment for the vehicle 1200, and / or other functions.

[0275] In at least one embodiment, the vehicle 1200 may include any number of SoCs 1204. In at least one embodiment, each of the SoCs 1204 may include, without limitation, central processing units ("CPU(s)") 1206, graphics processing units ("GPU(s)") 1208, processor(s) 1210, cache(s) 1212, one or more accelerators 1214, one or more data memories 1216, and / or other components and features not illustrated. In at least one embodiment, the SoC(s) 1204 may be used to control the vehicle 1200 in a variety of platforms and systems. For example, in at least one embodiment, the SoC(s) 1204 may be combined in a system (e.g., system of the vehicle 1200) with a high definition (“HD”) map 1222 that receives map refreshes and / or updates via the network interface 1224 from one or more servers (in Fig. 12C not shown).

[0276] In at least one embodiment, the CPU(s) 1206 may include a CPU cluster or CPU complex (alternatively referred to herein as a "CCPLEX"). In at least one embodiment, the CPU(s) 1206 may include multiple cores and / or level two ("L2") caches. For example, in at least one embodiment, the CPU(s) 1206 may include eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU(s) 1206 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 megabyte (MB) L2 cache). In at least one embodiment, the CPU(s) 1206 (e.g., CCPLEX) may be configured to support simultaneous cluster operations, such that any combination of clusters of the CPU(s) 1206 may be active at any given time.

[0277] In at least one embodiment, one or more of the CPU(s) 1206 may implement power management capabilities, including, without limitation, one or more of the following features: individual hardware blocks may be automatically clock controlled when they are idle to conserve dynamic power; each core clock may be controlled when such core is not actively executing instructions due to the execution of wait-for-interrupt ("WFI") / wait-for-event ("WFE") instructions; each core may be independently power controlled; each core cluster may be independently clock controlled if all cores are clock controlled or power controlled; and / or each core cluster may be independently power controlled if all cores are power controlled.In at least one embodiment, the CPU(s) 1206 may further implement an enhanced power state management algorithm, where allowable power states and expected wake-up times are specified, and the hardware / microcode determines the best power state to enter for a core, a cluster, and a CCPLEX. In at least one embodiment, the processing cores may support simplified power state entry sequences in software, offloading work to microcode.

[0278] In at least one embodiment, the GPU(s) 1208 may include an integrated GPU (alternatively referred to herein as an "iGPU"). In at least one embodiment, the GPU(s) 1208 may be programmable and efficient for parallel workloads. In at least one embodiment, the GPU(s) 1208 may use an enhanced Tensor instruction set. In at least one embodiment, the GPU(s) 1208 may include one or more streaming microprocessors, where each streaming microprocessor may include a level one ("L1") cache (e.g., an L1 cache with a memory capacity of at least 96 KB), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with a memory capacity of 512 KB). In at least one embodiment, the GPU(s) 1208 may include at least eight streaming microprocessors.In at least one embodiment, the GPU(s) 1208 may utilize computational application programming interface(s) (API(s)). In at least one embodiment, the GPU(s) 1208 may utilize one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA model).

[0279] In at least one embodiment, one or more of the GPU(s) 1208 may be power-optimized for best computing performance in automotive and embedded use cases. For example, in at least one embodiment, the GPU(s) 1208 may be fabricated on a fin field-effect transistor ("FinFET") circuit. In at least one embodiment, each streaming microprocessor may include a number of mixed-precision processing cores partitioned into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores may be partitioned into four processing blocks. In at least one embodiment, each processing block may be assigned 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two NVIDIA mixed-precision Tensor Cores for deep learning matrix arithmetic, a level zero ("L0") instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file.In at least one embodiment, streaming microprocessors may include independent parallel integer and floating-point datapaths to enable efficient execution of workloads with a mix of computation and addressing calculations. In at least one embodiment, streaming microprocessors may include independent thread scheduling to enable finer-grained synchronization and collaboration between parallel threads. In at least one embodiment, streaming microprocessors may include a combined L1 data cache and shared memory unit to improve performance and simplify programming.

[0280] In at least one embodiment, one or more of the GPU(s) 1208 may include high bandwidth memory (“HBM”) and / or a 16 GB HBM2 memory subsystem to provide, in some examples, a peak memory bandwidth of approximately 900 GB / second. In at least one embodiment, in addition to or alternatively to HBM memory, a synchronous graphics random-access memory (“SGRAM”) may be used, such as a graphics double data rate type five (“GDDR5”) synchronous random-access memory.

[0281] In at least one embodiment, GPU(s) 1208 may include unified memory technology. In at least one embodiment, address translation services ("ATS") support may be used to enable GPU(s) 1208 to directly access page tables of CPU(s) 1206. In at least one embodiment, if the memory management unit (MMU) of a GPU of GPU(s) 1208 experiences a failure, an address translation request may be transmitted to CPU(s) 1206. In response, two CPUs of CPU(s) 1206, in at least one embodiment, may look in their page tables for a virtual-to-physical mapping for an address and transmit the translation back to GPU(s) 1208.In at least one embodiment, the unified memory technology may enable a single unified virtual address space for memory of both the CPU(s) 1206 and the GPU(s) 1208, thereby simplifying programming of the GPU(s) 1208 and porting of applications to the GPU(s) 1208.

[0282] In at least one embodiment, the GPU(s) 1208 may include any number of access counters that may track the frequency of access by the GPU(s) 1208 to memory of other processors. In at least one embodiment, the access counter(s) may help ensure that memory pages are moved to physical memory of a processor that accesses pages most frequently, thereby improving efficiency for memory regions shared by multiple processors.

[0283] In at least one embodiment, one or more of the SoC(s) 1204 may include any number of cache(s) 1212, including those described herein. In at least one embodiment, the cache(s) 1212 may include, for example, a level three ("L3") cache available to both the CPU(s) 1206 and the GPU(s) 1208 (e.g., connected to the CPU(s) 1206 and GPU(s) 1208). In at least one embodiment, the cache(s) 1212 may include a write-back cache that may track the states of lines, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, an L3 cache may include 4 MB of memory or more, depending on the embodiment, although smaller cache sizes may also be used.

[0284] In at least one embodiment, one or more of the SoC(s) 1204 may include one or more accelerators 1214 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoC(s) 1204 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, the large on-chip memory (e.g., 4 MB SRAM) may enable a hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, a hardware acceleration cluster may be used to supplement the GPU(s) 1208 and offload some tasks of the GPU(s) 1208 (e.g., free up more cycles of the GPU(s) 1208 to perform other tasks). In at least one embodiment, the accelerator(s) 1214 could be used for targeted workloads (e.g.,Perception, convolutional neural networks (CNNs), recurrent neural networks (RNNs), etc.) that are robust enough to be suitable for acceleration may be used. In at least one embodiment, a CNN may include region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (such as those used for object detection), or another type of CNN.

[0285] In at least one embodiment, the accelerator(s) 1214 (e.g., hardware acceleration clusters) may include one or more deep learning accelerators (“DLAs”). In at least one embodiment, DLA(s) may include, without limitation, one or more tensor processing units (“TPUs”) that may be configured to provide an additional tens of trillion operations per second for deep learning applications and inference. In at least one embodiment, TPUs may be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, the DLA(s) may further be optimized for a particular set of neural network types and floating-point operations, as well as for inferencing.In at least one embodiment, the design of the DLAs can provide more performance per millimeter than a typical general-purpose GPU and typically significantly exceeds the performance of a CPU. In at least one embodiment, the TPUs can perform multiple functions, including a single-instance convolution function, using, for example, INT8, INT16, and FP16 data types for features and weights, as well as post-processing functions.In at least one embodiment, the DLAs can quickly and efficiently execute neural networks, particularly CNNs, on processed or unprocessed data for a variety of functions, including, for example and without limitation: a CNN for object identification and recognition using data from camera sensors; a CNN for distance estimation using data from camera sensors; a CNN for emergency vehicle detection and identification and recognition using data from microphones; a CNN for facial recognition and vehicle owner identification using data from camera sensors; and / or a CNN for safety and / or security-related events.

[0286] In at least one embodiment, the DLA(s) may perform any function of the GPU(s) 1208, and by using an inference accelerator, a designer may, for example, target either the DLA(s) or the GPU(s) 1208 for any function. For example, in at least one embodiment, a designer may focus on processing CNNs and floating-point operations on the DLA(s) and leave other functions to the GPU(s) 1208 and / or accelerator(s) 1214.

[0287] In at least one embodiment, the accelerator(s) 1214 may include a programmable vision accelerator (“PVA”), which may alternatively be referred to herein as a machine vision accelerator. In at least one embodiment, the PVA may be designed and configured to accelerate machine vision algorithms for advanced driver assistance systems (“ADAS”) 1238, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. In at least one embodiment, the PVA may provide a balance between computational power and flexibility. In at least one embodiment, each PVA may include, for example and without limitation, any number of reduced instruction set (“RISC”) cores, direct memory access (“DMA”) cores, and / or any number of vector processors.

[0288] In at least one embodiment, RISC cores may interact with image sensors (e.g., the image sensors of the cameras described herein), image signal processors, etc. In at least one embodiment, each RISC core may include any amount of memory. In at least one embodiment, the RISC cores may use any number of protocols, depending on the embodiment. In at least one embodiment, RISC cores may execute a real-time operating system ("RTOS"). In at least one embodiment, RISC cores may be implemented with one or more integrated circuit devices, application-specific integrated circuits ("ASICs"), and / or memory devices. In at least one embodiment, the RISC cores could include, for example, an instruction cache and / or tightly coupled RAM.

[0289] In at least one embodiment, DMA may enable components of the PVA to access system memory independently of the CPU(s) 1206. In at least one embodiment, DMA may support any number of features used to provide optimization of a PVA, including, but not limited to, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, DMA may support up to six or more dimensions of addressing, which may include, without limitation, block width, block height, block depth, horizontal block gradation, vertical block gradation, and / or depth gradation.

[0290] In at least one embodiment, vector processors may be programmable processors that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing functions. In at least one embodiment, a PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, a PVA core may include a processor subsystem, DMA engine(s) (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, a vector processing subsystem may act as a primary processing unit of a PVA and include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, a VPU core may include a digital signal processor, such as aa Single Instruction Multiple Data ("SIMD"), Very Long Instruction Word ("VLIW") digital signal processor. In at least one embodiment, a combination of SIMD and VLIW can increase throughput and speed.

[0291] In at least one embodiment, each vector processor may include an instruction cache and be connected to dedicated memory. As a result, in at least one embodiment, each of the vector processors may be configured to operate independently of other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute a general computer vision algorithm, but for different image regions. In at least one embodiment, vector processors included in a particular PVA may concurrently execute different image processing algorithms for an image, or even different algorithms for successive images or portions of an image.In at least one embodiment, among other things, there may be any number of PVAs in a hardware acceleration cluster and any number of vector processors in each PVA. In at least one embodiment, the PVA may include additional memory for error correcting code (ECC) to increase overall system security.

[0292] In at least one embodiment, the accelerator(s) 1214 may include an on-chip computer vision network and static random-access memory (“SRAM”) to provide high-bandwidth, low-latency SRAM for the accelerator(s) 1214. In at least one embodiment, on-chip memory may include at least 4 MB of SRAM, including, for example, and without limitation, eight field-configurable memory blocks accessible by both a PVA and a DLA. In at least one embodiment, each pair of memory blocks may include an extended peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used.In at least one embodiment, a PVA and a DLA may access memory via a backbone that provides high-speed access to memory to a PVA and a DLA. In at least one embodiment, a backbone may include an on-chip computer vision network that interconnects a PVA and a DLA with memory (e.g., using an APB).

[0293] In at least one embodiment, an on-chip computer vision network may include an interface that determines that both a PVA and a DLA are providing ready and valid signals before transmitting any control signal / address / data. In at least one embodiment, an interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data delivery. In at least one embodiment, an interface may conform to International Organization for Standardization ("ISO") 26262 or International Electrotechnical Commission ("IEC") 61508 standards, although other standards and protocols may be used.

[0294] In at least one embodiment, one or more of the SoC(s) 1204 may include a real-time ray tracing hardware accelerator. In at least one embodiment, the real-time ray tracing hardware accelerator may be used to quickly and efficiently determine positions and extents of objects (e.g., within a world model) to generate real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for simulation of sonar systems, for general wave propagation simulation, for comparison with lidar data for localization, and / or for other functions, and / or for other uses.

[0295] In at least one embodiment, the accelerator(s) 1214 may have a wide range of uses for autonomous driving. In at least one embodiment, a PVA may be used for critical processing stages in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of a PVA are well suited to algorithmic domains that require predictable processing with low power and low latency. In other words, a PVA demonstrates good computational performance for semi-dense or dense regular computations, even on small datasets, which may require predictable runtimes with low latency and low power. In at least one embodiment, such as in the vehicle 1200, the PVAs could be configured to execute classical computer vision algorithms, as they can be efficient at object detection and operating on integer mathematics.

[0296] For example, according to at least one embodiment of the technology, a PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching-based algorithm may be used in some examples, although this is not intended to be limiting. In at least one embodiment, Level 3-5 autonomous driving applications utilize motion estimation / stereo matching on the fly (e.g., structure from motion, pedestrian detection, lane detection, etc.). In at least one embodiment, a PVA may perform computer stereo vision functions on inputs from two monocular cameras.

[0297] In at least one embodiment, a PVA may be used to perform dense optical flow. For example, in at least one embodiment, a PVA may process raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar data. In at least one embodiment, a PVA is used for depth-of-flight processing, processing raw time-of-flight data to provide, for example, processed time-of-flight data.

[0298] In at least one embodiment, a DLA may be used to power any type of network to improve control and driving safety, including, for example and without limitation, a neural network that outputs a confidence measure for each object detection. In at least one embodiment, the confidence may be represented or interpreted as a probability, or as providing a relative "weight" of each detection compared to other detections. In at least one embodiment, a confidence measure further allows the system to make decisions about which detections should be considered true positives and which should be considered false positives. In at least one embodiment, a system may set a confidence threshold and consider only detections that exceed the threshold to be true positives.In an embodiment where an automatic emergency braking ("AEB") system is used, false positive detections would cause the vehicle to automatically perform emergency braking, which is clearly undesirable. In at least one embodiment, high-confidence detections may be considered triggers for AEB. In at least one embodiment, a DLA may execute a neural network to regress the confidence value. In at least one embodiment, the neural network may use as its input at least a subset of parameters, such as the bounding box dimensions, the ground plane estimate obtained (e.g., from another subsystem), the output of IMU sensor(s) 1266 correlating with the orientation of the vehicle 1200, the distance, the 3D location estimates of the object obtained from the neural network and / or other sensors (e.g., LIDAR sensor(s) 1264 or RADAR sensor(s) 1260), and others.

[0299] In at least one embodiment, one or more of the SoC(s) 1204 may include one or more data stores 1216 (e.g., memory). In at least one embodiment, the data stores 1216 may be on-chip memory of the SoC(s) 1204, which may store neural networks to be executed on the GPU(s) 1208 and / or a DLA. In at least one embodiment, the capacity of the data stores 1216 may be large enough to store multiple instances of neural networks for redundancy and security. In at least one embodiment, the data stores 1216 may include L2 or L3 cache(s).

[0300] In at least one embodiment, one or more of the SoC(s) 1204 may include any number of processor(s) 1210 (e.g., embedded processors). In at least one embodiment, the processor(s) 1210 may include a booting and power management processor, which may be a dedicated processor and subsystem to handle booting power and management functions and associated security enforcement. In at least one embodiment, the booting and power management processor may be part of a booting sequence of the SoC(s) 1204 and provide runtime power management services. In at least one embodiment, a booting power and management processor may provide clock and voltage programming, assistance with system transitions to a low-power state, management of thermal and temperature sensors of the SoC(s) 1204, and / or management of power states of the SoC(s) 1204.In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC(s) 1204 may use ring oscillators to detect temperatures of CPU(s) 1206, GPU(s) 1208, and / or accelerator(s) 1214. If temperatures are determined to exceed a threshold, in at least one embodiment, a booting and power management processor may then enter a temperature fault routine and place the SoC(s) 1204 into a lower power state and / or place the vehicle 1200 into a drive-to-safe-stop mode (e.g., bring the vehicle 1200 to a safe stop).

[0301] In at least one embodiment, processor(s) 1210 may further include a set of embedded processors that may serve as an audio processing engine, which may be an audio subsystem enabling full hardware support for multi-channel audio across multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In at least one embodiment, an audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0302] In at least one embodiment, the processor(s) 1210 may further include an always-on processor engine that may provide necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, the always-on processor engine may include, without limitation, a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O control peripherals, and routing logic.

[0303] In at least one embodiment, processor(s) 1210 may further include a security cluster engine, including, without limitation, a dedicated processor subsystem for handling security management for automotive applications. In at least one embodiment, a security cluster engine may include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, an interrupt controller, etc.), and / or routing logic. In a security mode, two or more cores, in at least one embodiment, may operate in a lockstep mode, functioning as a single core with comparison logic to detect any differences between their operations.In at least one embodiment, processor(s) 1210 may further include a real-time camera engine, which may include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, processor(s) 1210 may further include a high dynamic range signal processor, which may include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.

[0304] In at least one embodiment, processor(s) 1210 may include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate a final image for a playback program window. In at least one embodiment, a video image compositor may perform lens distortion correction on the wide-view camera(s) 1270, surround camera(s) 1274, and / or in-cabin surveillance camera sensor(s). In at least one embodiment, the in-cabin surveillance camera sensor(s) are preferably monitored by a neural network running on another instance of SoC 1204 and configured to detect and respond to events in the cabin.In at least one embodiment, an in-cabin system may perform, without limitation, lip reading to activate cellular service and place a call, dictate emails, change a vehicle destination, activate or change a vehicle infotainment system and its settings, or provide voice-activated web browsing. In at least one embodiment, certain features are available to a driver when a vehicle is operating in an autonomous mode and are disabled otherwise.

[0305] In at least one embodiment, a video image compositor may include advanced temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, when motion occurs in a video, the noise reduction appropriately weights the spatial information and reduces weights of information provided by neighboring frames. In at least one embodiment, where an image or a portion of an image has no motion, the temporal noise reduction performed by the video image compositor may use information from a previous frame to reduce noise in the current frame.

[0306] In at least one embodiment, a video image compositor may also be configured to perform stereo dewarping on the input stereo lens frames. In at least one embodiment, a video image compositor may further be used for user interface composition when an operating system desktop is in use and the GPU(s) 1208 are not required to continuously render new surfaces. When the GPU(s) 1208 are powered on and actively performing 3D rendering, in at least one embodiment, a video image compositor may be used to offload the GPU(s) 1208 to improve computational performance and responsiveness.

[0307] In at least one embodiment, one or more of the SoC(s) 1204 may further include a Mobile Industry Processor Interface ("MIPI") serial camera interface for receiving video and input from cameras, a high-speed interface, and / or a video input block that may be used for a camera and related pixel input functions. In at least one embodiment, one or more of the SoC(s) 1204 may further include input / output controller(s) that may be controlled by software and may be used to receive I / O signals that are not assigned to a specific role.

[0308] In at least one embodiment, one or more of the SoC(s) 1204 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio encoders / decoders ("codecs"), power management, and / or other devices. In at least one embodiment, the SoC(s) 1204 may be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet channels), sensors (e.g., LIDAR sensor(s) 1264, RADAR sensor(s) 1260, etc., which may be connected via Ethernet channels), data from the bus 1202 (e.g., speed of the vehicle 1200, steering wheel position, etc.), data from GNSS sensor(s) 1258 (e.g., connected via an Ethernet bus or a CAN bus), etc.In at least one embodiment, one or more of the SoC(s) 1204 may further include dedicated high-performance mass storage controllers, which may include their own DMA engines and which may be used to offload routine data management tasks from the CPU(s) 1206.

[0309] In at least one embodiment, the SoC(s) 1204 may be an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture, exploiting and efficiently deploying computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible, reliable driving software stack along with deep learning tools. In at least one embodiment, the SoC(s) 1204 may be faster, more reliable, and even more power and space efficient than conventional systems. For example, in at least one embodiment, the accelerator(s) 1214, when combined with the CPU(s) 1206, GPU(s) 1208, and data storage(s) 1216, may provide a fast, efficient platform for Levels 3-5 autonomous vehicles.

[0310] In at least one embodiment, computer vision algorithms may be executed on CPUs that can be configured using a high-level programming language, such as C, to perform a wide variety of processing algorithms on a wide variety of visual data. However, in at least one embodiment, the CPUs are often unable to meet the computational performance requirements of many computer vision applications, such as execution time and power consumption. In at least one embodiment, many CPUs are unable to execute complex, real-time object detection algorithms used in in-vehicle ADAS applications and practical Level 3-5 autonomous vehicles.

[0311] The embodiments described herein enable multiple neural networks to be executed simultaneously and / or sequentially and the results combined to enable Levels 3-5 autonomous driving functionality. For example, in at least one embodiment, a CNN executing on a DLA or a discrete GPU (e.g., GPU(s) 1220) may include text and word recognition that enables the reading and understanding of traffic signs, including signs for which a neural network has not been specifically trained. In at least one embodiment, a DLA may further include a neural network capable of identifying, interpreting, and providing semantic understanding of a sign and passing this semantic understanding to path planning modules running on a CPU complex.

[0312] In at least one embodiment, multiple neural networks may be executed simultaneously, such as for driving at Level 3, 4, or 5. For example, in at least one embodiment, a warning sign reading "Caution: Flashing lights indicate icing" along with an electric light may be interpreted independently or jointly by multiple neural networks. In at least one embodiment, such a warning sign may itself be identified as a traffic sign by a first deployed neural network (e.g., a neural network that has been trained), and a text "Flashing lights indicate icing" may be interpreted by a second deployed neural network, which informs vehicle path planning software (preferably executing on a CPU complex) that, when flashing lights are detected, icy conditions exist.In at least one embodiment, a flashing light may be identified by running a third deployed neural network across multiple frames, informing vehicle path planning software of the presence (or absence) of flashing lights. In at least one embodiment, all three neural networks may run simultaneously, such as within a DLA and / or on GPU(s) 1208.

[0313] In at least one embodiment, a CNN for facial recognition and vehicle owner identification may use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 1200. In at least one embodiment, an always-on sensor processing engine may be used to unlock a vehicle when an owner approaches a driver-side door and turns on lights, and to disable such a vehicle in a security mode when an owner exits such a vehicle. In this way, the SoC(s) 1204 provide security against theft and / or carjacking.

[0314] In at least one embodiment, a CNN for detecting and identifying emergency vehicles may use data from microphones 1296 to detect and identify emergency vehicle sirens. In at least one embodiment, the SoC(s) 1204 use a CNN to classify ambient and urban noise, as well as to classify visual data. In at least one embodiment, a CNN running on a DLA is trained to identify a relative approach speed of an emergency vehicle (e.g., by using a Doppler effect). In at least one embodiment, a CNN may also be trained to identify emergency vehicles specific to a local area in which a vehicle is operating, as identified by the GNSS sensor(s) 1258.In at least one embodiment, when operating in Europe, a CNN attempts to detect European sirens, and in North America, a CNN attempts to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program may be used to execute an emergency vehicle safety routine with the aid of the ultrasonic sensor(s) 1262 to slow a vehicle, pull over to the side of the road, park a vehicle, and / or idle a vehicle until the emergency vehicles have passed.

[0315] In at least one embodiment, the vehicle 1200 may include CPU(s) 1218 (e.g., discrete CPU(s) or dCPU(s)) that may be coupled to the SoC(s) 1204 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, the CPU(s) 1218 may include, for example, an x86 processor. The CPU(s) 1218 may be used to perform any of a variety of functions, including, for example, mediating potentially inconsistent results between ADAS sensors and SoC(s) 1204 and / or monitoring the status and health of the controller(s) 1236 and / or an infotainment system on a chip ("infotainment SoC") 1230.

[0316] In at least one embodiment, the vehicle 1200 may include GPU(s) 1220 (e.g., discrete GPU(s) or dGPU(s)) that may be coupled to the SoC(s) 1204 via a high-speed interconnect (e.g., NVIDIA's NVLINK channel). In at least one embodiment, the GPU(s) 1220 may provide additional functionality for artificial intelligence, such as by executing redundant and / or distinct neural networks, and may be used to train and / or update neural networks based at least in part on inputs (e.g., sensor data) from sensors of a vehicle 1200.

[0317] In at least one embodiment, vehicle 1200 may further include network interface 1224, which may include, without limitation, wireless antenna(s) 1226 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, network interface 1224 may be used to enable wireless connectivity to Internet cloud services (e.g., to server(s) and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). In at least one embodiment, a direct link may be established between vehicle 120 and another vehicle and / or an indirect link may be established (e.g., via networks and over the Internet) to communicate with other vehicles.In at least one embodiment, direct links may be provided using a vehicle-to-vehicle communication link. In at least one embodiment, a vehicle-to-vehicle communication link may provide information to the vehicle 1200 about vehicles in the vicinity of the vehicle 1200 (e.g., vehicles in front of, beside, and / or behind the vehicle 1200). In at least one embodiment, such aforementioned functionality may be part of a cooperative adaptive cruise control functionality of the vehicle 1200.

[0318] In at least one embodiment, the network interface 1224 may include an SoC that provides modulation and demodulation functionality and enables the controller(s) 1236 to communicate over wireless networks. In at least one embodiment, the network interface 1224 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. In at least one embodiment, frequency conversions may be performed in any technically feasible manner. For example, frequency conversions may be performed by known methods and / or using superheterodyne techniques. In at least one embodiment, the radio frequency front end functionality may be provided by a separate chip.In at least one embodiment, the network interfaces may include wireless functionality for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0319] In at least one embodiment, the vehicle 1200 may further include one or more data stores 1228, which may include, without limitation, off-chip memory (e.g., located outside of the SoC(s) 1204). In at least one embodiment, the data stores 1228 may include, without limitation, one or more memory elements, including RAM, SRAM, dynamic random-access memory (“DRAM”), video random-access memory (“VRAM”), flash memory, hard drives, and / or other components and / or devices capable of storing at least one bit of data.

[0320] In at least one embodiment, the vehicle 1200 may further include GNSS sensors 1258 (e.g., GPS and / or assisted GPS sensors) to assist with mapping, sensing, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1258 may be used, including, for example, and without limitation, a GPS using a USB connector with an Ethernet-to-serial bridge (e.g., RS-232 bridge).

[0321] In at least one embodiment, the vehicle 1200 may further include RADAR sensor(s) 1260. In at least one embodiment, the RADAR sensor(s) 1260 may be used by the vehicle 1200 for long-range vehicle detection, even in darkness and / or adverse weather conditions. In at least one embodiment, the RADAR functional safety level may be ASIL B. In at least one embodiment, the RADAR sensors 1260 may use a CAN bus and / or the bus 1202 (e.g., for transmitting the data generated by the RADAR sensors 1260) to control and access object tracking data, with access to Ethernet channels for accessing raw data in some examples. In at least one embodiment, a wide variety of RADAR sensor types may be used. For example, and without limitation, the RADAR sensor(s) 1260 may be suitable for use as front, rear, and side RADAR.In at least one embodiment, one or more of the RADAR sensor(s) 1260 is a pulse Doppler RADAR sensor.

[0322] In at least one embodiment, the RADAR sensor(s) 1260 may include different configurations, such as long range and narrow field of view, short range and wide field of view, short range side coverage, etc. In at least one embodiment, the long range RADAR may be used for adaptive cruise control functionality. In at least one embodiment, long range RADAR systems may provide a wide field of view realized by two or more independent scans, such as within a range of 250 m (meters). In at least one embodiment, the RADAR sensor(s) 1260 may help distinguish between static and moving objects and may be used by the ADAS system 1238 for emergency braking and forward collision warning.In at least one embodiment, the sensor(s) 1260 included in a long-range radar system may include, without limitation, a monostatic multimodal radar with multiple (e.g., six or more) fixed radar antennas and a high-speed CAN and FlexRay interface. In at least one embodiment with six antennas, four central antennas may create a focused beam pattern configured to record the surroundings of the vehicle 1200 at higher speeds with minimal interference from traffic in adjacent lanes. In at least one embodiment, two additional antennas may expand the field of view, allowing for rapid detection of vehicles entering or exiting a lane of the vehicle 1200.

[0323] For example, in at least one embodiment, medium-range radar systems may include a range of up to 160 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, short-range radar systems may include, without limitation, any number of radar sensors 1260 configured for installation at both ends of a rear bumper. When installed at both ends of a rear bumper, the radar sensor system may, in at least one embodiment, generate two beams that constantly monitor blind spots in a rearward direction and adjacent to a vehicle. In at least one embodiment, short-range radar systems may be used in the ADAS system 1238 for blind spot detection and / or lane change assistance.

[0324] In at least one embodiment, the vehicle 1200 may further include ultrasonic sensor(s) 1262. In at least one embodiment, the ultrasonic sensor(s) 1262, which may be positioned at a front, rear, and / or side location of the vehicle 1200, may be used for parking assistance and / or for creating and updating an occupancy grid. In at least one embodiment, a wide variety of ultrasonic sensor(s) 1262 may be used, and different ultrasonic sensor(s) 1262 may be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, the ultrasonic sensor(s) 1262 may operate at functional safety levels of ASIL B.

[0325] In at least one embodiment, the vehicle 1200 may include LIDAR sensor(s) 1264. In at least one embodiment, the LIDAR sensor(s) 1264 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LIDAR sensors 1264 may operate at the ASIL B functional safety level. In at least one embodiment, the vehicle 1200 may include multiple LIDAR sensors 1264 (e.g., two, four, six, etc.) that may use an Ethernet channel (e.g., to provide data to a Gigabit Ethernet switch).

[0326] In at least one embodiment, the LIDAR sensor(s) 1264 may be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, commercially available LIDAR sensor(s) 1264 may, for example, have an advertised range of approximately 100 m, with an accuracy of 2 cm to 3 cm, and with support for a 100 Mbps Ethernet connection. In at least one embodiment, one or more non-protruding LIDAR sensors may be used. In such an embodiment, the LIDAR sensor(s) 1264 may include a small device that may be embedded in a front, rear, side, and / or corner location of the vehicle 1200.In at least one embodiment, the LIDAR sensor(s) 1264 in such an embodiment may provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees with a range of 200 m, even for low-reflectivity objects. In at least one embodiment, the front-mounted LIDAR sensor(s) 1264 may be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0327] In at least one embodiment, LIDAR technologies, such as 3D flash LIDAR, may also be used. In at least one embodiment, 3D flash LIDAR uses a laser flash as a transmission source to illuminate the surroundings of the vehicle 1200 up to approximately 200 m. In at least one embodiment, a flash LIDAR unit includes, without limitation, a receptor that records the laser pulse time of flight and reflected light at each pixel, which in turn corresponds to a range from the vehicle 1200 to objects. In at least one embodiment, flash LIDAR may enable highly accurate and distortion-free images of the surroundings to be generated with each laser flash. In at least one embodiment, four flash LIDAR sensors may be deployed, one on each side of the vehicle 1200.In at least one embodiment, 3D flash lidar systems include, without limitation, a solid-state 3D staring array lidar camera with no moving parts other than a fan (e.g., a non-scanning lidar device). In at least one embodiment, the flash lidar device may use a 5-nanosecond Class I (eye-safe) laser pulse per image and collect the reflected laser light as a 3D range point cloud and co-registered intensity data.

[0328] In at least one embodiment, the vehicle 1200 may further include IMU sensor(s) 1266. In at least one embodiment, the IMU sensor(s) 1266 may be located at a center of a rear axle of the vehicle 1200. In at least one embodiment, the IMU sensor(s) 1266 may include, for example, and without limitation, an accelerometer, a magnetometer, gyroscope(s), a magnetic compass, magnetic compasses, and / or other types of sensors. In at least one embodiment, such as in six-axis applications, the IMU sensor(s) 1266 may include, without limitation, accelerometers and gyroscopes. In at least one embodiment, such as in nine-axis applications, the IMU sensor(s) 1266 may include, without limitation, accelerometers, gyroscopes, and magnetometers.

[0329] In at least one embodiment, the IMU sensor(s) 1266 may be implemented as a miniaturized, high-performance GPS-Aided Inertial Navigation System (“GPS / INS”) that combines micro-electro-mechanical systems (“MEMS”) inertial sensors, a highly sensitive GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and attitude. In at least one embodiment, the IMU sensor(s) 1266 may enable the vehicle 1200 to estimate its heading without requiring input from a magnetic sensor by directly observing changes in velocity from a GPS and correlating them to the IMU sensor(s) 1266. In at least one embodiment, the IMU sensor(s) 1266 and GNSS sensor(s) 1258 may be combined into a single integrated unit.

[0330] In at least one embodiment, the vehicle 1200 may include microphone(s) 1296 placed in and / or around the vehicle 1200. In at least one embodiment, the microphone(s) 1296 may be used, among other things, for detecting and identifying emergency vehicles.

[0331] In at least one embodiment, the vehicle 1200 may further include any number of camera types, including stereo camera(s) 1268, wide-view camera(s) 1270, infrared camera(s) 1272, surround camera(s) 1274, long-range camera(s) 1298, medium-range camera(s) 1276, and / or other camera types. In at least one embodiment, cameras may be used to capture image data around the entire periphery of the vehicle 1200. The types of cameras used depend on the vehicle 1200 in at least one embodiment. In at least one embodiment, any combination of camera types may be used to provide the necessary coverage around the vehicle 1200. In at least one embodiment, the number of cameras employed may vary depending on the embodiment.For example, in at least one embodiment, vehicle 1200 could include six cameras, seven cameras, ten cameras, twelve cameras, or another number of cameras. In at least one embodiment, cameras can support, for example, and without limitation, Gigabit Multimedia Serial Link ("GMSL") and / or Gigabit Ethernet communications. In at least one embodiment, each camera could be as previously described herein with respect to . Fig. 12A and Fig. 12B described in more detail.

[0332] In at least one embodiment, the vehicle 1200 may further include vibration sensor(s) 1242. In at least one embodiment, the vibration sensor(s) 1242 may measure vibrations of components of the vehicle 1200, such as axle(s). For example, in at least one embodiment, changes in the vibrations may indicate a change in the road surface. When two or more vibration sensors 1242 are used, in at least one embodiment, the differences between the vibrations may be used to determine the friction or slippage of the road surface (e.g., when there is a difference in vibration between a powered axle and a freely rotating axle).

[0333] In at least one embodiment, the vehicle 1200 may include the ADAS system 1238. In at least one embodiment, the ADAS system 1238 may include, without limitation, an SoC in some examples.In at least one embodiment, the ADAS system 1238 may include, without limitation, any number and combination of an autonomous / adaptive / automatic cruise control (“ACC”) system, a cooperative adaptive cruise control (“CACC”) system, a forward crash warning (“FCW”) system, an automatic emergency braking (“AEB”) system, a lane departure warning (“LDW”) system, a lane keep assist (“LKA”) system, a blind spot warning (“BSW”) system, a rear cross-traffic warning (“RCTW”) system, a collision warning (“CW”) system, a lane centering (“LC”) system, and / or other systems, features, and / or functions.

[0334] In at least one embodiment, the ACC system may use RADAR sensor(s) 1260, LIDAR sensor(s) 1264, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, a longitudinal ACC system monitors and controls the distance to another vehicle immediately in front of the vehicle 1200 and automatically adjusts the speed of the vehicle 1200 to maintain a safe distance from preceding vehicles. In at least one embodiment, a lateral ACC system performs follow-through and advises the vehicle 1200 to change lanes when necessary. In at least one embodiment, lateral ACC is related to other ADAS applications, such as LC and CW.

[0335] In at least one embodiment, a CACC system utilizes information from other vehicles that may be received via the network interface 1224 and / or the wireless antenna(s) 1226 from other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet). In at least one embodiment, direct links may be provided by a vehicle-to-vehicle ("V2V") communication link, while indirect links may be provided by an infrastructure-to-vehicle ("I2V") communication link. Generally, V2V communication provides information about immediately preceding vehicles (e.g., vehicles immediately ahead of and in the same lane as vehicle 1200), while I2V communication provides information about more distantly preceding traffic.In at least one embodiment, a CACC system may include either one or both I2V and V2V information sources. In at least one embodiment, a CACC system may be more reliable given information about vehicles ahead of vehicle 1200 and has the potential to improve traffic flow smoothness and reduce congestion on the road.

[0336] In at least one embodiment, an FCW system is configured to warn a driver of a hazard so that such a driver can take corrective action. In at least one embodiment, an FCW system uses a forward-facing camera and / or RADAR sensor(s) 1260 coupled, i.e., electrically coupled, to a dedicated processor, DSP, FPGA, and / or ASIC to provide driver feedback, such as a display, speaker, and / or vibrating component. In at least one embodiment, an FCW system may provide a warning, such as in the form of a sound, a visual warning, a vibration, and / or a rapid braking pulse.

[0337] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object, and may automatically apply the brakes if a driver does not take corrective action within a predetermined time or distance parameter. In at least one embodiment, the AEB system may utilize forward-facing camera(s) and / or RADAR sensor(s) 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When an AEB system detects a hazard, in at least one embodiment, it typically first alerts a driver to take corrective action to avoid a collision, and if that driver does not take corrective action, that AEB system may automatically apply the brakes in an effort to prevent or at least mitigate the impact of a predicted collision.In at least one embodiment, the AEB system may include techniques such as dynamic brake assist and / or braking due to an impending collision.

[0338] In at least one embodiment, an LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle 1200 crosses lane markings. In at least one embodiment, an LDW system is not activated when a driver indicates an intentional lane departure, such as by activating a turn signal. In at least one embodiment, an LDW system may utilize forward- and side-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC, i.e., electrically coupled, to provide driver feedback, such as a display, speaker, and / or vibrating component. In at least one embodiment, an LKA system is a variation of an LDW system.In at least one embodiment, an LKA system provides a steering input or braking to correct the vehicle 1200 if the vehicle 1200 begins to depart from its lane.

[0339] In at least one embodiment, a BSW system detects and warns a driver of vehicles in a blind spot of an automobile. In at least one embodiment, the BSW system may provide a visual, audible, and / or tactile alert to indicate that merging into or changing lanes is unsafe. In at least one embodiment, a BSW system may provide an additional warning when the driver activates the turn signal. In at least one embodiment, a BSW system may utilize rear-facing cameras and / or RADAR sensors 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.

[0340] In at least one embodiment, an RCTW system may provide visual, audible, and / or tactile notification when an object is detected outside the range of a rear-view camera when the vehicle 1200 is reversing. In at least one embodiment, an RCTW system includes an AEB system to ensure that vehicle brakes are applied to avoid a collision. In at least one embodiment, an RCTW system may utilize one or more rear-facing RADAR sensors 1260 coupled, i.e., electrically coupled, to a dedicated processor, DSP, FPGA, and / or ASIC to provide driver feedback, such as a display, speaker, and / or vibrating component.

[0341] In at least one embodiment, conventional ADAS systems may be prone to false positives, which may be annoying and distracting for the driver, but are typically not catastrophic because conventional ADAS systems warn a driver and allow that driver to decide whether a safety condition truly exists and act accordingly. In at least one embodiment, in the event of conflicting results, the vehicle 1200 itself decides whether to consider the result of a primary computer or a secondary computer (e.g., a first controller or a second controller of the controllers 1236). In at least one embodiment, the ADAS system 1238 may, for example, be a backup and / or secondary computer that provides perception information to a rationality module of a backup computer.In at least one embodiment, a rationality monitor of a backup computer may execute redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. In at least one embodiment, the outputs from the ADAS system 1238 may be provided to a supervisory MCU. If outputs from a primary computer and outputs from a secondary computer conflict, a supervisory MCU, in at least one embodiment, determines how to resolve the conflict to ensure safe operation.

[0342] In at least one embodiment, a primary computer may be configured to provide a monitoring MCU with a confidence score indicating the primary computer's confidence in a selected result. In at least one embodiment, the monitoring MCU may follow the primary computer's instruction if the confidence score exceeds a threshold, regardless of whether the secondary computer provides a conflicting or inconsistent result. In at least one embodiment, if a confidence score does not meet a threshold and primary and secondary computers indicate different results (e.g., a conflict), a monitoring MCU may arbitrate between the computers to determine an appropriate result.

[0343] In at least one embodiment, a monitoring MCU may be configured to execute neural network(s) trained and configured to determine, based at least in part on outputs from a primary computer and outputs from a secondary computer, the conditions under which that secondary computer provides false alarms. In at least one embodiment, the neural networks in a parent MCU may learn when the output of a secondary computer is trustworthy and when it is not. For example, in at least one embodiment, if the secondary computer is, for example, a RADAR-based FCW system, neural networks in that monitoring MCU may detect when an FCW system identifies metallic objects that do not actually pose hazards, such as a drainage grate or manhole cover, that trigger an alarm.In at least one embodiment, a neural network in a monitoring MCU may learn to override lane departure warning when a secondary computing system is a camera-based lane departure warning system and lane departure is actually the safest maneuver. In at least one embodiment, a monitoring MCU may include at least one of a DLA or a GPU capable of executing neural network(s) with associated memory. In at least one embodiment, a monitoring MCU may comprise and / or be included as a component of one or more SoC(s) 1204.

[0344] In at least one embodiment, the ADAS system 1238 may include a secondary computer that performs ADAS functionality using traditional computer vision rules. In at least one embodiment, this secondary computer may use classic computer vision rules (if-then), and the presence of a neural network(s) in a supervisory MCU may improve reliability, safety, and computational performance. In at least one embodiment, the different implementation and intentional non-identity make the overall system more fault-tolerant, particularly against errors caused by software functions (or software-hardware interfaces).For example, in at least one embodiment, if there is a software bug or error in the software running on a primary computer and non-identical software code running on a secondary computer provides a consistent overall result, then a monitoring MCU may have greater confidence that an overall result is correct and a bug in the software or hardware on that primary computer does not cause a significant error.

[0345] In at least one embodiment, an output of the ADAS system 1238 may be fed to a perception block of a primary computer and / or a dynamic driving task block of a primary computer. For example, if the ADAS system 1238 indicates a forward collision warning due to an immediately ahead object, a perception block in at least one embodiment may use this information in identifying objects. In at least one embodiment, a secondary computer may have its own neural network trained, thus reducing the risk of false positives, as described herein.

[0346] In at least one embodiment, the vehicle 1200 may further include an infotainment SoC 1230 (e.g., an in-vehicle infotainment system (IVI system)). Although illustrated and described as an SoC, in at least one embodiment, the infotainment SoC 1230 may not be an SoC and may include, without limitation, two or more discrete components. In at least one embodiment, the infotainment SoC 1230 may include, without limitation, a combination of hardware and software that may be used to provide the vehicle 1200 with audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g.,Navigation systems, reverse parking assistance, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, oil level, door open / close, air filter information, etc.). The infotainment SoC 1230 could include, for example, radios, record players, navigation systems, video playback devices, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, a hands-free voice control, a heads-up display (HUD), an HMI display 1234, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, the infotainment SoC 1230 may be further used to provide information (e.g.,visual and / or audible), such as information from the ADAS system 1238, autonomous driving information such as planned vehicle maneuvers, trajectories, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0347] In at least one embodiment, the infotainment SoC 1230 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 1230 may communicate with other devices, systems, and / or components of the vehicle 1200 via the bus 1202. In at least one embodiment, the infotainment SoC 1230 may be coupled to a supervisory MCU so that a GPU of an infotainment system may perform some self-driving functions in the event that the primary controller(s) 1236 (e.g., primary and / or backup computers of the vehicle 1200) fail. In at least one embodiment, the infotainment SoC 1230 may place the vehicle 1200 in a drive-to-safe-stop mode, as described herein.

[0348] In at least one embodiment, the vehicle 1200 may further include an instrument cluster 1232 (e.g., a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). In at least one embodiment, the instrument cluster 1232 may include, without limitation, a controller and / or a supercomputer (e.g., a discrete controller or a discrete supercomputer). In at least one embodiment, the instrument cluster 1232 may include, without limitation, any number and combination of a set of instruments, such as speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift position indicator, seat belt warning light(s), parking brake warning light(s), engine malfunction light(s), supplemental restraint system information (e.g., airbag), lighting controls, safety system controls, navigation information, etc.In some examples, the information may be displayed and / or shared between the infotainment SoC 1230 and the instrument cluster 1232. In at least one embodiment, the instrument cluster 1232 may be included as part of the infotainment SoC 1230, or vice versa.

[0349] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the system of Fig. 12C for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0350] Fig. 12D is an illustration of a system for communication between cloud-based server(s) and the autonomous vehicle 1200 of Fig. 12A according to at least one embodiment. In at least one embodiment, the system may include, without limitation, the server(s) 1278, the network(s) 1290, and any number and type of vehicles, including the vehicle 1200. In at least one embodiment, the server(s) 1278 may include, without limitation, a plurality of GPUs 1284(A)-1284(H) (collectively referred to herein as GPUs 1284), PCIe switches 1282(A)-1282(D) (collectively referred to herein as PCIe switches 1282), and / or CPUs 1280(A)-1280(B) (collectively referred to herein as CPUs 1280). In at least one embodiment, the GPUs 1284, CPUs 1280, and PCIe switches 1282 may be interconnected with high-speed interconnects, such as, without limitation, the NVLink interfaces 1288 developed by NVIDIA and / or PCIe interconnects 1286.In at least one embodiment, the GPUs 1284 are connected via an NVLink and / or NVSwitch SoC, and the GPUs 1284 and the PCIe switches 1282 are connected via PCIe interconnects. Although eight GPUs 1284, two CPUs 1280, and four PCIe switches 1282 are illustrated, this is not intended to be limiting. In at least one embodiment, each of the server(s) 1278 may include, without limitation, any number of GPUs 1284, CPUs 1280, and / or PCIe switches 1282 in any combination. For example, in at least one embodiment, the server(s) 1278 could each include eight, sixteen, thirty-two, and / or more GPUs 1284.

[0351] In at least one embodiment, the server(s) 1278 may receive, via the network(s) 1290 and from vehicles, image data representative of images depicting unexpected or changed road conditions, such as recently commenced roadwork. In at least one embodiment, the server(s) 1278 may transmit, via the network(s) 1290 and to the vehicles, neural networks 1292, updated or otherwise, and / or map information 1294, including without limitation information regarding traffic and road conditions. In at least one embodiment, updates to the map information 1294 may include, without limitation, updates to the HD map 1222, such as information regarding construction, potholes, detours, flooding, and / or other obstacles.In at least one embodiment, the neural networks 1292 and / or map information 1294 may have resulted from new training and / or experience represented in data received from any number of vehicles in an environment and / or may be based at least in part on training performed in a data center (e.g., using the server(s) 1278 and / or other servers).

[0352] In at least one embodiment, the server(s) 1278 may be used to train machine learning models (e.g., neural networks) based at least in part on training data. In at least one embodiment, the training data may be generated by vehicles and / or generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is tagged (e.g., if the associated neural network benefits from supervised learning) and / or subjected to other preprocessing. In at least one embodiment, any amount of training data is not tagged and / or preprocessed (e.g., if the associated neural network does not require supervised learning).In at least one embodiment, once the machine learning models are trained, the machine learning models may be used by vehicles (e.g., transmitted to vehicles via network(s) 1290) and / or the machine learning models may be used by server(s) 1278 to remotely monitor vehicles.

[0353] In at least one embodiment, the server(s) 1278 may receive data from vehicles and apply the data to real-time neural networks for intelligent real-time inference. In at least one embodiment, the server(s) 1278 may include deep learning supercomputers and / or dedicated AI computers powered by the GPU(s) 1284, such as the DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, the server(s) 1278 may include a deep learning infrastructure using CPU-powered data centers.

[0354] In at least one embodiment, the deep learning infrastructure of the server(s) 1278 may be capable of rapid, real-time inference and may use this capability to evaluate and verify the state of processors, software, and / or associated hardware in the vehicle 1200. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from the vehicle 1200, such as a sequence of images and / or objects that the vehicle 1200 has located in that sequence of images (e.g., via machine vision and / or other machine learning techniques for object classification).In at least one embodiment, the deep learning infrastructure may execute its own neural network to identify objects and compare them to objects identified by the vehicle 1200, and if the results do not match and the deep learning infrastructure concludes that the AI in the vehicle 1200 is malfunctioning, then the server(s) 1278 may transmit a signal to the vehicle 1200 instructing a fail-safe computer of the vehicle 1200 to take over control, notify the passengers, and perform a safe parking maneuver.

[0355] In at least one embodiment, the server(s) 1278 may include GPU(s) 1284 and one or more programmable inference accelerators (e.g., TensorRT-3 devices from NVIDIA). In at least one embodiment, a combination of GPU-powered servers and inference acceleration may enable real-time responsiveness. In at least one embodiment, such as when compute power is less critical, servers powered by CPUs, FPGAs, and other processors may be used for inference. In at least one embodiment, the hardware structure(s) 915 are used to perform one or more embodiments. Details regarding the hardware structure(s) 915 are described herein in connection with Fig. 9A and / or 9B provided.

[0356] Fig. 13 is a block diagram illustrating an example computer system, which may be a system of interconnected devices and components, a system on a chip (SOC), or a combination thereof, formed with a processor that may include execution units for executing an instruction, according to at least one embodiment. In at least one embodiment, a computer system 1300 may include, without limitation, a component, such as a processor 1302, for utilizing execution units including logic for performing algorithms on process data according to the present disclosure, such as in the embodiment described herein.In at least one embodiment, computer system 1300 may include processors such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessors available from Intel Corporation of Santa Clara, California, although other systems (including personal computers having other microprocessors, engineering workstations, set-top boxes, and the like) may be used. In at least one embodiment, computer system 1300 may run a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0357] Embodiments may be used in other devices such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and portable PCs. In at least one embodiment, embedded applications may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions according to at least one embodiment.

[0358] In at least one embodiment, computer system 1300 may include, without limitation, processor 1302, which may include, without limitation, one or more execution units 1308 to perform training and / or inference of a machine learning model according to the techniques described herein. In at least one embodiment, computer system 1300 is a single-processor desktop or server system, but in another embodiment, computer system 1300 may be a multiprocessor system. In at least one embodiment, processor 1302 may include, without limitation, a Complex Instruction Set Computer ("CISC") microprocessor, a Reduced Instruction Set Computing ("RISC") microprocessor, a Very Long Instruction Word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor.In at least one embodiment, the processor 1302 may be coupled to a processor bus 1310 that may transmit data signals between the processor 1302 and other components in the computer system 1300.

[0359] In at least one embodiment, processor 1302 may include, without limitation, an internal Level 1 ("L1") cache ("cache") 1304. In at least one embodiment, processor 1302 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache may be external to processor 1302. Other embodiments may include a combination of both internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, a register bank 1306 may store different types of data in various registers, including, without limitation, integer registers, floating-point registers, status registers, and an instruction pointer register.

[0360] In at least one embodiment, execution unit 1308, including without limitation logic for performing integer and floating-point operations, is also located within processor 1302. In at least one embodiment, processor 1302 may also include read-only memory ("ROM") for microcode ("ucode") that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 1308 may include logic for handling a packed instruction set 1309. In at least one embodiment, by including packed instruction set 1309 in an instruction set of a general-purpose processor, along with associated instruction execution circuitry, operations used by many multimedia applications may be performed using packed data within processor 1302.In one or more embodiments, many multimedia applications may be accelerated and executed more efficiently by using a full width of a processor's data bus to perform operations on packed data, which may eliminate the need to transfer smaller units of data across that processor's data bus to perform one or more operations on one data item at a time.

[0361] In at least one embodiment, execution unit 1308 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 1300 may include, without limitation, memory 1320. In at least one embodiment, memory 1320 may be a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 1320 may store instruction(s) 1319 and / or data 1321 represented by data signals that may be executed by processor 1302.

[0362] In at least one embodiment, a system logic chip may be coupled to processor bus 1310 and memory 1320. In at least one embodiment, a system logic chip may include, without limitation, a memory controller hub (“MCH”) 1316, and processor 1302 may communicate with MCH 1316 via processor bus 1310. In at least one embodiment, MCH 1316 may provide a high-bandwidth memory path 1318 to memory 1320 for instruction and data storage, as well as for storing graphics commands, data, and textures. In at least one embodiment, MCH 1316 may route data signals between processor 1302, memory 1320, and other components in computer system 1300, and may bridge data signals between processor bus 1310, memory 1320, and a system I / O interface 1322.In at least one embodiment, a system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 1316 may be coupled to memory 1320 through a high-bandwidth memory path 1318, and a graphics / video card 1312 may be coupled to MCH 1316 through an Accelerated Graphics Port ("AGP") interconnect 1314.

[0363] In at least one embodiment, computer system 1300 may use system I / O interface 1322 as a proprietary hub interface bus to couple MCH 1316 to an I / O controller hub (ICH) 1330. In at least one embodiment, ICH 1330 may provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, a local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 1320, a chipset, and processor 1302. Examples may include, without limitation, an audio controller 1329, a firmware hub (“Flash BIOS”) 1328, a wireless transceiver 1326, a data store 1324, a legacy I / O controller 1323 containing user input and keyboard interfaces 1325, a serial expansion port 1327, such as a Universal Serial Bus (“USB”) port, and a network controller 1334.In at least one embodiment, data storage 1324 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0364] In at least one embodiment, Fig. 13 a system that includes interconnected hardware devices or ‘chips’, whereas Fig. 13 may illustrate an exemplary SoC in other embodiments. In at least one embodiment, the Fig. 13 may be interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of computer system 1300 are interconnected using Compute Express Link (CXL) interconnects.

[0365] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the system of Fig. 13 be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0366] Fig. 14 is a block diagram illustrating an electronic device 1400 for utilizing a processor 1410 according to at least one embodiment. In at least one embodiment, the electronic device 1400 may be, for example and without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0367] In at least one embodiment, electronic device 1400 may include, without limitation, processor 1410 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1410 is coupled using a bus or interface, such as an I 2C-bus, a system management bus (SMBus), a low-pin count (LPC) bus, a serial peripheral interface (SPI), a high-definition audio (HDA) bus, a serial advance technology attachment (SATA) bus, a universal serial bus (USB) (version 1, 2, 3, etc.), or a universal asynchronous receiver / transmitter (UART) bus. In at least one embodiment, Fig. 14 a system that includes interconnected hardware devices or ‘chips’, whereas Fig. 14 may illustrate an exemplary SoC in other embodiments. In at least one embodiment, the Fig. 14 may be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of Fig. 14 interconnected using Compute Express Link (CXL) interconnects.

[0368] In at least one embodiment, Fig. 14 a display 1424, a touchscreen 1425, a touchpad 1430, a near field communications (NFC) unit 1445, a sensor hub 1440, a thermal sensor 1446, an express chipset (EC) 1435, a trusted platform module (TPM) 1438, BIOS / firmware / flash memory (BIOS, FW Flash) 1422, a DSP 1460, a drive 1420, such as a solid state disk (SSD) or a hard disk drive (HDD), a wireless local area network (WLAN) unit 1450, a Bluetooth unit 1452, a wireless wide area network (WWAN) unit 1456, a global positioning system (GPS) unit 1455, a camera (“USB 3.0 camera”) 1454, such as a USB 3.0 camera, and / or a low-power double data rate ("LPDDR") memory unit ("LPDDR3") 1415, implemented, for example, in an LPDDR3 standard. These components may each be implemented in any suitable manner.

[0369] In at least one embodiment, other components may be communicatively coupled to processor 1410 through components described herein. In at least one embodiment, an accelerometer 1441, an ambient light sensor (ALS) 1442, a compass 1443, and a gyroscope 1444 may be communicatively coupled to sensor hub 1440. In at least one embodiment, a thermal sensor 1439, a fan 1437, a keyboard 1436, and a touchpad 1430 may be communicatively coupled to EC 1435. In at least one embodiment, speakers 1463, headphones 1464, and a microphone ("Mic") 1465 may be communicatively coupled to an audio unit ("Audio Codec and Class D Amplifier") 1462, which in turn may be communicatively coupled to DSP 1460. In at least one embodiment, the audio unit 1462 may include, for example and without limitation, an audio encoder / decoder ("codec") and a Class D amplifier.In at least one embodiment, a SIM card ("SIM") 1457 may be communicatively coupled to the WWAN unit 1456. In at least one embodiment, components such as the WLAN unit 1450 and the Bluetooth unit 1452, as well as the WWAN unit 1456, may be implemented in a Next Generation Form Factor ("NGFF").

[0370] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the system of Fig. 14 for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0371] Fig. Figure 15 illustrates a computer system 1500 according to at least one embodiment. In at least one embodiment, the computer system 1500 is configured to implement various processes and methods described in this disclosure.

[0372] In at least one embodiment, computer system 1500 includes, without limitation, at least one central processing unit ("CPU") 1502 coupled to a communications bus 1510 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communications protocol(s). In at least one embodiment, computer system 1500 includes, without limitation, main memory 1504 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data is stored in main memory 1504, which may take the form of random access memory ("RAM").In at least one embodiment, a network interface subsystem (“network interface”) 1522 provides an interface to other computing devices and networks to receive and transmit data from and to other systems with computer system 1500.

[0373] In at least one embodiment, computer system 1500 includes, without limitation, input devices 1508, a parallel processing system 1512, and display devices 1506, which may be implemented with a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light-emitting diode ("LED"), a plasma display, or other suitable display technology. In at least one embodiment, user input is received from input devices 1508 such as a keyboard, mouse, touchpad, microphone, etc. In at least one embodiment, each module described herein may be located on a single semiconductor platform to form a processing system.

[0374] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the system of Fig. 15 be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0375] Fig. 16 illustrates a computer system 1600 according to at least one embodiment. In at least one embodiment, the computer system 1600 includes, without limitation, a computer 1610 and a USB flash drive 1620. In at least one embodiment, the computer 1610 may include, without limitation, any number and type of processor(s) (not shown) and memory (not shown). In at least one embodiment, the computer 1610 includes, without limitation, a server, a cloud instance, a laptop, and a desktop computer.

[0376] In at least one embodiment, USB flash drive 1620 includes, without limitation, a processing unit 1630, a USB interface 1640, and USB interface logic 1650. In at least one embodiment, processing unit 1630 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, processing unit 1630 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, processing unit 1630 comprises an application-specific integrated circuit ("ASIC") optimized to perform any number and type of operations associated with machine learning.For example, in at least one embodiment, processing unit 1630 is a tensor processing unit ("TPC") optimized for performing machine learning inference operations. In at least one embodiment, processing unit 1630 is a vision processing unit ("VPU") optimized for performing machine vision and machine learning inference operations.

[0377] In at least one embodiment, USB interface 1640 may be any type of USB plug or receptacle. For example, in at least one embodiment, USB interface 1640 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 1640 is a USB 3.0 Type-A plug. In at least one embodiment, USB interface logic 1650 may include any amount and type of logic that enables processing unit 1630 to interface with devices (e.g., computer 1610) via USB plug 1640.

[0378] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the system of Fig. 16 for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0379] Fig. 17A illustrates an example architecture in which a plurality of GPUs 1710(1)-1710(N) are communicatively coupled to a plurality of multi-core processors 1705(1)-1705(M) via high-speed links 1740(1)-1740(N) (e.g., buses, point-to-point interconnects, etc.). In at least one embodiment, the high-speed links 1740(1)-1740(N) support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or higher. In at least one embodiment, various interconnect protocols may be used, including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0. In various figures, "N" and "M" represent positive integers, the values of which may vary from figure to figure.

[0380] Additionally, and in at least one embodiment, two or more of the GPUs 1710 are interconnected via high-speed links 1729(1)-1729(2), which may be implemented using similar or different protocols / links than those used for the high-speed links 1740(1)-1740(N). Similarly, two or more of the multi-core processors 1705 may be interconnected via a high-speed link 1728, which may be symmetric multi-processor (SMP) buses operating at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, all communication between the various Fig. 17A can be achieved using similar protocols / links (e.g., via a common interconnection structure).

[0381] In at least one embodiment, each multi-core processor 1705 is communicatively coupled to a processor memory 1701(1)-1701(M) via memory interconnects 1726(1)-1726(M), respectively, and each GPU 1710(1)-1710(N) is communicatively coupled to GPU memory 1720(1)-1720(N) via GPU memory interconnects 1750(1)-1750(N), respectively. In at least one embodiment, memory interconnects 1726 and 1750 may utilize similar or different memory access technologies. The processor memories 1701(1)-1701(M) and the GPU memories 1720 may be, for example and without limitation, volatile memories such as dynamic random access memories (DRAMs) (including stacked DRAMs), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or non-volatile memories such as 3D XPoint or Nano-Ram.In at least one embodiment, a portion of the processor memory 1701 may be volatile memory and another portion may be non-volatile memory (e.g., using a two-level memory (2LM) memory hierarchy).

[0382] As described herein, various multi-core processors 1705 and GPUs 1710 may be physically coupled to a specific memory 1701 or 1720, respectively, and / or implement a unified memory architecture in which a virtual system address space (also referred to as "effective address space") is distributed across different physical memories. For example, processor memories 1701(1)-1701(M) may each comprise 64 GB of system memory address space, and GPU memories 1720(1)-1720(N) may each comprise 32 GB of system memory address space, resulting in a total addressable memory of 256 GB when M=2 and N=4. Other values for N and M are possible.

[0383] Fig. 17B illustrates additional details for interconnection between a multi-core processor 1707 and a graphics acceleration module 1746 according to an example embodiment. In at least one embodiment, the graphics acceleration module 1746 may include one or more GPU chips integrated on a line card coupled to the processor 1707 via a high-speed link 1740 (e.g., a PCIe bus, NVLink, etc.). Alternatively, in at least one embodiment, the graphics acceleration module 1746 may be integrated on a package or die with the processor 1707.

[0384] In at least one embodiment, processor 1707 includes a plurality of cores 1760A-1760D, each having a translation lookaside buffer (“TLB”) 1761A-1761D and one or more caches 1762A-1762D. In at least one embodiment, cores 1760A-1760D may include various other components for executing instructions and processing data, not illustrated. In at least one embodiment, caches 1762A-1762D may include Level 1 (L1) and Level 2 (L2) caches. Additionally, one or more shared caches 1756 may be included within caches 1762A-1762D and shared between sets of cores 1760A-1760D. For example, one embodiment of processor 1707 includes 24 cores, each with its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches.In this embodiment, one or more L2 and L3 caches are shared between two adjacent cores. In at least one embodiment, processor 1707 and graphics acceleration module 1746 are coupled to system memory 1714, which includes processor memories 1701(1)-1701(M) of FIG. Fig. 17A may include.

[0385] In at least one embodiment, coherency for data and instructions stored in various caches 1762A-1762D, 1756, and system memory 1714 is maintained via inter-core communication over a coherency bus 1764. For example, in at least one embodiment, each cache may have cache coherency logic / circuitry associated with it to communicate over the coherency bus 1764 in response to detected reads or writes to specific cache lines. In at least one embodiment, a cache snooping protocol is implemented over the coherency bus 1764 to control cache accesses via snooping.

[0386] In at least one embodiment, a proxy circuit 1725 communicatively couples the graphics acceleration module 1746 to the coherency bus 1764, enabling the graphics acceleration module 1746 to participate in a cache coherency protocol as a peer of the cores 1760A-1760D. Specifically, in at least one embodiment, an interface 1735 provides connectivity to the proxy circuit 1725 via a high-speed link 1740, and an interface 1737 connects the graphics acceleration module 1746 to the high-speed link 1740.

[0387] In at least one embodiment, an accelerator integration circuit 1736 provides cache management, memory access, context management, and interrupt management services on behalf of a plurality of graphics processing engines 1731(1)-1731(N) of the graphics acceleration module 1746. In at least one embodiment, the graphics processing engines 1731(1)-1731(N) may each comprise a separate graphics processing unit (GPU). Alternatively, in at least one embodiment, the graphics processing engines 1731(1)-1731(N) may comprise different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines.In at least one embodiment, the graphics acceleration module 1746 may be a GPU with a plurality of graphics processing engines 1731(1)-1731(N), or the graphics processing engines 1731(1)-1731(N) may be individual GPUs integrated on a common package, line card, or die.

[0388] In at least one embodiment, accelerator integration circuit 1736 includes a memory management unit (MMU) 1739 for performing various memory management functions, such as virtual-to-physical memory translations (also referred to as effective-to-real memory translations) and memory access protocols for accessing system memory 1714. MMU 1739 may also include an address translation buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations, in at least one embodiment. In at least one embodiment, a cache 1738 may store instructions and data for efficient access by graphics processing engines 1731(1)-1731(N).In at least one embodiment, the data stored in cache 1738 and graphics memories 1733(1)-1733(M) is kept coherent with core caches 1762A-1762D, 1756, and system memory 1714, possibly using a fetch unit 1744. As mentioned, this may be accomplished via proxy circuitry 1725 on behalf of cache 1738 and memories 1733(1)-1733(M) (e.g., sending updates to cache 1738 regarding modifications / accesses to cache lines in processor caches 1762A-1762D, 1756 and receiving updates from cache 1738).

[0389] In at least one embodiment, a set of registers 1745 stores context data for threads executed by graphics processing engines 1731(1)-1731(N), and a context management circuit 1748 manages thread contexts. For example, context management circuit 1748 may perform save and restore operations to save and restore contexts of different threads during context switches (e.g., when a first thread is saved and a second thread is saved so that a second thread can be executed by a graphics processing engine). For example, upon a context switch, context management circuit 1748 may save current register values to a designated region in memory (e.g., identified by a context pointer). It may then restore the register values upon return to a context.In at least one embodiment, an interrupt management circuit 1747 receives and processes interrupts received from system devices.

[0390] In at least one embodiment, virtual / effective addresses from a graphics processing engine 1731 are translated by the MMU 1739 into real / physical addresses in the system memory 1714. In at least one embodiment, the accelerator integration circuit 1736 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 1746 and / or other accelerator devices. The graphics accelerator module 1746 may, in at least one embodiment, be dedicated to a single application executing on the processor 1707 or shared among multiple applications. In at least one embodiment, a virtualized graphics execution environment is depicted in which the resources of the graphics processing engines 1731(1)-1731(N) are shared with multiple applications or virtual machines (VMs).In at least one embodiment, the resources may be divided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with VMs and / or applications.

[0391] In at least one embodiment, accelerator integration circuitry 1736 acts as a bridge to a system for graphics acceleration module 1746 and provides address translation and system memory caching services. Furthermore, in at least one embodiment, accelerator integration circuitry 1736 may provide virtualization facilities for a host processor to manage the virtualization of graphics processing engines 1731(1)-1731(N), interrupts, and memory management.

[0392] Because, in at least one embodiment, the hardware resources of graphics processing engines 1731(1)-1731(N) are explicitly mapped to a real address space seen by host processor 1707, any host processor can directly address these resources using an effective address value. In at least one embodiment, a function of accelerator integration circuit 1736 is to physically separate graphics processing engines 1731(1)-1731(N) so that they appear to a system as independent entities.

[0393] In at least one embodiment, one or more graphics memories 1733(1)-1733(M) are each coupled to each of the graphics processing engines 1731(1)-1731(N), and N=M. In at least one embodiment, the graphics memories 1733(1)-1733(M) store instructions and data processed by each of the graphics processing engines 1731(1)-1731(N). In at least one embodiment, the graphics memories 1733(1)-1733(M) may be volatile memories, such as DRAMs (including stacked DRAMs), GDDR memories (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memories, such as 3D XPoint or Nano-Ram.

[0394] In at least one embodiment, to reduce data traffic over high-speed link 1740, warping techniques may be used to ensure that the data stored in graphics memories 1733(1)-1733(M) is data most frequently used by graphics processing engines 1731(1)-1731(N) and preferably not used (at least not frequently) by cores 1760A-1760D. Similarly, in at least one embodiment, a warping mechanism attempts to keep data needed by the cores (and preferably not by graphics processing engines 1731(1)-1731(N)) within caches 1762A-1762D, 1756, and system memory 1714.

[0395] Fig. 17C illustrates another exemplary embodiment in which accelerator integration circuit 1736 is integrated with processor 1707. In this embodiment, graphics processing engines 1731(1)-1731(N) communicate directly via high-speed link 1740 with accelerator integration circuit 1736 via interface 1737 and interface 1735 (which may again be any form of bus or interface protocol). In at least one embodiment, accelerator integration circuit 1736 may perform similar operations to those described with respect to Fig. 17B, but possibly with higher throughput due to its close proximity to the coherence bus 1764 and caches 1762A-1762D, 1756. In at least one embodiment, an accelerator integration circuit supports different programming models, including a dedicated process programming model (without virtualization of the graphics acceleration module) and shared programming models (with virtualization), which may include programming models controlled by the accelerator integration circuit 1736 and programming models controlled by the graphics acceleration module 1746.

[0396] In at least one embodiment, the graphics processing engines 1731(1)-1731(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can direct other application requests to the graphics processing engines 1731(1)-1731(N), thus providing virtualization within a VM / partition.

[0397] In at least one embodiment, the graphics processing engines 1731(1)-1731(N) may be shared between multiple VM / application partitions. In at least one embodiment, shared models may use a system hypervisor to virtualize the graphics processing engines 1731(1)-1731(N) and provide access by any operating system. For single-partition systems without a hypervisor, the graphics processing engines 1731(1)-1731(N) are owned by an operating system in at least one embodiment. In at least one embodiment, an operating system may virtualize the graphics processing engines 1731(1)-1731(N) to provide access to any process or application.

[0398] In at least one embodiment, the graphics acceleration module 1746 or an individual graphics processing engine 1731(1)-1731(N) selects a process element using a process identifier. In at least one embodiment, the process elements are stored in system memory 1714 and are addressable using the effective address to real address translation technique described herein. In at least one embodiment, a process identifier may be an implementation-specific value provided to a host process when it registers its context with the graphics processing engine 1731(1)-1731(N) (i.e., calls system software to add a process element to a list linked to the process element). In at least one embodiment, the lower 16 bits of a process identifier may be an offset of a process element within a list linked to the process element.

[0399] Fig. 17D illustrates an exemplary accelerator integration slice 1790. In at least one embodiment, a "slice" comprises a predetermined portion of the processing resources of accelerator integration circuit 1736. In at least one embodiment, an application is effective address space 1782 within system memory 1714 that stores process elements 1783. In at least one embodiment, process elements 1783 are stored in response to GPU calls 1781 from applications 1780 executing on processor 1707. In at least one embodiment, a process element 1783 contains the process state for the corresponding application 1780. In at least one embodiment, a work descriptor (WD) 1784 contained in process element 1783 may be a single task requested by an application or may contain a pointer to a queue of tasks.In at least one embodiment, the WD 1784 is a pointer to a task request queue in the effective address space 1782 of an application.

[0400] In at least one embodiment, the graphics acceleration module 1746 and / or the individual graphics processing engines 1731(1)-1731(N) may be shared by all or a subset of the processes in a system. In at least one embodiment, an infrastructure for establishing process states and sending a WD 1784 to a graphics acceleration module 1746 to start a task in a virtualized environment may be included.

[0401] In at least one embodiment, a dedicated process programming model is implementation-specific. In at least one embodiment, in this model, a single process owns the graphics acceleration module 1746 or a single graphics processing engine 1731. When the graphics acceleration module 1746 is owned by a single process, in at least one embodiment, a hypervisor initializes the accelerator integration circuit 1736 for an owning partition, and an operating system initializes the accelerator integration circuit 1736 for an owning process when the graphics acceleration module 1746 is allocated.

[0402] In operation, in at least one embodiment, a WD fetch unit 1791 in accelerator integration slice 1790 fetches the next WD 1784, which includes an indication of work to be performed by one or more graphics processing engines of graphics acceleration module 1746. In at least one embodiment, data from WD 1784 may be stored in registers 1745 and used by MMU 1739, interrupt management circuitry 1747, and / or context management circuitry 1748, as illustrated. For example, one embodiment of MMU 1739 includes segment / page walkup circuitry for accessing segment / page tables 1786 within virtual address space 1785 of an OS. In at least one embodiment, interrupt management circuitry 1747 may process interrupt events 1792 received from graphics acceleration module 1746.When performing graphics operations, in at least one embodiment, an effective address 1793 generated by a graphics processing engine 1731(1)-1731(N) is translated into a real address by the MMU 1739.

[0403] In at least one embodiment, registers 1745 are duplicated for each graphics processing engine 1731(1)-1731(N) and / or each graphics acceleration module 1746, and they may be initialized by a hypervisor or an operating system. Each of these duplicated registers may be included in an accelerator integration slice 1790 in at least one embodiment. Example registers that may be initialized by a hypervisor are shown in Table 1. Table 1 - Registers initialized by hypervisor Register-Nr. Beschreibung 1 Slice-Steuerregister 2 Bereichszeiger geplante Prozesse reale Adresse (RA) 3 Autoritätsmasken-Überschreibungsregister 4 Unterbrechungsvektor-Tabelleneintragsversatz 5 Unterbrechungsvektor-Tabelleneintragsbegrenzung 6 Zustandsregister 7 Logische Partitions-ID 8 Datensatzzeiger Hypervisor-Beschleuniger-Nutzung reale Adresse (RA) 9 Speicherbeschreibungsregister

[0404] Example registers that can be initialized by an operating system are shown in Table 2. Table 2 - Registers initialized by the operating system Register-Nr Beschreibung 1 Process and thread identification 2 Context Store / Restore Pointer Effective Address (EA) 3 Record pointer accelerator usage virtual address (VA) 4 Memory segment table pointer virtual address (VA) 5 Authority mask 6 Work descriptor

[0405] In at least one embodiment, each WD 1784 is specific to a particular graphics acceleration module 1746 and / or the graphics processing engines 1731(1)-1731(N). In at least one embodiment, it contains all the information required for a graphics processing engine 1731(1)-1731(N) to perform work, or it may be a pointer to a memory location where an application has established a command queue of work to be completed.

[0406] Fig.17E illustrates additional details for an exemplary embodiment of a shared model. This embodiment includes a real hypervisor address space 1798 in which a process element list 1799 is stored. In at least one embodiment, the real hypervisor address space 1798 may be accessed via a hypervisor 1796 that virtualizes the graphics acceleration engine for the operating system 1795.

[0407] In at least one embodiment, shared programming models enable all or a subset of processes from all or a subset of partitions in a system to use a graphics acceleration module 1746. In at least one embodiment, there are two programming models where the graphics acceleration module 1746 is shared among multiple processes and partitions: shared via time slices and shared via directed graphics.

[0408] In at least one embodiment, in this model, the system hypervisor 1796 has the graphics acceleration module 1746 and makes its functionality available to all operating systems 1795.For a graphics acceleration module 1746 to support virtualization by the system hypervisor 1796, in at least one embodiment, the graphics acceleration module 1746 must adhere to certain requirements, such as (1) an application's task request must be autonomous (i.e., state does not need to be maintained between tasks), or the graphics acceleration module 1746 must provide a mechanism for saving and restoring context, (2) the graphics acceleration module 1746 guarantees that an application's task request will complete within a predetermined amount of time, including any translation errors, or the graphics acceleration module 1746 provides an ability to preempt processing of a task, and (3) the graphics acceleration module 1746 must be guaranteed inter-process fairness when operating in a directed shared programming model.

[0409] In at least one embodiment, the application 1780 is required to make a system call to the operating system 1795 with a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the graphics acceleration module type describes a targeted acceleration function for a system call. In at least one embodiment, the graphics acceleration module type may be a system-specific value.In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 1746 and may be in the form of an instruction of the graphics acceleration module 1746, an effective address pointer to a user-defined structure, an effective address pointer to an instruction queue, or any other data structure to describe work to be performed by the graphics acceleration module 1746.

[0410] In at least one embodiment, an AMR value is an AMR state to be used for a current process. In at least one embodiment, a value passed to an operating system is similar to an application setting an AMR. In at least one embodiment, if implementations of accelerator integration circuit 1736 (not shown) and graphics acceleration module 1746 do not support a User Authority Mask Override Register (UAMOR), an operating system may apply a current UAMOR value to an AMR value before passing an AMR in a hypervisor call. In at least one embodiment, hypervisor 1796 may optionally apply a current Authority Mask Override Register (AMOR) value before placing an AMR in process element 1783.In at least one embodiment, CSRP is one of the registers 1745 that contain an effective address of a region in an application's effective address space 1782 for the graphics acceleration module 1746 to save and restore context state. In at least one embodiment, this pointer is optional if no state needs to be saved between tasks or upon task preemption. In at least one embodiment, the context save / restore region may be pinned system memory.

[0411] Upon receiving a system call, the operating system 1795 may verify whether the application 1780 is registered and has been granted authority to use the graphics acceleration module 1746. In at least one embodiment, the operating system 1795 then calls the hypervisor 1796 with the information shown in Table 3. Table 3 - OS-to-Hypervisor call parameters Parameter No. Description 1 A work descriptor (WD) 2 An Authority Mask Register (AMR) value (possibly masked) 3 A context save / restore area pointer (CSRP) with effective address (EA) 4 A process ID (PID) and optional thread ID (TID) 5 An accelerator utilization record pointer (AURP) with virtual address (VA) 6 Virtual address of a storage segment table pointer (SSTP) 7 A logical interrupt service number (LISN)

[0412] In at least one embodiment, upon receiving a hypervisor call, the hypervisor 1796 verifies that the operating system 1795 is registered and has been granted authority to use the graphics acceleration module 1746. In at least one embodiment, the hypervisor 1796 then places the process element 1783 in a process element-linked list for a corresponding type of graphics acceleration module 1746. In at least one embodiment, a process element may include the information shown in Table 4. Table 4 - Process element information Element No. Description 1 A work descriptor (WD) 2 An Authority Mask Register (AMR) value (possibly masked) 3 A context save / restore area pointer (CSRP) with effective address (EA) 4 A process ID (PID) and optional thread ID (TID) 5 An accelerator utilization record pointer (AURP) with virtual address (VA) 6 Virtual address of a storage segment table pointer (SSTP) 7 A logical interrupt service number (LISN) 8 Interrupt vector table derived from hypervisor call parameters 9 A status register (SR) value 10 A logical partition ID (LPID) 11 A record pointer hypervisor accelerator usage with real address (RA) 12 Storage Descriptor Register (SDR)

[0413] In at least one embodiment, the hypervisor initializes a plurality of registers 1745 of the accelerator integration slice 1790.

[0414] As in Fig.17F, in at least one embodiment, a unified memory addressable via a common virtual memory address space is used to access the physical processor memories 1701(1)-1701(N) and the GPU memories 1720(1)-1720(N). In this implementation, operations performed on the GPUs 1710(1)-1710(N) use a same virtual / effective memory address space to access the processor memories 1701(1)-1701(M) and vice versa, which simplifies programmability. In at least one embodiment, a first portion of a virtual / effective address space is assigned to the processor memory 1701(1), a second portion is assigned to the second processor memory 1701(N), a third portion is assigned to the GPU memory 1720(1), and so on.In at least one embodiment, this distributes an entire virtual / effective memory space (sometimes referred to as effective address space) across each of processor memory 1701 and GPU memory 1720, allowing any processor or GPU to access any physical memory with a virtual address mapped to that memory.

[0415] In at least one embodiment, the warp / coherence management circuit 1794A-1794E within one or more MMUs 1739A-1739E ensures cache coherence between caches of one or more host processors (e.g., 1705) and GPUs 1710 and implements warp techniques that indicate physical memories in which certain types of data should be stored. Although in at least one embodiment, multiple instances of the warp / coherence management circuit 1794A-1794E in Fig.17F, the distortion / coherence circuitry may be implemented within an MMU of one or more host processors 1705 and / or within the accelerator integration circuitry 1736.

[0416] One embodiment enables GPU memory 1720 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology, but without incurring computational performance penalties associated with full system cache coherence. In at least one embodiment, the ability to access GPU memory 1720 as system memory without burdensome cache coherence overhead provides a favorable operating environment for GPU offloading. In at least one embodiment, this arrangement enables host processor 1705 software to set up operands and access computation results without the overhead of traditional I / O DMA data copies.In at least one embodiment, such traditional copies involve driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are inefficient relative to simple memory accesses. In at least one embodiment, the ability to access GPU memory 1720 without cache coherence overheads may be critical to the execution time of an offloaded computation. For example, in cases with significant streaming write memory traffic, the cache coherence overhead may significantly reduce an effective write bandwidth seen by a GPU 1710, in at least one embodiment. In at least one embodiment, operand setup efficiency, result access efficiency, and GPU computation efficiency may play a role in determining the effectiveness of GPU offloading.

[0417] In at least one embodiment, the selection of GPU skew and host processor skew is driven by a skew tracker data structure. For example, in at least one embodiment, a skew table may be used, which may be a page-granular structure (e.g., controlled at a memory page granularity) that includes 1 or 2 bits per GPU-bound memory page. In at least one embodiment, a skew table may be implemented in a stolen memory region of one or more GPU memories 1720, with or without a skew cache in a GPU 1710 (e.g., to cache frequently / recently used skew table entries). Alternatively, in at least one embodiment, an entire skew table may be maintained within a GPU.

[0418] In at least one embodiment, prior to actually accessing GPU memory, a warp table entry associated with each access to GPU-bound memory 1720 is accessed, causing the following operations. In at least one embodiment, local requests from a GPU 1710 that find their page in the GPU warp are forwarded directly to a corresponding GPU memory 1720. In at least one embodiment, local requests from a GPU that find its page in the host warp are forwarded to processor 1705 (e.g., via a high-speed link as described herein). In at least one embodiment, requests from processor 1705 that find a requested page in the host processor warp complete a request like a normal memory read.Alternatively, requests directed to a GPU warp page may be forwarded to a GPU 1710. In at least one embodiment, a GPU may then convert a page to a host processor warp if it is not currently using a page. In at least one embodiment, a warp state of a page may be changed either by a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, a purely hardware-based mechanism.

[0419] A mechanism for changing the warp state, in at least one embodiment, employs an API call (e.g., OpenCL), which in turn invokes a GPU's device driver, which in turn sends a message to a GPU (or queues a command descriptor) instructing it to change a warp state and, on some transitions, perform a cache flush operation in a host. In at least one embodiment, a cache flush operation is used for a transition from the host processor 1705 warp to the GPU warp, but not for an opposite transition.

[0420] In at least one embodiment, cache coherence is maintained by causing GPU-skewed pages to be temporarily uncacheable by host processor 1705. To access these pages, in at least one embodiment, processor 1705 may request access from GPU 1710, which may or may not grant access immediately. Therefore, to reduce communication between processor 1705 and GPU 1710, it is advantageous in at least one embodiment to ensure that GPU-skewed pages are those needed by a GPU but not by host processor 1705, and vice versa.

[0421] The hardware structure(s) 915 are used to perform one or more embodiments. Details regarding hardware structure(s) 915 may be described herein in connection with Fig. 9A and / or 9B must be provided.

[0422] Fig.Figure 18 illustrates exemplary integrated circuits and associated graphics processors that may be fabricated using one or more IP cores, according to various embodiments described herein. In addition to what is illustrated, other logic and circuitry may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0423] Fig.18 is a block diagram illustrating an exemplary system-on-a-chip integrated circuit 1800 that may be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, the integrated circuit 1800 includes one or more application processors 1805 (e.g., CPUs), at least one graphics processor 1810, and may additionally include an image processor 1815 and / or a video processor 1820, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 1800 includes peripheral or bus logic, including a USB controller 1825, a UART controller 1830, an SPI / SDIO controller 1835, and an I 2 2S / I 22C controller 1840. In at least one embodiment, the integrated circuit 1800 may include a display device 1845 coupled to one or more of a High Definition Multimedia Interface (HDMI) controller 1850 and a Mobile Industry Processor Interface (MIPI) display interface 1855. In at least one embodiment, storage may be provided by a flash memory subsystem 1860 including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1865 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits additionally include an embedded security engine 1870.

[0424] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the integrated circuit 1800 may be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0425] Fig.19A-19B illustrate example integrated circuits and associated graphics processors that may be fabricated using one or more IP cores, according to various embodiments described herein. In addition to what is illustrated, other logic and circuitry may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0426] Fig. 19A-19B are block diagrams illustrating example graphics processors for use within an SoC according to embodiments described herein. Fig. 19A illustrates an exemplary graphics processor 1910 of an integrated circuit as a system on a chip that may be manufactured using one or more IP cores, according to at least one embodiment. Fig.19B illustrates an additional exemplary graphics processor 1940 of an integrated circuit system on a chip that may be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, the graphics processor 1910 is Fig. 19A, a low-performance graphics processor core. In at least one embodiment, the graphics processor 1940 is Fig. 19B, a graphics processor core with higher computing power. In at least one embodiment, each of the graphics processors 1910, 1940 may be a variant of the graphics processor 1810 of Fig. be 18.

[0427] In at least one embodiment, graphics processor 1910 includes a vertex processor 1905 and one or more fragment processors 1915A-1915N (e.g., 1915A, 1915B, 1915C, 1915D through 1915N-1, and 1915N). In at least one embodiment, graphics processor 1910 may execute different shader programs via separate logic, such that vertex processor 1905 is optimized to perform operations for vertex shader programs, while one or more fragment processors 1915A-1915N perform shading operations on fragments (e.g., pixels) for fragment or pixel shader programs. In at least one embodiment, vertex processor 1905 performs a vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data.In at least one embodiment, the fragment processor(s) 1915A-1915N use primitive and vertex data generated by the vertex processor 1905 to produce a frame buffer displayed on a display device. In at least one embodiment, the fragment processor(s) 1915A-1915N are optimized to execute fragment shader programs, such as those provided in an OpenGL API, which can be used to perform similar operations as a pixel shader program, such as those provided in a Direct 3D API.

[0428] In at least one embodiment, graphics processor 1910 additionally includes one or more memory management units (MMUs) 1920A-1920B, cache(s) 1925A-1925B, and circuit interconnect(s) 1930A-1930B. In at least one embodiment, one or more MMU(s) 1920A-1920B provide virtual to physical address mapping for graphics processor 1910, including vertex processor 1905 and / or fragment processor(s) 1915A-1915N, which may reference vertex or image / texture data stored in memory, in addition to the vertex or image / texture data stored in one or more cache(s) 1925A-1925B. In at least one embodiment, one or more MMU(s) 1920A-1920B may be synchronized with other MMUs within a system, including one or more MMUs associated with one or more application processor(s) 1805, image processors 1815, and / or video processors 1820 of Fig.18, so that each processor 1805-1820 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1930A-1930B enable the graphics processor 1910 to interface with other IP cores within the SoC, either via an internal bus of the SoC or via a direct connection.

[0429] In at least one embodiment, the graphics processor 1940 includes one or more shader cores 1955A-1955N (e.g., 1955A, 1955B, 1955C, 1955D, 1955E, 1955F through 1955N-1 and 1955N), as shown in Fig.19B, which provides a unified shader core architecture where a single core or type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, a number of shader cores may vary. In at least one embodiment, the graphics processor 1940 includes an inter-core task manager 1945 acting as a thread dispatcher to dispatch execution threads to one or more shader cores 1955A-1955N, and a tiling unit 1958 for accelerating tiling operations for tile-based rendering, in which rendering operations for a scene are divided in image space, for example, to exploit local spatial coherence within a scene or to optimize the use of internal caches.

[0430] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the integrated circuit 19A and / or 19B may be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0431] Fig. 20A-20B illustrate additional example graphics processor logic according to embodiments described herein. Fig.20A illustrates a graphics core 2000 that, in at least one embodiment, is included within the graphics processor 1810 of Fig. 18 and in at least one embodiment, a unified shader core 1955A-1955N as in Fig. 19B can be. Fig. 20B illustrates a highly parallel general-purpose graphics processing unit (“GPGPU”) 2030 suitable for use on a multi-chip module in at least one embodiment.

[0432] In at least one embodiment, the graphics core 2000 includes a shared instruction cache 2002, a texture unit 2018, and a cache / shared memory 2020 that are common to the execution resources within the graphics core 2000. In at least one embodiment, the graphics core 2000 may include multiple slices 2001A-2001N or a partition for each core, and a graphics processor may include multiple instances of the graphics core 2000. In at least one embodiment, the slices 2001A-2001N may include support logic including a local instruction cache 2004A-2004N, a thread scheduler 2006A-2006N, a thread dispatcher 2008A-2008N, and a set of registers 2010A-2010N.In at least one embodiment, slices 2001A-2001N may include a set of additional function units (AFUs) 2012A-2012N, floating-point units (FPUs) 2014A-2014N, integer arithmetic logic units (ALUs) 2016A-2016N, address computational units (ACUs) 2013A-2013N, double-precision floating-point units (DPFPUs) 2015A-2015N, and matrix processing units (MPUs) 2017A-2017N.

[0433] In at least one embodiment, the FPUs 2014A-2014N may perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPUs 2015A-2015N may perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 2016A-2016N may perform 8-bit, 16-bit, and 32-bit variable-precision integer operations and may be configured for mixed-precision operations. In at least one embodiment, the MPUs 2017A-2017N may also be configured for mixed-precision matrix operations, including floating-point and 8-bit half-precision integer operations.In at least one embodiment, the MPUs 2017A-2017N may perform a variety of matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFUs 2012A-2012N may perform additional logical operations not supported by floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0434] The inference and / or training logic 915 is used to perform inference and / or training operations in connection with one or more embodiments. Details regarding the inference and / or training logic 915 are described below in connection with the Fig.9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 in the graphics core 2000 may be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0435] Fig.20B illustrates, in at least one embodiment, a general-purpose processing unit (GPGPU) 2030 that may be configured to perform highly parallel computational operations by an array of graphics processing units. In at least one embodiment, the GPGPU 2030 may be directly linked to other instances of the GPGPU 2030 to create a multi-GPU cluster to improve training speed for deep neural networks. In at least one embodiment, the GPGPU 2030 includes a host interface 2032 to enable connection to a host processor. In at least one embodiment, the host interface 2032 is a PCI Express interface. In at least one embodiment, the host interface 2032 may be a vendor-specific communication interface or communication fabric.In at least one embodiment, GPGPU 2030 receives instructions from a host processor and uses a global scheduler 2034 to dispatch execution threads associated with those instructions to a set of compute clusters 2036A-2036H. In at least one embodiment, compute clusters 2036A-2036H share a cache 2038. In at least one embodiment, cache 2038 may serve as a master cache for caches within compute clusters 2036A-2036H.

[0436] In at least one embodiment, GPGPU 2030 includes memory 2044A-2044B coupled to compute clusters 2036A-2036H via a set of memory controllers 2042A-2042B. In at least one embodiment, memory 2044A-2044B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.

[0437] In at least one embodiment, the compute clusters 2036A-2036H each include a set of graphics cores, such as the graphics core 2000 of Fig.20A, which may include multiple types of integer and floating-point logic units capable of performing computational operations with a range of precisions, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each of compute clusters 2036A-2036H may be configured to perform 16-bit or 32-bit floating-point operations, while another subset of the floating-point units may be configured to perform 64-bit floating-point operations.

[0438] In at least one embodiment, multiple instances of GPGPU 2030 may be configured to operate as a compute cluster. In at least one embodiment, the communication used by compute clusters 2036A-2036H for synchronization and data exchange varies depending on the embodiment. In at least one embodiment, multiple instances of GPGPU 2030 communicate via host interface 2032. In at least one embodiment, GPGPU 2030 includes an I / O hub 2039 that couples GPGPU 2030 to a GPU ...

Claims

[1] A computer system comprising one or more processors and a computer-readable memory storing instructions executable by the one or more processors to cause the computer system to at least: Identifying a strategy (404) to cause a machine (106, 206) to perform at least one movement, the strategy including at least a plurality of strategy levels comprising: a first strategy level (108) for causing the machine (106, 206) to execute a first movement sequence that reaches a neutral state, the first movement sequence being limited by at least a first parameter associated with the machine and a second parameter associated with a range in which the machine is to operate, and a second strategy level (110) for causing the machine (106, 206) to execute a second movement sequence without affecting the neutral state associated with the first strategy level; and Executing the strategy (404) to cause the machine to perform the at least one movement, wherein the at least one movement comprises at least the first movement sequence and the second movement sequence. [2] The computer system of claim 1, wherein the first strategy level (108) comprises at least a first second-order nonlinear differential equation and the second strategy level (110) comprises at least a second second-order nonlinear differential equation, wherein the first second-order nonlinear differential equation is neutral to cause the first motion sequence to come to a standstill and the second second-order nonlinear differential equation is neutral to cause the second motion sequence to come to a standstill. [3] The computer system of claim 1, wherein the first strategy level (108) is a first geometric structure comprising a second-order nonlinear differential equation, and the second strategy level (110) is a second geometric structure comprising another second-order differential equation. [4] The computer system of claim 1, wherein the first parameter comprises at least first data comprising one or more limit values associated with a joint of the machine (106, 206), and the second parameter comprises at least second data comprising a target position to be reached by the machine (106, 206), the target position being in a Euclidean space associated with the region in which the machine is to operate. [5] The computer system of claim 1, wherein the first strategy level (108) is energized by a first Finsler energy and the second strategy level (110) is energized by a second Finsler energy, the first Finsler energy being homogeneous of degree two and the second Finsler energy being homogeneous of degree two. [6] The computer system of claim 1, wherein the machine (106, 206) is an articulated robot comprising at least one arm, wherein the first strategy level causes the at least one arm to perform the first motion sequence comprising a rectilinear motion limited by the first parameter and the second parameter, and wherein the first parameter corresponds to a joint of the at least one arm and the second parameter corresponds to a coordinate in the region in which the articulated robot is to operate. [7] The computer system of claim 1, wherein the second movement sequence is to occur subsequent to the first movement sequence, the second movement sequence being to cause a gripper of the machine to perform a movement based on at least one task that the machine is to perform. [8] The computer system of claim 1, wherein the plurality of strategy levels comprises a third strategy level, the third strategy level causing the machine to execute a third movement sequence without affecting the neutral state associated with the first strategy level and a neutral state associated with the second strategy level, and the third movement sequence is intended to cause the machine to avoid at least one obstacle in the area in which the machine is intended to operate. [9] Device comprising: one or more processors and memory storing executable instructions that, as a result of execution by the one or more processors, cause the device to: Generating a first strategy level (108) to cause a machine (106, 206) to execute a first motion sequence that causes the machine to accelerate to reach a neutral state, and generating a second strategy level (110) to cause the machine to execute a second movement sequence without affecting the neutral state to be reached by the machine; and Execute the first strategy level and the second strategy level to cause the machine to execute the first motion sequence and the second motion sequence. [10] The apparatus of claim 9, wherein the one or more processors and the memory store executable instructions that, as a result of execution by the one or more processors, further cause the apparatus to generate a strategy comprising the first strategy level (108) and the second strategy level (110), the strategy to cause the machine to perform at least one action. [11] The apparatus of claim 10, wherein the first strategy level (108) and the second strategy level are generated sequentially, starting with the first strategy level and followed by the second strategy level. [12] The apparatus of claim 9, wherein the first strategy level (108) comprises at least a first second-order nonlinear differential equation and the second strategy level (110) comprises at least a second second-order nonlinear differential equation, wherein the first second-order nonlinear differential equation is neutral to cause the first motion sequence to come to a standstill and the second second-order nonlinear differential equation is neutral to cause the second motion sequence to come to a standstill. [13] The apparatus of claim 9, wherein the first strategy level (108) is a first geometric structure comprising a second-order nonlinear differential equation, and the second strategy level (110) is a second geometric structure comprising another second-order differential equation. [14] The apparatus of claim 9, wherein the first strategy level (108) comprises at least first data comprising one or more limit values associated with a joint of the machine, and the second strategy level comprises at least second data comprising a target position to be reached by the machine, the target position being in a Euclidean space associated with the region in which a portion of the machine is to operate. [15] The apparatus of claim 9, wherein the first strategy level (108) is energized by a first Finsler energy and the second strategy level (110) is energized by a second Finsler energy. [16] The apparatus of claim 9, wherein the machine is an articulated robot comprising at least one arm, wherein the first strategy level causes the at least one arm to perform the first motion sequence comprising a substantially rectilinear motion. [17] The apparatus of claim 9, wherein the second movement sequence is to occur subsequent to the first movement sequence, the second movement sequence being to cause a gripper of the machine to perform a movement based on at least one task that the machine is to perform. [18] The apparatus of claim 9, wherein the one or more processors and the memory store executable instructions that, as a result of execution by the one or more processors, further cause the apparatus to generate a third strategy level, the third strategy level causing the machine to execute a third movement sequence without affecting the neutral state associated with the first strategy level and a neutral state associated with the second strategy level, and the third movement sequence is to cause the machine to avoid at least one obstacle in the area in which the machine is to operate. [19] Computer-implemented method comprising: Generating a first strategy level (108) to cause a machine to execute a first motion sequence that causes the machine to accelerate to reach a neutral state, and generating a second strategy level (110) to cause the machine to execute a second movement sequence without affecting the neutral state to be reached by the machine; and Executing the first strategy level (108) and the second strategy level to cause the machine to execute the first movement sequence and the second movement sequence. [20] The computer-implemented method of claim 19, further comprising a third strategy level, wherein the third strategy level causes the machine to execute a third movement sequence without affecting the neutral state associated with the first strategy level and a neutral state associated with the second strategy level, and wherein the third movement sequence is intended to cause the machine to avoid at least one obstacle in the area in which the machine is intended to operate. [21] The computer-implemented method of claim 19, wherein the first strategy level (108) and the second strategy level (110) are generated sequentially, starting with the first strategy level and followed by the second strategy level. [22] The computer-implemented method of claim 19, wherein the first strategy level comprises at least a first second-order nonlinear differential equation and the second strategy level (110) comprises at least a second second-order nonlinear differential equation, wherein the first second-order nonlinear differential equation is neutral to cause the first motion sequence to come to a standstill and the second second-order nonlinear differential equation is neutral to cause the second motion sequence to come to a standstill. [23] The computer-implemented method of claim 19, wherein the first strategy level (108) is a first geometric structure comprising a second-order nonlinear differential equation, and the second strategy level (110) is a second geometric structure comprising another second-order differential equation. [24] The computer-implemented method of claim 19, wherein the first strategy level (108) comprises at least first data comprising one or more limit values associated with a joint of the machine, and the second strategy level (110) comprises at least second data comprising a target position to be reached by the machine, the target position being in a Euclidean space associated with the region in which a portion of the machine is to operate.

Citation Information

Patent Citations

  • Robot device, robot device action control method, external force detecting device and external force detecting method

    EP1195231A1

  • Artificial intelligence system for efficiently learning robotic control policies

    US10926408B1

  • Control apparatus and control method for robot arm, robot, control program for robot arm, and robot arm control-purpose integrated electronic circuit

    US20120173021A1

  • Hierarchical and interpretable skill acquisition in multi-task reinforcement learning

    US20190130312A1

  • Viewpoint invariant visual servoing of robot end effector using recurrent neural network

    US20200114506A1