Method and apparatus for operating a robot
Through the combination of deep deterministic strategy gradient method and artificial neural network, the robot learning process is stabilized by using resultless actions and boundaries, divergence problems in the robot learning process are solved and safe and efficient manipulation is achieved.
Patent Information
- Application Number
- CN202110612407.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-03
- Filing Date
- 2021-06-02
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2041-06-02
AI Technical Summary
The prior art is prone to divergence in the robot learning process, and it is difficult to effectively control the robot to achieve the target state, and there are also safety hazards.
The deep deterministic strategy gradient method is adopted, combined with the first and second artificial neural networks, and by limiting the vector and quality levels of the adjustment quantity, the learning process is stabilized by using resultless actions and boundaries to ensure that the robot safely reaches the target state.
The stability and safety of robot control are achieved, the divergence in the learning process is avoided, and the efficiency and safety of robots in industrial applications are improved.
Smart Images

Figure CN113752251B_ABST
Abstract
Description
Technical Field
[0001] The present invention is based on a method and a device for operating a robot. Background Art
[0002] Robots are used in many industrial applications. The movement strategy that the robot implements in the application can be specified by a controller in a closed control loop or by an agent that learns and specifies the strategy based on a model or independently of the model. Summary of the Invention
[0003] By means of the device and the method according to the independent claims, applications which are improved compared to conventional applications can be realized.
[0004] A method for operating a robot, wherein a first portion of a manipulated variable is determined for controlling the robot for a transition from the first state to a second state based on a first state of the robot and / or its environment and based on outputs of a first model; wherein a second portion of the manipulated variable is determined based on the first state and independently of the first model; wherein a quality level is determined using a second model based on the first state and based on outputs of the first model; wherein at least one parameter of the first model is determined based on the quality level; wherein at least one parameter of the second model is determined based on the quality level and a target value; wherein a target value is determined based on a reward, the reward being assigned to the transition from the first state to the second state. As a result, a remaining strategy for controlling the robot that is particularly effective in reaching a target is adopted without divergence that interferes with the learning process.
[0005] It may be provided that at least one force and at least one torque acting on an end effector of the robot are determined, wherein the first state and / or the second state is determined as a function of the at least one force and the at least one torque.
[0006] Preferably, the first state and / or the second state are defined relative to an axis, wherein a force causes the end effector to move in the direction of the axis, and wherein a torque causes the end effector to rotate about the axis. This manipulation is particularly effective for exploration and industrial applications. This makes exploration (i.e., especially random experimentation with new maneuvers) safer, ensuring that neither the robot nor the manipulated object, nor anyone nearby, is injured.
[0007] It may be provided that a vector is determined that defines a constant portion of the manipulated variable, wherein the vector defines a first force, a second force, a third force, a first moment, a second moment, and a third moment, wherein different axes are defined for these forces, and wherein each moment is assigned another of the different axes. Vectors are particularly well-suited for describing states and for control.
[0008] The first model can include a first function approximator, in particular a first Gaussian process or a first artificial neural network, wherein the first part of the vector defines the input thereto, wherein the input is defined independently of the second part of the vector.
[0009] The second model can include a second function approximator, in particular a second Gaussian process or a second artificial neural network, wherein the vector defines the input thereto.
[0010] It can be provided that a vector defining a manipulated variable is determined; wherein the vector defines a first force, a second force, a third force, a first torque, a second torque, and a third torque; wherein different axes are defined for the forces, wherein each torque is assigned another of the different axes; wherein a first part of the vector is defined independently of an output of a first artificial neural network, in particular a constant one, which the first model includes; wherein a second part of the vector is defined as a function of the output of the first artificial neural network. As a result, the robot can be controlled in a predefined manner using constant variables, and thus, depending on the task, the robot can be moved more quickly into a final state.
[0011] The end effector preferably comprises at least one finger having a section complementary to the workpiece, facilitating manual gripping or self-centering of the surface of the section.
[0012] The theoretical value can be determined based on a bound, wherein the bound is determined based on a graph in which nodes define states of the robot; a subgraph of the graph is determined based on the first state, the subgraph including a first node representing the first state; and the bound is determined based on a value assigned to a node of the subgraph, where a path from the first node to a second node includes the node, the path representing a final state for the robot. Within the subgraph, a Q value assigned to the subgraph can be analytically determined. The Q value can be used as a lower bound.
[0013] Preferably, the graph is determined based on at least one state of the robot, wherein edges defining fruitless actions are assigned to nodes that are leaves in the graph and are not assigned to a final state of the robot. This prevents causes of divergence during the learning process.
[0014] Ineffective actions can be assigned to a particularly constant value for the first portion of the manipulated variable. This allows particularly well-suited limits to be determined to avoid divergence during the learning process. In certain cases, the inclusion of ineffective actions can completely define the limits initially; in other cases, a higher lower limit can be determined than without ineffective actions.
[0015] The target value is preferably determined according to predefined limits. In this way, domain knowledge can be taken into account depending on the task.
[0016] For training the first artificial neural network, a cost function can be determined from the output of the second artificial neural network, wherein parameters of the first artificial neural network are learned during training for which the cost function has a smaller value than for the other parameters.
[0017] In the training of the second artificial neural network, a cost function can be defined for the output of the second artificial neural network based on the output of the second artificial neural network and a theoretical value, wherein the parameters of the second artificial neural network are learned, and the cost function has a smaller value for the parameters of the second artificial neural network than for other parameters.
[0018] A device for operating a robot, characterized in that the device is designed to carry out the method according to one of the preceding claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Further advantageous embodiments can be found in the following description and the accompanying drawings, in which:
[0020] Figure 1 A schematic diagram showing a robot and equipment for operating the robot,
[0021] Figure 2 A schematic diagram showing parts of the device,
[0022] Figure 3 Steps in a method for operating a robot are shown. DETAILED DESCRIPTION
[0023] Figure 1A robot 102 and a device 104 for operating the robot 102 are schematically shown. The robot 102 is configured to grip a first workpiece 106 using an end effector (in this example, a gripping device 108). The robot 102 can move in a workspace 110 in a variety of different postures p. In this example, the postures p can be described using a three-dimensional Cartesian coordinate system. The origin 112 of the coordinate system is centrally located between two fingers 114 of the gripping device 108 in this example, with which the first workpiece 106 can be gripped. Other arrangements of the Cartesian coordinate system are also possible.
[0024] The Cartesian coordinate system defines positions for motion in the workspace 110 using coordinates x, y, z.
[0025] exist Figure 1 , a second workpiece 116 is shown in the workspace 110 . In this example, an opening 118 is provided in the second workpiece 116 , which is configured to receive the first workpiece 106 .
[0026] For example, the first workpiece 106 is a shaft, in particular a motor shaft. For example, the second workpiece 116 is a ball bearing configured to receive the shaft. The ball bearing can be arranged in a motor housing.
[0027] The robot 102 is designed to move the first workpiece 106 in the workspace 110 along a strategically defined trajectory such that the first workpiece 106 is accommodated in the second workpiece 116 at the end of the trajectory.
[0028] The device 104 comprises at least one processor 120 and at least one memory 122 for instructions, the method described below being carried out when the instructions are executed by the at least one processor 120. At least one graphics processing unit (GPU) can also be provided, with which function approximators, in particular artificial neural networks, can be trained particularly efficiently. The at least one processor 120 and the at least one memory 122 can be implemented as one or more microprocessors. The device 104 can be arranged outside the robot 102 or can be arranged in such a way that it is integrated into the robot 102. Data lines can be provided for communication between the processor, the memory, the control device and the robot 102. These data lines are Figure 1 Not shown in the figure.
[0029] The device 104 may include an output device 124, which is designed to control the robot 102. The output device 124 may include an output stage or a communication interface for controlling one or more actuators of the robot 102.
[0030] Figure 2 Portions of the device 104 are shown schematically.
[0031] The device 104 comprises an agent 202 , which is in particular autonomous, and which is configured to interact with its environment.
[0032] At each discrete time step t, agent 202 can observe the state , and take actions according to the policy π , the action Limits the next state After each action, agent 202 receives a reward .
[0033] The environment is illustrated in this example by a Markov decision process having states, actions, transition dynamics, and a reward function. The transition dynamics can be stochastic.
[0034] Future Rewards The expected value of the sum of By results To define, where the factor , when from the state When the tracking strategy π is launched, the future reward is achieved .
[0035] For actions that are independent of the policy π in the first step, we can consider the Q value. The Q value can be considered as the expected value of the sum of future rewards:
[0036] ,
[0037] When the action is performed at the instantaneous time step t And when following policy π from then on, achieve said future reward.
[0038] The goal of agent 202 is to determine the optimal strategy for reaching the final state . You can set and confirm the following instructions (Angabe) : This instruction Indicates whether the robot 102 has completed its task, that is, whether it has reached the final state. , for each current state , select an action , the action Make all relations to the corresponding state Maximize the expected reward for the future state. In this example, the expected reward for the future state is considered in the sum of the future states in a manner weighted by the discount γ, which sum defines the expected reward.
[0039] Agent 202 can do this by determining the Q function for the environment and choosing an action , the action Maximize the Q value of the Q function at each time step t.
[0040] Q-function is used to optimize the function approximator and, for example, the Q function can be determined by estimating the Q value based on experience and an instantaneous estimate of the Q function. :
[0041] .
[0042] This is called temporal difference learning. To confirm, status Whether it is a final state.
[0043] These functions can be approximated by artificial neural networks. For example, two artificial neural networks are used. The first artificial neural network 204 represents the deterministic policy π. The first artificial neural network 204 is called an actor network. The second artificial neural network 206 is based on the state at its input. and actions The second artificial neural network 206 is called a critic network. This behavior is called Deep Deterministic Policy Gradients and is described, for example, in "Continuous control with deep reinforcement learning" by T.P. Lillicrap, J.J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra (arXiv preprint arXiv:1509.02971, 2015).
[0044] Deep Deterministic Policy Gradient is a model-free method for continuous state and action spaces. To avoid instabilities in learning with nonlinear neural networks, bounds on the Q-values can be determined. These bounds are applied during the learning process, ensuring that the learning process remains stable.
[0045] During the learning process, data about the agent's interactions with the environment are stored in a recurrent memory (Wiedergabespeicher). Instead of a list of transitions, rewards, and an indication of whether the final state has been reached, These transitions, rewards and instructions are stored in the replay memory and are processed as follows. It is also possible to store this list and then continuously derive new graphs from it. The transitions in this example include the state ,action , by in the state Implementation Action The state achieved The reward achieved through this and instructions Instead of a list of transitions, a data graph is provided in which the first node defines the state , and the second node is defined by the state Implementation Action The state achieved In the graph, an edge between a first node and a second node defines an action in this case .award and instructions Assigned to this edge. Different transitions in the data graph can lead to divergence of the learning process with different probabilities. For temporal difference learning, the probability is related to the structure of the data graph.
[0046] The probability of a transition that reaches the final state to produce a divergence is minimized relative to the probability of other transitions that do not reach the final state. The probability of forming a state from which the final state can be reached via the path in the graph The probability of causing divergence in the learning process is lower.
[0047] The lower bound can be determined based on the data graph. For example, a subgraph of the data graph can be determined from which all Q-values assigned to the subgraph can be analytically determined, assuming the subgraph is complete. These Q-values can be used as the lower bound for the data graph. Other lower and upper bounds can be defined a priori (i.e., based on domain knowledge about the robot 102, the first workpiece 106, the second workpiece 116, and / or the task to be solved and / or the reward function used).
[0048] One possibility of using limits is to restrict the Q function according to a lower limit LB and an upper limit UB. This results in the Q function:
[0049] .
[0050] In the training of the second artificial neural network 206, the output of the second artificial neural network 206 The cost function can be restricted to mean squared error:
[0051] .
[0052] In this example, the goal of the training is to learn the parameters of the second artificial neural network 206 for which the cost function has a smaller value than for the other parameters. The cost function is minimized, for example, using a gradient descent method.
[0053] It can be set that the following nodes in the data graph are equipped with work without results, that is, zero actions: Limited Action No edges originate from this node. Other nodes can also be assigned to fruitless tasks. In this example, fruitless actions are actions that do not change the state of the robot 102, for example by presetting an acceleration to zero.
[0054] This avoids leaves in the data graph from which no further actions originate. A lower bound can be determined for each transition that ends in an infinite loop. More lower bounds can be determined due to inconclusive actions than if there were no inconclusive actions. This makes it possible for the lower bound to be narrower overall.
[0055] As input to the first artificial neural network 204, in this example, the defined state The first part of the variables of the first artificial neural network 204 is the input and the defined state The second part of the variable is irrelevant.
[0056] The variable can be the variable being estimated These variables may also be measured variables, or they may be calculated or estimated based on measured variables.
[0057] In this example, the estimated force in the x-direction , the estimated force in the y direction , the estimated force in the z direction , the estimated moment of rotation about an axis extending in the x-direction , the estimated moment of rotation about an axis extending in the y direction and the estimated moment of rotation about an axis extending in the z direction , limited status , these forces and moments appear on the gripping device 108 at the instantaneous pose p. In this example, the estimated forces , the estimated force Independently, and with the estimated torque Independently, the inputs to the first artificial neural network 204 are determined. In this example, the second portion of the variables is not used as input to the first artificial neural network 204. This reduces the dimensionality and thus ensures that the first artificial neural network 204 becomes smaller and can therefore be trained more quickly or more simply. In other applications, these variables could be determined and used for this purpose, without using the additional variables. This is particularly advantageous for a movement of the robot 102 in which the first workpiece 106 is inserted into the opening 118 extending in the z-direction.
[0058] In this example, the action is defined by the first part, which is in particular constant, and by the second part In the example of the movement of the gripping device 108, the first part defines the force in the x-direction , force in the y direction , force in the z direction and the moment for rotation about an axis extending in the z direction .force Can be In this example, this means that the robot 102 moves the gripping device 108 continuously along the z-axis of the gripping device 108. The second part defines in this example the torque for rotation about an axis extending in the x-direction and / or defines a moment about an axis extending in the y direction. In this example, the first theoretical value is determined based on the first output variable of the first artificial neural network 204. In this example, the second theoretical value is determined based on the second output variable of the first artificial neural network 204. In this example, each of these output variables is scaled to an interval of [-3, 3]Nm.
[0059] In summary, in this example, as an action Determine the adjustment amount In this example, in the first part is , , The first part can also be set to another value. When the robot 102 is to be controlled differently, the adjustment amount can also be constructed differently. In this example, the adjustment amount is output to the regulator 208 , for adjusting the new posture p of the gripping device 108, the controller 208 controls the robot 102 until the new posture p is reached. In this example, if the controller 208 for the gripping device 108 has reached a stable state (for example, a stable state has been reached at a speed lower than a predetermined critical speed), then the action End. It can be configured that the regulator 208 is configured to determine the state of the robot 102 It may be provided that the controller 208 is configured to indicate that the robot 102 has reached the final state. Set to the first value (for example, 1). You can set it to initialize the indicator with another value , or otherwise indicate Set to another value (such as 0). Can be measured to determine the indication The speed can be calculated or estimated from the measured variables to determine the indication speed.
[0060] As a result, the robot 102 exerts a constant force in the z-direction on the first workpiece 106 , using which the workpiece 106 is moved in the z-direction. The robot 102 is moved relative to the gripping device 108 by moments in the x-direction and the y-direction.
[0061] An example of another application is the following movement of the robot 102: in which the first workpiece 108 is to be screwed into the opening 118. In this case, a continuous rotational movement about an axis extending in the z direction may be meaningful. This can be taken into account by: Independently of the estimated force in the y direction Independently of the estimated force in the z direction Regardless, according to the estimated torque , according to the estimated moment And with the estimated moment Independently, the input of the first artificial neural network 204 is determined. The first output of the first artificial neural network 204 can in this case be the torque The second output of the first artificial neural network 204 in this case can be a first theoretical value for the torque In this case, the second theoretical value for determining the control variable is determined independently of the first artificial neural network 204. Another variable.
[0062] award It can be predetermined by a first reward function that assigns a value for the reward to the transition to the final state , and each additional transfer is assigned a value for the reward .
[0063] award This can be predetermined by a second reward function that assigns a value to the transition depending on the distance between the momentary position p of the gripping device 108 and its position p in the final state. In this example, the final state is reached when the second workpiece 116 accommodates the first workpiece 106 in the opening 118 provided for this purpose.
[0064] In this example, the position error , determine the first reward In this example, the position error Determined as the Euclidean difference vector Norm. For example, the difference vector is determined from the position of the instantaneous posture p from the gripping device 108 and the position of the posture p from the gripping device 108 in the final state. In this example, for this purpose, the orientation error , determine the second reward In this example, the orientation error is determined as the angular error for rotation about the x-axis and the angular error for rotation about the y-axis between the orientation of the momentary pose p from the gripping device 108 and the orientation of the pose p from the gripping device 108 in the final state In this example, rotations about the z axis are not considered. Rotations about the z axis can be considered in other tasks.
[0065] In this example, the following formula is used:
[0066]
[0067] To determine the second reward function, the formula has an adjustable first parameter =0.015 and adjustable second parameter =0.7. After this, the reward In this example it remains in the interval [-1, 0].
[0068] In such an arrangement, it is easily possible for divergence to occur during the learning process, since long path lengths occur between the initial state and the final state in the associated data graph.
[0069] In order to avoid divergence during the learning process (i.e., during training), it is provided that in the learning process, actions without results are taken and additional upper and lower limits UB and LB are used. In this example, the upper and lower limits UB and LB are defined based on the minimum and maximum rewards. In this example, for the Q function Set lower limit and an upper bound of 0, for the Q function Applicable:
[0070] .
[0071] In this example, for the second part of the adjustment amount, a no-result action is defined. Defined as an action without result. In other scenarios, an action without result can also be defined for another part of the adjustment amount.
[0072] In this example, a learning device 210 is provided, which is configured to determine a reward function based on at least one of the reward functions. The learning device 210 is constructed in this example to determine the Q function It can be provided that the learning device 210 is configured to evaluate the indication , and according to the instructions The value of either Arrival status The Q function is determined independently of the Q value when it is the final state The value of, or otherwise determine the Q function according to the Q value The value of . It can be set that the learning device is constructed to use the lower limit LB or the upper limit UB to limit the Q function The value of .
[0073] In order to train the first artificial neural network 204, the Q value at the output of the second artificial neural network 206 may be used. To constrain the cost function:
[0074] .
[0075] The goal of the training in this example is to learn the parameters of the first neural network 204 for which the cost function has a smaller value than for the other parameters. , the Q value The cost function is thus minimized in this example using a gradient descent method.
[0076] In the following, it is described how the first artificial neural network 204 and the second artificial neural network 206 are trained. For the training, the Adam optimizer may be employed.
[0077] The first artificial neural network 204 (ie, the actor network) comprises three fully connected layers in this example, wherein the two hidden layers each comprise 100 neurons. In this example, the estimated variables are used. The forces and moments can be linearly mapped to values in the interval [-1, +1] describing the state in this example. The first artificial neural network 204 includes a set of parameters for the defined action The output layers of the forces and moments of the action are defined in this example. These layers include tanh activation functions. The weights for these layers can be initialized randomly, in particular by using a Glorot uniform distribution. The output of the output layer defines the weights for the action in this example. Adjustment amount In this example, the output layer is two-dimensional. The first output in this example defines the moment in the x-direction The second output defines the moment in the y direction in this example In this example, the control variable is constantly predetermined, independently of the first artificial neural network 204. The first part ,For example , , , The first and second outputs limit the adjustment amount. The second part.
[0078] The second artificial neural network 206 (i.e., the critic network) comprises three fully connected layers in this example, wherein the two hidden layers each comprise 100 neurons. The second artificial neural network 206 comprises an input layer for the following forces and torques: and actions These forces and moments can be linearly mapped to values in the interval [-1, +1] describing the state in this example. The second artificial neural network 206 includes a one-dimensional output layer. The output layer does not include any nonlinearity in this example, and in particular does not include an activation function. At the output of the second artificial neural network 206, the Q value is output. The other layers comprise ReLU activation functions in this example. The weights for these layers may be initialized randomly, in particular by using a He uniform distribution.
[0079] During training, it can be provided that gripping device 108 is moved from a starting position. The starting position can be, for example, one of eight possible predetermined starting positions. In this example, a training episode ends either when the final state is reached or after a predetermined number of calculation steps t=T (e.g., T=1000). In this example, training was performed over a predetermined number of episodes (e.g., 40 episodes). In this example, the weights of first artificial neural network 204 and / or second artificial neural network 206 were determined after a predetermined number of calculation steps t=N (e.g., N=20). After training, a testing phase can be performed. In this testing phase, a predetermined number of episodes can be performed, in this example, eight episodes. This starting position can be distinguished from one or more starting positions from training during the testing phase.
[0080] During training, different actions are performed by the first artificial neural network 204 according to the weights of the first artificial neural network 204. During training, the Q function is implemented by the second artificial neural network 206 according to the weights of the second artificial neural network 206. In this example, training is performed by adapting the weights of the first artificial neural network 204 and / or the second artificial neural network 206. The goal of training is, for example, to adapt the weights of the first artificial neural network 204 and / or the second artificial neural network 206 so that the action determined by the first artificial neural network 204 is , Q function The value determined by the second artificial neural network 206 takes a larger value than the value for the other weights. For example, the following weights are determined for the first artificial neural network 204 and / or the second artificial neural network 206: for the weights, the Q function The value of takes the maximum value. In order to adapt the weights, the following cost function is used in this example: the cost function is defined according to the weights of the first artificial neural network 204 and / or the second artificial neural network 206.
[0081] In this example, a cost function is defined based on the outputs of the first artificial neural network 204 and the second artificial neural network 206. In this example, the weights of the first artificial neural network 204 and / or the second artificial neural network 206 are adapted based on the value of the gradient of the cost function and based on a learning rate, which defines how the value of the gradient of the cost function affects each of the weights. Different learning rates can be selected for the first artificial neural network 204 and the second artificial neural network. Preferably, the first learning rate of the first artificial neural network 204 is less than the second learning rate of the second artificial neural network 206. Advantageously, the first learning rate is one-tenth the second learning rate.
[0082] In addition to the first learning rate and the second learning rate, the above-mentioned predetermined number of calculation steps and / or rounds can also be varied as a hyperparameter of the training.
[0083] The first reward function is preferably used because it requires less computing memory. This is particularly advantageous in embedded systems. The first reward function is a sparse reward function, which is simpler to define generally than the second reward function. This allows the user to complete training more quickly during operation. This is particularly advantageous in industrial environments where robot 102 is to be trained for a task.
[0084] This method achieves the following advantages in training compared to methods based on deep deterministic policy gradients:
[0085] Greater robustness to changes in hyperparameters (Robustheit),
[0086] reaching the final state more reliably,
[0087] Smaller variance in initializing weights using different random seeds,
[0088] Higher robustness to changes in the reward function,
[0089] Higher robustness with respect to limited memory (e.g. in embedded systems),
[0090] Safer exploration.
[0091] Figure 3 The steps in the method for operating the robot 102 are schematically shown.
[0092] In step 300, at least one force and at least one torque are determined, the at least one force and the at least one torque acting on the end effector 108 of the robot 102. It can be provided that the estimated variable is determined , or determine the corresponding variables as described.
[0093] In step 302, a first state is determined based on at least one force and at least one moment. .
[0094] First state In this example, the definition is with respect to an axis, wherein the force causes a movement of end effector 108 in the direction of the axis, and wherein the torque causes a rotation of end effector 108 about the axis.
[0095] In this example, the following vector is determined: , wherein the vector defines a first force, a second force, a third force, a first moment, a second moment and a third moment, wherein different axes are defined for the forces, wherein each moment is assigned a further one of the different axes. For example, by the estimated variable , constraining the vector.
[0096] In step 304, based on a first state of the robot 102 and / or the environment of the robot 102 , and according to the output of the first model, determine the first part of the adjustment amount for adjusting the robot 102 from the first state To the second state transfer to control robot 102.
[0097] The first model includes, for example, a first artificial neural network 204. In this example, the first part of the vector defines the input to the first artificial neural network 204 for determining the first part of the control variable. In this example, the second part of the vector defines the input to the first artificial neural network 204 independently of the second part of the vector.
[0098] In step 306, according to the first state , and independently of the first model, determine the second part of the manipulated variable.
[0099] In this example, a vector defining the control variable is determined. This vector includes a first force, a second force, a third force, a first moment, a second moment, and a third moment, wherein different axes are defined for these forces. Each moment is assigned another of the different axes.
[0100] In this example, the first part of the vector is defined independently, in particular constantly, of the output of the first artificial neural network 204. In this example, the second part of the vector is defined as a function of the output of the first artificial neural network 204.
[0101] In this example, the vector is as follows for the adjustment amount Determined as described.
[0102] In step 308, a data graph is determined based on at least one state of the robot 102. In this example, the last transition performed is used to complete the data graph. Edges defining inconclusive actions can be assigned to nodes that are leaves in the data graph and are not assigned to a final state of the robot 102. Optionally, inconclusive actions are assigned to, in particular, constant values for the first part of the manipulated variable.
[0103] In step 310, a quality level is determined using a second model based on the first state and based on the output of the first model.
[0104] The second model comprises, in this example, a second artificial neural network 206. The vector y defines, in this example, an input to the second artificial neural network 206. The output of the second artificial neural network 206 defines, in this example, a quality level.
[0105] In step 312 , the parameters of the first model are determined based on the quality level. To this end, training is performed as described above for the first artificial neural network 204 .
[0106] In step 314 , a theoretical value is determined as a function of the reward assigned to the transition from the first state to the second state.
[0107] In this example, the theoretical value is determined based on the bounds. The bounds are determined based on the data graph. In this example, a subgraph of the graph is determined based on the first state, and the bounds are determined based on the Q values assigned to the nodes of the subgraph, as described.
[0108] In optional step 314 , it is provided that a target value is determined according to predefined limits which take into account domain knowledge in a task-dependent manner.
[0109] In step 316 , at least one parameter of the second model is determined based on the quality level and the theoretical value. To this end, training is performed as described above for the second artificial neural network 206 .
[0110] The steps of the method can be repeated in this or another sequence for multiple rounds and / or cycles in order to train the first artificial neural network 204 according to the optimal strategy and maximize the Q value of the second artificial neural network 206. It can be provided that the robot 102 is controlled according to the control variable ζ, which is determined by the optimal strategy.
[0111] End effector 108 may include at least one finger 114 having a section complementary to first workpiece 110 and a self-centering surface. This allows for particularly good support under constant downward pressure. Self-centering is also particularly important in the case of high moments that do not act vertically downward relative to workspace 110. This prevents twisting of the object that would otherwise occur.
Claims
1. A method for operating a robot (102), characterized in that A first part of a regulating variable is determined (304) for controlling the robot (102) for a transition of the robot (102) from the first state to a second state based on a first state of the robot (102) and / or an environment of the robot (102) and based on an output of a first model; wherein a second part of the regulating variable is determined (306) based on the first state and independently of the first model; wherein a quality level is determined (310) using a second model based on the first state and based on an output of the first model; wherein at least one parameter of the first model is determined (312) based on the quality level; wherein at least one parameter of the second model is determined (316) based on the quality level and a theoretical value; wherein the theoretical value is determined (314) based on a reward, the reward being assigned to the transition from the first state to the second state.
2. The method according to claim 1, characterized in that At least one force and at least one torque are determined (300), the at least one force and the at least one torque acting on an end effector (108) of the robot (102), wherein the first state and / or the second state are determined (302) based on the at least one force and the at least one torque.
3. The method according to claim 2, characterized in that The first state and / or the second state are defined (302) with respect to an axis, wherein a force causes the end effector (108) to move in the direction of the axis, and wherein a torque causes the end effector (108) to rotate about the axis.
4. The method according to claim 3, characterized in that Determine (302) the following vector: the vector defines a constant part of the adjustment quantity, wherein the vector defines a first force, a second force, a third force, a first moment, a second moment and a third moment, wherein different axes are defined for the forces, and wherein each moment is assigned a further axis from the different axes.
5. The method according to claim 4, characterized in that The first model comprises a first function approximator comprising a first Gaussian process or a first artificial neural network (204), wherein a first portion of the vector defines (304) an input to the first function approximator, wherein the input is defined (306) independently of a second portion of the vector.
6. The method according to claim 5, characterized in that The second model comprises a second function approximator comprising a second Gaussian process or a second artificial neural network (206), wherein the vector defines an input to the second function approximator.
7. The method according to any one of claims 5 to 6, characterized in that A vector defining (304, 306) the regulating variable is determined, wherein the vector defines a first force, a second force, a third force, a first moment, a second moment and a third moment, wherein different axes are defined for the forces, wherein each moment is assigned a further axis of the different axes, wherein a first part of the vector is defined (306) constantly regardless of an output of the first artificial neural network (204), the first model comprising the first artificial neural network (204), wherein a second part of the vector is defined (304) as a function of the output of the first artificial neural network (204).
8. The method according to any one of claims 2 to 6, characterized in that The end effector (108) includes at least one finger (114) having a section complementary to the workpiece (110) to facilitate manual or self-centering implementation of the surface of the section.
9. The method according to any one of claims 1 to 6, characterized in that The theoretical value is determined (310) based on a limit, wherein the limit is determined based on a graph, in which nodes define states of the robot (102), wherein based on the first state, a subgraph of the graph is determined, the subgraph including a first node, the first node representing the first state, wherein the limit is determined based on values assigned to the following nodes of the subgraph: a path from the first node to a second node including the node, the path representing a final state for the robot (102).
10. The method according to claim 9, characterized in that Based on at least one state of the robot (102), the graph is determined (308), wherein edges defining actions without consequences are assigned to nodes that are leaves in the graph and are not assigned to the final state of the robot (102).
11. The method according to claim 10, characterized in that The fruitless action is assigned ( 308 ) to a constant value for the first portion of the manipulated variable, thereby keeping the changing portion of the remaining strategy for determining the limit out of account.
12. The method according to any one of claims 1 to 6, characterized in that The target value is determined (314) based on predefined limits.
13. The method according to claim 6, characterized in that To train the first artificial neural network (204), a cost function is determined as a function of the output of the second artificial neural network (206), wherein parameters of the first artificial neural network (204) are learned during training, for which the cost function has a smaller value than for other parameters.
14. The method according to claim 6, characterized in that In the training of the second artificial neural network (206), a cost function is defined for the output of the second artificial neural network (206) based on the output of the second artificial neural network (206) and the theoretical value, wherein the parameters of the second artificial neural network (206) are learned, and the cost function has a smaller value for the parameters of the second artificial neural network (206) than for other parameters.
15. A device (104) for operating a robot (102), characterized in that The device (104) comprises a processor (120) and a memory (122), wherein a computer program is stored on the memory (122), the computer program comprising instructions configured to implement the method according to any one of claims 1 to 14 when executed on the processor.
16. A computer program product comprising a computer program, characterized in that The computer program includes instructions that, when executed by a computer, implement the steps of the method according to any one of claims 1 to 14 .
Citation Information
Patent Citations
Spherical joint double-arm robot coordination moving method based on geometric projection
CN107584474A
Methods and apparatus for reinforcement learning
US20150100530A1