Method and device for automatically influencing an actuator

Through the exploration strategy of multiple artificial neural networks and upper confidence boundary optimization, the problem of limited resources in model-free reinforcement learning is solved, and the efficient automatic impact and rapid learning of robot actuators are realized.

CN111984000BActive Publication Date: 2025-07-11ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010429169.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-21
Filing Date
2020-05-20
Publication Date
2025-07-11
Estimated Expiration
2040-05-20

AI Technical Summary

Technical Problem

Existing model-free reinforcement learning requires a large number of examples in robotic actuator control to find task objectives, with limited resources and high cost, and the exploration process is time-consuming and labor-intensive.

Method used

The advantages of multiple artificial neural networks are defined by exploring strategies and road point sequences, using upper confidence boundaries to maximize action selection, combining gradient descent methods to train the network, optimize the exploration process, adapt to the environment and robot dynamics uncertainty.

Benefits of technology

In the absence of environmental models, the automatic impact efficiency and learning speed of the robot actuator are significantly improved, resource consumption is reduced, and quick and effective task completion is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111984000B_ABST
    Figure CN111984000B_ABST
Patent Text Reader

Abstract

The present invention relates to a device and a method for automatically influencing an actuator, wherein at least one state of the actuator or its environment is provided by an exploration strategy for learning a policy, wherein an action for automatically influencing the actuator is defined according to the state by the policy, wherein the state value is defined as the expected value of the sum of rewards achieved starting from the state under following the policy, wherein the state-action value is defined as the expected value of the sum of rewards achieved when first performing an arbitrary action in the state and then following the policy, wherein an advantage is defined according to the difference between the state value and the state-action value, wherein, according to the action and the state, a plurality of advantages are defined by a plurality of artificial neural networks independent of each other, wherein the policy for the state defines an action that maximizes the empirical average of the distribution of the plurality of advantages, wherein the exploration strategy pre-gives at least one state that locally maximizes the upper confidence bound, and wherein the upper confidence bound is defined according to the empirical average and variance of the distribution of the plurality of advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is based on a computer-implemented method and a device for automatically influencing an actuator, in particular by means of reinforcement learning. Background Art

[0002] Model-free reinforcement learning enables an agent to learn a task without prior knowledge. In model-free reinforcement learning, first, a premise is set for exploration: the goal of the task is known. However, the path, i.e., the actions required to reach the goal, is unknown. Then, the agent learns to reach the goal. In order to transfer conventional model-free reinforcement learning, for example, in robotics to influence an actuator, millions of examples are required until the robot finds the goal of the task.

[0003] One possibility to avoid this laborious exploration is learning from demonstrations. When learning from demonstrations, a human expert explains the task, i.e., the sequence of actions leading to the goal. Then, the agent also learns to reach the goal starting from a new starting position.

[0004] However, the resources available for this, i.e., for example, time, cost, demonstrations, or monitoring by a person, are usually scarce.

[0005] Therefore, it is desirable to further improve the influence on an actuator by means of reinforcement learning. Summary of the Invention

[0006] This is achieved by the subject matter according to the invention.

[0007] A method for automatically influencing an actuator, in particular a robot, a machine, at least partially autonomous vehicle, a tool or a part thereof, provides at least one state of the actuator or the environment of the actuator by means of a learning-based exploration strategy. An action for automatically influencing the actuator is defined according to the state by the rule. A state value is defined as the expected value of the sum of rewards achieved starting from the state when following the rule. A state-action value is defined as the expected value of the sum of rewards achieved when first performing an arbitrary action in the state and then following the rule. An advantage is defined according to the difference between the state value and the state-action value. A plurality of advantages are defined by a plurality of mutually independent artificial neural networks according to the action and the state. The rule for the state defines an action that maximizes the empirical mean of the distribution of the plurality of advantages. The exploration strategy pre-specifies at least one state that locally maximizes the upper confidence bound. The upper confidence bound is defined according to the empirical mean and variance of the distribution of the plurality of advantages. These advantages are determined or approximated by artificial neural networks, and these artificial neural networks are trained independently of the state value, and independently of the targets of the state value and the state-action value. The action that maximizes the upper confidence bound is the action performed at the instantaneous state where exploration should occur, with the highest probability of achieving a high long-term return. Thus, the (locally) optimal action for the exploration at the instantaneous state is executed.

[0008] Preferably, to automatically influence the actuator, a sequentially arranged sequence of waypoints is provided, which waypoints are defined by the state of the actuator or the state of its environment.

[0009] In one aspect, the actuator is moved to a waypoint, and when the waypoint is reached, an action for the waypoint is determined or executed. Thus, the optimal action for the waypoint is determined and executed.

[0010] In one aspect, it is checked whether the actuator collides with an obstacle in the environment of the actuator when moving from one waypoint in the sequence to the next waypoint. If a collision is identified, the movement to the next waypoint is interrupted, and instead of moving to the next waypoint, a movement to the waypoint following the next waypoint in the sequence, in particular directly following the next waypoint, is started. In one aspect, the exploration is performed under the assumption that there are no obstacles in the environment of the actuator. However, as long as a collision occurs with a certain waypoint during the movement, that waypoint is excluded. This improves the process of the exploration and enables the method to be used in situations where there is no correct environmental model and / or robot dynamics model.

[0011] Preferably, it is checked whether the next waypoint can be reached when moving from one waypoint in the sequence to the next waypoint, wherein if it is determined that the next waypoint cannot be reached, the movement to the next waypoint is interrupted, and wherein instead of moving to the next waypoint, a movement to the waypoint following the next waypoint in the sequence is started. In one aspect, the exploration is performed under the assumption that each waypoint can be reached. However, as long as a waypoint cannot be reached, that waypoint is excluded. This improves the flow of the exploration and enables the method to be used in situations where there is no correct environmental model and / or robot dynamics model.

[0012] Preferably, the sum over multiple upper confidence bounds is determined or approximated, wherein the sequence of waypoints is provided with states whose upper confidence bounds maximize the sum. Thereby, the state that is most uncertain for the exploration is determined as a waypoint. This enables particularly efficient model-free reinforcement learning based on model-based automatic demonstration. In a training method for reinforcement learning, the manipulation of the robot is set, wherein a signal for the manipulation is determined and the state is determined by manipulating the robot.

[0013] Preferably, it is provided that the distribution of the plurality of advantages is determined based on the actions and states, wherein the upper confidence bounds are determined based on the empirical mean and variance of the distribution of the plurality of advantages. In model-free reinforcement learning, the upper confidence bounds are determined based on data points that pair the actions and states with each other. For example, a set of artificial neural networks is used, which determine various advantages based on the actions and the states. The actions and the states for the artificial neural networks are stored, for example, in the replay memory of the artificial neural network. The parameters of these artificial neural networks are preferably initialized with different values, for example randomly. Thereby, different advantages are determined based on the actions and states of different artificial neural networks.

[0014] Preferably, it is provided that a target advantage of the advantages and a target state value of the state values are provided, wherein a gradient descent method is executed, and at least one parameter of the artificial neural network that determines the advantages based on the actions and the states is determined using the gradient descent method based on the target advantage and the advantages, and / or at least one parameter of the artificial neural network that determines the state values based on the states is determined using the gradient descent method based on the target state value and the state values. The data for the gradient descent method, for example from a previous manipulation of the actuator or a simulation of these manipulations, is stored in the replay memory.

[0015] A device for automatically influencing an actuator provides that the device includes a processor and a memory for a plurality of artificial neural networks, the processor and the memory being configured to execute the method.

[0016] The device preferably comprises a plurality of artificial neural networks, which are configured to respectively provide one of the plurality of advantages.

[0017] Preferably, a computer-implemented method for training an artificial neural network to determine an advantage in reinforcement learning comprises: determining digital data regarding states, actions, and rewards for manipulating an actuator according to the upper confidence bound on the advantage distribution, the advantage distribution being provided by a plurality of artificial neural networks respectively for determining one advantage; storing the data in a database; collecting from the database digital data regarding states, actions, and rewards and digital data regarding the target advantages for these states and actions; generating a training data set including the data; training the plurality of artificial neural networks for determining the advantage in a pre-given set of waypoints, wherein the sum over a plurality of upper confidence bounds is determined or approximated, wherein a sequentially ordered sequence of waypoints is provided, which waypoints are defined by the states of the actuator or its environment, wherein the directly successive waypoints can be approached along a straight line in the sequentially ordered sequence, and wherein the following states are provided for the sequence of waypoints, the upper confidence bounds of which states maximize the sum. Description of the Drawings

[0018] Other advantageous embodiments result from the following description and the drawings. In the drawings:

[0019] Figure 1 A schematic diagram of a reinforcement learning system is shown,

[0020] Figure 2 A schematic diagram of an agent for reinforcement learning is shown,

[0021] Figure 3 Steps in a method for influencing an actuator are shown,

[0022] Figure 4 Steps in a method for reinforcement learning are shown,

[0023] Figure 5 A schematic diagram of a device for reinforcement learning is shown,

[0024] Figure 6 A method for training an artificial neural network to determine an advantage in reinforcement learning is shown. Detailed Description

[0025] Figure 1A schematic diagram of a reinforcement learning system is shown. The learning system uses an exploration strategy to explore in order to learn regulations that, although there is still uncertainty in reinforcement learning, have a high probability of success. For example, an artificial neural network as described below is used to approximate the uncertainty. For reinforcement learning, first, the exploration strategy is used to navigate to a particularly uncertain state, and there, the so-called waypoint regulation is used to determine and execute the most uncertain action. The action to be executed can have the following purposes: solving a manipulation task or moving the robot to a navigation target. For example, a vacuum cleaner robot can be set, whose goal is to approach its charging station. Hereinafter, the robot or a part of the robot is regarded as an actuator. An actuator also includes a machine, a vehicle, a tool, or a part of them that is provided with an action. In this article, providing an action means determining the action and / or outputting a control signal according to the action, and the control signal controls the actuator to execute the action.

[0026] In the following description, based on Figure 1 A system 100 for reinforcement learning using an agent 102 is described, and the agent interacts with its environment 104 at time steps t = 1,..., T. In this example, the environment 104 includes a robot 106, and the robot 106 includes an actuator 108. The agent 102 observes the state s at time step t t . In this example, a set S of n states is set as S = {s0, s1,..., s n-1}}. In this example, the state s is defined by the posture of the robot 106 or the actuator 108 according to DIN EN ISO 8373 i ∈S. For example, the position description in a plane can be set by two position coordinates x, y in a Cartesian coordinate system and the rotational position of the actuator, and the rotational position has an angle around the y-axis relative to the plane within between In this example, the origin 110 of the coordinate system is located at the end of the actuator 108, and a workpiece 112 is arranged at this end. The workpiece 112 includes a base 114, a first pin 116, and a second pin 118. In this example, these pins extend from the base 114 at an immutable angle of 90° and parallel to each other in the same direction.

[0027] In this example, the task that the agent 102 should solve is to insert two pins into two corresponding holes 120 of a container 122.

[0028] The agent 102 is configured to control the robot 106 to execute an action a at time step t t . In this example, the robot 106 includes a control device 124, and the control device is configured to according to the action a tManipulate the actuator 108. In this example, the action a t defines the x and y coordinates of the rated position of the workpiece 112 and the angle of the rated orientation of the workpiece 112 The control device 124 is configured to move the actuator 108 such that the workpiece 112 is moved to the rated position in the rated orientation. The control device 124 is configured to move the actuator 108 according to the exploration strategy and / or regulations described below. The robot 106 may include a sensor 126, which is configured to monitor the collision of the robot 106 with an obstacle. In this example, the sensor 126 is configured to measure the momentum acting on the actuator 108. The control device 124 is configured to interrupt the movement when a collision is recognized. In this example, the control device 124 is configured to recognize the collision when the momentum exceeds a threshold. In this example, the actuator 108 is a robotic arm, which is configured to move the workpiece 112 to a position pre-given by the x and y coordinates in the plane according to the action a t and rotate the workpiece 112 to an orientation pre-given by the angle In this example, the state s t is defined by the x and y coordinates of the actual position of the workpiece 112 and the angle of the actual orientation of the workpiece 112 The actual position and the actual orientation are determined by the control device 124 according to the posture of the robotic arm. In this example, the action a t and the state s t are both normalized to a value range between -1 and 1.

[0029] In this example, a set of m actions is set The agent 102 is configured to determine the action a t at time step t according to the state s t observed at time step t. The agent 102 is configured to determine the action a t according to the regulation π(s t ), which associates the action a t with the state s t .

[0030] The trajectory τ π ={s0, a0, s1, a1,..., s T} defines a sequence of states s t and actions a t for t = 1,..., T, which is defined by the regulation π(s t ) for T time steps t. Along the trajectory τ π , a reward r t is defined for time step t, and the reward is obtained by performing the action a at time step tt is achieved. Along the trajectory τ π The sum of all the rewards obtained defines the return where γ ∈ (0, 1] is the discount factor.

[0031] The state s t value of the state is defined by the expected future return that occurs when executing the policy π(s t ). For executing any action a t and then executing the policy π(s t ), the expected future return is defined by the state-action value . In the case where an arbitrary action a is first selected by the agent 102 t and then the policy π(s t ) is executed, the difference between the expected future return in this case and the expected future return in the case of following the policy π(s t ) defines the advantage A π (s t , a t ) = Q π (s t , a t ) - V π (s t ). denotes the corresponding expected value.

[0032] The goal of the agent 102 is to determine the optimal policy π* that maximizes the expected future return for each state s t .

[0033] This is described below based on Figure 2 the agent 102 based on an artificial neural network. The agent 102 includes an input for the state s t and the action a t . The agent 102 includes an artificial neural network. The agent 102 includes a first multi-layer perceptron 202 with four fully-connected layers. The input of the first multi-layer perceptron 202 is defined by the state s t . In this example, the input is defined by the triple x s,t , y s,t , φ s,t , which describes the actual position and actual orientation in the state s t . The agent 102 includes a second multi-layer perceptron 204 with four layers. The input of the second multi-layer perceptron 204 is defined by the output of the first multi-layer perceptron 202. The output of the second multi-layer perceptron 204 defines the state value V π (s t ). More precisely, the state value V π (st ) estimate. The first multi-layer perceptron 202 and the second multi-layer perceptron 204 define a state value network. The agent 102 includes a number K of advantage networks 206-1,..., 206-K. Each of the advantage networks 206-1,..., 206-K includes a third multi-layer perceptron 208-1, 208-2,..., 208-K having four layers. Each of the advantage networks 206-1,..., 206-K includes a fourth multi-layer perceptron 210-1, 210-2,..., 210-K having four layers.

[0034] Instead of a multi-layer perceptron, other artificial neural networks, such as a deep convolutional neural network, can also be used. The structure of a multi-layer perceptron with four fully connected layers can also be implemented with other numbers of layers or only partially connected layers.

[0035] The action a at the input of the agent 102 t defines the input of each of the advantage networks 206-1,..., 206-K. More precisely, the action a t defines the input of its third artificial neural network 208-1, 208-2,..., 208-K in each of the advantage networks 206-1, 206-2,..., 206-K. In this example, the input is defined by the triple x a,t , y a,t , φ a,t which illustrates the actual position and actual orientation when performing the action a t . In each of the advantage networks 206-1,..., 206-K, the output of the first artificial neural network 202 and the output of the third artificial neural network 208-1, 208-2,..., 208-K in the same advantage network 206-1,..., 206-K define the input of the fourth artificial neural network 210-1, 210-2,..., 210-K in the same advantage network 206-1,..., 206-K. In this example, the elements at the input of the fourth artificial neural network 210-1, 210-2,..., 210-K are determined according to the elements of these outputs in the element-wise addition 212. The output of the fourth artificial neural network 210-1, 210-2,..., 210-K defines its advantage for each of the advantage networks 206-1,..., 206-K where k = 1,..., K. More precisely, the advantage is output at the output of the fourth artificial neural network 210-1, 210-2,..., 210-K estimate.

[0036] This means predicting a single state value V π(s t ) and a distribution with K advantages and k = 1,..., K.

[0037] In this example, each artificial neural network contains four fully connected layers. Looking at the input, each artificial neural network in this example uses leaky ReLU activation for the first three layers and linear activation for the last layer. Other network architectures and activations can also be set.

[0038] Agent 102 is constructed to pre - give an action a for a determined state s t according to a regulation π(s t ). In this example, the regulation π(s ) is defined according to the distribution of advantages t :

[0039]

[0040] μ A is determined, for example, by the following formula:

[0041]

[0042] Agent 102 is constructed to use the action a t = π(s t ) to control the robot 106 to reach a pre - given number (for example, T = 200 steps), and update each advantage network separately and distinctly according to the data (s t , a t , r t , s t+1 ) i in the separate replay memories of each of the advantage networks 206 - 1,..., 206 - K. For example, the data (s t , a t , r t , s t+1 ) i is sampled. Agent 102 can be constructed to train different advantage networks 206 - 1,..., 206 - K with different data, for example, by pre - giving different starting states randomly especially at the beginning.

[0043] Agent 102 is constructed to update the state - value network 204 according to the data sampled from all the replay memories.

[0044] Agent 102 is constructed to initialize different artificial neural networks with random parameters, especially random weights.

[0045] In addition, the agent 102 is configured to apply the stochastic gradient descent method to train all artificial neural networks. To this end, the estimates of these artificial neural networks of the advantage are compared with the iteratively computed target advantage and the state value V π (s t ) is compared with the iteratively computed target state value , where the iteratively computed target advantage and the iteratively computed target state value are based on the data (s t , a t , r t , s t+1 ) i stored in the replay memory. For each of the k advantage networks 206-1, ..., 206-K, the gradient descent method is performed according to the comparison result by the following formula:

[0046]

[0047] where,

[0048]

[0049] where is an indicator function indicating whether the state s t+1 is a final state. For example, the indicator function assigns the value 1 of the indicator function to the said final state. For example, the value 0 is assigned to all other states.

[0050] defines the mean μ for K advantages A (s t , a t ) with k = 1, ..., K. The mean μ t of the pre-given inputs s t , a A at each of the advantage networks 206-1, ..., 206-K is, for example:

[0051]

[0052] In this example, the argmax function for calculating the specified π(s t ) is approximated by evaluating a pre-given number (e.g., K = 100) of sampled actions. More precisely, the data uniformly distributed in the replay memory by means of the index i is sampled. In this example, and k denote the indices of the advantage networks with 1, …, K. In this example, In this example, the agent 102 is configured to perform a pre-given number of iterations of the gradient descent method, e.g., 50 iterations, after each data collection, and train the artificial neural network during these iterations. Alternatively, it can be stipulated that in each step, the optimal action for evaluating π(s t ) is performed by a separate gradient descent method. Thus, better and more precise optimization can be achieved while spending more time.

[0053] The described network architecture includes two separate data streams. This reduces the variance in the gradient descent method. Thus, the training can be performed faster and more stably.

[0054] K advantage networks 206-1,..., 206-K form a complete set of artificial neural networks in one of the data streams. This data stream represents an approximate calculation of model uncertainty. There is model uncertainty especially for regions with states that have not been explored.

[0055] In one aspect, the agent 102 is configured to use a rough approximation of the dynamics of the actuator 108 or the robot 106 to select a trajectory for learning an exploration strategy that maximizes the UCB sum of the advantage A for a set of waypoints. The agent 102 is, for example, configured to systematically determine these waypoints by determining a sequence of h waypoints that approximately maximizes the following sum:

[0056]

[0057] For example, the sequence of waypoints is defined by:

[0058]

[0059] UCB A For example, it is determined by:

[0060]

[0061] where

[0062]

[0063] κ represents a tunable hyperparameter, and the tunable hyperparameter is, for example, selected as κ = 1.96.

[0064] The road punctuation points can be determined, for example, using Powell's method: Powell M.J.D. “An efficient method for finding the minimum of a function of several variables without calculating derivatives”, The Computer Journal, Vol. 7, No. 2, 1964, pp. 155 - 162.

[0065] The agent 102 is configured to manipulate the robot 106 according to the exploration strategy so as to sequentially approach the path points. If a point is empirically unreachable (e.g., due to obstacles or because assumptions regarding dynamics are violated), then the approach to that road punctuation point is interrupted and that road punctuation point is skipped. The agent 102 is configured, for example, to detect a collision with the obstacle by means of the sensor 126 and to skip that road punctuation point by interrupting the movement towards that road punctuation point and sequentially approaching the next road punctuation point.

[0066] For example, to manipulate the road punctuation points, it is assumed that each road punctuation point can be approached along a straight line. In this case, the agent 102 is configured to manipulate the next road punctuation point, i.e., in this example, to manipulate the next rated position of the actuator 108 along a straight line starting from the instantaneous actual position of the actuator 108.

[0067] The actuator 102 is configured to determine and execute the optimal action for that road punctuation point once it reaches the road punctuation point. In one aspect, the optimal action is determined by the following road punctuation point specification:

[0068]

[0069] This generates data that can be used to very quickly learn especially contact - intensive manipulation tasks. These trajectories can be used as demonstrations for the reinforcement learning to quickly converge to the specified π * (s t ) during learning. Thus, the exploration is preferably carried out in regions with large uncertainties in the state, but there is also a great chance of obtaining good long - term rewards with respect to the specified π(s t ).

[0070] The following is based on Figure 3 A computer - implemented method for influencing the actuator 108 is described. The agent 102 is configured to automatically influence the actuator 108 using this method.

[0071] After the start of this method, step 300 is executed.

[0072] In step 300, the state s of the actuator 108 or the environment of the actuator 108 is provided. i . In order to automatically affect the actuator 108, in this example, a sequence of waypoints arranged in the order provided according to an exploration strategy for learning is provided, and these waypoints are defined by the state s of the actuator 108 or its environment. h In this example, the directly successive waypoints in the arranged sequence can be approached along a straight line.

[0073] For example, determine or approximate the following sum on multiple upper confidence bounds:

[0074]

[0075] Provide the following states for the sequence of waypoints, and the upper confidence bounds of these states maximize this sum.

[0076] In step 302, the actuator 108 is moved to the waypoint. In this example, the actuator 108 is controlled by the agent to the following waypoint, which is the next waypoint to be manipulated in the arranged sequence of waypoints.

[0077] In the next step 304, it is checked whether the actuator 108 collides with an obstacle in the environment of the actuator 108 when moving from one waypoint in the sequence to the next waypoint.

[0078] If there is a collision, the movement to the next waypoint is interrupted and step 302 is executed. Thus, instead of moving to the next waypoint, the movement to the waypoint in the sequence that follows the next waypoint, especially the waypoint directly following the next waypoint, is started.

[0079] In principle, the exploration is performed under the assumption that there are no obstacles in the environment of the actuator 108. However, as long as a collision occurs during the movement to a certain waypoint, that waypoint is excluded.

[0080] If there is no collision, step 306 is executed.

[0081] In step 306, it is checked whether the next waypoint can be reached when moving from one waypoint in the sequence to the next waypoint. If the next waypoint cannot be reached, the movement to the next waypoint is interrupted and step 302 is executed. Thus, instead of moving to the next waypoint, the movement to the waypoint in the sequence that follows the next waypoint is started. In principle, the exploration is performed under the assumption that each waypoint can be reached. However, as long as a certain waypoint cannot be reached, that waypoint is excluded.

[0082] If the waypoint can be reached, step 308 is executed.

[0083] Steps 304 and 306 are optional. Step 308 can also be directly executed after step 302.

[0084] In step 308, according to the state s i and according to the waypoint regulation π W (s t ) provides an action a for automatically influencing the actuator 108. The state s i corresponds to the waypoint at arrival. When arriving at the said waypoint, step 308 is preferably executed.

[0085] In this example, the state value V(s t ) is defined according to the expected value of the reward r t ), and the said reward is realized based on the state s t when following the regulation π(s i ).

[0086] In this example, the state-action value Q t is defined according to the expected value of the reward r π (s t , a t ), and the said reward is realized when first performing an arbitrary action a i in the state s t and then executing the regulation π(s t ).

[0087] In this example, the advantage is defined according to the difference between the state value V(s t ) and the state-action value Q π (s t , a t ). The regulation π(s t ) defines an action a for the state s i , and predicts the maximum empirical average value of multiple advantages for this action:

[0088] π(s t ) = arg max a∈A μ A (s t , a)

[0089] Multiple advantages are defined according to the action a and the state s i

[0090] The upper confidence bound is determined according to the empirical average value μ A (s i , a) and the factor κ. The factor κ is among multiple advantages ​is defined on the distribution. κ can be, for example, 1.96.

[0091] The advantages are determined or approximated by an artificial neural network, which is trained independently of the state value and the targets of the state-action values. Multiple advantages are predicted for it The action with the maximum empirical mean value is the most uncertain action at the instantaneous state where the exploration should be performed. Thus, the optimal action for the exploration at the instantaneous state is executed.

[0092] Then step 310 is executed. In step 310, the action a is executed. Thus, the optimal action a for this waypoint is determined and executed.

[0093] Then in step 312, it is checked whether the next waypoint in the sequence is defined by following the controlled waypoint. If so, step 302 is executed. Otherwise, the method ends.

[0094] It can be stipulated that, in order to further explore, a certain number of random actions are to be executed after reaching a certain waypoint and executing the optimal action. For example, fewer than 10 (e.g., 5) random actions are executed.

[0095] Below is based on Figure 4 Describe a method for reinforcement learning.

[0096] After starting, step 400 is executed.

[0097] In step 400, multiple advantages are determined according to the action a and the state s i to determine multiple advantages of the distribution.

[0098] In this example, according to the empirical mean value μ of the distribution of multiple advantages A (s i , a) the upper confidence bound is determined.

[0099] For example, in step 400, according to the method described previously based on Figure 3 the exploration is performed by using a specified exploration strategy for multiple iterations.

[0100] In step 402 following step 400, it is checked whether the exploration has been completed. If the exploration has been completed, step 404 is executed. Otherwise, step 400 is executed for the next iteration of the exploration.

[0101] In step 404, the target advantage of the advantage and the target state value V^ of the state value V(s t ) are provided.

[0102] In the next step 406, the gradient descent method is executed, and using this gradient descent method, at least one parameter of one of the artificial neural networks is determined according to the target advantage and the advantage where the artificial neural network determines the advantage according to the action a and the state s i to determine the advantage Alternatively or additionally, using the gradient descent method, at least one parameter of the following artificial neural network is determined according to the target state value V^ and the state value V(s t ) where the artificial neural network determines the state value V(s i ) according to the state (s t ).

[0103] The gradient descent method is executed in a multiple-iteration manner. In this example, in step 408 following step 406, it is checked whether the training of the agent 102 using the gradient descent method has been iterated a pre-given number of repetitions. If so, step 410 is executed. Otherwise, step 400 is executed.

[0104] In step 410, the agent 102 trained in this way is used to control the actuator 108. Then the method ends.

[0105] In Figure 5 the actuator 102 is exemplarily shown. The agent 102 includes a processor 500 and a memory 502 for the described artificial neural network. The processor 500 is, for example, a microprocessor or a graphics processing unit of a graphics card. Multiple processors can be provided. The agent 102 is configured to execute the described method. More precisely, the actuator 102 includes artificial neural networks 210-1,..., 210-K, which are configured to respectively provide one of multiple advantages among

[0106] In Figure 6 a method for training an artificial neural network to determine the advantage in reinforcement learning is shown. The method, for example, includes the following steps:

[0107] 602: Determine digital data on the state, action, and reward for controlling the actuator 108 according to the upper confidence bound of the distribution regarding the advantage where the distribution is provided by multiple artificial neural networks respectively used to determine one advantage .

[0108] 604: Store the data in a database, for example, in the replay memory of the corresponding artificial neural network.

[0109] 606: Collect digital data on states, actions, and rewards from the database, as well as digital data on the target advantages of these states and actions. The target advantage is determined according to a calculation rule, for example. For example, randomly initialize an artificial neural network for this purpose.

[0110] 608: Generate a training data set including the data.

[0111] 610: Use the training data set to train the multiple artificial neural networks in the pre-given waypoints to determine the advantages.

Claims

1. A method for automatically influencing an actuator (108), characterized in that, Providing (300) at least one state of the actuator (108) or the environment of the actuator (108) by means of a learning prescribed exploration strategy, wherein an action for automatically influencing the actuator (108) is defined (308) according to the state by the prescription, wherein a state value is defined as the expected value of the sum of rewards achieved starting from the state when following the prescription, wherein a state-action value is defined as the expected value of the sum of rewards achieved when first performing an arbitrary action in the state and then following the prescription, wherein an advantage is defined according to the difference between the state value and the state-action value, wherein a plurality of advantages are defined by a plurality of mutually independent artificial neural networks according to the action and the state, wherein the prescription for the state defines an action that maximizes the empirical average of the distribution of the plurality of advantages, wherein the exploration strategy pre-specifies at least one state that locally maximizes the upper confidence bound, and wherein the upper confidence bound is defined according to the empirical average and variance of the distribution of the plurality of advantages.

2. The method according to claim 1, wherein The actuator (108) is a robot or a part thereof, or the actuator (108) is a machine or a part thereof, or the actuator (108) is a tool or a part thereof, or the actuator (108) is at least partially autonomous vehicle or a part thereof.

3. The method according to claim 1, characterized in that, To automatically influence the actuator (108), a sequence of ordered waypoints is provided (300) via the exploration strategy, the waypoints being defined by the actuator (108), or the state (s h ) of the environment of the actuator.

4. The method according to claim 3, characterized in that, Moving (302) the actuator (108) to a waypoint, wherein when the waypoint is reached, an action (a) for the waypoint is determined or executed.

5. The method according to claim 3 or 4, characterized in that Checking whether the actuator (108) collides with an obstacle in the environment of the actuator (108) when moving from one waypoint in the sequence to the next waypoint, wherein if a collision is identified, the movement to the next waypoint is interrupted, and wherein instead of moving to the next waypoint, the movement to the waypoint directly following the next waypoint in the sequence is started.

6. The method according to claim 3 or 4, characterized in that, Checking whether the next waypoint is reachable when moving from one waypoint in the sequence to the next waypoint, wherein if it is determined that the next waypoint is not reachable, the movement to the next waypoint is interrupted, and wherein instead of moving to the next waypoint, the movement to the waypoint following the next waypoint in the sequence is started.

7. The method according to claim 3 or 4, characterized in that Determining or approximating the sum over a plurality of upper confidence bounds, wherein the following states are provided for the sequence of waypoints, the upper confidence bounds of which maximize the sum.

8. The method according to any one of claims 1 to 4, characterized in that Determining (400) the distribution of the plurality of advantages according to the action and the state, wherein the upper confidence bound is determined (400) according to the empirical average and variance of the distribution of the plurality of advantages.

9. The method according to claim 8, wherein The target advantage that provides the advantages (404) and the target state value of the state value, wherein a gradient descent method is executed, and at least one parameter of an artificial neural network that determines the advantage according to the action and the state is determined by using the gradient descent method based on the target advantage and the advantage, and / or at least one parameter of an artificial neural network that determines the state value according to the state is determined by using the gradient descent method based on the target state value and the state value.

10. A device (102) for automatically influencing an actuator (108), characterized in that, The device (102) includes a processor (500) and a memory for multiple artificial neural networks, and the processor and the memory are configured to execute the method according to any one of claims 1 to 9.

11. The device (102) according to claim 10, characterized in that, The processor is a processor of a graphics card.

12. The device (102) according to claim 10 or 11, characterized in that, The device (102) includes multiple artificial neural networks (210-1,..., 210-K), and the multiple artificial neural networks are configured to respectively provide one of the multiple advantages.

13. A machine-readable storage medium having a computer program stored thereon, characterized in that, The computer program includes computer-readable instructions that, when executed on a computer, run the method according to any one of claims 1 to 9.

14. A computer program product, characterized in that, The computer program product includes computer-readable instructions that, when executed on a computer, run the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Robot reinforced learning initialization method based on neural network

    CN102402712A

  • Reinforcement learning using advantage estimates

    CN108701251A