Device and method for training control strategy by means of reinforcement learning

A technology of reinforcement learning and control strategy, applied in neural learning methods, general control systems, program control, etc., can solve problems such as slow convergence to the optimal control strategy

Pending Publication Date: 2022-05-27
ROBERT BOSCH GMBH
View PDF0 Cites 0 Cited by
  • Summary
  • Abstract
  • Description
  • Claims
  • Application Information

AI Technical Summary

Problems solved by technology

However, in practice, the convergence to the optimal control strategy can be very slow

Method used

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more

Image

Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
  • Device and method for training control strategy by means of reinforcement learning
  • Device and method for training control strategy by means of reinforcement learning
  • Device and method for training control strategy by means of reinforcement learning

Examples

Experimental program
Comparison scheme
Effect test

Embodiment approach

[0030] Various implementations, particularly the embodiments described below, may be implemented by means of one or more circuits. In one embodiment, a "circuit" may be understood as any type of logic implementing entity, which may be hardware, software, firmware, or a combination thereof. Thus, in one embodiment, a "circuit" may be a hard-wired logic circuit or a programmable logic circuit, such as a programmable processor, eg, a microprocessor. A "circuit" may also be software, such as any type of computer program, implemented or implemented by a processor. According to an alternative embodiment, any other type of implementation of the corresponding functions, which are described in more detail below, may be understood as a "circuit".

[0031] figure 1 Robotic device 100 is shown.

[0032]The robotic device 100 has a robot 101 such as an industrial robotic arm for manipulating or mounting a workpiece or one or more other objects 114 . The robot 101 has manipulators 102 ,...

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

PUM

No PUM Login to View More

Abstract

A method for training a control strategy by means of reinforcement learning is described, having performing a plurality of reinforcement learning training traversal, in each traversal, selecting an action to be performed starting from an initial state of the control traversal for each state of a sequence of states of an agent, selecting an action by specifying a planned range for at least some states, and training the selected action by specifying the planned range for at least some states. The planned range specifies the number of states; determining a plurality of sequences of reachable states starting from respective states having a specified number of states by applying an answer set programming solver to an answer set programming program modeling a relationship between actions and subsequent states reached by the actions; selecting a sequence that provides the maximum reward among the sequences, wherein the reward provided by the determined sequence is the sum of the rewards obtained when the state of the sequence is reached; and selecting, as an action for the respective state, an action usable to start from the respective state to a first state of the selected sequence.

Description

technical field [0001] In general, various embodiments relate to an apparatus and method for training a control policy with reinforcement learning. Background technique [0002] Control devices for machines like robots can be trained by so-called reinforcement learning, English Reinforcement Learning (RL), to perform specific tasks, eg in a production process. The implementation of a task usually involves choosing an action for each state of a sequence of states, that is to say it can be viewed as a sequential decision problem. According to the state reached by the selected action, especially the final state, these actions are rewarded (English return), for example, according to whether the action is allowed to reach it (for example, for achieving the goal of the task). to the final state of the . [0003] Reinforcement learning enables an agent (such as a robot) to learn from experience by adapting its behavior to maximize the rewards it earns over time. Simple trial-and...

Claims

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

Application Information

Patent Timeline
no application Login to View More
Patent Type & AuthorityApplications(China)
IPC IPC(8): B25J9/16B25J13/00G06N3/02G06N3/08
CPCG06N3/02G06N3/08B25J13/00B25J9/1664B25J9/163G05B2219/33056G06N5/04G06N7/01G06F18/217G06F18/24G05B15/02G06N20/00G06F18/24765
InventorD·斯捷潘诺娃J·厄施N·穆斯里乌T·艾特尔F·M·里希特
OwnerROBERT BOSCH GMBH