Inverse reinforcement learning management apparatus, inverse reinforcement learning management method, and inverse reinforcement learning management system
The inverse reinforcement learning management device addresses the challenge of generating accurate reward functions with limited learning information by identifying and constraining unsafe state-action pairs, resulting in safer and more optimal behavioral choices.
Patent Information
- Application Number
- JP2023185137
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-14
AI Technical Summary
Existing inverse reinforcement learning methods struggle to generate highly accurate reward functions, especially in scenarios with limited learning information, which can lead to unsafe behaviors being rewarded.
The proposed solution involves an inverse reinforcement learning management device that includes a processor, memory, and storage unit. This device generates an initial reward function based on available learning information, identifies invalid state-action pairs, and applies constraints to the initial reward function to generate a modified reward function, ensuring safer and more accurate behavioral choices.
The approach enables the generation of highly accurate reward functions even with limited learning information, preventing unsafe behaviors and ensuring optimal action selection in various states.
Smart Images

Figure 2025074384000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an inverse reinforcement learning management device, an inverse reinforcement learning management method, and an inverse reinforcement learning management system. [Background technology]
[0002] Reinforcement learning (RL) is used in a variety of fields, including scheduling problems, resource management, and organizational operations.
[0003] "Reinforcement learning" is a type of machine learning method in which an agent (a learning system) learns optimal actions while interacting with the environment. More specifically, the agent observes the state, selects an appropriate action according to that state, and is given a reward according to the selected action. Through trial and error, the agent learns a policy (hereafter referred to as "policy") that specifies the rules for selecting actions that maximize rewards.
[0004] The reward given to an agent is calculated by a so-called reward function. In order to obtain a reinforcement learning agent that can select the optimal action for various states, it is important to design an appropriate reward function. However, this reward function is difficult to design manually because it is necessary to assign appropriate weights to a large number of possible actions in various states.
[0005] In recent years, inverse reinforcement learning (IRL), which automatically designs reward functions, has been attracting attention. In inverse reinforcement learning, when there is a wealth of learning information created by experts that shows appropriate behavior in various states, it is possible to automatically generate a reward function that gives an appropriate reward to an agent by using the learning information.
[0006] One of the conventional proposals for generating reward functions using inverse reinforcement learning is the work of Anwar, Usman et al. (Non-Patent Document 1). In Non-Patent Document 1, they write: "In real-world settings, there are many constraints that are difficult to describe mathematically. However, in the real-world deployment of reinforcement learning (RL), it is important for RL agents to recognize these constraints in order to act safely. In this study, we consider the problem of learning constraints from demonstrations of the agent's behavior that abides by the constraints. We experimentally validate our approach and show that our framework can learn the constraints that the agent is most likely to respect." [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] Anwar, Usman, Shehryar Malik, Alireza Aghasi and Ali Ahmed. “Inverse Constrained Reinforcement Learning.” International Conference on Machine Learning (2020). https: / / arxiv.org / abs / 2011.09999 Summary of the Invention [Problem to be solved by the invention]
[0008] Patent Document 1 describes a means for determining constraints to be applied to a reward function using learning information related to the field to be studied. Patent Document 1 is based on the premise that abundant learning information related to the field to be studied is available. However, in reality, a large amount of highly accurate learning information is required to generate a reward function by inverse reinforcement learning that gives an appropriate reward according to the action selection of an agent, and in the case of a complex problem such as train scheduling, it is difficult to obtain learning information that takes into account the many possible states. If inverse reinforcement learning is performed based on insufficient learning information, the resulting reward function may give a high reward to, for example, dangerous actions, and an agent trained with such a reward function may learn a strategy that leads to an accident. Patent Document 1 does not consider an inverse reinforcement learning means capable of generating a highly accurate reward function even in a situation where there is little learning information.
[0009] Therefore, an object of the present disclosure is to provide an inverse reinforcement learning management means capable of generating a highly accurate reward function even in a situation where there is little learning information. [Means for solving the problem]
[0010] In order to solve the above problems, a representative inverse reinforcement learning device of the present invention includes a processor, a memory, and a storage unit, the storage unit stores learning information related to a predetermined scenario, The memory includes processing instructions for causing the processor to function as a function generation unit that generates an initial reward function that defines a reward to be given to a reinforcement learning agent based on the learning information, a transition management unit that generates an initial policy that defines the behavior of the reinforcement learning agent based on the initial reward function and generates transition information including transitions consisting of a sequence of state-action pairs that indicate specific actions to be performed for specific states in the scenario using the initial policy, a transition selection unit that determines, from the transition information, a subset of transitions that satisfy a predetermined dispersion criterion, and a constraint generation unit that identifies invalid state-action pairs from the subset of transitions, defines constraints on the initial reward function based on the invalid state-action pairs, and applies the constraints to the initial reward function to generate a modified reward function. Effect of the Invention
[0011] According to the present disclosure, it is possible to provide an inverse reinforcement learning management means capable of generating a highly accurate reward function even in a situation where there is little learning information. Other objects, configurations and effects will become apparent from the following description of the preferred embodiment of the invention. [Brief description of the drawings]
[0012] [Figure 1] FIG. 1 illustrates a computer system for implementing an embodiment of the present disclosure. [Diagram 2] FIG. 2 is a diagram illustrating an example of a configuration of an inverse reinforcement learning management system according to an embodiment of the present disclosure. [Diagram 3] FIG. 3 is a diagram illustrating an example of a logical configuration of an inverse reinforcement learning management system according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating an example of the flow of an initial reward function correction process according to an embodiment of the present disclosure. [Diagram 5] FIG. 5 is a diagram illustrating an example of a process flow for determining an invalid state-action pair according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating an example of the first invalid state-action pair determining means according to the embodiment of the present disclosure. [Figure 7] FIG. 7 is a diagram illustrating an example of the second invalid state-action pair determining means according to the embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the present invention is not limited to the embodiment. In addition, in the description of the drawings, the same parts are denoted by the same reference numerals. In addition, terms such as "first," "second," and "third" may be used to describe various elements or components in this disclosure, but it will be understood that these elements or components should not be limited by these terms. These terms are used only to distinguish one element or component from another element or component. Thus, a first element or component discussed below can also be referred to as a second element or component without departing from the teachings of the inventive concept.
[0014] Next, referring to Fig. 1, a computer system 100 for implementing an embodiment of the present disclosure will be described. The mechanisms and devices of the various embodiments disclosed herein may be applied to any suitable computing system. The main components of the computer system 100 include one or more processors 102, memory 104, a terminal interface 112, a storage interface 113, an I / O (input / output) device interface 114, and a network interface 115. These components may be interconnected via a memory bus 106, an I / O bus 108, a bus interface unit 109, and an I / O bus interface unit 110.
[0015] Computer system 100 may include one or more general purpose programmable central processing units (CPUs) 102A and 102B, collectively referred to as processors 102. In some embodiments, computer system 100 may include multiple processors, while in other embodiments, computer system 100 may be a single CPU system. Each processor 102 executes instructions stored in memory 104 and may include an on-board cache.
[0016] In one embodiment, memory 104 may include random access semiconductor memory, storage devices, or storage media (either volatile or non-volatile) for storing data and programs. Memory 104 may store all or a portion of the programs, modules, and data structures that implement the functions described herein. For example, memory 104 may store an inverse reinforcement learning executive application 150. In one embodiment, inverse reinforcement learning executive application 150 may include instructions or descriptions for executing on processor 102 the functions described below.
[0017] In some embodiments, the inverse reinforcement learning manager application 150 may be implemented in hardware via semiconductor devices, chips, logic gates, circuits, circuit cards, and / or other physical hardware devices instead of or in addition to a processor-based system. In some embodiments, the inverse reinforcement learning manager application 150 may include data other than instructions or descriptions. In some embodiments, cameras, sensors, or other data input devices (not shown) may be provided to communicate directly with the bus interface unit 109, the processor 102, or other hardware of the computer system 100.
[0018] Computer system 100 may include a bus interface unit 109 that provides communication between processor 102, memory 104, display system 124, and I / O bus interface unit 110. I / O bus interface unit 110 may couple to an I / O bus 108 for transferring data to and from various I / O units. I / O bus interface unit 110 may communicate via I / O bus 108 with multiple I / O interface units 112, 113, 114, and 115, also known as I / O processors (IOPs) or I / O adapters (IOAs).
[0019] Display system 124 may include a display controller, a display memory, or both. The display controller may provide video, audio, or both data to display device 126. Computer system 100 may also include one or more sensors or other devices configured to collect data and provide the data to processor 102.
[0020] For example, computer system 100 may include biometric sensors to collect heart rate data, stress level data, etc., environmental sensors to collect humidity data, temperature data, pressure data, etc., and motion sensors to collect acceleration data, movement data, etc. Other types of sensors may also be used. Display system 124 may be connected to a display device 126, such as a standalone display screen, a television, a tablet, or a handheld device.
[0021] The I / O interface unit provides the ability to communicate with various storage or I / O devices. For example, the terminal interface unit 112 may be attached to user I / O devices 116, such as user output devices, such as a video display, a television with speakers, and user input devices, such as a keyboard, a mouse, a keypad, a touchpad, a trackball, buttons, a light pen, or other pointing device. A user may use a user interface to enter input data or instructions to the user I / O devices 116 and the computer system 100, and receive output data from the computer system 100, by manipulating the user input devices. The user interface may be displayed on a display, played through speakers, or printed via a printer, for example, via the user I / O devices 116.
[0022] Storage interface 113 allows attachment of one or more disk drives or direct access storage device 117 (usually a magnetic disk drive storage device, but may be an array of disk drives or other storage devices configured to appear as a single disk drive). In some embodiments, storage device 117 may be implemented as any secondary storage device. Contents of memory 104 may be stored in storage device 117 and retrieved from storage device 117 as needed. I / O device interface 114 may provide an interface to other I / O devices such as printers, fax machines, etc. Network interface 115 may provide a communications path to allow computer system 100 and other devices to communicate with each other. This communications path may be, for example, network 130.
[0023] In one embodiment, computer system 100 may be a device that receives requests from other computer systems (clients) without a direct user interface, such as a multi-user mainframe computer system, a single-user system, or a server computer, etc. In other embodiments, computer system 100 may be a desktop computer, a portable computer, a laptop, a tablet computer, a pocket computer, a telephone, a smartphone, or any other suitable electronic device.
[0024] Next, an inverse reinforcement learning management system according to an embodiment of the present disclosure will be described with reference to FIG.
[0025] 2 is a diagram illustrating an example of a configuration of an inverse reinforcement learning management system 200 according to an embodiment of the present disclosure. The inverse reinforcement learning management system 200 is a system that can generate a highly accurate reward function even in a situation where there is little learning information. As shown in FIG. 2, the inverse reinforcement learning management system 200 includes an inverse reinforcement learning management device 210, a communication network 250, and a user terminal 260. In the inverse reinforcement learning management system 200, the inverse reinforcement learning management device 210 and the user terminal 260 may be connected to each other via the communication network 250.
[0026] The inverse reinforcement learning management device 210 is a device for generating a high-precision reward function, and mainly includes a memory 220, a storage unit 230, a processor 244, and an input / output unit 246, as shown in FIG. In one embodiment, the inverse reinforcement learning manager 210 may be implemented by the computer system 100 shown in FIG.
[0027] The memory 220 may be a memory for storing an inverse reinforcement learning management application 150 for implementing the functions of an inverse reinforcement learning management means according to an embodiment of the present disclosure. The inverse reinforcement learning management application 150 may include processing instructions for implementing the functions of software modules such as a function generator 221, a transition manager 225, a transition selector 229, and a constraint generator 232, as shown in FIG.
[0028] The function generating unit 221 is a functional unit that generates a reward function that defines a reward to be given to a reinforcement learning agent (not shown in FIG. 2). For example, the function generating unit 221 may generate an initial reward function that defines a reward to be given to the reinforcement learning agent based on the learning information stored in the learning information DB 236, or may generate a modified reward function by applying a constraint generated by the constraint generating unit 232 to the initial reward function. In one embodiment, the function generating unit 221 may be an existing inverse reinforcement learning means such as Maximum Entropy that generates a reward function based on the learning information. The details of the function generated by the function generating unit 221 will be described later, and therefore will not be described here.
[0029] The transition management unit 225 is a functional unit that generates an initial policy that specifies the behavior of the reinforcement learning agent based on the initial reward function generated by the function generation unit 221, and then generates transition information including transitions consisting of a sequence of state-action pairs that indicate specific actions to be performed for specific states in a predetermined scenario using the initial policy. In one embodiment, the transition management unit 225 may include a reinforcement learning unit (e.g., the reinforcement learning unit 226 shown in FIG. 3) that learns a policy using the reward function, and a transition generation unit (e.g., the transition generation unit 227 shown in FIG. 3) that generates transition information indicating transitions based on the learned policy. The details of the functions of the transition management unit 225 will be described later, and therefore will not be described here.
[0030] The transition selection unit 229 is a functional unit that determines, from the transition information generated by the transition management unit 225, a subset of transitions that meets a predetermined variance criterion. The function of the transition selection unit 229 will be described in detail later, and therefore will not be described here.
[0031] The constraint generation unit 232 is a functional unit that identifies invalid state-action pairs from the subset of transitions selected by the transition selection unit 229, generates constraints for the initial reward function based on the identified invalid state-action pairs, and applies the generated constraints to the initial reward function generated by the function generation unit 221, thereby generating a modified reward function. The function of the constraint generating unit 232 will be described in detail later, and therefore will not be described here.
[0032] The storage unit 230 is a storage area that accommodates a database (hereinafter, "DB") for storing various information according to an embodiment of the present disclosure, and may include a learning information DB 236 as shown in FIG.
[0033] The learning information DB236 is a database that stores learning information used to generate an initial reward function. This learning information is information indicating state-action pairs indicating states that occur in various scenarios and actions to be taken for the states. This learning information may also be manually created by a user such as an expert. As an example, the learning information DB236 may be information including state-action pairs indicating an action of "changing the route from track A to track B" for a state of "track A is flooded" in a scenario of "train operation in rainy weather."
[0034] The processor 244 is a processing unit for executing processing instructions that define the functionality of each functional unit of the inverse reinforcement learning management application 150 stored by the memory 220 .
[0035] The input / output unit 246 is a functional unit for receiving information input to the inverse reinforcement learning management device 210 and outputting data such as a subset of transitions generated by the inverse reinforcement learning management device 210. In one embodiment, the input / output unit 246 may include, for example, a keyboard, a mouse, a display that displays a GUI (Graphical User Interface), and the like. In one embodiment, the input / output unit 246 may provide the user terminal 260 with a GUI that inputs and outputs various information.
[0036] Communications network 250 may include, for example, a local area network (LAN), a wide area network (WAN), a satellite network, a cable network, a WiFi network, or any combination thereof.
[0037] The user terminal 260 is a terminal device that can be used by a user of the inverse reinforcement learning management device 210. By using the user terminal 260, the user can create learning information to be stored in the learning information DB 236, check a subset of transitions, and determine invalid state-action pairs. As an example, the user terminal 260 may include a smartphone, a smart watch, a tablet, a personal computer, etc., and is not particularly limited. For ease of explanation, FIG. 2 illustrates an example of a configuration including one user terminal 260. However, the number of user terminals 260 is not limited, and a configuration including a plurality of user terminals 260 is also possible.
[0038] According to the inverse reinforcement learning management system 200 described above, it is possible to provide an inverse reinforcement learning management means capable of generating a highly accurate reward function even in a situation where there is little learning information.
[0039] Next, a logical configuration of an inverse reinforcement learning management system according to an embodiment of the present disclosure will be described with reference to FIG.
[0040] 3 is a diagram illustrating an example of a logical configuration of an inverse reinforcement learning management system 200 according to an embodiment of the present disclosure. As described above, the inverse reinforcement learning management device 210 is a device for generating a highly accurate reward function, and includes a learning information DB 236, a function generation unit 221, a transition management unit 225 including a reinforcement learning unit 226 and a transition generation unit 227, a transition selection unit 229, a constraint generation unit 232, and a user terminal 260, as shown in FIG.
[0041] First, the function generating unit 221 generates an initial reward function R(s, a) that defines a reward to be given to the reinforcement learning agent included in the reinforcement learning unit 226, based on the learning information stored in the learning information DB 236. As described above, the learning information here is information indicating states that occur in various scenarios and state-action pairs that indicate actions to be taken for the states, and may be created in advance by a user such as an expert. Also, as described above, the function generating unit 221 may be, for example, an existing inverse reinforcement learning means such as Maximum Entropy that generates a reward function based on the learning information.
[0042] The initial reward function R(s,a) generated here is a reward function that specifies the reward given to the reinforcement learning agent, but since it is generated based on limited learning information, its accuracy may be limited. For this reason, this initial reward function R(s,a) may give a high reward to an inappropriate action selected by the reinforcement learning agent for a certain state, and the reinforcement learning agent may learn an inappropriate policy. As an example, when the learning information is information about train operation, the initial reward function R(s,a) may give a high reward to a dangerous operation that can arrive on time, or a conservative operation that excessively prioritizes safety, and the reinforcement learning agent may learn a policy that may lead to an accident or a policy that makes it difficult to arrive on time. In this disclosure, a state-action pair in which an action selected by the reinforcement learning agent for a state is inappropriate is called an "invalid state-action pair." In view of the above, the present disclosure relates to identifying such invalid state-action pairs, generating and applying negatively rewarded constraints to the initial reward function R(s,a), and thereby generating a modified reward function from which a good policy can be learned.
[0043] Next, the reinforcement learning unit 226 of the transition management unit 225 generates an initial policy that specifies the behavior of the reinforcement learning agent based on the initial reward function R(s, a) generated by the function generation unit 221. This initial policy is information indicating a rule for selecting an action to be performed for a given state. However, as described above, since this initial policy is generated using an initial reward function generated based on limited learning information, its accuracy is limited and it may allow invalid state-action pairs.
[0044] Next, the transition generation unit 227 of the transition management unit 225 generates transition information including multiple transitions in a given scenario (e.g., train scheduling) using the initial policy generated by the reinforcement learning unit 226. A transition here is a sequence of state-action pairs indicating a specific action to be taken for a specific state. In the following formula 1, as an example of a transition, a transition τ 1 Shows.
number
[0045] Next, the transition selection unit 229 selects a subset of transitions that satisfy a predetermined variance criterion from among the transitions included in the transition information generated by the transition generation unit 227. The subset of transitions selected here is a set of transitions that may be invalid state-action pairs that include inappropriate actions. The process of selecting a subset of transitions will be described in detail later, and therefore will not be described here.
[0046] Next, the constraint generating unit 232 identifies invalid state-action pairs from the subset of transitions selected by the transition selecting unit 229, defines constraints for the initial reward function generated by the function generating unit 221 based on the identified invalid state-action pairs, and generates a modified reward function by applying the generated constraints to the initial reward function. In one embodiment, the identification of invalid state-action pairs may be performed based on a user instruction input via the user terminal 260. The process of identifying invalid state-action pairs and defining constraints will be described in detail later, and therefore will not be described here.
[0047] According to the inverse reinforcement learning management system 200 shown in FIG. 3, by using the corrected reward function to train a reinforcement learning agent, for example, by the reinforcement learning unit 226, it is possible to obtain a reinforcement learning model that can select optimal actions for various states even in a situation where there is little learning information.
[0048] Next, with reference to FIG. 4, a flow of an initial reward function correction process according to an embodiment of the present disclosure will be described.
[0049] Fig. 4 is a diagram showing an example of the flow of an initial reward function correction process 400 according to an embodiment of the present disclosure. The initial reward function correction process 400 shown in Fig. 4 is a process for generating a corrected reward function by defining constraints based on invalid state-action pairs and applying them to the initial reward function, and is performed by the transition selection unit 229 and the constraint generation unit 232 shown in Figs. 2 and 3.
[0050] First, in step S405, the transition selection unit 229 acquires transition information generated for a given scenario (for example, train scheduling) from the above-mentioned transition generation unit 227. In one embodiment, this transition information is generated based on the above-mentioned initial policy by selecting state-action pairs (s i ,a i ) n (where "n" is a number ranging from 1 to N) generated by the transition generator 227. In this case, the transition information is represented by state-action pairs (s i ,a i ) n A set of transitions T(τ 1 ... τ N ) and a set of rewards R(r 1 ...r N ) may also be included.
[0051] Next, in step S410, the transition selection unit 229 selects a reward set R(r 1 ...r N For each state-action pair (s i ,a i ) n Reward for i , the set of rewards R(r 1 ...r N ) variance for the whole i Here, the variance is the reward r i , the set of rewards R(r 1 ...r N ) is information that quantitatively indicates the variation relative to the whole, and may be calculated by existing statistical analysis means.
[0052] Next, in step S415, the transition selection unit 229 selects each state-action pair (s i ,a i ) n Reward for i , the set of rewards R(r 1 ...r N ) the overall variance v i Based on this, we define a set of transitions T(τ 1 ... τ N ) a subset T of transitions whose variance satisfies a given variance criterion si This variance criterion defines an acceptable variance and may be freely set. As an example, this variance criterion may be "10%". Note that the subset of transitions T si For each state-action pair (s i ,a i ) n Since the transitions are generated based on different initial conditions of the same scenario, they should correspond to similar rewards. Therefore, a subset of transitions T si The state-action pairs (s i ,a i ) nBy identifying state-action pairs that satisfy a variance criterion (e.g., high variance), we can identify a subset T of transitions that may be associated with abnormally high or low rewards and contain the “invalid state-action pairs” mentioned above. si can be selected.
[0053] Next, in step S420, the constraint generation unit 232 selects the subset T si For each transition in , the transition τ n The final state of s F The final state here is the transition τ n The state-action pairs (s i ,a i ) n , and "F" is the state that corresponds to the last state of the sequence of τ n The state-action pairs (s i ,a i ) n As an example, the transition τ 1 For the final state, "s 1 5. By working backwards from this final state and evaluating the validity of preceding state-action pairs, we can determine invalid state-action pairs that contain inappropriate behaviors.
[0054] Next, in step S425 or step S430, the constraint generation unit 232 generates a constraint for each transition τ n The final state of s F Based on this, a process of determining a set U of invalid state-action pairs is performed. Note that since there are two or more means for determining invalid state-action pairs, they are referred to here as a "first invalid state-action pair determination means" and a "second invalid state-action pair determination means", and their details will be described later.
[0055] Next, in step S435, the constraint generation unit 232 obtains from the transition selection unit 229 the set U of invalid state-action pairs determined in step S425 or step S430.
[0056] Next, in step S440, the constraint generation unit 232 calculates the invalid state-action pairs (s iU ,a iU ) n Here, we generate constraints for invalid state-action pairs (s iU ,a iU ) n The constraint on the invalid state-action pair (s iU ,a iU ) n State s in iU Regarding the action a iU This constraint specifies that if an invalid state-action pair R(s iU ,a iU ) n For "R(s iU ,a iU ) n =-1".
[0057] Next, in step S445, the constraint generating unit 232 generates a modified reward function by applying the constraint generated in step S440 to the initial reward function generated by the function generating unit 221. This modified reward function is generated by applying the determined invalid state-action pair (s iU ,a iU ), this reward function imposes a constraint that specifies that a negative reward will be given to the reinforcement learning agent. Therefore, compared to the initial reward function described above, this prevents the selection of risky or overly conservative actions, resulting in a more accurate policy.
[0058] According to the above-described initial reward function modification process 400, a modified reward function can be generated by identifying transitions including invalid state-action pairs based on the variance of rewards for transitions generated based on the initial policy, defining constraints that stipulate that negative rewards are given to the reinforcement learning agent for the identified invalid state-action pairs, and applying the defined constraints to the initial reward function. Then, by using the modified reward function thus generated to train the reinforcement learning agent, for example, by the reinforcement learning unit 226, a reinforcement learning model that can select optimal actions for various states can be obtained even in a situation where there is little learning information.
[0059] Next, a process for determining an invalid state-action pair according to an embodiment of the present disclosure will be described with reference to FIG.
[0060] 5 is a diagram showing an example of a flow of a process 500 for determining an invalid state-action pair according to an embodiment of the present disclosure. The process 500 for determining an invalid state-action pair is a process for determining an invalid state-action pair including an inappropriate action from among a subset of transitions in which the reward satisfies a predetermined variance threshold. As described above, there are two or more means for determining an invalid state-action pair, and therefore in FIG. 5, a first invalid state-action pair determination means is described in steps S525 to S545, and a second invalid state-action pair determination means is described in steps S550 to S560.
[0061] First, in step S505, the constraint generation unit 232 calculates the reward r i A subset T of transitions that satisfies a given variance threshold si Get the.
[0062] Next, in step S510, the constraint generation unit 232 generates a subset T si For each transition in , the transition τ n The final state of s F As mentioned above, the final state s Fis the transition τ n The state-action pairs (s i ,a i ) n , and "F" is the state that corresponds to the last state of the sequence of τ n The state-action pairs (s i ,a i ) n As an example, the transition τ 1 For the final state, "s 1 5". This final state s F By working backwards from the beginning and evaluating the validity of preceding state-action pairs, invalid state-action pairs (s iU ,a iU ) n It is possible to determine the following.
[0063] Next, in step S515, the constraint generation unit 232 generates a constraint for each transition τ n The final state of s F Among them, the final state s corresponding to the invalid state is FU Here, an "invalid state" refers to a state that is not permissible in the scenario. As an example, if the scenario is "train scheduling," a state in which one or more trains have not arrived at the destination may be determined to be an invalid final state. In one embodiment, the constraint generator 232 performs a constraint generation process on each transition τ n The final state of s F is presented to the user via the user terminal 260, and the user is prompted to specify a final state corresponding to an invalid state, thereby generating an invalid final state s FU In one embodiment, the constraint generator 232 may determine the final state s of each transition identified in step S510. F , to the acceptable existing final states, and if it does not match the acceptable existing final state (does not meet a certain similarity criterion), the final state is marked as an invalid final state s FU It may be judged as follows.
[0064] Next, in step S520, the constraint generating unit 232 determines whether the invalid final state s FU The spatiotemporal coordinates c that specify i Here, we determine whether the spatiotemporal coordinate c i (also referred to in this disclosure as the “first spatiotemporal coordinate”) corresponds to the transition τ n The state-action pairs (s i ,a i ) n A particular state in i This is information to specify the geographical location, time and subject of the occurrence of the event. For example, in the case of a "train scheduling" scenario, the spatiotemporal coordinates c i is a specific state s i may include information about the train, station and time when the invalid final state s FU When the location, time and target information of the occurrence of the invalid final state s are recorded and understood in a log or the like, the constraint generation unit 232 FU The spatiotemporal coordinates c that specify i is identifiable, the process proceeds to step S525, and the determination of the set U of invalid state-action pairs is performed by a first invalid state-action pair determination means, which will be described in steps S525 to S545.
[0065] On the other hand, the invalid final state s FU If the location, time, and target information of the occurrence of the condition are not recorded in a log or the like and are unknown, the constraint generation unit 232 generates an invalid final state s FU The spatiotemporal coordinates c that specify i is not identifiable, the process proceeds to step S550, and the determination of the set U of invalid state-action pairs is performed by a second invalid state-action pair determination means, which will be described in steps S550 to S560.
[0066] In step S520, the invalid final state s FU The spatiotemporal coordinates c that specify i If it is determined that the invalid final state s can be identified, then in step S525, the constraint generation unit 232 determines the invalid final state s FUspatiotemporal coordinates c specifying the geographical place, time and subject of occurrence i Here, the constraint generating unit 232 determines whether the final state s FU The spatiotemporal coordinates of c i may be determined based on a log record, or based on a user input received via the user terminal 260. As an example, in the case of a scenario of “train scheduling”, the constraint generating unit 232 may determine whether a particular invalid final state s FU The spatiotemporal coordinates of c i For example, "Train 36, Station B, 14:44" may be determined.
[0067] Next, in step S530, the constraint generating unit 232 calculates the spatiotemporal coordinate c i and transition τ n In the spatiotemporal coordinate c i The state s corresponding to i (For example, an invalid final state s FU ) preceding state-action pair (s i-1 ,a i-1 ) n The spatiotemporal coordinates c that identify i-1 (also referred to as the “second spatiotemporal coordinate” in this disclosure), where the initial value of i is the transition τ n The state-action pairs (s i ,a i ) n may be set to a total number F of
[0068] Next, in step S535, the constraint generating unit 232 i (For example, an invalid final state s FU ) spatiotemporal coordinate c i and transition τ n In state s i The state-action pair (s i-1 ,a i-1 ) n The spatiotemporal coordinates of c i-1 As a result of comparing with the spatiotemporal coordinate c i-1 is the spatiotemporal coordinate c i If it is determined that the similarity criterion for the spatiotemporal coordinate c is not satisfied, the process proceeds to step S540.i-1 is the spatiotemporal coordinate c i If it is determined that the similarity criterion for τ is satisfied, the process returns to step S530, after which “i” is decremented by 1 and the transition τ n Previous state-action pairs (e.g., (s i-2 ,a i-2 ) n ) and compare the spatiotemporal coordinates.
[0069] The similarity criterion here is information for quantitatively evaluating the similarity of multiple spatiotemporal coordinates, and is used to determine whether a certain spatiotemporal coordinate has changed relative to another spatiotemporal coordinate. The evaluation of the similarity of spatiotemporal coordinates may be performed, for example, by natural language processing. As an example, the state s i The spatiotemporal coordinates of c i is "Train 36, Station B, 14:44", and state s i The state-action pair (s i-1 ,a i-1 ) n The spatiotemporal coordinates c that identify i-1 is "Train 36, Station C, 14:44", the constraint generation unit 232 i-1 is the spatiotemporal coordinate c i , and may determine that the similarity criteria are not met.
[0070] Next, in step S540, the constraint generation unit 232 calculates the spatiotemporal coordinate c i The spatiotemporal coordinate c that is judged not to satisfy the similarity criterion for i-1 The state-action pair (s i-1 ,a i-1 ) n It is judged whether or not satisfies a predetermined validity criterion. The validity criterion here is a criterion for judging whether a state-action pair is valid or invalid. i The spatiotemporal coordinate c that is judged not to satisfy the similarity criterion for i-1 The state-action pair (s i-1 ,a i-1 ) nIf it is determined that i satisfies the predetermined validity criterion, the process returns to step S530, after which "i" is decremented by 1 and the transition τ n Previous state-action pairs (e.g., (s i-2 ,a i-2 ) n ) and compare the spatiotemporal coordinates. In one embodiment, the constraint generator 232 calculates the spatiotemporal coordinate c i The spatiotemporal coordinate c that is judged not to satisfy the similarity criterion for i-1 The state-action pair (s i-1 ,a i-1 ) n is presented to the user via the user terminal 260, and the state s i-1 Actions against a i-1 By having the user decide whether the state-action pair (s i-1 ,a i-1 ) n In addition, in one embodiment, the constraint generating unit 232 may determine whether the spatiotemporal coordinate c i The spatiotemporal coordinate c that is judged not to satisfy the similarity criterion for i-1 The state-action pair (s i-1 ,a i-1 ) n Compare the state-action pair (s) to existing acceptable state-action pairs. If the state-action pair (s) does not match an acceptable state-action pair (does not satisfy the similarity criterion), i-1 ,a i-1 ) n may be determined not to meet the efficacy criteria.
[0071] Next, in step S545, the constraint generation unit 232 determines whether the state-action pairs (s i-1 ,a i-1 ) n is an invalid state-action pair, and adds the invalid state-action pair to the set U of invalid state-action pairs.
[0072] In step S520, the invalid final state s FUThe spatiotemporal coordinates c that specify i If it is determined that the transition τ n In the invalid final state s FU Specific states such as i The state-action pair that precedes the i-1 ,a i-1 ) n In one embodiment, the constraint generator 232 now identifies the transition τ n A graph showing each state-action pair in the state s is presented to the user via the user terminal 260. i By accepting user input indicating the region on the graph that corresponds to the state-action pair that precedes (s i-1 ,a i-1 ) n may be specified.
[0073] Next, the constraint generating unit 232 determines the state s i The state-action pair (s i-1 ,a i-1 ) n State s in i-1 It is judged whether the state s satisfies a predetermined validity criterion. i-1 If it is determined that the state s satisfies the validity criterion, the process proceeds to step S560. i-1 does not satisfy the validity criterion, the process returns to step S550, after which "i" is decremented by 1 and the transition τ n Identify the previous state-action pair in In one embodiment, the constraint generator 232 determines whether the state s i-1 is presented to the user via the user terminal 260, and the state s i-1 By having the user decide whether or not state s is appropriate, i-1 In one embodiment, the constraint generator 232 may also compare the state s to acceptable existing states and, if the state s does not match an acceptable existing state (does not meet the similarity criteria), deselect the state s . i-1 may be determined not to meet the efficacy criteria.
[0074] Next, in step S560, the constraint generation unit 232 determines whether the state s i-1 A state-action pair (s i-1 ,a i-1 ) n is determined to be an invalid state-action pair, and the invalid state-action pair (s i-1 ,a i-1 ) n to the set of invalid state-action pairs U. Here, we add the state s i-1 (also referred to as the second state in this disclosure) i-1 ,a i-1 ) n is judged to be an invalid state-action pair if the final state s FU Since is invalid, the invalid final state s FU The inappropriate behavior that caused the invalid final state s FU The valid state immediately before s i-1 Actions taken against a i-1 (In this disclosure, this is also referred to as the second behavior).
[0075] According to the process 500 for determining invalid state-action pairs described above, invalid state-action pairs including inappropriate actions can be determined from the subset of transitions whose rewards satisfy a predetermined variance threshold. Then, as described above, for each invalid state-action pair determined in this way, constraints that give a negative reward to the reinforcement learning agent are generated, and these constraints are applied to the initial reward function, thereby obtaining a reinforcement learning model that can select optimal actions for various states even in a situation where there is little learning information.
[0076] Next, an example of a means for determining an invalid state-action pair according to an embodiment of the present disclosure will be described with reference to FIGS. In the explanation of FIG. 6 and FIG. 7, the transition τ 1 is used as an example of a transition.
[0077] 6 is a diagram illustrating an example of a first invalid state-action pair determination means according to an embodiment of the present disclosure. As described above, the first invalid state-action pair determination means determines whether a transition τ n Invalid final state s in FU The spatiotemporal coordinates c that specify i This is a procedure for determining invalid state-action pairs that is implemented when it is possible to identify
[0078] First, the constraint generating unit 232 calculates the final state s FU As the transition τ 1 In the state "s5 1 " is determined. As described above, this state s5 1 is the transition τ 1 This is the last state in the set of states s5 and is an invalid state that is determined to be inappropriate. Next, the constraint generation unit 232 generates a set of invalid final states s5. 1 Here, the constraint generator 232 identifies the spatio-temporal coordinates c5 that specify the geographical location, time, and subject of the invalid final state s FU The spatiotemporal coordinates c5 of the vehicle may be determined based on a log record, or based on a user input received via the user terminal 260. For example, the user may input the invalid final state s5 1 The spatio-temporal coordinate c5 may be identified by selecting the region in which the event occurred. As an example, the spatio-temporal coordinate c5 may be "Train 36, Station B, 14:44".
[0079] Next, the constraint generating unit 232 calculates the spatiotemporal coordinate c5 and the transition τ 1 In the state s5 corresponding to the spatiotemporal coordinate c5, 1 The state-action pair (s4, a4) that precedes 1 If it is determined that the spatiotemporal coordinate c4 has changed with respect to the spatiotemporal coordinate c5 (i.e., does not satisfy the similarity criterion), the constraint generation unit 232 compares the state-action pair (s4, a4) 1 As described above, the state-action pair (s4, a4) is evaluated to see if it satisfies a certain validity criterion. 1Whether a given state-action pair satisfies a certain validity criterion may be determined by comparison with existing acceptable state-action pairs, or may be determined by a user.
[0080] State-action pair (s4,a4) 1 does not satisfy a predetermined validity criterion, the constraint generator 232 1 to the set of invalid state-action pairs U.
[0081] On the other hand, the state-action pair (s4, a4) 1 satisfies a predetermined validity criterion, or if it is determined that the spatiotemporal coordinate c4 has no change with respect to the spatiotemporal coordinate c5 (i.e., satisfies a similarity criterion), the constraint generator 232 1 The spatiotemporal coordinates c4 that specify the transition τ 1 In the state-action pair (s4, a4) 1 The state-action pair (s3, a3) that precedes 1 When it is determined that the spatio-temporal coordinate c3 has changed relative to the spatio-temporal coordinate c4, the constraint generation unit 232 compares the state-action pair (s4, a4) 1 If it is determined that the spatio-temporal coordinate c3 has no change with respect to the spatio-temporal coordinate c4, the constraint generation unit 232 determines whether the state-action pair (s3, a3) 1 The spatiotemporal coordinates c3 that specify the transition τ 1 In the state-action pair (s3, a3) 1 The state-action pair (s2, a2) that precedes 1 The spatiotemporal coordinate c2 that specifies By repeating this process, transition τ 1 The effectiveness of all state-action pairs can be evaluated.
[0082] In this way, the first invalid state-action pair determination means goes back from the final state in the transition, and determines the validity of state-action pairs that have a change in the spatiotemporal coordinates of the immediately succeeding state-action pair, and omits determining the validity of state-action pairs that have no change in the spatiotemporal coordinates of the immediately succeeding state-action pair, thereby making it possible to efficiently identify state-action pairs that include inappropriate actions.
[0083] 7 is a diagram illustrating an example of the second invalid state-action pair determination means according to the embodiment of the present disclosure. As described above, the second invalid state-action pair determination means determines whether or not the transition τ n Invalid final state s in FU The spatiotemporal coordinates c that specify i is not identifiable.
[0084] First, the constraint generating unit 232 calculates the final state s FU As the transition τ 1 In the state "s5 1 However, in the case of the state s5 1 Since the location, time, or object of occurrence is unknown, the spatiotemporal coordinate c5 cannot be specified. In this case, the constraint generation unit 232 1 The state-action pair (s4, a4) that precedes 1 In this case, the constraint generating unit 232 identifies the state-action pair (s4, a4) 1 If it is determined that the state s4 satisfies the validity criterion, the constraint generation unit 232 deletes the invalid state-action pair (s4, a4) 1 is judged to be an invalid state-action pair and added to the set U of invalid state-action pairs.
[0085] As described above, the state-action pair (s4, a4) including the state s4 determined to satisfy the validity criterion is 1 is judged to be an invalid state-action pair when the final state s5 1 Since is invalid, the final state s5 is invalid. 1The inappropriate behavior that caused the invalid final state s5 1 This is because it can be inferred that the action is a4 that was performed for the valid state s4 immediately preceding the action a4.
[0086] On the other hand, the state-action pair (s4, a4) 1 If state s4 in does not satisfy a predetermined validity criterion, the constraint generator 232 1 The state-action pair (s3, a3) that precedes 1 Identify the state-action pair (s3, a3) 1 If it is determined that the state s3 satisfies the validity criterion, the constraint generation unit 232 deletes the invalid state-action pair (s3, a3) 1 is judged to be an invalid state-action pair and added to the set U of invalid state-action pairs. By repeating this process, transition τ 1 The effectiveness of all state-action pairs can be evaluated.
[0087] In this way, according to the second invalid state-action pair determination means, even if the spatiotemporal coordinates of the final state are unknown, by tracing back from the final state in the transition and determining the validity of each state-action pair, it is possible to reliably identify state-action pairs that include inappropriate actions.
[0088] As described above, an inverse reinforcement learning management means according to an embodiment of the present disclosure generates an initial reward function based on learning information, generates an initial policy that specifies the behavior of a reinforcement learning agent based on the initial reward function, generates transition information using the initial policy, including transitions consisting of a sequence of state-action pairs that indicate specific actions to be performed for specific states in a specified scenario, determines a subset of transitions that satisfy a specified dispersion criterion from the transition information, identifies invalid state-action pairs from the subset of transitions, specifies constraints on the initial reward function based on the invalid state-action pairs, and generates a modified reward function by applying the constraints to the initial reward function.
[0089] In an inverse reinforcement learning management means according to an embodiment of the present disclosure, transition information including multiple transitions for different initial conditions of a given scenario is generated, and then a subset of transitions that satisfy a given dispersion criterion is determined from the transition information, thereby making it possible to identify a subset of transitions that correspond to abnormally high or abnormally low rewards and may include invalid state-action pairs.
[0090] After identifying a subset of transitions that may contain invalid state-action pairs, for each transition in the subset of transitions, we can determine the validity of each state-action pair by working backwards from the final state-action pair to identify invalid state-action pairs in that transition that contain inappropriate actions but are associated with high rewards.
[0091] More specifically, if the spatiotemporal coordinates indicating the location, object, time, etc. at which a state-action pair occurred are known, by comparing the spatiotemporal coordinates of adjacent state-action pairs, it is possible to judge the validity of only those state-action pairs that have a change relative to the adjacent state-action pairs, and to omit the validity judgment for state-action pairs that have a change relative to the adjacent state-action pairs. This shortens the processing time and makes it possible to efficiently judge invalid state-action pairs.
[0092] For the determined invalid state-action pairs, a constraint is generated that gives the reinforcement learning agent a negative reward, and by applying this to the initial reward function, a modified reward function can be generated. In this way, by training the reinforcement learning agent using the generated modified reward function, it is possible to avoid action selection that leads to inappropriate states and obtain a highly accurate policy. As an example, when an inverse reinforcement learning management means according to an embodiment of the present disclosure is applied to train scheduling, the reinforcement learning agent can generate a highly accurate operation control plan that balances safety and immediacy, avoiding risky operations or conservative operations that overly prioritize safety.
[0093] As described above, the inverse reinforcement learning management means according to the embodiment of the present disclosure includes the following aspects.
[0094] (Aspect 1) An inverse reinforcement learning management device, A processor, a memory, and a storage unit are provided, The storage unit is storing learning information for a given scenario; The memory includes: a function generating unit that generates an initial reward function that defines a reward to be given to the reinforcement learning agent based on the learning information; a transition management unit that generates an initial policy that specifies a behavior of the reinforcement learning agent based on the initial reward function, and generates transition information including transitions consisting of a sequence of state-action pairs that indicate specific actions to be taken for specific states in the scenario, using the initial policy; a transition selection unit that determines a subset of transitions that satisfies a predetermined variance criterion from the transition information; a constraint generator for identifying invalid state-action pairs from the subset of transitions, defining constraints on the initial reward function based on the invalid state-action pairs, and applying the constraints to the initial reward function to generate a modified reward function; The inverse reinforcement learning management device includes a processing instruction for causing the processor to function as a
[0095] (Aspect 2) The management unit generating a plurality of transitions corresponding to a plurality of different initial conditions as the transition information for the scenario; 2. The inverse reinforcement learning management device according to claim 1,
[0096] (Aspect 3) The constraint generation unit determining a first spatiotemporal coordinate identifying a first state-action pair that is a final state-action pair in a first transition in the subset of transitions; determining a second spatiotemporal coordinate identifying a second state-action pair that precedes the first state-action pair in the first transition; if the first spatio-temporal coordinate and the second spatio-temporal coordinate do not satisfy a predetermined similarity criterion, determining whether the second state-action pair satisfies a predetermined validity criterion; identifying the second state-action pair as the invalid state-action pair if it is determined that the second state-action pair does not meet a predetermined validity criterion; 3. The inverse reinforcement learning management device according to aspect 1 or 2.
[0097] (Aspect 4) The constraint generation unit determining a third spatio-temporal coordinate identifying a third state-action pair that precedes the second state-action pair in the first transition if the second state-action pair meets a predetermined validity criterion or if the first spatio-temporal coordinate and the second spatio-temporal coordinate meet a predetermined similarity criterion; if the second spatio-temporal coordinate and the third spatio-temporal coordinate do not satisfy a predetermined similarity criterion, determining whether the third state-action pair satisfies a predetermined validity criterion; identifying the third state-action pair as the invalid state-action pair if it is determined that the third state-action pair does not meet a predetermined validity criterion; 4. An inverse reinforcement learning management device according to claim 3.
[0098] (Aspect 5) The constraint generation unit determining whether a second state in a second state-action pair preceding a first state-action pair that is a final state-action pair in a first transition in the subset of transitions satisfies a predetermined validity criterion; if the second state satisfies a predetermined validity criterion, determining that a second action corresponding to the second state in the second state-action pair does not satisfy the validity criterion, and identifying the second state-action pair as the invalid state-action pair; 5. The inverse reinforcement learning management device according to any one of aspects 1 to 4.
[0099] (Aspect 6) The constraint generation unit if the second state does not satisfy a predetermined validity criterion, determining whether a third state in a third state-action pair preceding the second state-action pair satisfies a predetermined validity criterion; if the third state satisfies a predetermined validity criterion, determining that a third action in the third state-action pair corresponding to the third state does not satisfy the validity criterion, and identifying the third state-action pair as the invalid state-action pair; 6. An inverse reinforcement learning management device according to claim 5,
[0100] (Aspect 7) The transition management unit training the reinforcement learning agent using the modified reward function; 7. An inverse reinforcement learning management device according to any one of aspects 1 to 6.
[0101] Although the embodiment of the present invention has been described above, the present invention is not limited to the above-described embodiment, and various modifications are possible without departing from the gist of the present invention. [Explanation of symbols]
[0102] 150 Inverse Reinforcement Learning Management Applications 200 Reverse Reinforcement Learning Management System 210 Inverse Reinforcement Learning Management Device 220 Memory 221 Function Generator 225 Transition Management Department 229 Transition Selection Section 232 Constraint generator 236 Learning Information DB 244 processors 246 Input / output section 250 Communication Network 260 User Terminals
Claims
1. An inverse reinforcement learning management device, A processor, a memory, and a storage unit are provided, The storage unit is storing learning information for a given scenario; The memory includes: a function generating unit that generates an initial reward function that defines a reward to be given to the reinforcement learning agent based on the learning information; a transition management unit that generates an initial policy that specifies a behavior of the reinforcement learning agent based on the initial reward function, and generates transition information including transitions consisting of a sequence of state-action pairs that indicate specific actions to be performed for specific states in the scenario using the initial policy; a transition selection unit that determines a subset of transitions that satisfies a predetermined variance criterion from the transition information; a constraint generator for identifying invalid state-action pairs from the subset of transitions, defining constraints on the initial reward function based on the invalid state-action pairs, and applying the constraints to the initial reward function to generate a modified reward function; The inverse reinforcement learning management device includes a processing instruction for causing the processor to function as a
2. The transition management unit generating a plurality of transitions corresponding to a plurality of different initial conditions as the transition information for the scenario; The inverse reinforcement learning management device according to claim 1 .
3. The constraint generation unit determining a first spatiotemporal coordinate identifying a first state-action pair that is a final state-action pair in a first transition in the subset of transitions; determining a second spatiotemporal coordinate identifying a second state-action pair that precedes the first state-action pair in the first transition; if the first spatio-temporal coordinate and the second spatio-temporal coordinate do not satisfy a predetermined similarity criterion, determining whether the second state-action pair satisfies a predetermined validity criterion; if it is determined that the second state-action pair does not satisfy a predetermined validity criterion, identifying the second state-action pair as the invalid state-action pair; The inverse reinforcement learning management device according to claim 1 .
4. The constraint generation unit determining a third spatio-temporal coordinate identifying a third state-action pair that precedes the second state-action pair in the first transition if the second state-action pair meets a predetermined validity criterion or if the first spatio-temporal coordinate and the second spatio-temporal coordinate meet a predetermined similarity criterion; if the second spatio-temporal coordinate and the third spatio-temporal coordinate do not satisfy a predetermined similarity criterion, determining whether the third state-action pair satisfies a predetermined validity criterion; if it is determined that the third state-action pair does not satisfy a predetermined validity criterion, identifying the third state-action pair as the invalid state-action pair; The inverse reinforcement learning management device according to claim 3 .
5. The constraint generation unit determining whether a second state in a second state-action pair preceding a first state-action pair that is a final state-action pair in a first transition in the subset of transitions satisfies a predetermined validity criterion; If the second state satisfies a predetermined validity criterion, determining that a second action corresponding to the second state in the second state-action pair does not satisfy the validity criterion, and identifying the second state-action pair as the invalid state-action pair. The inverse reinforcement learning management device according to claim 1 .
6. The constraint generation unit if the second state does not satisfy a predetermined validity criterion, determining whether a third state in a third state-action pair preceding the second state-action pair satisfies a predetermined validity criterion; If the third state satisfies a predetermined validity criterion, determining that a third action corresponding to the third state in the third state-action pair does not satisfy the validity criterion, and identifying the third state-action pair as the invalid state-action pair. The inverse reinforcement learning management device according to claim 5 .
7. The transition management unit training the reinforcement learning agent using the modified reward function; The inverse reinforcement learning management device according to claim 1 .
8. An inverse reinforcement learning management method executed in an inverse reinforcement learning management device, The inverse reinforcement learning management device, A processor, a memory, and a storage unit are provided, The storage unit is storing learning information for a given scenario; The memory includes: generating an initial reward function that defines a reward to be given to the reinforcement learning agent based on the learning information; generating an initial policy governing the behavior of the reinforcement learning agent based on the initial reward function; generating transition information including transitions each consisting of a sequence of state-action pairs indicating a specific action to be taken for a specific state in the scenario using the initial policy; determining a subset of the transitions from the transition information that satisfies a predetermined variance criterion; determining whether a first spatiotemporal coordinate can be identified that identifies a first state-action pair that is a final state-action pair in a first transition in the subset of transitions; determining the first spatiotemporal coordinate if the first spatiotemporal coordinate is identifiable; determining a second spatiotemporal coordinate identifying a second state-action pair that precedes the first state-action pair in the first transition; if the first spatio-temporal coordinate and the second spatio-temporal coordinate do not satisfy a predetermined similarity criterion, determining whether the second state-action pair satisfies a predetermined validity criterion; identifying the second state-action pair as an invalid state-action pair if it is determined that the second state-action pair does not satisfy a predetermined validity criterion; determining a third spatiotemporal coordinate identifying a third state-action pair that precedes the second state-action pair in the first transition if the second state-action pair meets a predetermined validity criterion or if the first spatiotemporal coordinate and the second spatiotemporal coordinate meet a predetermined similarity criterion; if the second spatio-temporal coordinate and the third spatio-temporal coordinate do not satisfy a predetermined similarity criterion, determining whether the third state-action pair satisfies a predetermined validity criterion; identifying the third state-action pair as the invalid state-action pair if it is determined that the third state-action pair does not satisfy a predetermined validity criterion; if the first spatiotemporal coordinate is not identifiable, determining whether a second state in a second state-action pair that precedes the first state-action pair satisfies a predetermined validity criterion; if the second state satisfies a predetermined validity criterion, determining that a second action in the second state-action pair corresponding to the second state does not satisfy the validity criterion, and identifying the second state-action pair as the invalid state-action pair; if the second state does not satisfy a predetermined validity criterion, determining whether a third state in a third state-action pair preceding the second state-action pair satisfies a predetermined validity criterion; if the third state satisfies a predetermined validity criterion, determining that a third action in the third state-action pair does not satisfy the validity criterion, and identifying the third state-action pair as the invalid state-action pair; defining a constraint on the initial reward function based on the invalid state-action pairs; applying the constraint to the initial reward function to generate a modified reward function; The inverse reinforcement learning management method includes a processing instruction for causing the processor to execute the processing instruction.
9. Inverse reinforcement learning management device An inverse reinforcement learning management system connected to a user terminal via a communication network, The inverse reinforcement learning management device, A processor, a memory, and a storage unit are provided, The storage unit is storing learning information for a given scenario; The memory includes: a function generating unit that generates an initial reward function that defines a reward to be given to the reinforcement learning agent based on the learning information; a transition management unit that generates an initial policy that specifies a behavior of the reinforcement learning agent based on the initial reward function, and generates transition information including transitions consisting of a sequence of state-action pairs that indicate specific actions to be performed for specific states in the scenario using the initial policy; a transition selection unit that determines a subset of transitions that satisfies a predetermined variance criterion from the transition information; a constraint generator for identifying invalid state-action pairs from the subset of transitions, defining constraints on the initial reward function based on the invalid state-action pairs, and applying the constraints to the initial reward function to generate a modified reward function; and an input / output unit that outputs output information by the reinforcement learning agent trained using the modified reward function to the user terminal; and processing instructions for causing the processor to function as a
Citation Information
Cited By
Long text generation model optimization method and device based on adaptive constraint reward
CN121094110A