Agent control device, learning device, agent control method, learning method, agent control program, and learning program
Patent Information
- Application Number
- JP2022149607
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-09-20
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2042-09-20
Smart Images

Figure 0007920777000001 
Figure 0007920777000002 
Figure 0007920777000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an agent control device, a learning device, an agent control method, a learning method, an agent control program, and a learning program.
Background Art
[0002] Multi-agent navigation technology has been disclosed that enables each agent to reach a destination (goal) without colliding with each other in an environment where there are multiple agents such as autonomously traveling mobile robots, autonomous driving vehicles, and drones (see Non-Patent Documents 1 and 2). Non-Patent Documents 1 and 2 disclose technology for minimizing the sum of path lengths (required time) of all agents, or the path length (required time) of the agent that arrives at the goal latest.
Prior Art Literature
Non-Patent Literature
[0003]
Non-Patent Literature 1
Non-Patent Literature 2
Summary of Invention
Problem to be Solved by the Invention
[0004] When there are differences in routes and operations among multiple agents, minimizing the sum of the route lengths (travel times) of all agents, or the route length (travel time) of the agent that reaches the goal last, may result in a specific agent arriving at the destination extremely late. The former minimizes the overall time by using the differences in route lengths and travel times of individual agents, while the latter minimizes route lengths and travel times by increasing the priority of a specific agent and decreasing the priority of other agents.
[0005] For these reasons, when navigating multiple agents with different origins, destinations, and routes, minimizing the variation in route length and travel time for each agent is not considered because it would hinder the aforementioned minimization.
[0006] This disclosure is made in view of the above points, and aims to provide an agent control device, a learning device, an agent control method, a learning method, an agent control program, and a learning program for reducing the variation in the cost required for agent movement in an environment where multiple agents exist. [Means for solving the problem]
[0007] In one aspect of this disclosure, an agent control device is provided, comprising: a movement determination unit that inputs a first observation set of observed states of a controlled agent and a second observation set of observed states of at least one other agent surrounding the controlled agent to a first model and determines information relating to the movement of the controlled agent based on the output of the first model; a change amount determination unit that inputs the first observation set and the second observation set to a second model and determines a change amount for the information relating to the movement of the controlled agent based on the output of the second model; and an operation control unit that operates the controlled agent by applying the change amount determined by the change amount determination unit to the information relating to the movement determined by the movement determination unit.
[0008] This allows for the application of the output of the second model to the output of the first model, thereby controlling agent movement in environments with multiple agents to reduce variability in the costs associated with agent movement. These costs could include, for example, the time required to travel from start to finish, the distance traveled, or the amount of energy consumed during the journey. Furthermore, "cost" is synonymous with "value."
[0009] The change amount determination unit may determine the change amount based on the output of the second model in order to reduce the variation in the compensation required for the movement of the controlled agent and the other agents. This makes it possible to control the movement of the agents in a way that reduces the variation in the compensation borne by each agent.
[0010] The change amount determination unit may determine the change amount based on the output of the second model in order to reduce the variation in delays from the planned travel time of the controlled agent and the other agents. This makes it possible to control the movement of the agents in order to reduce the variation in delays from the planned travel time of each agent.
[0011] The change amount determination unit may determine the change amount based on the output of the second model to reduce the variation in the extension of the controlled agent and the other agents from their planned travel distances. This makes it possible to control the movement of the agents in a way that reduces the variation in the extension of each agent from their planned travel distance.
[0012] The change amount determination unit may determine the change amount based on the output of the second model to reduce the variation in the increase from the planned energy consumption of the controlled agent and the other agents. This makes it possible to control the movement of the agents to reduce the variation in the extension from the planned travel distance of each agent.
[0013] The change amount determination unit may determine the change amount based on the output of the second model in such a way as to reduce the average cost required for the movement of the controlled agent and the other agents. This makes it possible to control the movement of the agents in such a way as to reduce the average cost required for the movement of each agent.
[0014] The information relating to movement may include at least one of the following: the direction of movement of the controlled agent, the amount of movement, the time required for movement, the direction of rotation, and the amount of rotation. This allows the movement of the agent to be controlled by changing at least one of the direction of movement, the amount of movement, the time required for movement, the direction of rotation, and the amount of rotation.
[0015] The motion control unit may determine whether the controlled agent incurs a cost for movement based on the output of the second model when only the first observation set is input to the second model. This makes it possible to determine whether each agent incurs a cost for movement.
[0016] Furthermore, from another perspective of this disclosure, a learning device is provided, which includes a learning unit that takes a first observation set of observations of the state of a controlled agent and a second observation set of observations of the state of at least one other agent surrounding the controlled agent as inputs, inputs the first observation set and the second observation set into a first model, and performs learning of a second model whose output is the amount of change in information regarding the movement of the controlled agent determined by the output of the first model.
[0017] This allows for the training of a second model that applies the output of the first model to control agent movement in an environment with multiple agents, thereby reducing the variability in the costs incurred by agent movement.
[0018] Another aspect of this disclosure provides an agent control method in which a processor inputs a first observation set of the state of a controlled agent and a second observation set of the state of at least one other agent surrounding the controlled agent into a first model, determines information regarding the movement of the controlled agent based on the output of the first model, inputs the first and second observation sets into a second model, determines a change in the information regarding the movement of the controlled agent based on the output of the second model, and performs a process to operate the controlled agent by applying the determined change to the determined information regarding the movement.
[0019] This allows for the application of the output of the second model to the output of the first model, thereby controlling agent movement in environments with multiple agents to reduce variations in the costs associated with agent movement.
[0020] Another aspect of this disclosure provides a learning method in which a processor takes a first set of observations of the state of a controlled agent and a second set of observations of the state of at least one other agent surrounding the controlled agent as inputs, inputs the first set of observations and the second set of observations into a first model, and performs a process of learning a second model whose output is the amount of change in information regarding the movement of the controlled agent determined by the output of the first model.
[0021] This allows for the training of a second model that applies the output of the first model to control agent movement in an environment with multiple agents, thereby reducing the variability in the costs incurred by agent movement.
[0022] According to another aspect of the present disclosure, there is provided an agent control program that causes a computer to execute a process of: inputting a first observation set obtained by observing a state of a controlled agent and a second observation set obtained by observing a state of at least one other agent around the controlled agent into a first model, determining information relating to movement of the controlled agent based on an output of the first model; inputting the first observation set and the second observation set into a second model, determining an amount of change for the information relating to movement of the controlled agent based on an output of the second model; and causing the controlled agent to operate by applying the determined amount of change to the determined information relating to movement.
[0023] Accordingly, by applying the output of the second model to the output of the first model, movement of the agent can be controlled so as to reduce variation in the cost required for the agent to move in an environment where a plurality of agents exist.
[0024] According to another aspect of the present disclosure, there is provided a learning program that causes a computer to execute a process of training a second model, wherein the second model receives, as inputs, a first observation set obtained by observing a state of a controlled agent and a second observation set obtained by observing a state of at least one other agent around the controlled agent, the first observation set and the second observation set are input to a first model, and the second model outputs an amount of change for information relating to movement of the controlled agent determined based on the output of the first model.
[0025] Accordingly, the second model whose output is applied to the output of the first model for controlling agent movement can be trained to reduce variation in the cost required for agent movement in an environment where a plurality of agents exist. Effects of the Invention
[0026] According to this disclosure, it is possible to provide an agent control device, a learning device, an agent control method, a learning method, an agent control program, and a learning program for reducing the variability in the cost required for agent movement in an environment where multiple agents exist. [Brief explanation of the drawing]
[0027] [Figure 1] This diagram illustrates an example where multiple agents move to their respective target locations. [Figure 2] This figure shows a schematic configuration of an agent navigation system according to an embodiment of the disclosed technology. [Figure 3] This is a block diagram showing an example of a learning device hardware configuration. [Figure 4] This is a block diagram showing an example of the functional configuration of a learning device. [Figure 5] This is a block diagram showing an example of the hardware configuration of an agent control unit. [Figure 6] This is a block diagram showing an example of the functional configuration of an agent control device. [Figure 7] This is a flowchart showing the learning process flow by the learning device. [Figure 8] This flowchart shows the flow of agent control processing by the agent control unit. [Figure 9] This diagram illustrates a specific example of controlling agent movement using an agent control device. [Figure 10] This diagram illustrates a specific example of controlling agent movement using an agent control device. [Figure 11] This diagram illustrates a specific example of controlling agent movement using an agent control device. [Figure 12] This diagram illustrates a specific example of controlling agent movement using an agent control device. [Figure 13] This diagram illustrates a specific example of controlling agent movement using an agent control device. [Figure 14]This diagram illustrates a specific example of controlling agent movement using an agent control device. [Modes for carrying out the invention]
[0028] First, the Discloser will explain the circumstances that led to the conceiving of the embodiments of this disclosure. As mentioned above, there is a multi-agent navigation technology that enables multiple agents, such as self-driving mobile robots, autonomous vehicles, and drones, to reach their destination without colliding with each other in an environment where multiple agents exist.
[0029] However, if the goal is to minimize the sum of the path lengths (travel times) of all agents, or the path length (travel time) of the agent that reaches the goal last, certain agents may arrive at their destination extremely late. Each agent often has to take detours to avoid other agents or wait for other agents to pass, resulting in delays in arrival time compared to taking a selfish path while completely ignoring other agents.
[0030] Figure 1 shows an example of multiple agents moving to their respective destinations. Agents A and B are both represented as single circles at their starting point, moving straight through intersections, and then as double circles after reaching their destinations. For example, consider the case where there are two types of agents, A (1 unit) and B (4 units), moving near an intersection as shown in Figure 1. Here, we compare the case where all of Agent B's units move straight through the intersection before Agent A moves straight through the intersection (Pattern 1), and the case where all of Agent B's units move straight through the intersection before Agent A moves straight through the intersection (Pattern 2). In Pattern 1, Agent A must wait at the intersection for all four Agent B units to pass. In Pattern 2, Agent A passes through without waiting for Agent B, and Agent B can move after waiting for one unit. Comparing Pattern 1 and 2, it can be seen that Pattern 2 is more efficient. Thus, efficiency changes depending on how the agents are operated. When there are many agents and the intersecting situations become complex, efficiency can be evaluated by the variability of the delay, and if this variability is small, it can be evaluated as efficient. Note that "variability" refers to statistical variability and may include variance or standard deviation.
[0031] In light of the points mentioned above, the Disclosing Party has diligently considered technologies to achieve navigation with less variation in delays, such that arrival delays are not concentrated on specific agents but are instead experienced by all agents to a similar degree. As a result, the Disclosing Party has devised technologies to achieve navigation with less variation in delays, such that all agents are experienced by a similar degree, as described below.
[0032] Hereinafter, an example of an embodiment of this disclosure will be described with reference to the drawings. In each drawing, identical or equivalent components and parts are given the same reference numerals. Also, the dimensional ratios in the drawings are exaggerated for illustrative purposes and may differ from the actual ratios.
[0033] Figure 2 is a diagram showing the schematic configuration of the agent navigation system according to this embodiment. The agent navigation system shown in Figure 2 includes a learning device 1 and an agent control device 2.
[0034] Learning device 1 is a device that uses machine learning to train a collaborative model 10 that agent control device 2 uses to control the actions of agents A1 and A2. Agent control device 2 is a device that uses the collaborative model 10 to control the actions of agents A1 and A2. When controlling the actions of agents A1 and A2, agent control device 2 uses the output of the movement model 20, as well as the output of the collaborative model 10 to change the output of the movement model 20. The movement model 20 is a model that determines the next direction of movement or the amount of movement to be made based on information about obstacles in the environment in which agents A1 and A2 exist, their own destination, and surrounding agents. The movement model 20 may use classical algorithms such as the Dynamic Window Approach, or it may be a neural network trained by machine learning.
[0035] The collaborative model 10 is a machine learning model that modifies the output of the movement model 20. The collaborative model 10 can be a neural network that has been machine learning using a predetermined method.
[0036] Agents A1 and A2 are examples of controlled agents in this disclosure, and examples of such controlled agents include self-driving mobile robots, AMRs (Autonomous Mobile Robots), AGVs (Automated Guided Vehicles), and autonomous vehicles. Furthermore, cooperative model 10 is an example of a second model in this disclosure and is used when the agent control device 2 controls the operation of agents A1 and A2.
[0037] When training the cooperative model 10, the learning device 1 uses a first observation set that observes the state of a certain agent (e.g., agent A1) and a second observation set that observes the state of at least one other agent (e.g., agent A2) surrounding agent A1. Each observation set includes information observed by each agent, such as its own position, the relative position of the goal from its own position, the arrangement of surrounding objects, and the current time delay compared to the time it would take to reach the goal while ignoring other agents. Therefore, agents A1 and A2 may be equipped with sensors to observe their surroundings.
[0038] Furthermore, the observation timing and interval of each observation set are not limited to a specific pattern; for example, observations of the observation set may be performed in accordance with the agent's operating interval.
[0039] The learning device 1 then trains the collaborative model 10 to reduce the variation in delays among multiple agents reaching each agent's goal. In addition, the learning device 1 may train the collaborative model 10 to keep the delay for agents to reach the goal small. In other words, the learning device 1 trains the collaborative model 10 to reduce the variation in delays for agents A1 and A2, or to reduce the time to reach the goal.
[0040] In this embodiment, the learning device 1 and the agent control device 2 are separate devices, but this disclosure is not limited to this example, and the learning device 1 and the agent control device 2 may be the same device. Also, in this embodiment, the collaborative model 10 exists independently of the learning device 1 and the agent control device 2, but this disclosure is not limited to this example, and for example, the collaborative model 10 may be held in the learning device 1 or the agent control device 2. Furthermore, there may be multiple agents.
[0041] Next, we will explain an example of the configuration of the learning device 1.
[0042] Figure 3 is a block diagram showing an example of the hardware configuration of learning device 1.
[0043] As shown in Figure 3, the learning device 1 includes a CPU (Central Processing Unit) 11, ROM (Read Only Memory) 12, RAM (Random Access Memory) 13, storage 14, input unit 15, display unit 16, and communication interface (I / F) 17. Each component is connected to the others via a bus 19 so that they can communicate with each other.
[0044] The CPU 11 is a central processing unit that executes various programs and controls various components. Specifically, the CPU 11 reads a program from the ROM 12 or storage 14 and executes the program using the RAM 13 as a working area. The CPU 11 controls each of the above components and performs various calculations according to the program recorded in the ROM 12 or storage 14. In this embodiment, the ROM 12 or storage 14 stores a learning program for learning the cooperative model 10.
[0045] ROM12 stores various programs and data. RAM13 temporarily stores programs or data as a working area. Storage14 consists of a storage device such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory, and stores various programs, including the operating system, and various data.
[0046] The input unit 15 includes a pointing device such as a mouse and a keyboard, and is used for various types of input.
[0047] The display unit 16 is, for example, a liquid crystal display and displays various information. The display unit 16 may also function as an input unit 15 by employing a touch panel system.
[0048] The communication interface 17 is an interface for communicating with other devices, and standards such as Ethernet®, FDDI, and Wi-Fi® can be used.
[0049] When executing the above learning program, the learning device 1 uses the above hardware resources to implement various functions. The functional configuration implemented by the learning device 1 will be described below.
[0050] Figure 4 is a block diagram showing an example of the functional configuration of the learning device 1.
[0051] As shown in Figure 4, the learning device 1 has an acquisition unit 101 and a learning unit 102 as its functional configuration. Each functional configuration is realized by the CPU 11 reading and executing a learning program stored in the ROM 12 or storage 14.
[0052] The acquisition unit 101 acquires a first set of observations, which is used to train the collaborative model 10, which consists of observations of the state of a certain agent (e.g., agent A1) and a second set of observations, which consists of observations of the state of at least one other agent (e.g., agent A2) surrounding agent A1. The number of agents that generate the second set of observations is not limited to one.
[0053] The learning unit 102 performs machine learning on the collaborative model 10 using the first and second observation sets acquired by the acquisition unit 101. The learning unit 102 performs machine learning on the collaborative model 10 to reduce the variation in delays among multiple agents reaching the goal of each agent. In addition, the learning unit 102 may perform machine learning on the collaborative model 10 to keep the delay in agents reaching the goal as low as possible. The cost required for movement may be, for example, the time required to move from the start to the goal, the distance traveled, the amount of energy consumed during the movement, etc.
[0054] In other words, the learning unit 102 may train the cooperative model 10 so that the variation in delays from the planned travel time of the controlled agent and other agents is reduced based on the output of the cooperative model 10. The learning unit 102 may also train the cooperative model 10 so that the variation in the extension from the planned travel distance of the controlled agent and other agents is reduced based on the output of the cooperative model 10. The learning unit 102 may also train the cooperative model 10 so that the variation in the increase from the planned energy consumption of the controlled agent and other agents is reduced based on the output of the cooperative model 10. In addition, the learning unit 102 may machine-learn the cooperative model 10 so that at least one of the delay in the agent reaching the goal, the extension from the planned travel distance, or the increase from the planned energy consumption is kept low.
[0055] Specifically, the learning unit 102 trains the collaborative model 10 using machine learning to reduce the variation in latency among multiple agents. In addition, the learning unit 102 may train the collaborative model 10 using machine learning to keep the latency for agents to reach the goal low. In this embodiment, the learning unit 102 trains the collaborative model 10 by reinforcement learning.
[0056] An example of machine learning of the collaborative model 10 by the learning unit 102 is shown. When an agent reaches the goal, a reward is given that expresses how much the agent's actions, changed by the collaborative model 10, contributed to improving the time lag of other agents. Such rewards include, for example, the improvement in the amount of time lag, the improvement in the time of reaching the goal, and the improvement in the Q-value if the movement model 20 is acquired through reinforcement learning. The Q-value if the movement model 20 is acquired through reinforcement learning is a value determined based on factors such as how quickly the agent reached the goal and whether it collided with obstacles. Of course, the reward is not limited to the example shown, and the user may arbitrarily define other rewards. The learning unit 102 performs reinforcement learning of the collaborative model 10 in order to maximize such rewards.
[0057] Next, we will explain an example configuration of the agent control device 2.
[0058] Figure 5 is a block diagram showing an example of the hardware configuration of the agent control device 2.
[0059] As shown in Figure 5, the agent control device 2 includes a CPU 21, ROM 22, RAM 23, storage 24, input unit 25, display unit 26, and communication interface (I / F) 27. Each component is connected to the others via a bus 29 so that they can communicate with each other.
[0060] The CPU 21 is a central processing unit that executes various programs and controls each component. Specifically, the CPU 21 reads a program from the ROM 22 or storage 24 and executes the program using the RAM 23 as a workspace. The CPU 21 controls each of the above components and performs various calculations according to the program recorded in the ROM 22 or storage 24. In this embodiment, the ROM 22 or storage 24 stores an agent control program that controls the operation of the agent using the outputs of the cooperative model 10 and the mobile model.
[0061] ROM22 stores various programs and data. RAM23 temporarily stores programs or data as a working area. Storage24 consists of a storage device such as an HDD, SSD, or flash memory, and stores various programs, including the operating system, and various data.
[0062] The input unit 25 includes a pointing device such as a mouse and a keyboard, and is used for various types of input.
[0063] The display unit 26 is, for example, a liquid crystal display and displays various information. The display unit 26 may also function as an input unit 25 by employing a touch panel system.
[0064] The communication interface 27 is an interface for communicating with other devices, and standards such as Ethernet®, FDDI, and Wi-Fi® are used.
[0065] When executing the above learning program, the agent control unit 2 uses the above hardware resources to implement various functions. The functional configuration implemented by the agent control unit 2 will be described below.
[0066] Figure 6 is a block diagram showing an example of the functional configuration of the agent control device 2.
[0067] As shown in Figure 6, the agent control device 2 has an acquisition unit 201, a movement determination unit 202, a change amount determination unit 203, and an operation control unit 204 as its functional configuration. Each functional configuration is realized by the CPU 21 reading and executing an agent control program stored in the ROM 22 or storage 24.
[0068] The acquisition unit 201 acquires a first observation set, which is used to control the operation of agents A1 and A2, which is an observation set of the state of a certain agent (e.g., agent A1) and a second observation set, which is an observation set of the state of at least one other agent (e.g., agent A2) surrounding agent A1. The number of agents that generate the second observation set is not limited to one.
[0069] The movement determination unit 202 inputs the first and second observation sets acquired by the acquisition unit 201 into the movement model 20 and uses the output from the movement model 20 to determine information regarding the movement of the controlled agent. Here, information regarding the movement of the controlled agent may include the direction of movement, amount of movement, time required for movement, direction of rotation, amount of rotation, etc. For example, the movement determination unit 202 uses the output from the movement model 20 to determine that the information regarding the movement of the controlled agent is that it moves 1 meter in the west direction.
[0070] The change amount determination unit 203 inputs the first observation set and the second observation set to the cooperative model 10 and uses the output from the cooperative model 10 to determine the change amount for information regarding the movement of the controlled agent. More specifically, the change amount determination unit 203 determines the change amount such that the variation in the cost required for the movement of the controlled agent and other agents is reduced. The change amount determination unit 203 may also determine the change amount such that both the variation and the average of the required costs are reduced. The cost required for movement may be, for example, the time required to move from the start to the goal, the distance traveled, the amount of energy consumed during the movement, etc. Note that "cost" is an example of "compensation" in this disclosure. The output of the cooperative model 10 may include information for determining whether or not to bear the cost.
[0071] In other words, the change amount determination unit 203 may determine the change amount based on the output from the cooperative model 10 so as to reduce the variation in delays from the planned travel time of the controlled agent and other agents. Alternatively, the change amount determination unit 203 may determine the change amount based on the output from the cooperative model 10 so as to reduce the variation in extensions from the planned travel distance of the controlled agent and other agents. Alternatively, the change amount determination unit 203 may determine the change amount based on the output from the cooperative model 10 so as to reduce the variation in increases from the planned energy consumption of the controlled agent and other agents. Alternatively, the change amount determination unit 203 may determine the change amount based on the output from the cooperative model 10 so as to reduce the variation and average of delays from the planned travel time of the controlled agent and other agents. Alternatively, the change amount determination unit 203 may determine the change amount based on the output from the cooperative model 10 so as to reduce the variation and average of extensions from the planned travel distance of the controlled agent and other agents. The change amount determination unit 203 may also determine the change amount based on the output from the cooperative model 10 such that the variation and average of the increase from the planned energy consumption of the controlled agent and other agents are reduced.
[0072] The change amount determination unit 203 can arbitrarily select the cost to increase, but the user may specify which cost to prioritize, or a preferred cost may be specified for each agent.
[0073] The amount of change determined by the change amount determination unit 203 may be either 0 or 1, which determines whether to move the controlled agent or not, or it may be a weight applied to the amount of movement, which is a value between 0 and 1.
[0074] The motion control unit 204 controls the operation of the controlled agent by applying the change amount determined by the change amount determination unit 203 to the information regarding the movement of the controlled agent determined by the movement determination unit 202.
[0075] For example, if the change amount determination unit 203 determines that the change amount is 0, the operation control unit 204 controls the controlled agent to stop. Also, for example, if the movement determination unit 202 determines that the controlled agent will move 1 meter westward, and the change amount determination unit 203 determines that the change amount is 0.5, the operation control unit 204 controls the controlled agent to move 0.5 meters westward. Also, for example, if the movement determination unit 202 determines that the controlled agent will move 10 meters westward in 1 second, and the change amount determination unit 203 determines that the change amount is 0.1, the operation control unit 204 controls the controlled agent to move 10 meters westward in 10 seconds.
[0076] When the motion control unit 204 controls the motion of the controlled agent, it may determine whether or not the controlled agent will incur the cost of movement based on the output of the cooperative model 10 when only the first set of observations is input to the cooperative model 10.
[0077] The agent control device 2, with this configuration, can control the movement of agents in an environment where multiple agents exist, thereby avoiding situations where a particular agent experiences extreme delays.
[0078] Next, we will explain the operation of learning device 1.
[0079] Figure 7 is a flowchart showing the flow of the learning process by the learning device 1. The CPU 11 reads the learning program from the ROM 12 or storage 14, loads it into the RAM 13, and executes it, thereby performing the learning process.
[0080] In step S101, the CPU 11 inputs the first and second sets of observations, which were acquired for training the collaborative model 10, into the collaborative model 10.
[0081] Following step S101, in step S102, the CPU 11 reinforces the collaborative model 10 using the output from the collaborative model 10. An example of the learning process for the collaborative model 10 was explained in the processing of the learning unit 102 above, so the details are omitted here, but the CPU 11 reinforces the collaborative model 10 in a way that maximizes the reward obtained.
[0082] Next, we will explain the operation of the agent control device 2.
[0083] Figure 8 is a flowchart showing the flow of agent control processing by the agent control device 2. The agent control processing is performed when the CPU 21 reads the agent control program from the ROM 22 or storage 24, loads it into the RAM 23, and executes it.
[0084] In step S111, the CPU 21 acquires a first observation set and a second observation set from each agent, inputs the acquired first and second observation sets into the movement model 20, and uses the output from the movement model 20 to determine information regarding the movement of the controlled agent.
[0085] Following step S111, in step S112, the CPU 21 inputs the acquired first and second observation sets into the cooperative model 10 and uses the output from the cooperative model 10 to determine the amount of change in the information regarding the movement of the controlled agent.
[0086] Following step S112, in step S113, the CPU 21 applies the amount of change determined in step S112 to the information regarding the movement of the controlled agent determined in step S111, and controls the movement of the controlled agent.
[0087] Next, we will explain a specific example of controlling the movement of the agent using the agent control device 2.
[0088] Figures 9 to 14 illustrate specific examples of agent movement control by the agent control device 2. Figures 9 to 14 show examples of controlling the movement of agents A1 and A2 near an intersection. In Figures 9 to 14, code G1 represents the destination of agent A1, and code G2 represents the destination of agent A2. Note that while Figures 9 to 14 show only two agents, A1 and A2, for simplicity, the movement of three or more agents can be controlled similarly.
[0089] For example, as shown in Figure 9, if agents A1 and A2 are located near an intersection and move toward their respective destinations G1 and G2, there is a possibility that agents A1 and A2 will collide. Figure 10 shows agents A1 and A2 approaching each other. Therefore, the agent control device 2 needs to control the movements of agents A1 and A2 to avoid a collision.
[0090] Here, the agent control device 2 controls the operation of agents A1 and A2 by applying the output of the cooperation model 10 to the output of the movement model 20 so that agents A1 and A2 reach destination points G1 and G2 with minimal variation in delay, or so that they reach destination points G1 and G2 with minimal variation and average delay. For example, if agent A2 is made to wait in place and agent A1 is allowed to pass first, agent A1 will reach destination point G1 as scheduled, but agent A2 will be delayed in reaching destination point G2. However, if agent A1 is made to wait in place and agent A2 is allowed to pass first, both agents A1 and A2 will be delayed in reaching their destinations. In this case, the agent control device 2 controls the operation of agents A1 and A2 so that agent A1 is made to wait in place and agent A2 is allowed to pass first.
[0091] Figure 11 shows that, under the control of the agent control device 2, agent A1 remains stationary, while agent A2 moves towards destination point G2 to avoid agent A1.
[0092] Figure 12 shows how, under the control of the agent control device 2, agent A1 begins to move after agent A2 has passed, and moves towards the destination point G1.
[0093] Figure 13 shows agent A2 reaching destination G2 under the control of agent control device 2. At this point, let's assume that agent A2 is already behind schedule in arriving at destination G2.
[0094] Figure 14 shows that agent A1 has reached the destination G1 under the control of agent control device 2. At this point, let's assume that agent A1 is also behind schedule in arriving at the destination G1.
[0095] By using the outputs of the cooperation model 10 and the movement model 20 in this way, the agent control device 2 can control the movement of agents in such a way that the variation in the cost required for the movement of each agent is reduced, or the variation and average cost are reduced.
[0096] While embodiments of the present disclosure have been described in detail above with reference to the attached drawings, the technical scope of the present disclosure is not limited to these examples. It is clear that a person with ordinary skill in the art of the present disclosure may conceive of various modifications or alterations within the scope of the technical idea set forth in the claims, and these modifications or alterations are also understood to fall within the technical scope of the present disclosure.
[0097] Furthermore, the effects described in the above embodiments are descriptive or illustrative, and are not limited to those described in the above embodiments. In other words, the technology relating to this disclosure may produce other effects that would be obvious to a person of ordinary skill in the art of this disclosure from the descriptions in the above embodiments, in addition to or in lieu of the effects described in the above embodiments.
[0098] In addition, the learning process and agent control process that the CPU reads and executes in each of the above embodiments may be executed by various processors other than the CPU. Examples of such processors include PLDs (Programmable Logic Devices) such as FPGAs (Field-Programmable Gate Arrays) whose circuit configuration can be changed after manufacturing, and dedicated electrical circuits that are processors with circuit configurations specifically designed to execute specific processes, such as ASICs (Application Specific Integrated Circuits). Furthermore, the learning process and agent control process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (for example, multiple FPGAs, and a combination of a CPU and an FPGA). More specifically, the hardware structure of these various processors is an electrical circuit that combines circuit elements such as semiconductor elements.
[0099] Furthermore, while the above embodiments describe a configuration in which the learning process and agent control process programs are pre-stored (installed) in ROM or storage, the system is not limited to this. The programs may be provided in a form recorded on a non-transitory recording medium such as a CD-ROM (Compact Disk Read Only Memory), DVD-ROM (Digital Versatile Disk Read Only Memory), or USB (Universal Serial Bus) memory. Alternatively, the programs may be provided in a form that can be downloaded from an external device via a network.
[0100] The following are examples of applications of this disclosure. This disclosure is applicable to rescue robots, for example, in the event of a large-scale disaster when multiple rescue robots are used to rescue multiple victims, it is possible to control the rescue robots to reach the victims with minimal delay and with minimal variation in delay for multiple victims, thus proving effective in saving multiple lives in a time-sensitive situation. This disclosure is also applicable to automated warehouses using mobile robots, where it is possible to improve the efficiency of time and the energy consumption of the mobile robots by reducing delays and variations in the process of retrieving goods using multiple mobile robots. Furthermore, this disclosure is applicable to cooking robots and pharmaceutical compounding robots (especially dual-arm robots), where it is effective in reducing variations in the timing of compounding ingredients while avoiding collisions between the robot's arms. Finally, this disclosure is applicable to autonomously moving machines such as drones and self-driving vehicles. [Explanation of Symbols]
[0101] 1. Learning device 2 Agent control unit 10. Collaborative Model 20 Mobile Models
Claims
1. A movement determination unit inputs a first observation set, which observes the state of the controlled agent, and a second observation set, which observes the state of at least one other agent surrounding the controlled agent, into a first model, and determines information regarding the movement of the controlled agent based on the output of the first model. A change amount determination unit inputs the first observation set and the second observation set into a second model, and determines a change amount that is a weight between 0 and 1 applied to the amount of movement, which is a value of 0 or 1 that determines whether or not to move the controlled agent, or a change amount that is applied to the amount of movement, and which is determined to reduce the variation in the compensation required for the movement of the controlled agent and the other agents compared to when the change amount is not applied to the information on the movement determined by the movement determination unit. An operation control unit that operates the controlled agent by providing the change amount determined by the change amount determination unit to the information regarding the movement determined by the movement determination unit, An agent control device equipped with the following:
2. The agent control device according to claim 1, wherein the change amount determination unit determines the change amount based on the output of the second model, so as to reduce the variation in the delay from the planned movement time of the controlled agent and the other agents compared to the case where the change amount is not applied to the movement information determined by the movement determination unit.
3. The agent control device according to claim 1, wherein the change amount determination unit determines the change amount based on the output of the second model, so as to reduce the variation in the extension of the controlled agent and the other agents from the planned travel distance, compared to the case where the change amount is not applied to the information on the movement determined by the movement determination unit.
4. The agent control device according to claim 1, wherein the change amount determination unit determines the change amount based on the output of the second model, so as to reduce the variation in the increase from the planned energy consumption of the controlled agent and the other agents compared to the case where the change amount is not applied to the information on movement determined by the movement determination unit.
5. The agent control device according to claim 1, wherein the change amount determination unit determines the change amount based on the output of the second model, such that the average cost of the movement of the controlled agent and the other agents is reduced compared to a case where the change amount is not applied to the movement information determined by the movement determination unit.
6. The agent control device according to any one of claims 1 to 5, wherein the information relating to the movement is at least one of the direction of movement, amount of movement, time required for movement, direction of rotation, and amount of rotation of the controlled agent.
7. The agent control device according to any one of claims 1 to 5, wherein the operation control unit determines whether or not the controlled agent incurs a cost for movement based on the output of the second model when only the first observation set is input to the second model.
8. A learning device comprising a learning unit that takes as input a first observation set of the state of a controlled agent and a second observation set of the state of at least one other agent surrounding the controlled agent, inputs the first observation set and the second observation set to a first model, and outputs a change amount which is 0 or 1, or a weight between 0 and 1 applied to the amount of movement, that determines whether or not to move the controlled agent determined by the output of the first model, and that the change amount is determined to reduce the variation in the compensation required for the movement of the controlled agent and the other agents compared to when the change amount is not applied to the determined information regarding the movement.
9. The processor, A first observation set, which observes the state of the controlled agent, and a second observation set, which observes the state of at least one other agent surrounding the controlled agent, are input to the first model, and information regarding the movement of the controlled agent is determined from the output of the first model. The first observation set and the second observation set are input to the second model, and the output of the second model determines a change amount which is 0 or 1, or a weight between 0 and 1 applied to the amount of movement, that determines the change amount such that the variation in the compensation required for the movement of the controlled agent and the other agents is reduced compared to when the change amount is not applied to the determined information regarding the movement. The controlled agent is operated by applying the determined amount of change to the determined information regarding the movement. A method for controlling agents to execute processes.
10. The processor, The first model is trained by inputting a first observation set, which observes the state of the controlled agent, and a second observation set, which observes the state of at least one other agent surrounding the controlled agent. The first and second observation sets are input to the first model, and the output of the second model is a change amount that is either 0 or 1, or a weight between 0 and 1 applied to the amount of movement, which determines whether to move or not the controlled agent determined by the output of the first model, and is determined to reduce the variability in the compensation required for the movement of the controlled agent and the other agents compared to when the change amount is not applied to the determined information about the movement. A learning method for executing a process.
11. On the computer, A first observation set, which observes the state of the controlled agent, and a second observation set, which observes the state of at least one other agent surrounding the controlled agent, are input to the first model, and information regarding the movement of the controlled agent is determined from the output of the first model. The first observation set and the second observation set are input to the second model, and the output of the second model determines a change amount which is 0 or 1, or a weight between 0 and 1 applied to the amount of movement, that determines the change amount such that the variation in the compensation required for the movement of the controlled agent and the other agents is reduced compared to when the change amount is not applied to the determined information regarding the movement. The controlled agent is operated by applying the determined amount of change to the determined information regarding the movement. An agent control program that executes a process.
12. On the computer, The first model is trained by inputting a first observation set, which observes the state of the controlled agent, and a second observation set, which observes the state of at least one other agent surrounding the controlled agent. The first and second observation sets are input to the first model, and the output of the second model is a change in the information regarding the movement of the controlled agent, determined by the output of the first model, which is determined to reduce the variability in the compensation required for the movement of the controlled agent and the other agents compared to when the change is not applied to the determined information regarding the movement. A learning program that executes a process.
Citation Information
Patent Citations
Traffic signal lamp control method based on cooperative multi-agent reinforcement learning
CN115083174A
Deep learning based motion control of a group of autonomous vehicles
EP3832420A1
Area management system, method and program
JP2019007832A
Information processor, information processing method, program, and system
JP2019125345A
Route adjustment system, route adjustment device, route adjustment method, and route adjustment program
JP2022526546A