Control device, control system, control method, and program

The control device and system address task inefficiencies by calculating request and response parameters and importance to enhance agent cooperation, ensuring efficient task execution in unknown environments.

JP7715099B2Active Publication Date: 2025-07-30TOYOTA JIDOSHA KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022135849
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-07-30
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently achieving task goals when the number of agents required is unknown, leading to potential task stagnation.

Method used

A control device and system that calculates request and response parameters based on observation information, importance processing, and task selection to efficiently distribute tasks among multiple agents, using a learned policy to enhance cooperation and task execution.

Benefits of technology

Ensures efficient task achievement even in unknown environments by appropriately selecting and executing tasks based on agent cooperation and importance calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007715099000027
    Figure 0007715099000027
  • Figure 0007715099000028
    Figure 0007715099000028
  • Figure 0007715099000029
    Figure 0007715099000029
Patent Text Reader

Abstract

To provide a control apparatus capable of making it possible to, even in an environment in which tasks are not yet known, efficiently achieve the target for the tasks.SOLUTION: A request response processing unit 130 calculates, based on observation information about an agent, another agent around the agent and tasks, a request parameter as to whether or not to request help, and a response parameter as to whether or not to respond to a request from the other agent. An importance processing unit 140 performs processing for calculating, based on at least the request parameter of the other agent and the response parameter of the agent, importance of each of the tasks for the agent. A task selection unit 150 selects the task to be performed by the agent according to the importance. A task execution unit 160 performs the control so that the agent performs the selected task.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a control device, a control system, a control method, and a program.

Background Art

[0002] There is a technology for causing a plurality of agents (such as robots) to execute tasks. In relation to this technology, Patent Document 1 discloses a mobile agent having the ability to assemble general-purpose structures. In Patent Document 1, a plurality of mobile agents automatically operate components such as blocks on a work surface in order to execute operations such as the assembly of general-purpose structures. Also, various mobile agents may operate in cooperation with each other.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In an environment where the task is unknown, the number of agents required to execute the task may not be known. In such a case, with the technology of Patent Document 1, when a plurality of agents cooperate to execute a task, the task may stop progressing. Therefore, with the technology of Patent Document 1, there is a risk that the task target may not be efficiently achieved.

[0005] The present disclosure provides a control device, a control system, a control method, and a program that can efficiently achieve the task target even in an environment where the task is unknown.

Means for Solving the Problems

[0006] The control device according to the present disclosure is a control device that controls agents that execute tasks, wherein the more agents that execute the task, the higher the likelihood of achieving the goal of the task, and there are a plurality of the tasks in the environment. Based on observation information regarding the agent, other agents around the agent, and the task, a request response processing unit calculates a request parameter regarding whether or not to request support and a response parameter regarding whether or not to respond to a request from another agent, an importance processing unit performs processing for calculating the importance of each of the tasks regarding the agent based on at least the request parameter of other agents and the response parameter of the agent, a task selection unit selects the task that the agent should execute according to the importance, and a task execution unit controls the agent to execute the selected task.

[0007] Further, the control system according to the present disclosure is a control system that distributes and controls a plurality of agents that execute tasks, wherein the more agents that execute the task, the higher the likelihood of achieving the goal of the task, and there are a plurality of the tasks in the environment. The control system has a plurality of control devices that respectively control a plurality of agents, and each of the plurality of control devices includes a request response processing unit that calculates a request parameter regarding whether or not to request support and a response parameter regarding whether or not to respond to a request from another agent based on observation information regarding the agent related to the control device, other agents around the agent, and the task, an importance processing unit that performs processing for calculating the importance of each of the tasks regarding the agent based on at least the request parameter of other agents and the response parameter of the agent, a task selection unit that selects the task that the agent should execute according to the importance, and a task execution unit that controls the agent to execute the selected task.

[0008] Also, the control method according to the present disclosure is a control method for controlling an agent that executes a task. The task is such that the more agents that execute the task, the higher the likelihood of achieving the goal of the task. There are multiple tasks in the environment. Based on the observation information regarding the agent, other agents around the agent, and the task, a request parameter regarding whether to request support and a response parameter regarding whether to respond to a request from other agents are calculated. Based on at least the request parameter of other agents and the response parameter of the agent, a process for calculating the importance of each of the tasks regarding the agent is performed. According to the importance, the task to be executed by the agent is selected, and control is performed so that the agent executes the selected task.

[0009] Also, the program according to the present disclosure is a program for realizing a control method for controlling an agent that executes a task. The task is such that the more agents that execute the task, the higher the likelihood of achieving the goal of the task. There are multiple tasks in the environment. Based on the observation information regarding the agent, other agents around the agent, and the task, steps of calculating a request parameter regarding whether to request support and a response parameter regarding whether to respond to a request from other agents, a step of performing a process for calculating the importance of each of the tasks regarding the agent based on at least the request parameter of other agents and the response parameter of the agent, a step of selecting the task to be executed by the agent according to the importance, and a step of performing control so that the agent executes the selected task are executed by a computer.

[0010] In the present disclosure, even in an environment where the task is unknown, it is possible to efficiently achieve the goal of the task.

[0011] Preferably, the request response processing unit calculates the request parameter and the response parameter based on the policy learned for each agent. In the present disclosure, with such a configuration, it becomes possible to appropriately select a task to be executed for each agent.

[0012] Preferably, the request response processing unit calculates the request parameter and the response parameter based on the degree of request and the degree of response output from the policy by inputting the observation information into the policy. In the present disclosure, with such a configuration, it becomes possible to appropriately select a task to be executed for each agent.

[0013] Preferably, when the degree of request exceeds a predetermined threshold and the task that the agent is executing or attempting to execute is not in progress, the request response processing unit calculates the request parameter indicating that support is requested. In the present disclosure, with such a configuration, when support should be requested for the task that the agent is executing or attempting to execute, it is possible to appropriately calculate the request parameter indicating that support is requested.

[0014] Preferably, when the degree of response exceeds a predetermined threshold and the task that the agent is executing or attempting to execute is not in progress, the request response processing unit calculates the response parameter indicating that the request is responded to. In the present disclosure, with such a configuration, when the task that the agent is executing or attempting to execute is in progress, it is possible to continue executing the task.

[0015] Preferably, the importance processing unit calculates the importance of each of the tasks related to the agent based on the policy learned for each agent. In the present disclosure, with such a configuration, it becomes possible to appropriately calculate the importance of each task for each agent.

[0016] Further preferably, the importance processing unit inputs the observation information into the policy and calculates the importance of the task corresponding to the observation information regarding the agent based on the target value of the importance of the task corresponding to the observation information output from the policy. In the present disclosure, with such a configuration, for each agent, it is possible to calculate the importance of the task corresponding to the observation information so as to approach the target value. Thereby, it becomes possible to appropriately calculate the importance of the task.

Advantages of the Invention

[0017] According to the present disclosure, it is possible to provide a control device, a control system, a control method, and a program capable of efficiently achieving the target of a task even in an environment where the task is unknown.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0019] (Embodiment 1) Hereinafter, this embodiment will be described with reference to the drawings. For clarity of explanation, the following description and drawings are appropriately omitted and simplified. Also, in each drawing, the same reference numerals are assigned to the same elements, and duplicate explanations are omitted as necessary.

[0020] FIG. 1 is a diagram showing a control system 1 according to Embodiment 1. The control system 1 includes a control device 100 that controls each of a plurality of agents 10, and a monitoring device 60 that monitors each of a plurality of tasks 50. The agent 10 is, for example, a machine such as a robot, but is not limited thereto. Each agent 10 is disposed in the environment and autonomously operates in the environment under the control of the control device 100.

[0021] The control device 100 is, for example, a computer. The control device 100 may be built into an agent 10 that is a machine such as a robot. The control device 100 performs control to cause the corresponding agent 10 to execute the task 50. That is, the control system 1 controls a plurality of agents 10 in a distributed manner. Each control device 100 is communicably connected to other control devices 100 via a wired or wireless network. Also, each control device 100 is communicably connected to the monitoring device 60 via a wired or wireless network. The control device 100 will be described in detail later.

[0022] In the environment where the agents 10 exist, a plurality of tasks 50 exist. The agent 10 executes each of the plurality of tasks 50. A target (goal; end condition) is set for each task 50. By each agent 10 executing each task 50, the task 50 progresses, and when the target of each task 50 is achieved, each task 50 is realized (ended).

[0023] Here, for task 50, the more agents 10 that execute this task 50, the higher the feasibility (the possibility that the goal of task 50 is achieved). That is, even if one agent 10 tries to execute a certain task 50 and the task 50 does not progress, by having multiple agents 10 execute that task 50, the task 50 will progress, and the possibility that the task 50 is realized (the possibility that the goal of task 50 is achieved) will increase. That is, by having multiple agents 10 cooperate to execute task 50, the possibility that task 50 is realized increases. In other words, by having multiple agents 10 cooperate to execute task 50, the possibility that the goal of task 50 is achieved increases. However, the number of agents 10 required for the progress of task 50 is not known in advance. The number of agents 10 required for the progress of task 50 is determined by the agents 10 executing task 50. The control device 100 controls the agents 10 so that task 50 is executed to achieve the goal of task 50. Details will be described later.

[0024] The monitoring device 60 is, for example, a sensor or a camera, etc. The monitoring device 60 monitors (detects) the state of each task 50. Specifically, the monitoring device 60, for example, detects the position and speed of task 50. Also, the monitoring device 60 stores information on whether the task has ended. Further, the monitoring device 60 may store information regarding the goal of task 50. The monitoring device 60 may monitor whether the goal of task 50 has been achieved. Note that the monitoring device 60 may be provided for each task 50. Alternatively, one monitoring device 60 may monitor a plurality of tasks 50. Note that each agent 10 may detect the state of task 50. In this case, the monitoring device 60 may not be necessary. Also, the agent 10 may detect the state of task 50 and make an end determination.

[0025] Here, in Embodiment 1, Task 50 is the load to be transported. And for each Task 50, a goal (destination) of the load is set. Agent 10 transports the load (Task 50) so that the load reaches the goal. And the more the number of Agents 10 that transport the load (Task 50), the higher the possibility that the load (Task 50) reaches the goal. That is, depending on the load, it can be so large that it cannot be transported by a small number of Agents 10. That is, the size and weight can vary depending on the load. On the other hand, by having many Agents 10 cooperate to transport one load, the load can be moved. That is, by having many Agents 10 cooperate, the load can be transported (Task 50 can be advanced), and the load can be transported to its target position (the goal of Task 50 can be achieved). Note that the number of Agents 10 required to transport the load is unknown. The number of Agents 10 required for transportation is found only when an Agent 10 tries to execute the transportation of the load.

[0026] FIG. 2 is a diagram showing the configuration of the control device 100 according to Embodiment 1. As shown in FIG. 2, the control device 100 includes, as main hardware components, a control unit 102, a storage unit 104, a communication unit 106, and an interface unit 108 (IF; Interface). The control unit 102, the storage unit 104, the communication unit 106, and the interface unit 108 are interconnected via a data bus or the like. Note that Agent 10, which is a machine, may also have the hardware configuration of the control device 100 shown in FIG. 2. Further, the monitoring device 60 may also have the hardware configuration of the control device 100 shown in FIG. 2.

[0027] The control unit 102 is a processor such as a CPU (Central Processing Unit). The control unit 102 has a function as an arithmetic unit that performs control processing, arithmetic processing, and the like. Note that the control unit 102 may have a plurality of processors. The storage unit 104 is a storage device such as a memory or a hard disk. The storage unit 104 is, for example, a ROM (Read Only Memory) or a RAM (Random Access Memory). The storage unit 104 has a function of storing a control program, an arithmetic program, and the like executed by the control unit 102. That is, the storage unit 104 (memory) stores one or more instructions. Further, the storage unit 104 has a function of temporarily storing processing data and the like. The storage unit 104 may include a database. Also, the storage unit 104 may have a plurality of memories.

[0028] The communication unit 106 performs processing necessary for communicating with other devices such as another control device 100 or a monitoring device 60 via a network. The communication unit 106 may include a communication port, a router, a firewall, and the like. The interface unit 108 is, for example, a user interface (UI). The interface unit 108 has an input device such as a keyboard, a touch panel, or a mouse, and an output device such as a display or a speaker. The interface unit 108 may be configured such that the input device and the output device are integrated, for example, like a touch screen (touch panel). The interface unit 108 receives an operation of inputting data by a user (operator) and outputs information to the user.

[0029] The control device 100 according to Embodiment 1 includes, as components, an observation information acquisition unit 110, a policy storage unit 112, an action output unit 120, a request response processing unit 130, an importance processing unit 140, a task selection unit 150, and a task execution unit 160. Each of the above-described components can be realized, for example, by executing a program under the control of the control unit 102. More specifically, each component can be realized by the control unit 102 executing a program (instruction) stored in the storage unit 104. Further, necessary programs may be recorded on an arbitrary non-volatile recording medium and installed as needed to realize each component. Also, each component is not limited to being realized by software based on a program, and may be realized by any combination of hardware, firmware, and software. Further, each component may be realized using a user-programmable integrated circuit such as an FPGA (field-programmable gate array) or a microcomputer. In this case, a program composed of the above-described components may be realized using this integrated circuit. These are the same in other embodiments described later.

[0030] In the following description, the control device 100 to be described is referred to as "its own control device 100 (the said control device)". Further, a control device 100 other than its own control device 100 is referred to as "another control device 100". Also, an agent 10 controlled by its own control device 100 is referred to as "its own agent (the said agent)". Further, an agent 10 other than its own agent 10 is referred to as "another agent". Also, in the following description, the operation of its own control device 100 will be described, but another control device 100 performs the same operation.

[0031] The control device 100 controls its own agent 10 so that the task 50 is executed by the above-described components to achieve the goal of the task 50. That is, the control device 100 performs control for its own agent 10 to execute the task 50. The control device 100 calculates the request parameter and the response parameter of its own agent 10 based on the observation information regarding its own agent 10, other agents 10 around its own agent 10, and the task 50. Here, the "request parameter" is a parameter regarding whether to request support from other agents 10. Also, "requesting support" corresponds to having other agents 10 execute the task 50 in cooperation with its own agent 10. Further, the "response parameter" is a parameter regarding whether to respond to a request from other agents 10. Also, "responding to a request" corresponds to having its own agent 10 execute the task 50 in cooperation with other agents 10.

[0032] In addition, the control device 100 performs a process for calculating the importance of each task 50 regarding its own agent 10 based on the request parameter of other agents 10 and the response parameter of its own agent 10. Here, the "importance" is used to determine which task 50 the agent 10 selects and executes. The higher the importance of the task 50, the more likely it is to be selected by the agent 10 and executed by the selected agent 10.

[0033] Further, the control device 100 selects a task 50 to be executed by its own agent 10 according to the importance. The control device 100 controls so that its own agent 10 executes the selected task 50. Then, the control device 100 repeats the above processing for each control cycle. The importance of the task 50 being executed by the agent 10 that has calculated the request parameter indicating the need for support is likely to increase in another agent 10 that has calculated the response parameter indicating the response to the request. Therefore, the possibility that the agent 10 will come to support the task 50 increases. This will be described in detail below.

[0034] The observation information acquisition unit 110 acquires observation information from the surrounding environment. The observation information is information regarding its own agent 10, other agents 10 around its own agent 10, and the task 50. Therefore, the observation information includes information regarding its own agent 10. Further, the observation information includes information regarding other agents 10 and information regarding the task 50 around its own agent 10.

[0035] FIG. 3 is a diagram illustrating an environment in which the agent 10 and the task 50 according to the first embodiment exist. Let the number of agents 10 be M and the number of tasks 50 be N. Also, let its own agent 10 be "agent #i". i is an index indicating its own agent 10. Also, let other agents 10 be "agent #j". j is an index indicating other agents 10.

[0036] In addition, other agents 10 in the vicinity of agent #i are referred to as "neighboring agents". The neighboring agents may be, for example, a predetermined number of other agents 10 within a predetermined range of the distance from agent #i (indicated by the dashed circle in FIG. 3). Alternatively, the neighboring agents may be a predetermined number of other agents 10 that are closest to agent #i. In Embodiment 1, the "predetermined number" is set to 2. These two neighboring agents are denoted as agent #j1 and agent #j2, respectively. Also, FIG. 3 shows agent #1 and agent #M, which are agents 10 other than the neighboring agents. In actuality, there are (M - 3) agents 10 other than the neighboring agents, excluding the agent 10 itself and the two neighboring agents from the total number M of agents 10.

[0037] Also, let the index of task 50 be "l" (l ∈ {1, ···, N}). The task 50 in the vicinity of agent #i is referred to as a "neighboring task". The neighboring task may be, for example, a predetermined number of tasks 50 within a predetermined range of the distance from agent #i (indicated by the dashed circle in FIG. 3). Note that this "predetermined range" may be a range different from that defining the above neighboring agents. Alternatively, the neighboring task may be a predetermined number of tasks 50 that are closest to agent #i. In Embodiment 1, the "predetermined number" is set to 2. These two neighboring tasks are denoted as task #l1 and task #l2, respectively. Also, FIG. 3 shows task #1, task #2, and task #N, which are tasks 50 other than the neighboring tasks. In actuality, there are (N - 2) tasks 50 other than the neighboring tasks, excluding the neighboring tasks from the total number N of tasks 50.

[0038] Also, x indicates the position (current position) of agent 10. x i indicates the position of agent # i . x j indicates the position of agent # j . Also, z indicates the position (current position) of task 50. Also, z * indicates the target position (goal) of task 50. zl indicates the position of task #l. z l * indicates the target position of task #l. Note that the "position" of task 50 is not limited to indicating where task 50 is in the real space, and may also indicate the state of task 50. In this case, the "position" of task 50 may indicate a point in the virtual space representing the state of task 50. For example, the state of task 50 may indicate the progress of task 50, and the "position" of task 50 may indicate a point in the virtual space representing the progress of task 50.

[0039] Also, φ indicates the importance of each task 50 for each agent 10. φ i indicates the importance of each task 50 for agent #i. φ j indicates the importance of each task 50 for agent #j. Note that φ has a number of components corresponding to the number N of tasks 50, and indicates the importance of each of tasks #1, ···, #l, ···, #N. For example, the importance φ i l indicates the importance of task #l for agent #i. The importance for each agent 10 is calculated by the control device of each agent 100 and transmitted (broadcast) to the surrounding agents 10 (control devices 100). Details will be described later.

[0040] The observation information acquisition unit 110 acquires the positions of the surrounding agents 10 and tasks 50. Specifically, the observation information acquisition unit 110 acquires information about the agent 10 from other control devices 100 related to the surrounding agents 10. The information about the agent 10 indicates, for example, the position of the agent 10 and the importance of each task 50 related to the agent 10 (the importance of each task 50 for the agent 10). Also, the observation information acquisition unit 110 acquires information about each task 50 from the monitoring device 60. The information about the task 50 indicates, for example, the state of the task 50 and the target of the task 50. The state of the task 50 may include, for example, the position and speed of the task 50.

[0041] The observation information acquisition unit 110 calculates the distance D between agent #i and agent #j from the acquired position of agent #j. ij Here, D ij =||x i -x j ||2. Also, the observation information acquisition unit 110 calculates the distance between its own agent 10 (agent #i) and each task #l from the acquired position of task 50. Specifically, the observation information acquisition unit 110 calculates the distance D between agent #i and task #l using the following formula (1). il to calculate.

Equation

[0042] In Equation (1), "0.05" is a threshold value for determining whether task #l has reached the target position (i.e., whether task #l has achieved the target). If the distance between z l and z l * is 0.05 or less, task #l is considered to have reached the target position. Also, "1.0e4" is a value large enough not to be regarded as the vicinity of agent #i. That is, from Equation (1), for task 50 that has reached the target position, D il is calculated as a distance much larger than the actual distance. Therefore, for task 50 that has reached the target position, it can be ignored in subsequent processing.

[0043] The observation information acquisition unit 110 acquires the observation information o ij of its own agent 10 (agent #i) using D il and D i . The observation information acquisition unit 110 determines a predetermined number of neighboring agents using D ij and includes information about the neighboring agents as part of the observation information o i . Also, the observation information acquisition unit 110 determines a predetermined number of neighboring tasks using D il and includes information about the neighboring tasks as part of the observation information o iMake it a part of

[0044] Here, the neighboring tasks in Embodiment 1 will be described. The conditions for the neighboring tasks related to Agent #i are represented by the following formula (2). Note that, as shown in formula (2), the number of neighboring tasks related to Agent #i is two.

Number

[0045] Here, l i 1 is the neighboring task #l1 related to Agent #i and is defined by the following formula (3). That is, the neighboring task #l i 1 (neighboring task #l1) is the task 50 closest to Agent #i. Note that this neighboring task #l1 can be the task 50 that Agent #i is currently executing.

Number

[0046] Also, l i 2 is the neighboring task #l2 related to Agent #i and is defined by the following formula (4). That is, the neighboring task #l i 2 (neighboring task #l2) is the second-closest task 50 to Agent #i.

Number

[0047] The observation information acquisition unit 110 acquires observation information o as shown in the following formula (5). Note that in formula (5), T on the upper right indicates transpose. Also, formula (5) represents the observation information at a certain point in time (for example, time t). i is acquired. Note that in formula (5), T on the upper right indicates transpose. Also, formula (5) represents the observation information at a certain point in time (for example, time t).

Number

[0048] Here, in Equation (5), the following Equation (6) is information about its own agent 10 (agent #i). Note that Equation (6) shows, in order from the left, the position of agent #i, the importance of the neighboring task #l1 related to agent #i, and the importance of the neighboring task #l2 related to agent #i.

Equation

[0049] Also, in Equation (5), the following Equation (7) is information about neighboring agent #j1. Note that neighboring agent #j1 may be the other agent 10 closest to its own agent 10 (agent #i) in the same way as neighboring task #l1. Note that Equation (7) shows, in order from the left, the position of neighboring agent #j1, the importance of the neighboring task #l1 related to neighboring agent #j1, and the importance of the neighboring task #l2 related to neighboring agent #j1.

Equation

[0050] Also, in Equation (5), the following Equation (8) is information about neighboring agent #j2. Note that neighboring agent #j2 may be the other agent 10 second closest to its own agent 10 (agent #i) in the same way as neighboring task #l2. Note that Equation (8) shows, in order from the left, the position of neighboring agent #j2, the importance of the neighboring task #l1 related to neighboring agent #j2, and the importance of the neighboring task #l2 related to neighboring agent #j2.

Equation

[0051] Also, in Equation (5), the second o_(l1)^task from the right is information regarding neighboring task #l1. o_(l1)^task may indicate the state and target of neighboring task #l1. As described above, in Embodiment 1, task 50 is the load to be transported. In this case, o_(l1)^task can be defined as in the following Equation (9). Note that the right side of Equation (9) indicates, in order from the left, the position of neighboring task #l1, the speed of neighboring task #l1, and the target position (goal, i.e., the destination) of neighboring task #l1. The position and speed of neighboring task #l1 correspond to the state of the neighboring task. Note that the speed of the neighboring task can be calculated from the difference in position for each control cycle of the neighboring task. The same applies to the speeds of other objects described later.

Number

[0052] Similarly, in Equation (5), the first o_(l2)^task from the right is information regarding neighboring task #l2. o_(l2)^task may indicate the state and target of neighboring task #l2. Also, in Embodiment 1, o_(l2)^task can be defined as in the following Equation (10). Note that the right side of Equation (10) indicates, in order from the left, the position of neighboring task #l2, the speed of neighboring task #l2, and the target position (goal, i.e., the destination) of neighboring task #l2. The position and speed of neighboring task #l2 correspond to the state of the neighboring task.

Number

[0053] From Equation (5), Equation (9), and Equation (10), in Embodiment 1 where the task is the load to be transported, the observation information o i is represented as in the following Equation (11).

Number

[0054] The policy storage unit 112 stores a policy π (learned model) learned by reinforcement learning. The policy π is learned for each agent 10. Therefore, the learned policy π (parameters of the network (such as a neural network) constituting the policy π) can be different for each agent 10.

[0055] The policy π of agent #i NN,i takes the above-described observation information o i as input and outputs an action a shown by the following equation (12). i Therefore, the action a i is the output value of the policy π NN,i .

Equation

[0056] Here, c i l is the target value of the importance φ of the neighboring task #l regarding agent #i i l and corresponds to an index (intention) indicating how much agent #i values the neighboring task #l. Note that the importance φ i l can take values up to the target value c i l . In other words, the importance φ i l can grow up to the target value c i l . c ^(l1), which is the first component of the first right side (the second equation from the left) of equation (12), is the target value of the importance of the neighboring task #l1 regarding agent #i. Similarly, c ^(l2), which is the second component of the first right side of equation (12), is the target value of the importance of the neighboring task #l2 regarding agent #i.

[0057] Also, a i d , which is the third component of the first right side of equation (12), indicates the degree of request of agent #i. Also, a i σ which is the fourth component of the first right side of equation (12),​i σ indicates the degree of responsiveness of agent #i. Here, as shown in the following formula (13), a i d and a i σ can take values in the range of 0 or more and 1 or less.

Equation

[0058] The degree of requirement a i d indicates the degree to which agent #i requests support. That is, the higher the value of a i d , the higher the possibility that another agent 10 will cooperate with its own agent 10 to execute task 50. In other words, the higher the value of a i d , the higher the possibility that a request parameter indicating a request for support from another agent 10 will be calculated.

[0059] The degree of responsiveness a i σ indicates the degree to which agent #i responds to a request. That is, the higher the value of a i σ , the higher the possibility that its own agent 10 will cooperate with another agent 10 to execute task 50. In other words, the higher the value of a i σ , the higher the possibility that a response parameter indicating a response to a request from another agent 10 will be calculated.

[0060] Also, the policy π NN,i is learned to maximize the reward r i (t) shown in the following formula (14). That is, in the learning stage, the policy π NN,i takes the observation information o i as input and outputs the action a i . And when the action a i is output, the reward r i(t) is calculated. Then, the parameters (weights) of the network within the policy π are updated at any time so that the reward (cumulative reward) increases. Note that, as can be seen from the fact that there is no index i on the right side of Equation (14), the reward is common regardless of the agent. Then, from the obtained common reward, the network of Q-values different for each agent and the network of the policy are updated. Thereby, the policy π is learned. NN,i Since there is no index i on the right side of Equation (14), the reward is common regardless of the agent. Then, from the obtained common reward, the network of Q-values different for each agent and the network of the policy are updated. Thereby, the policy π NN,i is learned.

Number

[0061] t indicates time. Also, P l (t) indicates the degree of achievement of task #l at time t. Therefore, the first term on the right side of Equation (14) is the sum (Summation) of the degrees of achievement P l (t) of each task #l (#1, ···, #N) at time t. Note that the degree of achievement of task 50 may indicate the degree of achievement with respect to the goal of task 50. Alternatively, the degree of achievement of task 50 may indicate whether the goal of task 50 has been achieved.

[0062] Also, Q l (t) indicates the progress of task #l at time t. Therefore, the second term on the right side of Equation (14) is the sum of the progress Q l (t) of each task #l (#1, ···, #N) at time t. λ is a predetermined coefficient. Note that the progress of task 50 indicates how much task 50 has progressed. That is, the progress of task 50 indicates the progress of task 50. Therefore, if task 50 is progressing, the progress can be high, and if task 50 is stalled, the progress can be low.

[0063] Here, as described above, in Embodiment 1, task 50 is the load to be transported. And in Embodiment 1, that the goal of task 50 is achieved means that task 50 reaches the target position. Therefore, in Embodiment 1, P lDefine (t) as in the following formula (15). Formula (15) indicates that if the task #l, which is a piece of luggage, has reached the target position at time t, then P l (t) = 1; otherwise, P l (t) = 0. [Number] ···(15)

[0064] Also, in Embodiment 1, it can be said that if the luggage is moving fast, the transportation of that luggage is proceeding smoothly, and if the luggage is not moving very fast, the transportation of that luggage is stalled. That is, in Embodiment 1, the higher the speed of task 50, which is a piece of luggage, the higher the progress of that task 50 can be. Therefore, in Embodiment 1, Q l (t) is defined as in the following formula (16). As shown in formula (16), in Embodiment 1, the progress Q l (t) of task #l corresponds to the moving speed of task #l, which is a piece of luggage. That is, in Embodiment 1, the higher the speed of task #l, which is a piece of luggage, the higher the progress Q l (t) can be. [Number] ···(16)

[0065] [[ID=3,2]]From formula (16), in Embodiment 1, the reward r i (t) shown in the above formula (14) is expressed as in the following formula (17). [Number] ···(17)

[0066] The action output unit 120 outputs the action a i corresponding to the observation information o i using the above-described policy π. Specifically, the action output unit 120 inputs the observation information o i to the policy π NN,i . As a result, the policy π NN,i outputs the action ai Output.

[0067] The request response processing unit 130 calculates the request parameters and response parameters for its own agent 10. The request response processing unit 130 outputs the policy π NN,i Action a output from i Calculate the request parameters and response parameters for agent #i based on the i is the observation information i Therefore, it can be said that the request response processing unit 130 calculates the request parameters and response parameters based on the observation information.

[0068] The request response processing unit 130 outputs the policy π NN,i The degree of demand a output from i d Based on the request parameter d for agent #i i Specifically, the request response processing unit 130 calculates the request degree a i d A request parameter d indicates that assistance should be requested when i On the other hand, the request response processing unit 130 may calculate the request degree a i d is below the threshold, the request parameter d indicates that no support is requested. i Alternatively, the request response processing unit 130 may calculate the request degree a i d exceeds the threshold and the task 50 that the agent 10 is executing or about to execute is not progressing, a request parameter d i On the other hand, if the above condition is not met, the request response processing unit 130 may calculate a request parameter d i The request response processing unit 130 may calculate the calculated request parameter d i to the control device 100 for the other agent 10.

[0069] For example, the request response processing unit 130 calculates the request parameter d of agent #i using the following formula (18). i Here, d i = 1 indicates that agent #i requests support. d i = 0 indicates that agent #i does not request support. Therefore, the request parameter d i can function as a trigger for an event of a support request.

Equation

[0070] In formula (18), "0.5" is a predetermined threshold value. The threshold value is not limited to 0.5. Also, l i * indicates the currently selected task 50 for agent #i. In other words, l i * indicates the task 50 selected by the task selection unit 150 described later in the previous control cycle. To put it more simply, l i * indicates the task 50 that its own agent 10 (agent #i) is executing or about to execute. Also, Q_(l i * )(t) is the progress of task #l i * . And Q_(l i * )(t) = 0 indicates that task #l i * is not in progress.

[0071] Therefore, formula (18) means that when the request degree a i d ]>exceeds the threshold value "0.5" and the progress of the currently selected task #l for agent #i i * is 0 (that is, task #l i * is not in progress), the request parameter is di indicates that it is equal to 1. In other words, Equation (18) represents the degree of request a i d exceeds the threshold value of "0.5" and, when the currently selected task #l for agent #i i * is not in progress, it indicates that a request parameter d for requesting support is calculated. Also, Equation (18) represents that when the above conditions are not satisfied, d i is equal to 0. That is, Equation (18) represents that when the above conditions are not satisfied, a request parameter d for not requesting support is calculated. Note that in this embodiment, just because a request parameter indicating a request for support is calculated, it does not necessarily mean that another agent #j will actually come to support the currently selected task #l for agent #i i Whether another agent #j will actually come to support for task #l i can be determined according to the importance of that task #l. That is, whether another agent #j will actually come to support for task #l i * depends on the importance of task #l that agent #j has. In other words, whether another agent #j will actually come to support for task #l i * depends on the importance of task #l for agent #j. i * Note that when the currently selected task #l for its own agent #i is in progress, even if its own agent #i and another agent #j do not cooperate, task #l i * is in progress. In such a case, requesting support for task #l i * does not necessarily lead to another agent #j actually coming to support. Whether another agent #j will actually come to support for task #l i * depends on the importance of that task #l. That is, whether another agent #j will actually come to support for task #l i * depends on the importance of task #l for agent #j.

[0072] When the currently selected task #l for its own agent #i is in progress, even if its own agent #i and another agent #j do not cooperate, task #l i * is in progress. In such a case, requesting support for task #l i * is in progress.i * It may be wasteful to execute in cooperation with another agent #j. Therefore, in the above formula (18), even if the degree of request a i d is high, if the progress of the currently selected task #l for agent #i i * is not 0, then d i = 0. This can suppress making unnecessary requests.

[0073] Also, the request response processing unit 130 calculates a response parameter σ NN,i for agent #i based on the response degree a i σ output from the action output unit 120 according to the policy π. The request response processing unit 130 may calculate a response parameter σ i indicating that it responds to the request when the response degree a i σ exceeds a predetermined threshold. On the other hand, the request response processing unit 130 may calculate a response parameter σ i indicating that it does not respond to the request when the response degree a i σ is below the threshold. Alternatively, the request response processing unit 130 may calculate a response parameter σ i indicating that it responds to the request when the response degree a i σ exceeds the threshold and the task 50 that its own agent 10 is executing or about to execute is not in progress. On the other hand, the request response processing unit 130 may calculate a response parameter σ i indicating that it does not respond to the request when the above conditions are not met. i i

[0074] For example, the request response processing unit 130 calculates the response parameter σ i of agent #i using the following formula (19). Here, σ i = 1 indicates that agent #i responds to the request. σ i= 0 indicates that agent #i does not respond to the request. Therefore, the response parameter σ i can function as a trigger for the event of responding to the request.

Number

[0075] In Equation (19), "0.5" is a predetermined threshold value. The threshold value is not limited to 0.5. Also, this threshold value does not have to be the same as the threshold value in Equation (18). Also, as described above, l i * indicates the currently selected task 50 for agent #i. Also, Q_(l i * )(t) is the progress of task #l i * . And Q_(l i * )(t) = 0 indicates that task #l i * is not in progress.

[0076] Therefore, Equation (19) means that the response degree a i σ exceeds the threshold value "0.5", and the progress of the currently selected task #l for agent #i i * is 0 (that is, task #l i * is not in progress), then the response parameter is σ i = 1. In other words, Equation (19) means that when the response degree a i σ exceeds the threshold value "0.5", and the currently selected task #l for agent #i i * is not in progress, it indicates that the response parameter σ i indicating a response to the request is calculated. Also, Equation (19) means that when the above conditions are not met, σ iindicates that it is equal to 0. That is, when the above conditions are not satisfied, the response parameter σ i is calculated to indicate not responding to the request. Note that in this embodiment, just because a response parameter indicating responding to the request is calculated, it does not necessarily mean that the agent #i actually goes to support the task #l of another agent #j. Whether the agent #i actually goes to support the task #l of another agent #j can be determined according to the importance of the task #l. That is, whether the agent #i actually goes to support the task #l of another agent #j is determined by the importance of the task #l that the agent #j has. In other words, whether the agent #i actually goes to support the task #l of another agent #j is determined by the importance of the task #l for the agent #j.

[0077] Regarding the task #l selected for its own agent #i i * if it responds to the request and actually goes to support the task #l of another agent #j when the task #l is in progress, the agent #i will i * stop executing the task #l. However, stopping the execution of the task that the agent #i is in progress with can be wasteful. That is, for the task #l currently selected for the agent #i (that is, the task that the agent #i is executing) i * when it is in progress, it is preferable to continue executing that task #l i * Therefore, in the above formula (19), even if the response degree a i σ is high, if the progress degree of the task #l currently selected for the agent #i i * is not 0, then σ i = 0. This can suppress making a wasteful response.

[0078] Here, as described above, in Embodiment 1, task #l is the load to be transported. And in Embodiment 1, as shown in the above formula (16), the progress of task #l corresponds to the moving speed of task #l which is a load. Therefore, in Embodiment 1, for task #l i * the progress is expressed as in the following formula (20).

Equation

[0079] Therefore, in Embodiment 1, the request parameter d shown in the above formula (18) i is expressed as in the following formula (21).

Equation

[0080] Also, in Embodiment 1, the response parameter σ shown in the above formula (19) i is expressed as in the following formula (22).

Equation

[0081] The importance processing unit 140 updates (calculates) the importance of each surrounding task #l (l ∈ 1, ···, N) related to its own agent #i. Specifically, the importance processing unit 140 performs processing for calculating the importance of each task related to its own agent #i based on the request parameter of another agent #j and the response parameter of its own agent #i (the said agent).

[0082] Specifically, the importance processing unit 140 acquires the request parameter d j related to agent #j from the control device 100 of the surrounding agent #j. Also, the importance processing unit 140 acquires the importance φ of each task #l related to each surrounding agent #j from the control device 100 of each surrounding agent #jj l to obtain. As described above, the observation information acquisition unit 110 has obtained the importance levels of the neighboring tasks #l1, #l2 for the neighboring agents #j1, #j2. On the other hand, the importance processing unit 140 obtains, from the control devices 100 of all the (acquirable) agents #j in the vicinity, not only the neighboring agents, but also the importance level φ of each task #l for each agent #j j l to obtain. Note that, if the importance processing unit 140 fails to obtain the request parameter d j and the importance level φ j l from the control device 100 of the agent #j in the vicinity due to reasons such as communication being impossible, for this agent #j, d j = 0, φ j l may be set to 0.

[0083] Also, the importance processing unit 140 updates the importance level φ i l using the importance level φ j l for the task #l, the request parameter d j for other agents #j, and the response parameter σ i for its own agent #i. Also, when the task #l is a neighboring task, the importance processing unit 140 further updates the importance level φ i l for the task #l using the target value c i l of the importance level φ i l for the neighboring task #l of its own agent #i. Note that the importance level φ i l is the importance level of the task #l for its own agent #i. i l is the importance level of the task #l for its own agent #i.

[0084] Specifically, the importance processing unit 140 uses the following formula (23) to calculate the importance level φ of the task #l for the agent #i i lCalculate the amount of change (change). In Equation (23), k is a predetermined coefficient. For convenience of notation, the left side of Equation (23) may be expressed as “φ i l (dot)”.

Number

[0085] The importance processing unit 140 adds the change amount φ i l of the importance φ shown in Equation (23) i l to update the importance φ i l (dot) of task #l for agent #i. That is, the importance processing unit 140 updates the importance φ i l of task #l for agent #i according to the following Equation (24). Note that Δt represents the control period. The importance processing unit 140 updates the importance for all tasks #l. Note that the initial value φ i l of the importance for all agents #i (i = 1, ···, M) and all tasks #l (l = 1, ···, N) is predetermined. i l (0) is assumed to be predetermined.

Number

[0086] When task #l is not a neighboring task (when l does not satisfy the condition of the neighboring task shown in Equation (2)), the change amount φ i l of the importance φ i l (dot) is expressed as the following formula on the right side of Equation (23). The following formula on the right side of Equation (23) is the request parameter d j for agent #j and the importance φ j l from the importance φ i lThe sum, for all agents #j, of the product of the subtracted value and the coefficient k, is multiplied by the response parameter σ for agent #i i This corresponds to the multiplication result. Note that the lower formula on the right side of Equation (23) is the response parameter σ of the agent #i itself i If it is 0, it becomes 0. That is, when task #l is not a neighboring task, if the response parameter σ i is 0, the importance φ of task #l for the agent #i itself i l is not updated (does not change). Also, when the response parameter σ i is 1, for other agents #j (requesting agents) whose request parameter is 1, the sum of the differences obtained by subtracting the importance φ j l from the importance φ i l corresponds to the change amount φ i l (dot). Therefore, the more there are requesting agents #j whose importance φ j l is much larger than the importance φ i l , or the more there are requesting agents #j whose importance φ j l is larger than the importance φ i l , the more likely the importance φ of task #l for agent #i j l is to increase.

[0087] On the other hand, when task #l is a neighboring task (when task #l satisfies the conditions of the neighboring tasks shown in Equations (2) to (4)), the change amount φ i l of the importance φ i l (dot) is expressed as in the upper formula on the right side of Equation (23). The upper formula on the right side of Equation (23) subtracts the importance φ of task #l for agent #i from the target value c i l of the importance of task #l for agent #i in the lower formula i lIt corresponds to the value obtained by adding the product of the value obtained by subtracting and the coefficient k. Note that the second term in the upper expression on the right side of Equation (23) is the same as the lower expression on the right side of Equation (23). Therefore, the second term in the upper expression on the right side of Equation (23) is the response parameter σ of its own agent #i i becomes 0 if is 0. That is, when the task #l is a neighboring task, if the response parameter σ i is 0, according to the first term in the upper expression on the right side of Equation (23), the importance φ of the task #l with respect to its own agent #i i l can be updated to approach the target value c i l . Also, when the response parameter σ i is 1, the sum of the differences obtained by subtracting the importance φ j l from the importance φ i l , and the value obtained by adding the product of the value obtained by subtracting the importance φ i l from the target value c i l and the coefficient k corresponds to the change amount φ i l (dot). Therefore, when the target value c i l is large, the more there is a requesting agent #j whose importance φ j l is much larger than the importance φ i l , or the more there are requesting agents #j whose importance φ j l is larger than the importance φ i l , the importance φ of the task #l with respect to the agent #i j l can become larger.

[0088] In addition, the importance processing unit 140 performs processing on the importance of the task #l that has achieved the target. Specifically, the importance processing unit 140 processes the importance φ of the task #l that has achieved the target i lSet it to 0. Here, as described above, in Embodiment 1, when the goal of task #l is achieved, it means that the task #l, which is a piece of luggage, reaches the target position. Therefore, the importance processing unit 140 sets the importance φ of task #l that has reached the target position according to the following formula (25). i l Set it to 0. Note that δ is a threshold value for determining whether task #l has reached the target position (that is, whether task #l has achieved the goal). In the example of formula (1) and the like, it is 0.05. As a result, the processing for task #l that has achieved the goal is no longer performed, and each agent 10 will execute other tasks 50.

Number

[0089] In addition, the importance processing unit 140 transmits the calculated importance φ of each task #l regarding agent #i to the control device 100 of another agent #j. That is, the importance of each task 50 regarding each agent 10 is shared among each agent 10 (control device 100). As a result, the control device 100 of another agent #j performs the above processing for that agent #j. That is, the control device 100 of another agent #j calculates (updates) the importance φ of each task #l regarding that agent #j. i l j l

[0090] The task selection unit 150 selects the task #l to be executed by its own agent #i. i * Specifically, the task selection unit 150 selects, according to the following formula (26), among all tasks #l (l = 1, ···, N), the task #l with the maximum importance φ as the task #l to be executed by its own agent #i. i l i * ​​​Select as. In Embodiment 1, since Task 50 is the load to be transported, the Task Selection Unit 150 selects the load with the highest importance (Task #l i * ) for its own agent #i.

Number

[0091] The Task Execution Unit 160 performs processing for its own agent #i to execute Task #l. Specifically, the Task Execution Unit 160 controls the agent #i to execute the Task #l i * selected by the Task Selection Unit 150. More specifically, the Task Execution Unit 160 obtains the position of Task #l i * and the goal (end condition) to be achieved. The Task Execution Unit 160 moves the agent #i to the position of Task #l i * . At this time, the Task Execution Unit 160 may calculate the speed command value of the agent #i. Then, the Task Execution Unit 160 controls the agent #i to execute Task #l i * and achieve the goal of Task #l i * . When Task #l is a load, the Task Execution Unit 160 controls the arm of the agent #i to grip the load. At this time, the Task Execution Unit 160 may calculate the force and torque command values of the tip of the arm (end effector, etc.). The Task Execution Unit 160 controls the agent #i to transport Task #l i * to the target position of Task #l i * .

[0092] Note that when the request parameter d i of its own agent #i is 0, in the processing in the control device 100 of another agent #j, in the second term of the upper formula on the right side of Equation (23) and the lower formula on the right side, d jk(φ j l -φ i l ) = 0. Note that since this is the processing in the control device 100 of another agent #j, it should be noted that j corresponds to its own agent #i and i corresponds to another agent #j. Therefore, when the request parameter d i of its own agent #i is 0, even if the importance of task #l is high in its own agent #i in the processing in the control device 100 of another agent #j, it is highly likely that it will not affect the importance of task #l regarding another agent #j. Here, task #l with high importance in its own agent #i may include task #l i * selected for its own agent #i. From the above, when the request parameter d i is 0, in the processing in the control device 100 of another agent #j, the possibility that task #l i * selected for its own agent #i is selected is low. Therefore, the possibility that another agent #j will come to support for task #l i * decreases.

[0093] On the other hand, when the request parameter d i of its own agent #i is 1, in the processing in the control device 100 of another agent #j, the possibility that the importance (significance) of task #l with high importance in its own agent #i increases. Therefore, when the request parameter d i is 1, the possibility that task #l i * selected for its own agent #i is also selected in the processing in the control device 100 of another agent #j increases. That is, the possibility that another agent #j will come to support for task #l i * increases. Therefore, when task #l i * is not in progress, the request parameter d ibecomes 1, which increases the likelihood that the task #l will be executed in cooperation between the agent #i itself and another agent #j. As a result, the likelihood of achieving the goal of the task #l increases. i * This increases the likelihood that the goal of task #l will be achieved. i *

[0094] Also, when the response parameter σ i is 0, in the processing of the control device 100 of the agent #i itself, the second term of the upper formula and the lower formula on the right side of Equation (23) become 0. Therefore, when the response parameter σ i is 0, in the processing of the control device 100 of the agent #i itself, it is difficult for the importance of task #l with a high importance in another agent #j (requesting agent) to increase. Here, task #l with a high importance in another agent #j may include the task #l selected for the other agent #j j * From the above, when the response parameter σ i is 0, in the processing of the control device 100 of the agent #i itself, the task #l selected for the other agent #j j * has a lower likelihood of being selected. Therefore, the likelihood that the agent #i itself will go to support the task #l j * decreases.

[0095] i l <ElementName>φ< / ElementName> of the nearby task #l approaches the target value c i l from the first term on the right side of Equation (23). Here, the nearby task #l is the task #l selected in the previous control cycle of the agent #i and is the task #l that the agent #i is currently executing i * i and as long as the response parameter σ is 0, the likelihood that the nearby task #l will continue to be selected is high.

[0096] In contrast, for the response parameter σ i ​​When it is 1, in the process of the control device 100 of its own agent #i, the second term of the upper formula on the right side of formula (23) and the lower formula on the right side do not become 0. That is, the response parameter σ i When it is 1, in the process of the control device 100 of its own agent #i, the importance of task #l with a high importance in the requesting agent #j is likely to increase. From the above, the response parameter σ i When it is 1, the task #l selected for the requesting agent #j j * is likely to be selected also in the process of the control device 100 of its own agent #i. That is, its own agent #i is likely to go to support task #l j * Therefore, when the task #l selected by another requesting agent #j j * is not in progress, the response parameter σ i becoming 1 makes it highly likely that task #l j * will be executed cooperatively between its own agent #i and another requesting agent #j. As a result, the possibility of achieving the goal of task #l j * becomes high.

[0097] FIG. 4 is a flowchart showing a control method executed by the control device 100 according to Embodiment 1. As described above, the observation information acquisition unit 110 calculates the distances between the surrounding agents 10 and the luggage (task 50) and its own agent 10 (step S102). As described above, the observation information acquisition unit 110 acquires the observation information o i of its own agent 10 (agent #i) (step S110).

[0098] As described above, the action output unit 120 uses the policy π NN,i to obtain the action a i corresponding to the observation information o iOutput it (step S120). The request response processing unit 130 performs request response processing (step S130). Specifically, as described above, the request response processing unit 130 calculates the request parameter d i and the response parameter σ i for its own agent #i.

[0099] The importance processing unit 140 updates the importance of the task 50, which is the package (step S140). Specifically, as described above, the importance processing unit 140 updates (calculates) the importance of the surrounding task #l (package) related to its own agent #i. Also, the importance processing unit 140 processes the importance of the package that has reached the goal (step S142). Specifically, as described above, the importance processing unit 140 sets the importance of the task #l (package) that has reached the goal and achieved the target to 0.

[0100] Also, as described above, the task selection unit 150 selects the package with the highest importance for its own agent 10 (step S150). The task execution unit 160 performs processing so that its own agent 10 transports the selected package (step S160).

[0101] The control device 100 determines whether the distance between the positions of all packages and their target positions is less than a certain value (step S170). Here, the fact that the distance between the position of a package and its target position is less than a certain value means that it can be regarded that the package has reached the target position. Therefore, the control device 100 determines whether all packages have reached the target position. Note that the "certain value" corresponds to δ in formula (25) (for example, δ = 0.05). If the distance between the positions of all packages and their target positions is less than a certain value (YES in S170), the processing flow ends. On the other hand, if the distance between the positions of all packages and their target positions is not less than a certain value (NO in S170), the processing flow returns to S102. Then, the processing from S102 to S170 is repeated. This repetition of the processing is performed for each of the above-described control cycles.

[0102] As described above, the control device 100 according to Embodiment 1 calculates a request parameter and a response parameter based on observation information, and performs processing for calculating the importance of each task related to the agent based on the request parameter and the response parameter. Then, the control device 100 according to Embodiment 1 selects a task to be executed by the agent according to the importance, and controls the agent to execute the selected task. By being configured in this way, the control device 100 according to Embodiment 1 can appropriately select a task to be executed by the agent according to the importance calculated according to the observation information, the request parameter, and the response parameter. Thereby, even in an environment where the task is unknown, it is possible to suppress many agents from concentrating on one task, and it is possible to realize an operation in which an agent goes to support a task that does not progress. As a result, the task will surely progress. Therefore, even in an environment where the task is unknown, it is possible to make the task target be efficiently achieved. Therefore, overall, it is possible to reduce the execution time (total execution time) of the task.

[0103] Further, the control device 100 according to Embodiment 1 is configured to calculate a request parameter and a response parameter based on a policy learned for each agent. Further, preferably, the control device 100 according to Embodiment 1 is configured to calculate a request parameter and a response parameter based on the degree of request and the degree of response output from the policy by inputting the observation information into the policy. Thereby, it is possible to calculate the importance so that the importance of a task that requires support can be increased for each agent. Therefore, it is possible to appropriately select a task to be executed for each agent.

[0104] Further, when the degree of requirement of the control device 100 according to Embodiment 1 exceeds a predetermined threshold value and the task that the agent is executing or about to execute is not in progress, a request parameter indicating that support is requested is calculated. With such a configuration, when support should be requested for the task that the agent is executing or about to execute, a request parameter indicating that support is requested can be appropriately calculated. Therefore, it is possible to suppress making unnecessary requests.

[0105] Further, when the degree of response of the control device 100 according to Embodiment 1 exceeds a predetermined threshold value and the task that the agent is executing or about to execute is not in progress, a response parameter indicating that the request is responded to is calculated. With such a configuration, when the task that the agent is executing or about to execute is in progress, the task can continue to be executed. Therefore, it is possible to suppress making unnecessary responses.

[0106] Further, the control device 100 according to Embodiment 1 calculates the importance of each task related to the agent based on the policy learned for each agent. With such a configuration, it becomes possible to appropriately calculate the importance of each task for each agent.

[0107] Further, the control device 100 according to Embodiment 1 calculates the importance of the task corresponding to the observation information related to the agent based on the target value of the importance of the task corresponding to the observation information output from the policy by inputting the observation information into the policy. With such a configuration, for each agent, the importance of the task corresponding to the observation information can be calculated so as to approach the target value. Thereby, it becomes possible to appropriately calculate the importance of the task.

[0108] (Embodiment 2) Next, Embodiment 2 will be described. Note that the configuration of the control system 1 according to Embodiment 2 is substantially the same as the configuration of the control system 1 according to Embodiment 1 shown in FIG. 1, and thus the description thereof will be omitted. Also, the configuration of the control device 100 according to Embodiment 2 is substantially the same as the configuration of the control device 100 according to Embodiment 1 shown in FIG. 2, and thus the description thereof will be omitted. In Embodiment 2, Task 50 is different from that in Embodiment 1.

[0109] In Embodiment 2, Task 50 is a location where a large number of packages to be transported exist. And for each package, a goal (target position) which is the destination is set. And the goal of Task 50 according to Embodiment 2 is that all the packages existing at that location reach their respective goals. Note that, unlike Embodiment 1, the packages to be transported in Embodiment 2 can be small enough to be transported by one agent 10. Note that the larger the number of agents 10 that transport the packages existing at that location (Task 50), the higher the possibility that that location (Task 50) achieves the goal (the possibility that all the packages at that location reach the goal).

[0110] Task 50 according to Embodiment 2 may be, for example, each room in a hospital. And it may be assumed that there are medical records, medicines, specimens, etc., which are packages to be transported, in each room which is Task 50. Also, Task 50 according to Embodiment 2 may be, for example, a location where relief supplies are placed during a disaster. And the relief supplies may be packages to be transported.

[0111] In Embodiment 2, the observation information acquisition unit 110, similar to Embodiment 1, acquires observation information o as shown in Equation (5). iObtain it. At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and the location of task #l. Specifically, similar to Embodiment 1, the observation information acquisition unit 110 acquires the position of the location of task #l. The observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (location #l) from the acquired position of task #l (location #l). Then, the observation information acquisition unit 110 determines a predetermined number of neighboring tasks close to its own agent #i from the distances between its own agent #i and each task #l (location #l), and uses the information on the neighboring tasks as part of the observation information. In the examples of the above-described formulas (2) to (5), the neighboring tasks are the location #l1 closest to its own agent #i and the second-closest location #l2.

[0112] Also, in Embodiment 2, \(o_{(l1)}^{task}\) in formula (5) may indicate the state of each piece of luggage and the goal of each piece of luggage at location #l1 which is a neighboring task. The state of each piece of luggage may be the position and velocity of each piece of luggage. Note that \(o_{(l1)}^{task}\) may indicate the average of the states (position and velocity) of each piece of luggage at location #l1 which is a neighboring task and the average position of the goals of each piece of luggage. The average position may be the centroid (geometric center) of each position (goal). Alternatively, \(o_{(l1)}^{task}\) may indicate the number of pieces of luggage at location #l1 which is a neighboring task and the target number of pieces of luggage (that is, 0 pieces). The same applies to \(o_{(l2)}^{task}\).

[0113] In Embodiment 2, the policy storage unit 112 stores the policy π (learned model) learned by reinforcement learning in the same manner as in Embodiment 1. The policy π is learned for each agent 10. The policy π of agent #i NN,i takes the above-described observation information o i as input and outputs the action a i shown in the above formula (12). Also, similar to Embodiment 1, the policy π NN,i is learned to maximize the reward r i (t) shown in the above formula (14).

[0114] Here, in the second embodiment, P l (t) may indicate the degree of transport achievement of the luggage at location #l at time t. The degree of transport achievement may correspond to, for example, the ratio of the number of luggage transported to the goal to the number of luggage initially present at location #l. In addition, in the second embodiment, Q regarding the progress of task #l l (t) may be the amount of luggage lost at time t. The amount of luggage lost may correspond to, for example, the number of luggage lost (transported) from location #l per unit time.

[0115] 5 is a flowchart showing a control method executed by the control device 100 according to the second embodiment. As described above, the observation information acquisition unit 110 calculates the distance between the agent 10 itself and the surrounding agents 10 and the location (task 50) (step S202). As described above, the observation information acquisition unit 110 calculates the distance between the agent 10 itself and the surrounding agents 10 and the location (task 50) (step S203). i (Step S210). As in the first embodiment, the action output unit 120 obtains the measure π NN,i Using the observation information i Action a corresponding to i is output (step S220).

[0116] The request response processing unit 130 performs the request response process in the same manner as in the first embodiment (step S230). Specifically, the request response processing unit 130 uses the above equation (18) to calculate the request parameter d i The request response processing unit 130 may also calculate the response parameter σ i may be calculated.

[0117] The importance processing unit 140 updates the importance of the location, which is the task 50, in the same manner as in the first embodiment (step S240). That is, the importance processing unit 140 receives request parameters d related to the agent #j from the control devices 100 of other agents #j in the vicinity, in the same manner as in the first embodiment. jand the importance φ of each task #l (location #l) j l are obtained. Then, the importance processing unit 140 updates (calculates) the importance φ of the surrounding tasks #l (location #l) related to its own agent #i i l The importance processing unit 140 may update the importance φ of each task #l (location #l) related to its own agent #i by using the above formulas (23) and (24). i l

[0118] The importance processing unit 140 performs processing on the importance of the location where all the packages have reached the goal (step S242). Specifically, similar to Embodiment 1, the importance processing unit 140 sets the importance φ of the location #l where all the packages related to its own agent #i have reached the goal i l to 0.

[0119] Similar to Embodiment 1, the task selection unit 150 selects the location with the highest importance for its own agent 10 (step S250). That is, similar to Embodiment 1, the task selection unit 150 may use the above formula (26) to select the location #l with the highest importance φ for its own agent #i i l

[0120] The task execution unit 160 performs control so that its own agent 10 moves to the selected location and performs the package transportation process (step S260). Specifically, the task execution unit 160 controls the agent #i to move to the selected location l i * In addition, when the agent #i moves to the location l i * the task execution unit 160 controls to transport the packages existing at the location l i * to the goal. The method of transporting the packages is the same as that in Embodiment 1 described above.

[0121] ​​The control device 100 determines whether the distance between the location of all luggage and the goal is less than a certain value at all locations (step S270). Here, if the distance between the location of all luggage and the goal is less than the certain value, it means that the luggage can be considered to have reached the goal. Therefore, the control device 100 determines whether all luggage has reached the goal at all locations. Note that the "certain value" may correspond to δ in equation (25). If the distance between the location of all luggage and the goal at all locations is less than the certain value (YES in S270), the processing flow ends. On the other hand, if the distance between the location of all luggage and the goal at all locations is not less than the certain value (NO in S270), the processing flow returns to S202. Then, the processing of S202 to S270 is repeated. This processing is repeated at each control cycle described above.

[0122] As in the first embodiment, the control device 100 according to the second embodiment calculates request parameters and response parameters based on observation information and performs processing to calculate the importance of each task related to the agent based on the request parameters and response parameters. The control device 100 according to the second embodiment then selects a task to be executed by the agent based on the importance and controls the agent to execute the selected task. Therefore, as in the first embodiment, the control device 100 according to the second embodiment can appropriately select a task to be executed by the agent based on the importance calculated based on the observation information, request parameters, and response parameters. This prevents many agents from concentrating on a single task, even in an environment where tasks are unknown, and enables agents to provide support for tasks that are not progressing. This ensures that tasks progress. Therefore, even in an environment where tasks are unknown, it is possible to efficiently achieve task goals. Therefore, it is possible to reduce the overall task execution time (total execution time).

[0123] (Modification of Embodiments 1 and 2) In addition, in Embodiment 1 and Embodiment 2, the agent 10 is assumed to be a machine such as a robot. However, the agent 10 does not have to be a machine. The agent 10 may include machines and humans. That is, a robot and a human may cooperate to carry a plurality of pieces of luggage. At this time, the human may carry a communication terminal capable of communicating with the control device 100 related to the agent 10. Note that the agent 10 that is a machine can be controlled by substantially the same processing as in Embodiment 1 and Embodiment 2 described above.

[0124] At that time, the control device 100 of the agent 10 that is a machine may transmit its request parameters and the importance of each task to the communication terminal carried by the human. Based on the request parameters and the importance of each task obtained from other agents 10, the human may, at his or her own discretion, respond to the agent 10 that has transmitted the request parameters indicating a request for support, and may provide support for the task being executed by that agent 10. Note that the human can independently determine the luggage to be carried based on the sense of accomplishment obtained by the completion of the luggage transportation as the principle of action. Note that the human does not request support from other agents 10. That is, the human does not have to transmit request parameters indicating a request for support to other agents 10. This is because the timing of a human's request cannot be simulated in the learning of the strategy regarding the agent 10 that is a robot, and if a human requests support, there is a possibility that an undesirable result may occur in the behavior of the agent 10 that is a robot.

[0125] (Embodiment 3) Next, Embodiment 3 will be described. Regarding the configuration of the control system 1 according to Embodiment 3, since it is substantially the same as the configuration of the control system 1 according to Embodiment 1 shown in FIG. 1, the description thereof will be omitted. Also, regarding the configuration of the control device 100 according to Embodiment 3, since it is substantially the same as the configuration of the control device 100 according to Embodiment 1 shown in FIG. 2, the description thereof will be omitted. In Embodiment 3, the task 50 is different from the above-described embodiments. Here, in the above-described embodiments, the target of the task 50 is achieved by the agent 10 transporting the luggage. On the other hand, in Embodiment 3, the agent 10 does not have to transport the luggage when executing the task 50. A specific example of the task 50 in Embodiment 3 will be described later. Similar to Embodiment 1, the agent 10 autonomously operates in the environment under the control of the control device 100. Also, in Embodiment 3, the monitoring device 60 does not have to monitor the task 50. Each agent 10 may monitor (detect) the state of the task 50.

[0126] FIG. 6 is a flowchart showing a control method executed by the control device 100 according to Embodiment 3. Similar to S102 and S202, the observation information acquisition unit 110 calculates the distances between the surrounding agents 10 and the task 50 and its own agent 10 (step S302). Similar to S110 and S210, the observation information acquisition unit 110 acquires the observation information o i of its own agent 10 (agent #i) (step S310). The observation information o i will be described later. Similar to S120 and S220, the action output unit 120 uses the policy π NN,i to output the action a i corresponding to the observation information o i (step S320). The reward r NN,i regarding the policy π i (t) will be described later.

[0127] The request response processing unit 130 performs the request response process in the same manner as in S130 and S230 (step S330). Specifically, the request response processing unit 130 uses the above equation (18) to calculate the request parameter d i The request response processing unit 130 may also calculate the response parameter σ i may be calculated.

[0128] The importance processing unit 140 updates the importance of the task 50 in the same manner as in S140 and S240 (step S340). Specifically, the importance processing unit 140 receives request parameters d related to the agent #j from the control devices 100 of other nearby agents #j, in the same manner as in the above-described embodiment. j and the importance of each task #l φ j l Then, the importance processing unit 140 obtains the importance φ of the peripheral task #l related to the agent #i of the agent #i, as in the above-described embodiment. i l The importance processing unit 140 updates (calculates) the importance φ of each task #l related to its own agent #i using the above formulas (23) and (24). i l may be updated.

[0129] The importance processing unit 140 processes the importance of the completed task in the same manner as in S142 and S242 (step S342). Specifically, as in the above-described embodiment, the importance processing unit 140 calculates the importance φ of the task #l for which the goal has been achieved, for its own agent #i. i l Set to 0.

[0130] As in S150 and S250, the task selection unit 150 selects the task 50 with the highest importance for its own agent 10 (step S350). Specifically, as in the above-described embodiment, the task selection unit 150 uses the above formula (26) to calculate the importance φ i lYou may select the task #l with the largest

[0131] Similar to S160 and S260, the task execution unit 160 performs control so as to execute the task 50 selected by its own agent 10 (step S360). Specifically, the task execution unit 160 controls the agent #i to move to the position of the selected task l i * Also, when the agent #i moves to the position of the task l i * the task execution unit 160 performs control so as to execute the task l i * A specific example of the task 50 will be described later.

[0132] The control device 100 determines whether or not all the tasks 50 have been completed (step S370). If all the tasks 50 have been completed (YES in S370), the processing flow ends. On the other hand, if not all the tasks 50 have been completed (NO in S370), the processing flow returns to S302. Then, the processing of S302 to S370 is repeated. This repetition of the processing is performed for each of the above-described control cycles.

[0133] Similar to Embodiment 1, the control device 100 according to Embodiment 3 performs processing for calculating a request parameter and a response parameter based on observation information, and calculating the importance of each task related to the agent based on the request parameter and the response parameter. Then, the control device 100 according to Embodiment 3 selects a task to be executed by the agent according to the importance, and controls the agent to execute the selected task. Therefore, similar to Embodiment 1, the control device 100 according to Embodiment 3 can appropriately select a task to be executed by the agent according to the importance calculated according to the observation information, the request parameter, and the response parameter. As a result, even in an environment where the task is unknown, it is possible to suppress unnecessary concentration such that a plurality of agents concentrate on one task, and realize an operation in which the agent supports a task that is not progressing. As a result, the task will surely progress. Therefore, even in an environment where the task is unknown, it is possible to efficiently achieve the task target. Therefore, overall, it is possible to reduce the execution time (total execution time) of the task.

[0134] <Specific Example 1 of Embodiment 3> In Specific Example 1, the method according to this embodiment is applied to maintenance. In Specific Example 1, a plurality of agents 10, which are machines such as robots, perform maintenance (inspection) of a structure. Also, in Specific Example 1, a plurality of agents 10 perform inspection of a wide range of structures. Here, in Specific Example 1, the task 50 is each inspection location. Also, in Specific Example 1, the target of the task 50 is that a comprehensive inspection is performed for each inspection location. Note that each of the plurality of agents 10 may have different functions. Therefore, the plurality of agents 10 may be composed of different types of agents 10. A comprehensive inspection can be realized by a plurality of different types of agents 10. Therefore, the higher the number of agents 10, the higher the possibility that the target of the task 50 is achieved.

[0135] For example, one agent 10 may be a search robot that searches for an abnormality. Another agent 10 may be a response robot that responds to the abnormality. Another agent 10 (search robot) may have a camera that photographs the inspection location and determine the severity of the abnormality from the image obtained by the photography. Another agent 10 may have a function for performing a first non-destructive inspection (e.g., infrared inspection). Another agent 10 may have a function for performing a second non-destructive inspection (e.g., ultrasonic flaw detection). Another agent 10 may have a function for performing a third non-destructive inspection (e.g., radiographic inspection). Another agent 10 may have a function for performing a fourth non-destructive inspection (e.g., eddy current flaw detection).

[0136] In the specific example 1, the observation information acquisition unit 110 acquires the observation information o as shown in equation (5) in the same manner as in the above-described embodiment. i (S310). At that time, the observation information acquiring unit 110 calculates the distance between its own agent #i and task #l (inspection location #l) (S302). Specifically, as in the first embodiment, the observation information acquiring unit 110 acquires the position of the inspection location, which is task #l. From the acquired position of task #l (inspection location #l), the observation information acquiring unit 110 calculates the distance between its own agent #i and each task #l (inspection location #l). Then, from the distance between its own agent #i and each task #l (inspection location #l), the observation information acquiring unit 110 determines a predetermined number of neighboring tasks that are close to its own agent #i, and makes information about the neighboring tasks part of the observation information. In the example of the above-mentioned formulas (2) to (5), the neighboring tasks are the inspection location #l1 closest to its own agent #i and the second-closest inspection location #l2.

[0137] In addition, in Specific Example 1, \(o_{(l1)}^{task}\) in Equation (5) may indicate the state of the inspection point #l1, which is a neighboring task, and the end condition of the inspection. The state of each inspection point may be the position of the inspection point, the content of the executed inspection, and the severity of the abnormality. The end condition of the inspection may be that all types of agents 10 reach the inspection point and all types of inspections (comprehensive inspections) are executed. The same applies to \(o_{(l2)}^{task}\).

[0138] In Specific Example 1, the policy storage unit 112 stores a policy π (learned model) learned by reinforcement learning, in the same manner as in the above-described embodiment. The policy π is learned for each agent 10. The policy π of agent #i NN,i takes the above-described observation information o i as input and outputs an action a i shown in the above Equation (12). Also, similar to Embodiment 1, the policy π NN,i is learned to maximize the reward r i (t) shown in the above Equation (14).

[0139] Here, in Specific Example 1, \(P_{(l)}(t)\) regarding the achievement degree of task #l l may indicate the achievement degree of the inspection of inspection point #l at time t. The achievement degree of the inspection may correspond to, for example, the number of types of agents 10 that have reached a serious inspection point #l and executed processing. Alternatively, the achievement degree of the inspection may be the number of inspection items executed. Also, in Specific Example 1, \(Q_{(l)}(t)\) regarding the progress degree of task #l l may be the progress degree of the inspection of inspection point #l at time t. The progress degree of the inspection may correspond to the number of agents 10 that reach a serious inspection point #l per unit time. Alternatively, the progress degree of the inspection may be the number of inspection items executed per unit time.

[0140] In addition, in Specific Example 1, the request response processing unit 130, in the same manner as in the above-described embodiment, determines the request parameter d regarding agent #i from the action a NN,i output from the policy π i ​i and the response parameter σ i Then, the request response processing unit 130 calculates the request parameter d i In the first specific example, the importance processing unit 140 transmits the importance φ of the peripheral task #l (inspection location #l) related to its own agent #i to the control device 100 of the peripheral agent 10, as in the above-described embodiment. i l The importance processing unit 140 updates (calculates) the importance φ of each task #l (inspection location #l) related to its own agent #i using the above formulas (23) and (24). i l In the specific example 1, the task selection unit 150 may update the importance φ of all tasks #l by using the above formula (26), as in the above embodiment. i l The largest task #l (inspection point #l) is the task #l that should be executed by its own agent #i. i * (S350).

[0141] Here, in the specific example 1, the request response processing unit 130 may transmit request parameters to the control device 100 of an agent 10 of a type different from the agent #i of the request response processing unit 130 itself. This increases the possibility that an agent 10 of a type different from the agent #i of the request response processing unit 130 itself will reach the inspection point. Conversely, it decreases the possibility that an agent 10 of the same type as the agent #i of the request response processing unit 130 itself will reach the inspection point. In other words, when the agent #i, which is a search robot, detects a serious inspection point #l, it is assumed that the degree of request output from the policy in the control device 100 of the agent #i will increase. Then, the control device 100 of the agent #i, which is a search robot, i= 1 to the control device 100 of a different type of agent 10 (such as an agent 10 that performs non-destructive testing). As a result, it can be assumed that the importance of the inspection location #l will increase in the control device 100 of the different type of agent 10 (such as an agent 10 that performs non-destructive testing). Therefore, the control device 100 of the different type of agent 10 (such as an agent 10 that performs non-destructive testing) will select the inspection location #l, and the different type of agent 10 will be more likely to reach the inspection location #l. This increases the likelihood that the goal of task 50 will be achieved. Note that if information about which inspection locations have not yet been inspected is added to the observation information, the agent that obtained the observation information can determine whether to proactively respond to a request for assistance at the inspection location.

[0142] In the specific example 1, the task execution unit 160 controls the agent #i to execute the selected task #l (inspection point #l) (S360). i * (Inspection point #l i * ) position. The task execution unit 160 controls its own agent #i so that it executes an inspection that matches the function of its own agent #i. Then, the control device 100 of each agent 10 performs processing so that a multifaceted inspection is executed for all inspection points.

[0143] <Specific Example 2 of Embodiment 3> In Specific Example 2, the method according to this embodiment is applied to monitoring, patrol, and security. In Specific Example 2, a plurality of agents 10, which are machines such as robots, perform environmental conservation. Specifically, the plurality of agents 10 patrol the environment and perform processes related to monitoring, patrol, and security. The plurality of agents 10 patrol the environment to search for problems existing in the environment. Also, in Specific Example 2, when a problem is searched, the control device 100 of the agent 10 that has searched for the problem transmits information regarding the searched problem to the control devices 100 of the surrounding agents 10. Here, in Specific Example 2, the task 50 is the "searched problem". Also, in Specific Example 2, the goal of the task 50 is to solve the searched problem. The greater the number of agents 10, the higher the possibility of achieving the goal of the task 50.

[0144] Also, in Specific Example 2, the plurality of agents 10 may have a function to handle the searched problem. Also, as in Specific Example 1, the plurality of agents 10 may have different functions from each other. For example, when the searched problem is "removal of bulky waste", an agent 10 capable of transporting bulky waste may remove the bulky waste. Also, when the searched problem is "apprehension of a criminal", an agent 10 capable of apprehending a criminal may apprehend the criminal. Also, when the searched problem is "handling of a lost person", an agent 10 capable of providing directions may handle the lost person. In the following description, the "searched problem" may be simply referred to as "problem".

[0145] In Specific Example 2, the observation information acquisition unit 110, similar to the above-described embodiment, acquires observation information o as shown in Equation (5). iAcquire it (S310). At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and task #l (task #l) (S302). Specifically, similar to the first embodiment, the observation information acquisition unit 110 acquires the position of "searched task" which is task #l. The position of the "searched task" may be the position of the agent 10 that searched the searched task at the time of search. The observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (task #l) from the acquired position of task #l (task #l). Then, the observation information acquisition unit 110 determines a predetermined number of neighboring tasks close to its own agent #i from the distances between its own agent #i and each task #l (task #l), and uses the information on the neighboring tasks as part of the observation information. In the examples of the above-described formulas (2) to (5), the neighboring tasks are task #l1 which is the closest to its own agent #i and task #l2 which is the second closest.

[0146] Also, in Specific Example 2, o_(l1)^task in formula (5) may indicate the state of task #l1 which is a neighboring task and the end condition of the task. The state of each task may be the position of the task, the degree of solution of the task, the quality of the task, and the type of the task. The end condition of the task may be that the task is solved. The same applies to o_(l2)^task.

[0147] In Specific Example 2, the policy storage unit 112 stores the policy π (learned model) learned by reinforcement learning in the same manner as in the above-described embodiment. The policy π is learned for each agent 10. The policy π of agent #i NN,i takes the above-described observation information o i as input and outputs the action a i shown in the above formula (12). Also, similar to the first embodiment, the policy π NN,i is learned to maximize the reward r i (t) shown in the above formula (14).

[0148] Here, in Specific Example 2, P regarding the achievement degree of task #l l(t) may indicate the degree of achievement of problem-solving regarding problem #l at time t. The degree of achievement of problem-solving may correspond, for example, to the completion of the handling by agent 10 capable of handling the problem. Also, in Specific Example 2, Q regarding the progress of task #l l (t) may be the progress of handling the problem regarding problem #l at time t. The progress of handling the problem may correspond, for example, to the fact that agent 10 capable of handling the problem is performing the handling.

[0149] Also, in Specific Example 2, the request response processing unit 130, in the same manner as the above-described embodiment, the policy π NN,i from which the output action a i from which, for agent #i, the request parameter d i and the response parameter σ i are calculated (S330). Then, the request response processing unit 130 transmits the request parameter d i to the control device 100 of the surrounding agent 10. At this time, like the modification examples of Embodiment 1 and Embodiment 2, the control device 100 may transmit the request parameter to the terminal carried by a human. Thereby, a human may handle the problem. Also, like Specific Example 1, the control device 100 may transmit the request parameter to the control device 100 of an agent 10 of a type different from its own agent #i.

[0150] Also, in Specific Example 2, the importance processing unit 140, in the same manner as the above-described embodiment, updates (calculates) the importance φ i l of the surrounding task #l (problem #l) regarding its own agent #i. The importance processing unit 140 may update the importance φ i l of each task #l regarding its own agent #i using the above equations (23) and (24). Also, in Specific Example 2, the task selection unit 150, in the same manner as the above-described embodiment, according to the above equation (26), among all tasks #l, the importance φ i lThe task #l (issue #l) that has the largest number of tasks is the task #l that should be executed by its own agent #i. i * (S350).

[0151] If the agent #i that has searched for the problem #l does not have the function to deal with the problem, it is assumed that the degree of request output from the policy in the control device 100 of the agent #i will be large. i = 1 to the control devices 100 of the surrounding agents 10. Then, among the surrounding agents 10, the control devices 100 of the agents 10 that can handle the task acquire observation information indicating the quality and type of the task and the request parameters, and it is assumed that the importance of task #l will increase. Therefore, the control devices 100 of the agents 10 that can handle the task will select task #l, and the agents 10 that can handle the task will be more likely to reach the location of task #l. Note that if information about which tasks have not yet been handled is added to the observation information, the agent that acquired the observation information will be able to determine whether to proactively respond to the request for help with the task.

[0152] In the second specific example, the task execution unit 160 controls the agent #i to execute the selected task #l (task #l) (S360). i * (Assignment #l i * ) position. The task execution unit 160 controls its own agent #i so that it executes the tasks that it can execute. Then, the control device 100 of each agent 10 performs processing so that all the tasks that have been found are solved.

[0153] <Specific Example 3 of Embodiment 3> In Example 3, the method according to the present embodiment is applied to coexistence with nature. In Example 3, animals are monitored and their movements are controlled by a plurality of agents 10, which are machines such as robots. This prevents animals from invading farmland, thereby realizing a sustainable ecosystem and reducing agricultural damage.

[0154] In Example 3, multiple agents 10 detect animals by detecting objects moving on or around the farm. Each of the multiple agents 10 may have a different function. That is, as in Example 1, the multiple agents 10 may be composed of different types of agents. In this case, one agent 10 may have a function to search for animals. Another agent 10 may have a function to make animals leave the farm. Alternatively, each of the multiple agents 10 may have both a function to search for animals and a function to make animals leave the farm. That is, each of the multiple agents 10 may be the same type of agent 10. The control device 100 of an agent 10 that detects an animal transmits information about the animal (such as the animal's location) to the control devices 100 of the other agents 10. In Example 3, each detected animal is a task 50. The goal of the task 50 is to make the animals leave the farm. The greater the number of agents 10, the higher the likelihood that the goal of the task 50 will be achieved.

[0155] In the specific example 3, the observation information acquisition unit 110 acquires the observation information o as shown in equation (5) in the same manner as in the above-described embodiment. i(S310). At that time, the observation information acquisition unit 110 calculates the distance between its own agent #i and task #l (animal #l) (S302). Specifically, as in the first embodiment, the observation information acquisition unit 110 acquires the position of the animal, which is task #l. From the acquired position of task #l (animal #l), the observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (animal #l). Then, from the distance between its own agent #i and each task #l (animal #l), the observation information acquisition unit 110 determines a predetermined number of nearby tasks that are close to its own agent #i, and makes information about the nearby tasks part of the observation information. In the example of the above-mentioned formulas (2) to (5), the nearby tasks are animal #l1, which is closest to its own agent #i, and animal #l2, which is second closest.

[0156] In specific example 3, o_(l1)^task in equation (5) may indicate the state of animal #l1, which is a neighboring task, and the destination (goal) of animal #l1. The state of an animal may be the position and speed of the animal. The destination of animal #l1 may correspond to the original territory of the animal. The same applies to o_(l2)^task.

[0157] In the specific example 3, the policy storage unit 112 stores a policy π (learned model) that has been learned by reinforcement learning, as in the above-described embodiment. The policy π is learned for each agent 10. Policy π of agent #i NN,i is the above-mentioned observation information i is used as input, and the action a shown in the above equation (12) i As in the first embodiment, the policy π NN,i is the reward r shown in equation (14) above. i It is trained to maximize (t).

[0158] Here, in the specific example 3, P regarding the achievement level of task #l l (t) may be the distance to the boundary of the area to be protected from animals (the area to which animals should not be allowed to enter) at time t. Also, in specific example 3, Ql (t) may be the moving speed of animal #l towards the goal at time t.

[0159] Also, in Specific Example 3, similar to the above-described embodiment, the request response processing unit 130 calculates the request parameter d NN,i for agent #i and the response parameter σ i from the action a i output from π (S330). Then, the request response processing unit 130 transmits the request parameter d i to the control device 100 of the surrounding agent 10. At this time, similar to the modification examples of Embodiment 1 and Embodiment 2, the control device 100 may transmit the request parameter to the terminal carried by a human. Thereby, the human may cause the animal to withdraw. i

[0160] Also, in Specific Example 3, similar to the above-described embodiment, the importance processing unit 140 updates (calculates) the importance φ i l of the surrounding task #l (animal #l) related to its own agent #i (S340, S342). The importance processing unit 140 may update the importance φ i l of each task #l (animal #l) related to its own agent #i using the above formulas (23) and (24). Also, in Specific Example 3, similar to the above-described embodiment, the task selection unit 150 selects, according to the above formula (26), the task #l (animal #l) with the maximum importance φ i l among all tasks #l as the task #l i * that its own agent #i should execute (S350).

[0161] In addition, when the plurality of agents 10 are composed of agents 10 of different types, similar to Specific Example 1, the request response processing unit 130 may transmit request parameters to the control device 100 of an agent #i of a type different from its own agent #i. As a result, the possibility that an agent 10 of a type different from its own agent #i reaches the animal increases. Thus, similar to Specific Example 1, when an agent #i having a function of searching for an animal detects the animal, the possibility that an agent 10 having a function of driving the animal out of the farm reaches the animal increases.

[0162] Also, in Specific Example 3, the task execution unit 160 performs control so that its own agent #i executes the selected task #l (control of the animal) (S360). Specifically, the task execution unit 160 moves the agent #i to the position of the task #l i * (animal #l i * ). The task execution unit 160 controls its own agent #i so as to execute the departure of the animal (driving away of the animal). Then, the control device 100 of each agent 10 performs processing so that the departure (driving away) is executed for all animals.

[0163] <Specific Example 4 of Embodiment 3> In Example 4, the method according to the present embodiment is applied to the provision of various services. In Example 4, a plurality of agents 10, which are machines such as robots, support people living in an environment. This can improve the comfort level of the people. Specifically, the plurality of agents 10 patrol the environment and perform processing to meet people's requests. The agents 10 patrol the environment and solve tasks requested by people. Also, in Example 4, when a task is requested, the control device 100 of the agent 10 that requested the task transmits information about the requested task to the control devices 100 of nearby agents 10. Note that in Example 4, the task 50 is a "requested task." Also, in Example 4, the goal of the task 50 is to solve the requested task. The greater the number of agents 10, the higher the likelihood that the goal of the task 50 will be achieved.

[0164] In addition, in specific example 4, the multiple agents 10 may have the function of dealing with the requested task. Also, as in specific example 1, the multiple agents 10 may have different functions from each other. For example, if the requested task is "removal of bulky waste," an agent 10 capable of transporting bulky waste may remove the bulky waste. Also, if the requested task is "apprehending a criminal," an agent 10 capable of apprehending a criminal may apprehend the criminal. Also, if the requested task is "helping a lost person," an agent 10 capable of providing directions may help the lost person. In the following description, the "requested task" may be simply referred to as a "task."

[0165] In the specific example 4, the observation information acquisition unit 110 acquires the observation information o as shown in equation (5) in the same manner as in the above-described embodiment. iObtain it (S310). At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and task #l (task #l) (S302). Specifically, similar to the first embodiment, the observation information acquisition unit 110 acquires the position of "requested task" which is task #l. The position of the "requested task" may be the position of agent 10 that requested the task at the time of request. The observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (task #l) from the acquired position of task #l (task #l). Then, the observation information acquisition unit 110 determines a predetermined number of neighboring tasks close to its own agent #i from the distances between its own agent #i and each task #l (task #l), and uses the information about the neighboring tasks as part of the observation information. In the examples of the above-described formulas (2) to (5), the neighboring tasks are task #l1 which is the closest to its own agent #i and task #l2 which is the second closest.

[0166] Also, in Specific Example 4, o_(l1)^task in formula (5) may indicate the state of task #l1 which is a neighboring task and the end condition of the task. The state of each task may be the position of the task, the degree of solution of the task, the quality of the task, and the type of the task. The end condition of the task may be that the task is solved. The same applies to o_(l2)^task.

[0167] In Specific Example 4, the policy storage unit 112 stores the policy π (learned model) learned by reinforcement learning in the same manner as in the above-described embodiment. The policy π is learned for each agent 10. The policy π of agent #i NN,i takes the above-described observation information o i as input and outputs the action a i shown in the above formula (12). Also, similar to the first embodiment, the policy π NN,i is learned to maximize the reward r i (t) shown in the above formula (14).

[0168] Here, in Specific Example 4, P regarding the achievement degree of task #l l(t) may indicate the degree of achievement of problem solving for problem #l at time t. The degree of achievement of problem solving may correspond, for example, to the completion of the response by agent 10 capable of dealing with the problem. Also, in Specific Example 4, Q regarding the progress of task #l l (t) may be the progress of dealing with the problem for problem #l at time t. The progress of dealing with the problem may correspond, for example, to the fact that agent 10 capable of dealing with the problem is dealing with it.

[0169] Also, in Specific Example 4, the request response processing unit 130, in the same manner as in the above-described embodiment, from the policy π NN,i output action a i calculates request parameter d i and response parameter σ i for agent #i (S330). Then, the request response processing unit 130 transmits the request parameter d i to the control device 100 of the surrounding agent 10. At this time, like the modified examples of Embodiment 1 and Embodiment 2, the control device 100 may transmit the request parameter to the terminal carried by a human. Thereby, a human may deal with the problem. Also, like Specific Example 1, the control device 100 may transmit the request parameter to the control device 100 of an agent 10 of a type different from its own agent #i.

[0170] Also, in Specific Example 4, the importance processing unit 140, in the same manner as in the above-described embodiment, updates (calculates) the importance φ i l of the surrounding task #l (problem #l) regarding its own agent #i (S340, S342). The importance processing unit 140 may update the importance φ i l of each task #l regarding its own agent #i using the above formulas (23) and (24). Also, in Specific Example 4, the task selection unit 150, in the same manner as in the above-described embodiment, according to the above formula (26), among all tasks #l, the importance φ i lSelect the task #l with the highest task #l (task #l) as the task that the agent #i itself should execute i * as (S350).

[0171] In addition, when the agent #i that has searched for the task #l does not have a function to handle that task, it is assumed that the degree of request output from the policy will increase in the control device 100 of the agent #i. And the control device 100 of the agent #i, d i = 1 sends a request parameter to the control devices 100 of the surrounding agents 10. Then, among the control devices 100 of the surrounding agents 10 that can handle the task, when the observation information indicating the quality and type of the task and the above request parameter are acquired, it is possible that the importance of the task #l will increase. Therefore, in the control device 100 of the agent 10 that can handle the task, the task #l is selected, and the possibility that the agent 10 that can handle the task reaches the position of the task #l becomes high. If information on which tasks have not been completed is added to the observation information, it becomes possible for the agent side that has acquired the observation information to determine whether to actively respond to the request for support for that task.

[0172] Also, in Specific Example 4, the task execution unit 160 controls so that the selected task #l (task #l) is executed by the agent #i itself (S360). Specifically, the task execution unit 160 moves the agent #i to the position of the task #l i * (task #l i * ). The task execution unit 160 controls its own agent #i so that the agent #i executes the handling of tasks that the agent #i can execute. And the control device 100 of each agent 10 performs processing so that all the requested tasks are solved.

[0173] <Specific Example 5 of Embodiment 3> In Specific Example 5, the method according to this embodiment is applied to event handling. In Specific Example 5, a plurality of agents 10, which are machines such as robots, control the flow of people in an event. Specifically, the plurality of agents 10 search for the flow of people (crowd) to be guided and guide the flow of people to a position where they should be stopped, thereby organizing the flow of people (crowd). More specifically, for example, a plurality of agents 10 grip a single rope, and by moving each agent 10 to a predetermined position, the area can be divided by the rope. Thereby, the flow of people is organized by guiding the flow of people to the divided area. Also, in Specific Example 5, the control device 100 of the agent 10 that has searched for the flow of people to be guided may transmit information regarding the searched flow of people to the control devices 100 of the surrounding agents 10.

[0174] Here, in Specific Example 5, the task 50 is "a group of people (or simply 'the flow of people')". Also, in Specific Example 5, the goal of the task 50 is to guide the flow of people to a position where they should be stopped. Note that when many agents 10 organize the flow of people, the variations in the division increase and the size of the divided area increases. Therefore, the higher the number of agents 10, the higher the possibility that the goal of the task 50 is achieved.

[0175] In Specific Example 5, the observation information acquisition unit 110, similar to the above-described embodiment, acquires observation information o as shown in Equation (5). iObtain it (S310). At this time, the observation information acquisition unit 110 calculates the distance between its own agent #i and task #l (human flow #l) (S302). Specifically, similar to the first embodiment, the observation information acquisition unit 110 acquires the position of the human flow that is task #l. The observation information acquisition unit 110 calculates the distance between its own agent #i and each task #l (human flow #l) from the acquired position of task #l (human flow #l). Then, the observation information acquisition unit 110 determines a predetermined number of neighboring tasks close to its own agent #i from the distances between its own agent #i and each task #l (human flow #l), and uses the information on the neighboring tasks as part of the observation information. In the examples of the above-described formulas (2) to (5), the neighboring tasks are the human flow #l1 closest to its own agent #i and the human flow #l2 second closest.

[0176] Also, in Specific Example 5, \(o_{(l1)}^{task}\) in formula (5) may indicate the state of the human flow #l1, which is a neighboring task, and the goal (position to be stopped) of the human flow #l1. The state of the human flow may be the position and movement speed of the human flow. The same applies to \(o_{(l2)}^{task}\).

[0177] In Specific Example 5, the policy storage unit 112 stores the policy π (learned model) learned by reinforcement learning in the same manner as in the above-described embodiment. The policy π is learned for each agent 10. The policy π of agent #i NN,i takes the above-described observation information o i as input and outputs the action a i shown in the above formula (12). Also, similar to the first embodiment, the policy π NN,i is learned to maximize the reward r i (t) shown in the above formula (14).

[0178] Here, in Specific Example 5, \(P_{(l)}(t)\) regarding the degree of achievement of task #l l may indicate the presence or absence of reaching the goal (position to be stopped) of the human flow #l at time t. Also, in Specific Example 5, \(Q_{(l)}(t)\) regarding the progress of task #l l(t) may be the moving speed of the crowd flow #l towards the goal at time t.

[0179] Also, in Specific Example 5, similar to the above-described embodiment, the request response processing unit 130 calculates the request parameter d NN,i for the agent #i and the response parameter σ i from the action a i output from the policy π i (S330). Then, the request response processing unit 130 transmits the request parameter d i to the control device 100 of the surrounding agent 10. At this time, similar to the modification examples of Embodiment 1 and Embodiment 2, the control device 100 may transmit the request parameter to the terminal carried by a human. Thereby, a human may perform the arrangement of the crowd flow.

[0180] Also, in Specific Example 5, similar to the above-described embodiment, the importance processing unit 140 updates (calculates) the importance φ i l of the surrounding task #l (crowd flow #l) related to its own agent #i (S340, S342). The importance processing unit 140 may update the importance φ i l of each task #l (crowd flow #l) related to its own agent #i using the above equations (23) and (24). Also, in Specific Example 3, the task selection unit 150 selects, in the same manner as the above-described embodiment, the task #l (crowd flow #l) with the maximum importance φ i l from among all the tasks #l as the task #l i * that its own agent #i should execute (S350).

[0181] Also, in Specific Example 5, the task execution unit 160 controls so that its own agent #i executes the selected task #l (arrangement of the crowd flow) (S360). Specifically, the task execution unit 160 moves the agent #i to the task #l i * (crowd flow #l i *) position. The task execution unit 160 controls its own agent #i to execute people flow management (guiding people flow to a position where they should be stopped). Then, the control device 100 of each agent 10 performs processing so that management (guiding) is executed for all people flows.

[0182] (Variation) It should be noted that the present embodiment is not limited to the above embodiment, and can be appropriately modified within the scope of the gist. For example, the order of each step (process) in the above flowchart can be appropriately changed. Furthermore, one or more of each step (process) in the above flowchart can be omitted.

[0183] In the above-described embodiment, the agent 10 and the task 50 exist in a real space, but this is not limiting. The agent 10 and the task 50 may exist in a virtual space realized by a simulation, for example.

[0184] The above-mentioned program includes a set of instructions (or software code) that, when loaded into a computer, causes the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disk (DVD), Blu-ray® disk or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals. [Explanation of symbols]

[0185] 1 Control system 10 Agent 50 Task 60 Monitoring device 100 Control device 110 Observation information acquisition unit 112 Policy storage unit 120 Action output unit 130 Request response processing unit 140 Importance processing unit 150 Task selection unit 160 Task execution unit

Claims

1. A control device for controlling an agent that executes a task, wherein the task is such that the higher the number of agents executing the task, the higher the likelihood of achieving the goal of the task, and a plurality of the tasks exist in the environment, a request response processing unit that calculates a request parameter regarding whether to request assistance and a response parameter regarding whether to respond to a request from another agent based on observation information regarding the agent, other agents around the agent, and the task, performs processing for calculating the importance of each of the tasks regarding the agent based on at least the request parameter of other agents and the response parameter of the agent, inputs the observation information into a policy learned for each agent, and calculates the importance of the task corresponding to the observation information regarding the agent based on a target value of the importance of the task corresponding to the observation information output from the policy, when the response parameter indicates that the agent does not respond to a request from another agent, calculates the importance of the task corresponding to the observation information regarding the agent so that the importance of the task being executed by the agent approaches the target value, when the response parameter indicates that the agent responds to a request from another agent and the request parameter of the other agent indicates a request for assistance, calculates the importance of the task corresponding to the observation information regarding the agent so that the importance of the task being executed by the other agent increases based on the difference between the importance of the task being executed by the other agent and the importance of the task regarding the agent, an importance processing unit; a task selection unit that selects the task to be executed by the agent according to the importance; a task execution unit that controls the agent to execute the selected task; A control device having the above.

2. The request response processing unit calculates the request parameter and the response parameter based on the policy learned for each agent. The control device according to claim 1.

3. The request response processing unit calculates the request parameter and the response parameter based on the degree of request and the degree of response output from the policy by inputting the observation information into the policy. The control device according to claim 2.

4. When the degree of request exceeds a predetermined threshold value and the task being executed or about to be executed by the agent has not progressed, the request response processing unit calculates the request parameter indicating a request for support. The control device according to claim 3.

5. When the degree of response exceeds a predetermined threshold value and the task being executed or about to be executed by the agent has not progressed, the request response processing unit calculates the response parameter indicating a response to the request. The control device according to claim 3.

6. A control system for dispersedly controlling a plurality of agents that execute tasks, In the task, the higher the number of agents executing the task, the higher the possibility of achieving the goal of the task, and there are a plurality of the tasks in the environment. The control system includes a plurality of control devices that respectively control a plurality of agents. Each of the plurality of control devices A request response processing unit that calculates a request parameter regarding whether to request support and a response parameter regarding whether to respond to a request from another agent based on the agent related to the control device and observation information regarding other agents around the agent and the task; Performs processing for calculating the importance of each of the tasks related to the agent based on at least the request parameter of other agents and the response parameter of the agent. The observation information is input into the policy learned for each agent, and the importance of the task corresponding to the observation information is calculated based on the target value of the importance of the task corresponding to the observation information output from the policy. When the response parameter indicates that the agent does not respond to a request from another agent, the importance of the task corresponding to the observation information regarding the agent is calculated so that the importance of the task being executed by the agent approaches the target value. When the response parameter indicates a response to a request from another agent and the request parameter of the other agent indicates a request for support, based on the difference between the importance of the task being executed by the other agent and the importance of the task for the agent, calculate the importance of the task corresponding to the observation information regarding the agent so that the importance of the task being executed by the other agent increases. An importance processing unit; A task selection unit that selects the task to be executed by the agent according to the importance; A task execution unit that controls the agent to execute the selected task; Having A control system.

7. A control method for controlling an agent that executes a task, For the task, the higher the number of agents that execute the task, the higher the probability that the goal of the task is achieved. There are multiple tasks in the environment. Based on the observation information regarding the agent, other agents around the agent, and the task, calculate a request parameter regarding whether to request support and a response parameter regarding whether to respond to a request from another agent. Perform processing for calculating the importance of each task regarding the agent based on at least the request parameter of the other agent and the response parameter of the agent. Input the observation information into the policy learned for each agent, and calculate the importance of the task corresponding to the observation information regarding the agent based on the target value of the importance of the task corresponding to the observation information output from the policy. When the response parameter indicates that the agent does not respond to a request from another agent, calculate the importance of the task corresponding to the observation information regarding the agent so that the importance of the task being executed by the agent approaches the target value. When the response parameter indicates a response to a request from another agent, and when the request parameter of the other agent indicates a request for support, based on the difference between the importance of the task being executed by the other agent and the importance of the task regarding the agent, calculate the importance of the task corresponding to the observation information regarding the agent so that the importance of the task being executed by the other agent increases. Select the task to be executed by the agent according to the importance. Perform control so that the agent executes the selected task. A control method executed by a computer.

8. A program for realizing a control method for controlling an agent that executes a task, For the task, the higher the number of agents that execute the task, the higher the possibility that the goal of the task is achieved. There are multiple tasks in the environment. Based on the agent, the observation information regarding the agent, other agents around the agent, and the task, calculate a request parameter regarding whether to request support and a response parameter regarding whether to respond to a request from another agent. Perform processing for calculating the importance of each task regarding the agent based on at least the request parameter of the other agent and the response parameter of the agent. Input the observation information into the policy learned for each agent, and calculate the importance of the task corresponding to the observation information regarding the agent based on the target value of the importance of the task corresponding to the observation information output from the policy. When the response parameter indicates that the agent does not respond to a request from another agent, calculate the importance of the task corresponding to the observation information regarding the agent so that the importance of the task being executed by the agent approaches the target value. indicating that the response parameter responds to a request from another agent, and when the request parameter of the other agent indicates a request for support, calculating the importance of the task corresponding to the observation information regarding the agent so that the importance of the task being executed by the other agent is increased based on the difference between the importance of the task being executed by the other agent and the importance of the task regarding the agent; selecting the task to be executed by the agent according to the importance; controlling to execute the task selected by the agent; A program for causing a computer to execute.

Citation Information

Patent Citations

  • Mobile agents for manipulating, moving, and / or reorienting components

    JP2017094122A

  • Work personnel allocation system and work personnel allocation device

    JP2021051649A

  • Control device, control method, and control system

    WO2019058694A1